Data caching method and apparatus, electronic device, and readable storage medium
By introducing streaming buffers and basic caches into the system-level cache and adopting a bypass prediction mechanism, the hit rate of the system-level cache is improved, and the problem of low hit rate in the existing technology is solved, and the effect of reducing the number of memory accesses and reducing the latency and power consumption of the memory system is achieved.
Patent Information
- Application Number
- PCT/CN2024/093751
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-05-16
- Publication Date
- 2025-06-19
AI Technical Summary
In the prior art, the hit rate of system-level cache is low, resulting in the performance of computer systems being restricted by storage walls, and it is impossible to effectively reduce the number of accesses to memory and the power consumption of storage subsystems.
By introducing streaming buffers and basic caches into the system-level cache and using a bypass prediction mechanism, bypass prediction of memory access requests is written to the streaming buffer, and data blocks that do not need to be bypassed are written to the basic cache.
Improves the hit rate of the system-level cache, reduces the number of accesses to memory, and reduces the latency and power consumption of the memory system.
Smart Images

Figure CN2024093751_19062025_PF_FP_ABST
Abstract
Description
Data caching method, device, electronic device and readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 11, 2023, with application number 202311686437.0 and titled “A data caching method, device, electronic device and readable storage medium,” the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a data caching method, device, electronic device, and readable storage medium. Background Art
[0004] In recent years, the performance of computing components such as central processing units (CPUs) and graphics processing units (GPUs) has rapidly improved. Meanwhile, while memory speed has increased, the rate of improvement has far outstripped the performance gains of processors. This has led to a widening performance gap between computing and storage components. While waiting for data to be delivered by storage, computing components are forced to pause, wasting their performance. This phenomenon, known as the memory wall, has become a bottleneck restricting the performance of various computer systems.
[0005] To mitigate the negative impact of the memory wall on computer system performance, the concept of a memory hierarchy has been proposed in computer architecture. The memory hierarchy divides various storage components (including registers, cache, memory, and hard disks) into different tiers based on operating speed and unit cost. Storage components closer to the processor have faster operating speeds, smaller capacity, and higher unit cost. Storage components closer to the memory have larger capacity, slower operating speeds, and lower unit cost.
[0006] In the storage hierarchy, cache is an effective way to mitigate the impact of the memory wall. Cache typically operates slower than the processor but faster than memory. Thanks to the principle of program locality, copying recently accessed data to the faster cache effectively reduces memory access time, thereby masking the speed gap between the processor and memory. Furthermore, given that the memory subsystem accounts for a significant portion of total system power consumption, retrieving data from the cache reduces the number of memory accesses, thereby reducing memory power consumption. Therefore, cache has the ability to reduce both access latency and storage subsystem power consumption.
[0007] Current computer microarchitecture designs utilize a large, shared system cache (system cache) across multiple devices as the last line of defense for the entire memory hierarchy, hoping to reduce memory access latency and the number of memory accesses. However, experiments have shown that directly applying upper-level cache management methods to the system cache often results in a very low hit rate.
[0008] How to improve the hit rate of system-level cache and reduce memory access is an important issue for optimizing chip system energy consumption and improving user experience.
[0009] Summary of the Invention
[0010] Embodiments of the present application provide a data caching method, device, electronic device, and readable storage medium, which can improve the hit rate of system-level cache.
[0011] The present application discloses a data caching method, which is applied to a system-level cache, wherein the system-level cache includes a streaming buffer and a basic cache. The method includes:
[0012] receiving a memory access request sent by a processor, wherein the memory access request carries a memory access address;
[0013] When the memory access request satisfies a bypass prediction condition, performing bypass prediction on a target data block corresponding to the memory access address to obtain a prediction result;
[0014] If the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed, writing the target data block into the streaming buffer;
[0015] If the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed, writing the target data block into the basic cache;
[0016] The bypass prediction condition includes any one of the following:
[0017] The memory access request is a write request;
[0018] The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache;
[0019] The memory access request satisfies a prefetch condition; the prefetch condition is used to indicate a condition that needs to be satisfied in order to prefetch a data block in the memory into the system-level cache in advance.
[0020] On the other hand, an embodiment of the present application discloses a data cache device, which is applied to a system-level cache, wherein the system-level cache includes a streaming buffer and a basic cache; the device includes:
[0021] A request receiving module, configured to receive a memory access request sent by a processor, wherein the memory access request carries a memory access address;
[0022] A bypass prediction module, configured to perform bypass prediction on a target data block corresponding to the memory access address to obtain a prediction result when the memory access request satisfies a bypass prediction condition;
[0023] a first writing module, configured to write the target data block into the streaming buffer when the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed;
[0024] a second writing module, configured to write the target data block into the basic cache if the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed;
[0025] The bypass prediction condition includes any one of the following:
[0026] The memory access request is a write request;
[0027] The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache;
[0028] The memory access request satisfies a prefetch condition; the prefetch condition is used to indicate a condition that needs to be satisfied in order to prefetch a data block in the memory into the system-level cache in advance.
[0029] On the other hand, an embodiment of the present application further discloses an electronic device, which includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the aforementioned data caching method.
[0030] An embodiment of the present application further discloses a readable storage medium. When instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the aforementioned data caching method.
[0031] The embodiments of the present application include the following advantages:
[0032] The data caching method provided by the embodiment of the present application performs bypass prediction on the target data block corresponding to the memory access address when the received memory access request meets the bypass prediction condition, and writes the target data block that needs to be bypassed into the streaming buffer, so that the bypassed data block is cached in the streaming buffer for a short period of time. Even if the data block is bypassed by mistake, it will still exist in the streaming buffer and have a chance of being hit, thereby eliminating the impact of the mistaken bypass; the target data block that does not need to be bypassed is written into the basic high-speed cache, thereby avoiding the contention of the cache between the data block that will not be hit and other data blocks that may be hit, which is beneficial to improving the cache hit rate, thereby reducing the number of memory accesses and reducing the latency and power consumption of the memory system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] FIG1 is a flowchart of a data caching method according to an embodiment of the present invention;
[0035] FIG2 is a schematic diagram of a data caching process of the present application;
[0036] FIG3 is a schematic diagram of the structure of a hit history table of the present application;
[0037] FIG4 is a schematic structural diagram of a regional prefetcher of the present application;
[0038] FIG5 is a schematic structural diagram of a data caching device of the present application;
[0039] FIG6 is a structural block diagram of an electronic device provided by an example of this application. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0041] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of this application, the term "multiple" refers to two or more, and other quantifiers are similar.
[0042] Method Example
[0043] 1 , a flowchart of an embodiment of a data caching method of the present application is shown. The method may specifically include the following steps:
[0044] Step 101: Receive a memory access request sent by a processor, where the memory access request carries a memory access address.
[0045] Step 102: When the memory access request satisfies a bypass prediction condition, a bypass prediction is performed on the target data block corresponding to the memory access address to obtain a prediction result.
[0046] Step 103: When the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed, write the target data block into the streaming buffer.
[0047] Step 104 : When the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed, write the target data block into the basic cache.
[0048] The bypass prediction condition includes any one of the following:
[0049] A1. The memory access request is a write request;
[0050] A2: The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache;
[0051] A3. The memory access request satisfies a prefetch condition; the prefetch condition is used to indicate a condition that needs to be met to prefetch the data block in the memory into the system-level cache in advance.
[0052] It should be noted that the data caching method provided in the embodiments of the present application can be applied to a system-level cache, which includes a streaming buffer and a basic cache. The system-level cache is located in the memory controller and is the last cache for the processor to access memory. The system-level cache is shared by all devices on the chip, including the CPU, GPU, and various accelerator chips.
[0053] System-level caches feature low storage hierarchies, are shared across multiple devices, and have poor locality of incoming memory accesses, requiring unique management strategies that differ from traditional cache management methods. In low storage hierarchies, due to the filtering effect of upper-level caches, incoming memory access requests have very poor locality. As a result, many cache lines are removed from the system-level cache without being hit (re-referenced) after entering it. These cache lines that enter the system-level cache but are not hit are called long-reuse blocks. Long-reuse blocks, due to their large reuse distance, cannot be hit in the system-level cache. Furthermore, they exacerbate cache contention, causing some potentially hit blocks to be removed from the system-level cache, further reducing the cache hit rate. In order to solve this problem, the embodiment of the present application uses a bypass mechanism in the system-level cache to perform bypass prediction on the target data block, and only writes the target data block that does not need to be bypassed into the basic cache. The bypassed data block will not be inserted into the basic cache, thereby avoiding the contention of the long multiplexing block and other data blocks that may be hit for the cache, and avoiding other data blocks from being replaced out of the cache because the cache line is occupied by the long multiplexing block, which is beneficial to improving the cache hit rate.
[0054] Specifically, the embodiment of the present application divides the system-level cache into two parts: a stream buffer and a base cache. The stream buffer has a smaller capacity and is used to temporarily store bypassed data blocks; the base cache has a larger capacity and is used to cache data blocks that are not bypassed. For example, a small portion of the total capacity of the system-level cache is used as the stream buffer. In the stream buffer, a first-in-first-out replacement strategy can be used to manage bypassed data blocks, and all bypassed blocks enter and exit the stream buffer in order.
[0055] In an embodiment of the present application, if the system-level cache receives a memory access request sent by the processor, it can first determine whether the memory access request meets the bypass prediction condition, and if the memory access request meets the bypass prediction condition, bypass prediction is performed on the target data block corresponding to the memory access address carried in the memory access request, and the data block to be bypassed is written into the streaming buffer, so that the bypassed data block is cached in the streaming buffer for a short period of time. Even if the data block is bypassed by mistake, it will still exist in the streaming buffer and have a chance of being hit, thereby eliminating the impact of the mistaken bypass. Data blocks that do not need to be bypassed will be written to the basic cache, and long multiplexing blocks will not be inserted into the basic cache, thereby avoiding contention for the cache between long multiplexing blocks and other data blocks that may be hit, and improving the cache hit rate.
[0056] It will be understood that the processors in the embodiments of the present application may include, but are not limited to, processing modules or processing units in a CPU, a GPU, a data processing unit (DPU), a field programmable gate array (FPGA), and an application specific integrated circuit (ASIC), etc.
[0057] The bypass prediction condition in the embodiment of the present application includes any one of A1 to A3. In other words, as long as the memory access request received at the system level meets any one of A1 to A3, it can be determined that the memory access request meets the bypass prediction condition.
[0058] Referring to Figure 2, a flow chart of a data cache provided by an embodiment of the present application is shown. As shown in Figure 2, the system-level cache includes a bypass module, a streaming buffer and a basic cache. Among them, the bypass module is used to perform bypass prediction. If the memory access request is a write request, it means that the processor wants to write data to the system-level cache. In this case, the bypass module performs bypass prediction on the target data block written by the processor, and determines whether the target data block needs to be bypassed based on the prediction result. If the target data block needs to be bypassed, the target data block is written to the streaming buffer. If the target data block does not need to be bypassed, the target data block is written to the basic cache.
[0059] If the memory access request is a read request, the system-level cache is searched for a data block that matches the memory address carried in the memory access request. If a data block that matches the memory address exists in the system-level cache, it is directly returned to the processor. If a data block that matches the memory address does not exist in the system-level cache, the processor searches for a target data block that matches the memory address in memory, reads the target data block from memory, and writes it to the system-level cache.
[0060] Optionally, the bypass prediction condition includes: the memory access request is a read request, and no data block matching the memory access address exists in the system-level cache; and in step 102, if the memory access request satisfies the bypass prediction condition, performing bypass prediction on the target data block corresponding to the memory access address to obtain a prediction result, including:
[0061] Step S11: when the memory access request is a read request and there is no data block matching the memory access address in the system-level cache, obtaining a target data block matching the memory access address from the memory;
[0062] Step S12: perform bypass prediction on the target data block to obtain a prediction result.
[0063] When the memory access request is a read request and there is no data block matching the memory access address in the system-level cache, the system-level cache needs to read the data from the memory. As shown in FIG2 , the system-level cache can perform bypass prediction on the data read from the memory, that is, the target data block in the embodiment of the present application. If the target data block needs to be bypassed, the target data block is written to the streaming buffer; if the target data block does not need to be bypassed, the target data block is written to the base cache.
[0064] If the memory access request satisfies the prefetch condition, it means that the data block in the memory can be prefetched into the system-level cache in advance. In this case, the system-level cache needs to read data from the memory. As shown in Figure 2, the system-level cache can perform bypass prediction on the data read from the memory, that is, the target data block in the embodiment of the present application, and write the data block that needs to be bypassed into the streaming buffer, and write the data block that does not need to be bypassed into the basic cache.
[0065] If the system-level cache writes data to memory, no bypass prediction is required. Similarly, if the system-level cache returns data to the processor, no processing is required.
[0066] It is understandable that the main purpose of bypass prediction is to predict whether a data block is a long multiplexing block, that is, to predict whether the data block will be hit after entering the cache. If it is predicted that the data block will not be hit after entering the cache, it can be determined that the data block is a long multiplexing block and needs to be bypassed; conversely, if it is predicted that the data block may be hit after entering the cache, it can be determined that the data block is not a long multiplexing block and does not need to be bypassed. In an embodiment of the present application, it is possible to predict whether the data block corresponding to the memory access address needs to be bypassed based on each instruction in the program segment or code segment to which the memory access request belongs; it is also possible to predict whether the data block corresponding to the memory access address needs to be bypassed based on the historical hit situation of each data block in the cache. Other bypass prediction algorithms can also be used, and this embodiment of the present application does not make specific limitations.
[0067] The data caching method provided by the embodiment of the present application performs bypass prediction on the target data block corresponding to the memory access address when the received memory access request meets the bypass prediction condition, and writes the target data block that needs to be bypassed into the streaming buffer, so that the bypassed data block is cached in the streaming buffer for a short period of time. Even if the data block is bypassed by mistake, it will still exist in the streaming buffer and have a chance of being hit, thereby eliminating the impact of the mistaken bypass; the target data block that does not need to be bypassed is written into the basic high-speed cache, thereby avoiding the contention of the cache between the data block that will not be hit and other data blocks that may be hit, which is beneficial to improving the cache hit rate, thereby reducing the number of memory accesses and reducing the latency and power consumption of the memory system.
[0068] In an optional embodiment of the present application, in step 102, when the memory access request satisfies the bypass prediction condition, performing bypass prediction on the target data block corresponding to the memory access address to obtain a prediction result includes:
[0069] Step S21: when the memory access request satisfies the bypass prediction condition, determining the historical hit status of the target data block corresponding to the memory access address;
[0070] Step S22: Perform bypass prediction on the target data block according to the historical hit situation to obtain a prediction result.
[0071] In an embodiment of the present application, bypass prediction can be performed on a target data block based on the historical hit status of the target data block. The historical hit status can reflect whether the target data block was hit in the cache when the processor accessed the cache over a period of time. The time span of the historical data recorded in the historical hit status can be set according to actual needs. For example, a period can be pre-set to record the hit status of each data block corresponding to the memory access address in the cache when the processor accesses the cache during this period.
[0072] As an example, if the historical hit situation shows that the access to the target data block in the past period of time is all cache misses, that is, when the processor wants to access the target data block, the target data block is not in the cache, then it can be considered that the interval between the processor's access to the target data block is long, and the target data block is a long multiplexed block and should be bypassed.
[0073] On the contrary, if the historical hit situation shows that the target data block has been hit in the past period of time, it can be considered that the target data block is not a long-reuse block, has a chance of being hit in the cache, and should not be bypassed.
[0074] Optionally, the system-level cache further includes a hit history table, wherein the hit history table is used to record historical hit conditions of N consecutive accesses to the data block in the past, where N is a positive integer;
[0075] Step S22 performs bypass prediction on the target data block according to the historical hit situation to obtain a prediction result, including:
[0076] Sub-step S221, querying the hit history table to see whether there is an entry matching the memory access address;
[0077] Sub-step S222: If there is no entry matching the memory access address in the hit history table, it is determined that the target data block needs to be bypassed; or,
[0078] Sub-step S223: If there is an entry matching the memory access address in the hit history table, and the target data block has not been hit N times in the past, it is determined that the target data block needs to be bypassed.
[0079] In an embodiment of the present application, according to the principle of locality, only a hit history table with a relatively small capacity is used to record the hit history of the data block that has been recently accessed by the processor in the cache. The hit history table records the historical hits of the data block in the past N consecutive accesses. Exemplarily, referring to FIG3 , a schematic diagram of the structure of a hit history table provided by an embodiment of the present application is shown. As shown in FIG3 , the hit history table has a total of 32K sets, 4-way sets are connected, and each table entry corresponds to four memory blocks with consecutive addresses, which are used to record the past hit history of four blocks with consecutive physical addresses, and use the same address high bits corresponding to these four memory blocks as address tags for indexing to reduce the storage overhead for indexing.
[0080] If a block does not have a corresponding hit history, it is considered that this block is not the most recently accessed block and meets the characteristics of a long-reuse block, so this block needs to be bypassed.
[0081] It is understandable that a block has no corresponding hit history, including that there is no table entry matching the memory access address in the hit history table, and there is a table entry matching the memory access address in the hit history table, but the target data block has not been hit in the past N consecutive accesses. Therefore, for both cases, it can be considered that the target data block needs to be bypassed.
[0082] Optionally, the method further includes:
[0083] If there is a table entry matching the memory access address in the hit history table, and the target data block has been hit at least once in the past N consecutive accesses, it is determined that the target data block does not need to be bypassed.
[0084] If there is an entry matching the memory access address in the hit history table, and the target data block has been hit at least once in the past N consecutive accesses, it can be determined that the target data block has a chance of being hit in the system-level cache and should not be bypassed.
[0085] It should be noted that in the embodiment of the present application, a target data block hitting in the basic cache or hitting in the streaming buffer can be considered a hit.
[0086] As an example, assuming N=2, and the historical hit record of a data block in a certain table entry in the hit history table is "hit history 1=1, hit history 2=0", it means that this block was hit in the cache the last time it was accessed, but was not hit in the cache the last time it was accessed. If the access address carried in the received memory access request is the same as the address of the data block, it can be considered that the data block has a chance of hitting in the cache in the next access and should not be bypassed. The data block was hit in the cache the last time it was accessed, which may be a hit in the basic cache or a hit in the streaming buffer. In this embodiment of the application, no matter which part of the system-level cache the data block is hit in, it is considered a hit in the cache.
[0087] If the historical hit record of a data block in a table entry in the hit history table is "hit history 1 = 0, hit history 2 = 0", it means that this block did not hit in the cache when it was last accessed, and did not hit in the cache when it was accessed the previous time. If the memory access address carried in the received memory access request is the same as the address of the data block, it can be considered that the data block will not hit in the cache in the next access and should be bypassed.
[0088] Optionally, the method further includes:
[0089] Step S31: If there is no entry matching the memory access address in the hit history table, then determine a target entry in the hit history table; the hit count of the data block corresponding to the target entry is less than that of the data blocks corresponding to other entries;
[0090] Step S32: Clear the entry content of the target entry, and fill the hit status of the target data block into the target entry.
[0091] If there is no entry matching the access address in the hit history table, that is, no historical hit record of the data block corresponding to the access address has been recorded in the hit history table, then a hit record of the data block corresponding to the access address can be added to the hit history table. Specifically, a target entry in the hit history table whose hit count is less than the hit count of other data blocks is determined, the entry content of the target entry is cleared, and the hit record of the target data block corresponding to the access address is filled in the target entry.
[0092] It can be understood that in the hit history table shown in Figure 3, an item can be replaced in the same group of the hit history table. The replacement method is "replace the table item with the least number of '1's in the same group", reset all records of the item to 0, and then re-fill the records in the table item based on the hit situation of the target data block in this access. If this access hits in the cache, "Hit History 1" is updated to "1". If this access does not hit in the cache, "Hit History 1 = 0" is maintained.
[0093] In an optional embodiment of the present application, the system-level cache further includes a bypass history table, wherein the bypass history table is used to record address tags of data blocks bypassed in the system-level cache within a historical period; and the method further includes:
[0094] Step S41: If a data block matching the memory access address exists in the system-level cache, and / or an entry matching the memory access address exists in the bypass history table, it is determined that the memory access request hits the cache;
[0095] Step S42: If there is no data block matching the memory access address in the system-level cache, and there is no entry matching the memory access address in the bypass history table, it is determined that the memory access request misses the cache.
[0096] Step S43: Update the corresponding entry in the hit history table according to the memory access address.
[0097] It is understandable that the streaming buffer can only temporarily store some bypassed blocks. After the bypassed blocks are replaced in the streaming buffer, they cannot generate hits in the system-level cache. Therefore, the hits of these blocks cannot be recorded in the hit history table, resulting in these data blocks being bypassed and unable to enter the cache. To solve this problem, the embodiment of the present application uses a bypass history table to record the address tags of the data blocks that were bypassed in the historical period in the cache, where the historical period can be set according to actual needs. For example, the historical period is M clock cycles before the current moment, where M is a positive integer. Through the bypass history table, the most recently bypassed blocks can be determined.
[0098] The bypass history table has the same number of sets as the system-level cache, and their associativity is the same as the system-level cache. A set in the cache refers to a collection of cache lines, and an entry refers to the number of cache lines contained in each set, also known as associativity. The bypass history table records the address tags of the most recently bypassed blocks in each set. Using a first-in, first-out replacement method, the address tags of bypassed blocks are inserted at the end of the queue, removing the block at the head of the queue.
[0099] Each time the system-level cache receives a memory access request, it needs to update the hit history table. Specifically, the system-level cache and the bypass history table are checked to see whether there is a data block that matches the memory access address of the current memory access request. If a data block that matches the memory access address is found in the system-level cache and / or the bypass history table, it is determined that the memory access request hits the cache, and the hit status of the most recent access to the data block recorded in the hit history table can be updated to "hit". For example, in the hit history table shown in Figure 3, the "hit history 1" of the data block whose address tag matches the memory access address is updated to "1".
[0100] If a data block matching the memory access address is not found in the system-level cache and the bypass history table, it is determined that the memory access request did not hit the cache, and the hit status of the most recent access to the data block recorded in the hit history table can be updated to "miss". For example, in the hit history table shown in Figure 3, the "hit history 1" of the data block whose address tag matches the memory access address is updated to "0".
[0101] In an optional embodiment of the present application, the bypass prediction condition includes that the memory access request satisfies a prefetch condition; and in step 102, when the memory access request satisfies the bypass prediction condition, performing bypass prediction on the target data block corresponding to the memory access address to obtain a prediction result includes:
[0102] When the memory access request meets the prefetch condition, bypass prediction is performed on each prefetched target data block to obtain a prediction result for each data block; the target data block includes a data block in the preset memory area that matches the memory access address, and other data blocks in the preset memory area.
[0103] Prefetching involves learning the processor's memory access patterns to predict upcoming blocks. If the blocks are not already in the cache, they are prefetched from memory into the cache. This effectively increases cache hit rates and reduces memory access latency.
[0104] In the embodiments of the present application, when performing bypass prediction, prefetch blocks and actual memory access blocks are treated equally. If the memory access request meets the prefetch conditions, after prefetching the target data block in the preset memory area into the system-level cache, bypass prediction is also performed on each prefetched target data block. If the prediction result indicates that the prefetched target data block (hereinafter referred to as the "prefetch block") needs to be bypassed, the prefetch block is written to the streaming buffer; if the prediction result indicates that the prefetch block does not need to be bypassed, the prefetch block is written to the base cache.
[0105] In the embodiment of the present application, even if a prefetch block is bypassed, the prefetch block will enter the streaming buffer and will not be discarded. The prefetch block stored in the streaming buffer still has a chance to be used, thereby achieving coordinated work between prefetching and bypassing.
[0106] Furthermore, during prefetching, a further check can be performed to determine whether the block to be prefetched already exists in the base cache. If so, a prefetch request is not issued again. Instead, the base cache's replacement policy is notified to advance the block's position in the LRU (Least Recently Used) stack to achieve protection. If the prefetched block hits the base cache or the streaming buffer (the first use of the prefetched block), the replacement policy is notified not to update its position in the LRU stack. The LRU stack stores the page numbers of each currently used memory page. When a new process accesses a page, that page number is pushed to the top of the stack, and other page numbers are moved to the bottom. If the stack is full, the page number at the bottom is removed. This ensures that the top of the stack always contains the page number of the most recently accessed page, while the bottom contains the page number of the least recently accessed page.
[0107] In an optional embodiment of the present application, the bypass prediction condition includes that the memory access request satisfies a prefetch condition; and when the memory access request satisfies the bypass prediction condition, performing bypass prediction on the target data block corresponding to the memory access address to obtain a prediction result includes:
[0108] Step S51: if the memory access address belongs to a preset memory area, determining that the memory access request meets a prefetch condition;
[0109] Step S52: determining a first memory access bitmap of the preset memory area according to first historical memory access information of the preset memory area; the first memory access bitmap is used to record access conditions of each data block in the preset memory area within a historical period;
[0110] Step S53: if the first memory access bitmap satisfies a first condition, determining a target data block to be prefetched in the preset memory area according to the first memory access bitmap;
[0111] Step S54: perform bypass prediction on the target data block to obtain a prediction result.
[0112] The first condition includes at least one of the following:
[0113] The number of first flag bits included in the first memory access bitmap is greater than or equal to a first preset value; the first flag bit is used to indicate that the data block has been accessed within a historical period;
[0114] The credibility score of the first memory access bitmap is greater than or equal to a second preset value.
[0115] In an embodiment of the present application, a first memory access bitmap of a preset memory area may be determined based on first historical memory access information of the preset memory area, and then a target data block to be prefetched may be determined based on the first memory access bitmap.
[0116] The preset memory area may be any predetermined memory area, for example, the preset memory area may be a memory page. The first memory access bitmap is used to record the access status of each data block in the preset memory area in a historical period.
[0117] Due to the filtering effect of the upper multi-level cache, it is very difficult to find a certain stride pattern from the memory access requests reaching the system-level cache. Therefore, many existing stride-based prefetchers (including BOP, SPP, VLDP, etc.) cannot function well. There is a strong "regional" feature in the memory access requests reaching the system-level cache, that is, several specific data blocks in a memory area will be repeatedly accessed in a short period of time. Based on this, an embodiment of the present application provides a composite regional prefetcher that can prefetch multiple data blocks in a preset memory area into the system-level cache at the same time.
[0118] Specifically, the regional prefetcher may include an intra-page prefetcher and an inter-page prefetcher. The startup order of the two prefetchers is to start the intra-page prefetcher first. When the intra-page prefetcher does not issue a prefetch due to a lack of historical memory access information or the learned memory access features are not obvious, the inter-page prefetcher is started for prefetching. Exemplarily, the intra-page prefetcher can learn the memory access features on memory page A, and then use the features to generate a prefetch request for memory page A; the inter-page prefetcher learns the memory access features of memory pages B, C, D, etc. (memory pages B, C, D are adjacent to memory page A in physical address), and then uses them to generate a prefetch request for memory page A.
[0119] If the memory access address carried in the memory access request reaching the system-level cache belongs to a preset memory area, the in-page prefetcher in the system-level cache starts working: determining a first memory access bitmap based on the first historical memory access information of the preset memory area, and further determining whether the first memory access bitmap meets the first condition.
[0120] In an embodiment of the present application, if the number of first flag bits included in the first memory access bitmap is greater than or equal to a first preset value, it indicates that multiple data blocks in the preset memory area have been accessed within a historical period. The more blocks accessed in a memory area, the more accurate the memory access bitmap corresponding to the memory area, and the higher the accuracy of the prefetch issued based on this memory access bitmap. Therefore, in order to improve the accuracy of prefetching, the memory access request can be considered to meet the first condition and a prefetch can be issued only when the number of first flag bits included in the first memory access bitmap is greater than or equal to the first preset value.
[0121] Similarly, a larger credibility score for a memory access bitmap indicates a more accurate memory access bitmap, and a higher accuracy prefetch issued based on the memory access bitmap. Therefore, a memory access request may be considered to meet the prefetch condition and a prefetch may be issued only if the credibility score of the first memory access bitmap for the preset memory region is greater than or equal to a second preset value.
[0122] It is understood that both the first preset value and the second preset value can be set according to actual needs, and the embodiments of the present application do not specifically limit this. For example, if the preset memory area is a memory page, assuming that the memory page has four memory blocks, the first preset value can be 3 or 4. The second preset value can be 80%, 90%, etc.
[0123] In the embodiment of the present application, the target data block can be all the data blocks in the preset memory area, or can be the data block whose flag bit is the first flag bit in the first memory access bitmap. Optionally, when the first memory access bitmap satisfies the first condition, determining the target data block to be prefetched in the preset memory area according to the first memory access bitmap includes: when the first memory access bitmap satisfies the first condition, determining the data block whose flag bit is the first flag bit in the first memory access bitmap as the target data block.
[0124] Finally, bypass prediction is performed on each pre-fetched target data block, and the target data block that needs to be bypassed is inserted into the streaming buffer, and the target data block that does not need to be bypassed is inserted into the basic cache.
[0125] As an example, performing bypass prediction on the target data block to obtain a prediction result includes:
[0126] Step S61: determining the memory address of each target data block according to the starting address of the preset memory area and the address offset of the target data block; the address offset is used to indicate the offset value of the memory address of the target data block relative to the starting address of the preset memory area;
[0127] Step S62: Perform bypass prediction on the target data block according to the memory address of the target data block to obtain a prediction result.
[0128] When performing bypass prediction on each pre-fetched target data block, the memory address of each target data block can be determined based on the starting address of the preset memory area and the address offset of the target data block. For example, assuming that the preset memory area is a memory page, the page number of the memory page is PN, the size of a block in the memory page is 64 bytes, and the flag bit in the first memory access bitmap is 0, 1, or 4, then the memory addresses of the pre-fetched target data blocks are PN, PN+64, and PN+4×64, respectively.
[0129] Next, a bypass prediction is performed on the target data block based on the memory address of the target data block. For example, if there is no table entry matching the memory address of the target data block in the hit history table, it can be considered that the target data block needs to be bypassed; or, if there is a table entry matching the memory address of the target data block in the hit history table, but the target data block has not been hit N times in the past, it can be considered that the target data block needs to be bypassed. If there is a table entry matching the memory address of the target data block in the hit history table, and the target data block has been hit at least once in the past N consecutive accesses, it can be considered that the target data block should not be bypassed.
[0130] In the embodiment of the present application, each prefetch block can be accurately located by using the memory access bitmap, thereby improving the prefetch accuracy.
[0131] Optionally, the system-level cache also includes an accumulation table and a pattern table; wherein the accumulation table is used to record the access status of each data block in the preset memory area within a first historical period; the pattern table is used to record the access status of each data block in the preset memory area within a second historical period, and the second historical period is greater than the first historical period.
[0132] The step S51 of determining a first memory access bitmap of the preset memory area according to the first historical memory access information of the preset memory area includes:
[0133] Sub-step S511, querying whether there is a first table entry matching the memory access address in the accumulation table;
[0134] Sub-step S512: If a first entry matching the memory access address exists in the accumulation table, updating the flag bit of the target data block corresponding to the memory access address in the first entry to a first flag bit, and updating the timestamp of the first entry according to the access time of the memory access request;
[0135] Sub-step S513: determining a second entry in the accumulation table according to the timestamps corresponding to the respective entries, wherein the timestamp of the second entry is greater than the timestamps of the other entries in the accumulation table;
[0136] Sub-step S514: adding the content of the second entry to the pattern table, and clearing the content of the second entry from the accumulation table;
[0137] Sub-step S515: determining a first memory access bitmap of the preset memory area according to the accumulation table and the mode table.
[0138] The intra-page prefetcher in the system-level cache may include three parts: an accumulation table (AT), a pattern table (PT), and a filter table (FT). Among them, AGT is used to record the access status of each data block in the preset memory area in the first historical period, and PT is used to record the access status of each data block in the preset memory area in the second historical period, and the second historical period is greater than the first historical period. In an embodiment of the present application, AT is used to observe the memory access bitmap (Pattern) of a certain page, which is to complete the construction and learning of the Pattern. In this sense, AT is the source of PT. To be more precise, AT reflects the memory access information in the recent short period of time, while PT reflects the stable memory access information over a long period of time in the past. The filter table is used to record the memory access information of the data block, and then filter out the data that needs to be filled in the accumulation table based on the recorded memory access information.
[0139] Taking the preset memory area as a memory page as an example, referring to FIG4 , a schematic diagram of the structure of an intra-page prefetcher provided by an embodiment of the present application is shown. As shown in FIG4 , each preset memory area in the accumulation table corresponds to a row of table entries, which records the page number (PN) of the memory page, the memory access bitmap (Pattern), and the last access time (Last Access Time) of the memory page. The pattern table records the page number (PN) and the memory access bitmap (Pattern) of the memory page. The page number (PN) of the memory page, the address offset (offsets) of each block in the memory page, the last access time (Last Access Time) and the position (Position) recorded in the filtering table, wherein the address offset refers to the offset value of the memory block relative to the starting address of the memory page, and the position refers to the position of the last accessed memory block in the memory page.
[0140] In an embodiment of the present application, when a memory access request arrives at the system-level cache, the in-page prefetcher can first query whether there is a first table entry that matches the memory access address in the cumulative table based on the memory access address carried in the memory access request. If the first table entry that matches the memory access address is found in the cumulative table, the first table entry is updated, that is, the flag bit of the target data block corresponding to the memory access address in the first table entry is updated to the first flag bit, and the timestamp of the first table entry is updated according to the access time. Exemplarily, the flag bit corresponding to the target data block in the memory access bitmap of the first table entry in the cumulative table shown in Figure 4 can be updated to "1", and the last access time of the first table entry is updated to the access time corresponding to the memory access request, such as the time when the system-level cache receives the memory access request.
[0141] Next, based on the timestamps corresponding to the entries, such as the last access time recorded in the accumulation table shown in FIG4 , the second entry in the accumulation table having a timestamp greater than that of the other entries is determined. The second entry is the entry in the accumulation table that has not been accessed the longest.
[0142] The content of the second entry is added to the pattern table, and the content of the second entry is cleared from the accumulation table. As shown in FIG4 , the page number (PN) and access bitmap (Pattern) recorded in the second entry are added to the pattern table.
[0143] Finally, based on the accumulation table and the pattern table, the first memory access bitmap of the preset memory area can be determined. As an example, if there is no table entry corresponding to the preset memory area in the pattern table, that is, the access situation corresponding to the preset memory area has not been replaced from the accumulation table to the pattern table, in this case, the first memory access bitmap of the preset memory area can be determined based on the access situation recorded in the accumulation table, such as the memory access bitmap (Pattern). If there is a table entry corresponding to the preset memory area in the pattern table, the first memory access bitmap of the preset memory area can be determined directly based on the access situation recorded in the pattern table, such as the memory access bitmap (Pattern).
[0144] It is understandable that, in general, prefetching can only be used to improve hit rate and reduce latency, but cannot be used to optimize power consumption. However, the regional prefetcher provided in the embodiment of the present application has the potential to reduce power consumption, which is mainly due to the "aggregation effect" of the regional prefetcher. The aggregation effect refers to the regional prefetcher aggregating multiple scattered accesses to a continuous memory area and sending them all to the memory at once in a short period of time. Since a continuous memory area is often mapped to the same row of the same memory bank, the aggregation effect helps to improve the row hit rate of the memory bank, thereby reducing the number of row activations. Furthermore, the memory controller will enter a low power state after not receiving a memory access request for a period of time. The aggregation effect allows multiple requests that are dispersed in time to be processed in a concentrated short period of time, giving the memory controller more opportunities to enter low power mode.
[0145] Optionally, the system-level cache further includes a filter table; and the method further includes:
[0146] Step S71: If there is no first table entry matching the memory access address in the accumulation table, record the access information of the memory access address in the screening table;
[0147] Step S72: If there is at least one third entry in the filter table and the third entry has recorded M accesses, then add the entry content of the third entry to the accumulation table and clear the entry content of the third entry in the filter table; M is a positive integer.
[0148] In an embodiment of the present application, if there is no first table entry matching the memory access address in the accumulation table, the access information of the memory access address can be recorded in the filtering table, such as the page number of the memory page to which the memory access address belongs, the address offset of the memory block corresponding to the memory access address, the access time, the position of the accessed memory block in the memory page, and so on.
[0149] If there is a third table entry in the filter table that has recorded M accesses, for example, M memory blocks have been accessed in the memory page recorded by the third table entry, and / or the third table entry records the address offsets of the M accessed memory blocks, then the table content of the third table entry can be added to the accumulation table, and the table content of the third table entry can be cleared in the filter table. As an example, the table content of the third table entry in the filter table can be converted into the form of "address tag, memory access bitmap" and added to the accumulation table. As shown in Figure 4, the page number (PN) and the last access time recorded in the third table entry can be directly added to the accumulation table, and the accessed data blocks can be determined based on the address offset, and the corresponding memory access bitmap (Pattern) can be generated in the accumulation table, and the flag bit of the accessed data block is the first flag bit.
[0150] In an optional embodiment of the present application, the memory access bitmap includes a first bitmap of the preset memory area in the accumulation table and a second bitmap of the preset memory area in the mode table; the method further includes:
[0151] Step S81: Determine a first number of data blocks whose flag bits in the first bitmap and the second bitmap are both the first flag bits;
[0152] Step S82: Determine a second number of data blocks in the second bitmap whose flag bit is the first flag bit, and a third number of data blocks in the first bitmap whose flag bit is the first flag bit;
[0153] Step S83: Calculate the credibility score of the memory access bitmap according to a preset weight parameter, the first number, the second number, and the third number.
[0154] It can be understood that in the accumulation table and pattern table shown in Figure 4, each table entry records the memory access bitmap (Pattern) of the preset memory area. The memory access bitmap is used to reflect the access status of each data block in the preset memory area within a historical period. In an embodiment of the present application, the memory access bitmap recorded in the accumulation table and the pattern table can be used as the memory access bitmap of the preset memory area.
[0155] When both the accumulation table and the pattern table contain entries corresponding to a preset memory area, the credibility of the first memory access bitmap of the preset memory area can be calculated based on the first bitmap recorded in the accumulation table and the second bitmap recorded in the pattern table.
[0156] Specifically, assuming that in the preset memory area, the first number of data blocks that appear in both AT and PT and whose flag bits are all the first flag bit is X; the second number of data blocks that appear in PT and whose flag bit is the first flag bit is Y; and the third number of data blocks that appear in AT and whose flag bit is the first flag bit is Z, then the credibility score of the first memory access bitmap of the preset memory area can be expressed as: score = a × score coverage +b×score accuracy (1)
[0157] Among them, score coverage =X / Z, score accuracy =X / Y. a and b are both preset weight parameters.
[0158] When the credibility score is greater than or equal to the second preset value, it can be considered that the pre-fetch condition is met, and the memory addresses of each pre-fetched target data block are calculated according to the memory access bitmap in the PT, that is, the second bitmap.
[0159] In an optional embodiment of the present application, before performing bypass prediction on the target data block to obtain a prediction result, the method further includes:
[0160] Step S91: If the first memory access bitmap does not satisfy the first condition, obtain second historical memory access information of each memory area within a preset memory range; the preset memory range includes the preset memory area and other memory areas adjacent to the preset memory area;
[0161] Step S92: determining a second memory access bitmap based on the second historical memory access information; the second memory access bitmap is used to reflect the access status of each reference data block in the preset memory area within the preset memory range in a historical period;
[0162] Step S93: determining the data block whose flag bit in the second memory access bitmap is the first flag bit as the first data block;
[0163] Step S94: Determine a data block in the preset memory area that has the same address offset as the first data block as a target data block to be prefetched.
[0164] In an embodiment of the present application, if the first memory access bitmap of the preset memory area does not meet the first condition, it means that the preset memory area lacks historical memory access information, or the memory access characteristics in the first memory access bitmap are not obvious. In this case, the intra-page prefetcher cannot issue a prefetch based on the first memory access bitmap. At this time, the inter-page prefetcher can be started for prefetching. The basic idea of the inter-page prefetcher is: for memory areas that are adjacent in address, such as memory pages, their spatial memory access rules are similar, that is, you can refer to which blocks on the adjacent pages have been accessed, and then prefetch blocks with the same offset on the current page.
[0165] Specifically, when the first memory access bitmap does not meet the first condition, the inter-page prefetcher obtains the second historical memory access information of each memory area within the preset memory range, that is, the first historical memory access information of the preset memory area, and the historical memory access information of each memory area adjacent to the preset memory area.
[0166] Then, a second memory access bitmap is determined based on the second historical memory access information, and the second memory access bitmap is used to reflect the access status of each reference data block of the preset memory area within the preset memory range during the historical period. Exemplarily, the second memory access bitmap may include memory access bitmaps of each memory area adjacent to the preset memory area, and in the second memory access bitmap, the flag bit of the data block accessed during the historical period is the first flag bit. Alternatively, based on the memory access bitmaps of each memory area within the preset memory range, the sub-bitmap with the most first flag bits corresponding to the same address offset in these memory access bitmaps can be determined, and then the sub-bitmaps corresponding to each address offset can be combined to obtain the second memory access bitmap.
[0167] Next, the data block whose flag bit is the first flag bit in the second memory access bitmap is determined as the first data block, that is, the reference data block. Then, the data block with the same address offset as the first data block in the preset memory area is determined as the target data block to be prefetched. For example, assuming that the preset memory area is memory page A, the adjacent page of memory page A is memory page B, and there is a first data block b1 in the second memory access bitmap, which is the first data block in memory page B, and its address offset is 0, then the first data block a1 in memory page A can be determined as the target data block, and the address offset of data block a1 is 0. It should be noted that the address offset in this application refers to the offset value of the memory address of the data block relative to the starting address of the memory area to which it belongs.
[0168] Optionally, the second historical memory access information includes a historical page table, the historical page table being used to record page numbers and memory access bitmaps of each memory area within the preset memory range, the memory access bitmap being used to record access status of each data block in the memory area within a historical period. Determining the second memory access bitmap based on the second historical memory access information in step S92 includes:
[0169] Sub-step S921: querying, in the history page table, whether there is a first entry matching the first page number according to the first page number of the preset memory area;
[0170] Sub-step S922: If a first table entry matching the first page number exists in the history page table, set a flag bit of a data block matching the memory access address in a memory access bitmap of the first table entry to the first flag bit;
[0171] Sub-step S923: determining N first memory areas adjacent to the preset memory area in the historical page table according to the first page number;
[0172] Sub-step S924, performing a bitwise logical AND operation on the memory access bitmap of the preset memory area and the memory access bitmaps of the N first memory areas to obtain N third bitmaps;
[0173] Sub-step S925: Determine the bitmap containing the largest number of first flag bits among the N third bitmaps as the second memory access bitmap.
[0174] The main structure of the inter-page prefetcher is the Recent Page Table (RPT), which is used to record the page numbers and access bitmaps of each memory area within a preset memory range, such as the access features on a group of recently accessed memory pages, where each table entry can correspond to a memory page and record the page number and access bitmap of the memory page. When the access address of a memory access request belongs to the preset memory area, but the access features contained in the first access bitmap of the preset memory area do not support the intra-page prefetcher to issue a prefetch, the inter-page prefetcher can query the table entries of N memory areas adjacent to the preset memory area in the RPT table, and perform a logical AND operation on the first access bitmap of the preset memory area and the access bitmap of the adjacent area to obtain N third bitmaps, and determine the bitmap with the largest number of first flag bits contained in the N third bitmaps as the second access bitmap, and then determine the target data block to be prefetched in the preset memory area based on the first data block with the flag bit of the first flag bit in the second access bitmap.
[0175] Taking the memory area as a memory page as an example, the operation process of the inter-page prefetcher is as follows:
[0176] Step 1: When a memory access request arrives at the cache, it is first handled by the intra-page prefetcher. If no prefetch can be generated, it is then handled by the inter-page prefetcher. The inter-page prefetcher uses the page number (PN) of the memory page to which the current memory access request belongs to search the RPT table. If the page number is found, the second step is executed. If not, the inter-page prefetcher replaces an entry in the RPT table with a random one, records the PN in the page number field, sets all the pattern fields to 0, and then executes the second step.
[0177] Step 2: According to the block address currently being accessed, that is, the memory address carried in the memory access request, set the corresponding position in the pattern to 1.
[0178] Step 3: Search the RPT table for entries that are adjacent to the currently accessed page (i.e., the memory page to which the access address belongs). For example, the page number differences in the entries can be compared, and pages with differences less than a certain range are considered adjacent pages.
[0179] Step 4: Perform a bitwise AND operation on the pattern of the currently accessed page and the patterns of all N adjacent pages found, generating N new bitmaps. Find the entry with the most 1s in these N bitmaps and determine this entry as the second referenced memory access bitmap, ref_pattern.
[0180] Step 5: The data block with the flag bit set to "1" in the second access bitmap ref_pattern is determined as the first data block, and the data block in the currently accessed page with the same address offset as the first data block is determined as the target data block to be prefetched. Bypass prediction is performed on the target data blocks, and target data blocks that need to be bypassed are written to the streaming buffer, while target data blocks that do not need to be bypassed are written to the base cache.
[0181] In summary, the embodiment of the present application provides a data caching method, which, when a received memory access request meets the bypass prediction condition, performs bypass prediction on the target data block corresponding to the memory access address, and writes the target data block that needs to be bypassed into the streaming buffer, so that the bypassed data block is cached in the streaming buffer for a short period of time. Even if the data block is bypassed by mistake, it will still exist in the streaming buffer and have a chance of being hit, thereby eliminating the impact of the erroneous bypass; the target data block that does not need to be bypassed is written into the basic cache, thereby avoiding the contention of the cache for the data blocks that will not be hit and other data blocks that may be hit, which is beneficial to improving the cache hit rate, thereby reducing the number of memory accesses and reducing the latency and power consumption of the memory system.
[0182] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0183] Device embodiment
[0184] 5 , a block diagram of a data cache device of the present application is shown. The device may include:
[0185] A request receiving module 501 is configured to receive a memory access request sent by a processor, wherein the memory access request carries a memory access address;
[0186] A bypass prediction module 502 is configured to perform bypass prediction on a target data block corresponding to the memory access address to obtain a prediction result when the memory access request satisfies a bypass prediction condition;
[0187] A first writing module 503 is configured to write the target data block into the streaming buffer if the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed;
[0188] A second writing module 504 is configured to write the target data block into the basic cache if the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed;
[0189] The bypass prediction condition includes any one of the following:
[0190] The memory access request is a write request;
[0191] The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache;
[0192] The memory access request satisfies a prefetch condition; the prefetch condition is used to indicate a condition that needs to be satisfied in order to prefetch a data block in the memory into the system-level cache in advance.
[0193] Optionally, the bypass prediction module includes:
[0194] A first determining submodule is configured to determine a historical hit status of a target data block corresponding to the memory access address when the memory access request satisfies a bypass prediction condition;
[0195] The first prediction submodule is configured to perform bypass prediction on the target data block according to the historical hit situation to obtain a prediction result.
[0196] Optionally, the system-level cache further includes a hit history table, wherein the hit history table is used to record historical hit conditions of N consecutive accesses to the data block in the past, where N is a positive integer;
[0197] The first prediction submodule includes:
[0198] a first query unit, configured to query whether there is an entry in the hit history table that matches the memory access address;
[0199] A first determining unit is configured to determine that the target data block needs to be bypassed if there is no entry matching the memory access address in the hit history table; or
[0200] The second determining unit is configured to determine that the target data block needs to be bypassed if there is an entry matching the memory access address in the hit history table and the target data block has not been hit N times in the past.
[0201] Optionally, the first prediction submodule further includes:
[0202] The third determining unit is configured to determine that the target data block does not need to be bypassed if there is an entry matching the memory access address in the hit history table and the target data block has been hit at least once in the past N consecutive accesses.
[0203] Optionally, the device further comprises:
[0204] A first determining module is configured to determine a target entry in the hit history table if no entry matching the memory access address exists in the hit history table; the hit count of the data block corresponding to the target entry is less than that of the data blocks corresponding to other entries;
[0205] The first updating module is configured to clear the entry content of the target entry and fill the hit status of the target data block into the target entry.
[0206] Optionally, the system-level cache further includes a bypass history table, wherein the bypass history table is used to record address tags of data blocks bypassed in the system-level cache within a historical period; the apparatus further includes:
[0207] a second determining module, configured to determine that the memory access request hits the cache if a data block matching the memory access address exists in the system-level cache and / or an entry matching the memory access address exists in the bypass history table;
[0208] a third determining module, configured to determine that the memory access request misses the cache if no data block matching the memory access address exists in the system-level cache and no entry matching the memory access address exists in the bypass history table;
[0209] The second updating module is configured to update a corresponding entry in the hit history table according to the memory access address.
[0210] Optionally, the bypass prediction condition includes that the memory access request satisfies a prefetch condition; and the bypass prediction module includes:
[0211] A second determining submodule, configured to determine, when the memory access address belongs to a preset memory area, that the memory access request satisfies a prefetch condition;
[0212] A third determining submodule is configured to determine a first memory access bitmap of the preset memory area according to first historical memory access information of the preset memory area; the first memory access bitmap is configured to record access conditions of each data block in the preset memory area within a historical period;
[0213] a fourth determining submodule, configured to determine, if the first memory access bitmap satisfies a first condition, a target data block to be prefetched in the preset memory area according to the first memory access bitmap;
[0214] A second prediction submodule is used to perform bypass prediction on the target data block to obtain a prediction result;
[0215] The first condition includes at least one of the following:
[0216] The number of first flag bits included in the first memory access bitmap is greater than or equal to a first preset value; the first flag bit is used to indicate that the data block has been accessed within a historical period;
[0217] The credibility score of the first memory access bitmap is greater than or equal to a second preset value.
[0218] Optionally, the system-level cache further includes an accumulation table and a pattern table; wherein the accumulation table is used to record access status of each data block in a preset memory area within a first historical period; and the pattern table is used to record access status of each data block in the preset memory area within a second historical period, where the second historical period is greater than the first historical period.
[0219] The third determining submodule includes:
[0220] a second query unit, configured to query the accumulation table for a first entry that matches the memory access address;
[0221] a first updating unit, configured to update a flag bit of a target data block corresponding to the memory access address in the first entry to a first flag bit if a first entry matching the memory access address exists in the accumulation table, and update a timestamp of the first entry according to an access time of the memory access request;
[0222] a fourth determining unit, configured to determine, based on timestamps corresponding to the respective entries, a second entry in the accumulation table, wherein the timestamp of the second entry is greater than the timestamps of other entries in the accumulation table;
[0223] a second updating unit, configured to add the entry content of the second entry to the mode table and clear the entry content of the second entry from the accumulation table;
[0224] A fifth determining unit is configured to determine a first memory access bitmap of the preset memory area according to the accumulation table and the mode table.
[0225] Optionally, the system-level cache further includes a filter table; and the apparatus further includes:
[0226] An information recording module, configured to record access information of the memory access address in the screening table if there is no first table entry matching the memory access address in the accumulation table;
[0227] An information filtering module is used to add the entry content of the third entry to the accumulation table and clear the entry content of the third entry in the filtering table if there is at least one third entry in the filtering table and the third entry has recorded M visits; M is a positive integer.
[0228] Optionally, the first memory access bitmap includes a first bitmap of the preset memory area in the accumulation table and a second bitmap of the preset memory area in the mode table; the apparatus further includes:
[0229] A fourth determining module, configured to determine a first number of data blocks in which both flag bits of the first bitmap and the second bitmap are the first flag bits;
[0230] a fifth determining module, configured to determine a second number of data blocks in the second bitmap whose flag bit is the first flag bit, and a third number of data blocks in the first bitmap whose flag bit is the first flag bit;
[0231] A calculation module is configured to calculate a credibility score of the first memory access bitmap based on a preset weight parameter, the first number, the second number, and the third number.
[0232] Optionally, the fourth determining submodule includes:
[0233] The target data block determining unit is configured to determine, when the first memory access bitmap satisfies a first condition, a data block whose flag bit in the first memory access bitmap is a first flag bit as a target data block.
[0234] Optionally, the device further comprises:
[0235] a memory access information acquisition module, configured to acquire, when the first memory access bitmap does not satisfy the first condition, second historical memory access information of each memory area within a preset memory range; the preset memory range includes the preset memory area and other memory areas adjacent to the preset memory area;
[0236] a memory access bitmap determining module, configured to determine a second memory access bitmap based on the second historical memory access information; the second memory access bitmap being configured to reflect access conditions of each reference data block within the preset memory range in the preset memory area during a historical period;
[0237] A first data block determining module, configured to determine a data block whose flag bit in the second memory access bitmap is the first flag bit as a first data block;
[0238] The target data block determining module is configured to determine a data block in the preset memory area that has the same address offset as the first data block as a target data block to be prefetched.
[0239] Optionally, the second historical memory access information includes a historical page table, the historical page table is used to record the page number and memory access bitmap of each memory area within the preset memory range, and the memory access bitmap is used to record the access status of each data block in the memory area within a historical period;
[0240] The memory access bitmap determination module includes:
[0241] A page table query submodule, configured to query, according to the first page number of the preset memory area, in the history page table whether there is a first table entry that matches the first page number;
[0242] A setting submodule is configured to set a flag bit of a data block matching the memory access address in a memory access bitmap of the first entry to the first flag bit when a first entry matching the first page number exists in the history page table;
[0243] a memory region determination submodule, configured to determine, in the history page table according to the first page number, N first memory regions adjacent to the preset memory region;
[0244] A bitmap operation submodule, configured to perform bitwise logical AND operations on the memory access bitmap of the preset memory area and the memory access bitmaps of the N first memory areas to obtain N third bitmaps;
[0245] The bitmap determination submodule is configured to determine the bitmap containing the largest number of first flag bits among the N third bitmaps as the second memory access bitmap.
[0246] Optionally, the second prediction submodule includes:
[0247] a memory address determining unit, configured to determine a memory address of each target data block based on a starting address of the preset memory area and an address offset of the target data block; wherein the address offset is used to indicate an offset value of the memory address of the target data block relative to the starting address of the preset memory area;
[0248] The bypass prediction unit is used to perform bypass prediction on the target data block according to the memory address of the target data block to obtain a prediction result.
[0249] Optionally, the bypass prediction condition includes: the memory access request is a read request, and there is no data block matching the memory access address in the system-level cache;
[0250] The bypass prediction module includes:
[0251] a data acquisition submodule, configured to acquire a target data block matching the memory access address from a memory when the memory access request is a read request and no data block matching the memory access address exists in the system-level cache;
[0252] The third prediction submodule is used to perform bypass prediction on the target data block to obtain a prediction result.
[0253] In summary, the embodiment of the present application provides a data cache device, which, when a received memory access request meets the bypass prediction condition, performs bypass prediction on the target data block corresponding to the memory access address, and writes the target data block that needs to be bypassed into the streaming buffer, so that the bypassed data block is cached in the streaming buffer for a short period of time. Even if the data block is bypassed by mistake, it will still exist in the streaming buffer and have a chance of being hit, thereby eliminating the impact of the erroneous bypass; the target data block that does not need to be bypassed is written into the basic cache, thereby avoiding the contention of the cache for the data block that will not be hit and other data blocks that may be hit, which is beneficial to improving the cache hit rate, thereby reducing the number of memory accesses and reducing the latency and power consumption of the memory system.
[0254] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0255] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0256] Regarding the processor in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method and will not be elaborated here.
[0257] Referring to Figure 6, there is a block diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in Figure 6, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other via the communication bus. The memory is used to store executable instructions, and the executable instructions enable the processor to execute the data caching method of the aforementioned embodiment.
[0258] The processor may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0259] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus may be categorized as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG6 shows only one line, but this does not imply that there is only one bus or only one type of bus.
[0260] The memory may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only), a CD-ROM (Compact Disa Read Only), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0261] An embodiment of the present application also provides a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), enables the processor to execute the data caching method shown in Figure 1.
[0262] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0263] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0264] The present application embodiment is described with reference to the flow chart and / or block diagram of the method, terminal device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flow chart and / or block diagram and the combination of the process and / or box in the flow chart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes and / or one box or multiple boxes of the flow chart.
[0265] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0266] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0267] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0268] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0269] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0270] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0271] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0272] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0273] It is understood that the embodiments described in the present disclosure may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the modules, units, and subunits may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units or combinations thereof for performing the functions described in the present disclosure.
[0274] For software implementation, the techniques described in the embodiments of the present disclosure can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described in the embodiments of the present disclosure. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0275] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0276] The above is a detailed introduction to a data caching method, device, electronic device and readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A data caching method, applied to a system-level cache, wherein the system-level cache includes a streaming buffer and a basic cache; The method comprises: receiving a memory access request sent by a processor, wherein the memory access request carries a memory access address; When the memory access request satisfies the bypass prediction condition, bypass prediction is performed on the target data block corresponding to the memory access address to obtain a prediction result; When the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed, writing the target data block into the streaming buffer; When the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed, writing the target data block into the basic cache; The bypass prediction condition includes any one of the following: The memory access request is a write request; The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache; The memory access request satisfies a pre-fetch condition; the pre-fetch condition is used to indicate a condition that needs to be satisfied in order to pre-fetch a data block in the memory into the system-level cache in advance.
2. The method according to claim 1, wherein: When the memory access request satisfies the bypass prediction condition, bypass prediction is performed on the target data block corresponding to the memory access address to obtain a prediction result, including: In the case where the memory access request satisfies the bypass prediction condition, determining a historical hit situation of the target data block corresponding to the memory access address; Bypass prediction is performed on the target data block according to the historical hit situation to obtain a prediction result.
3. The method according to claim 2, wherein: The system-level cache also includes a hit history table, which is used to record the historical hit conditions of the data block for N consecutive accesses in the past, where N is a positive integer; The step of performing bypass prediction on the target data block according to the historical hit situation to obtain a prediction result includes: Querying the hit history table to see whether there is an entry matching the memory access address; If there is no table entry matching the memory access address in the hit history table, it is determined that the target data block needs to be bypassed; or, If there is a table entry matching the memory access address in the hit history table, and the target data block has not been hit N times in succession in the past, it is determined that the target data block needs to be bypassed.
4. The method according to claim 3, wherein: The method further comprises: If there is a table entry matching the memory access address in the hit history table, and the target data block has been hit at least once in the past N consecutive accesses, it is determined that the target data block does not need to be bypassed.
5. The method according to claim 3, wherein: The method further comprises: If there is no table entry matching the memory access address in the hit history table, determining a target table entry in the hit history table; the hit count of the data block corresponding to the target table entry is less than that of the data blocks corresponding to other table entries; The entry content of the target entry is cleared, and the hit status of the target data block is filled into the target entry.
6. The method according to claim 3, wherein: The system-level cache also includes a bypass history table, and the bypass history table is used to record the address tags of the data blocks bypassed in the system-level cache within the historical period; the method also includes: If a data block matching the memory access address exists in the system-level cache, and / or an entry matching the memory access address exists in the bypass history table, it is determined that the memory access request hits the cache; If there is no data block matching the memory access address in the system-level cache, and there is no table entry matching the memory access address in the bypass history table, determining that the memory access request misses the cache; The corresponding entry in the hit history table is updated according to the memory access address.
7. The method according to claim 1, wherein: The bypass prediction condition includes that the memory access request satisfies the pre-fetch condition; and when the memory access request satisfies the bypass prediction condition, bypass prediction is performed on the target data block corresponding to the memory access address to obtain a prediction result, including: In a case where the memory access address belongs to a preset memory area, determining that the memory access request satisfies a pre-fetch condition; Determine a first memory access bitmap of the preset memory area according to first historical memory access information of the preset memory area; the first memory access bitmap is used to record the access status of each data block in the preset memory area in a historical period; In a case where the first memory access bitmap satisfies a first condition, determining a target data block to be prefetched in the preset memory area according to the first memory access bitmap; Performing bypass prediction on the target data block to obtain a prediction result; The first condition includes at least one of the following: The number of first flag bits included in the first memory access bitmap is greater than or equal to a first preset value; the first flag bit is used to indicate that the data block is accessed within a historical period; The credibility score of the first memory access bitmap is greater than or equal to a second preset value.
8. The method according to claim 7, wherein: The system-level cache also includes an accumulation table and a mode table; wherein the accumulation table is used to record the access status of each data block in the preset memory area in a first historical period; the mode table is used to record the access status of each data block in the preset memory area in a second historical period, and the second historical period is greater than the first historical period; The determining a first memory access bitmap of the preset memory area according to first historical memory access information of the preset memory area includes: Querying whether there is a first table entry matching the memory access address in the accumulation table; In the case where there is a first table entry matching the memory access address in the accumulation table, updating a flag bit of a target data block corresponding to the memory access address in the first table entry to a first flag bit, and updating a timestamp of the first table entry according to an access time of the memory access request; Determine, according to the timestamps corresponding to the respective entries, a second entry in the accumulation table, wherein the timestamp of the second entry is greater than the timestamps of other entries in the accumulation table; Adding the entry content of the second entry to the mode table, and clearing the entry content of the second entry in the accumulation table; A first memory access bitmap of the preset memory area is determined according to the accumulation table and the mode table.
9. The method according to claim 8, wherein: The system-level cache also includes a filter table; the method also includes: If there is no first table entry matching the memory access address in the accumulation table, recording access information of the memory access address in the screening table; If there is at least one third entry in the filter table, and the third entry has recorded M accesses, the entry content of the third entry is added to the accumulation table, and the entry content of the third entry is cleared from the filter table; M is a positive integer.
10. The method according to claim 8, wherein: The first memory access bitmap includes a first bitmap of the preset memory area in the accumulation table and a second bitmap of the preset memory area in the mode table; the method further includes: Determine a first number of data blocks whose flag bits in the first bitmap and the second bitmap are both the first flag bits; Determine a second number of data blocks in the second bitmap whose flag bit is the first flag bit, and a third number of data blocks in the first bitmap whose flag bit is the first flag bit; A credibility score of the first memory access bitmap is calculated according to a preset weight parameter, the first number, the second number, and the third number.
11. The method according to claim 7, wherein: The determining, when the first memory access bitmap satisfies a first condition, a target data block to be prefetched in the preset memory area according to the first memory access bitmap comprises: When the first memory access bitmap satisfies a first condition, a data block whose flag bit in the first memory access bitmap is a first flag bit is determined as a target data block.
12. The method according to claim 7, wherein: Before performing bypass prediction on the target data block to obtain a prediction result, the method further includes: When the first memory access bitmap does not satisfy the first condition, obtaining second historical memory access information of each memory area within a preset memory range; the preset memory range includes the preset memory area and other memory areas adjacent to the preset memory area; Determine a second memory access bitmap according to the second historical memory access information; the second memory access bitmap is used to reflect the access status of each reference data block in the preset memory area within the preset memory range in the historical period; Determine a data block whose flag bit in the second memory access bitmap is the first flag bit as a first data block; A data block in the preset memory area having the same address offset as the first data block is determined as a target data block to be pre-fetched.
13. The method according to claim 12, wherein: The second historical memory access information includes a historical page table, the historical page table is used to record the page number and memory access bitmap of each memory area within the preset memory range, and the memory access bitmap is used to record the access status of each data block in the memory area within a historical period; The determining a second memory access bitmap according to the second historical memory access information includes: According to the first page number of the preset memory area, query in the historical page table whether there is a first table entry matching the first page number; If there is a first table entry matching the first page number in the history page table, setting a flag bit of a data block matching the memory access address in a memory access bitmap of the first table entry to the first flag bit; Determine N first memory areas adjacent to the preset memory area in the historical page table according to the first page number; Performing bitwise logical AND operations on the memory access bitmap of the preset memory area and the memory access bitmaps of the N first memory areas to obtain N third bitmaps; The bitmap including the largest number of first flag bits among the N third bitmaps is determined as the second memory access bitmap.
14. The method according to claim 11 or 12, wherein: The step of performing bypass prediction on the target data block to obtain a prediction result includes: Determine the memory address of each target data block according to the starting address of the preset memory area and the address offset of the target data block; the address offset is used to indicate the offset value of the memory address of the target data block relative to the starting address of the preset memory area; According to the memory address of the target data block, bypass prediction is performed on the target data block to obtain a prediction result.
15. The method according to claim 1, wherein: The bypass prediction condition includes: the memory access request is a read request, and there is no data block matching the memory access address in the system-level cache; When the memory access request satisfies the bypass prediction condition, bypass prediction is performed on the target data block corresponding to the memory access address to obtain a prediction result, including: When the memory access request is a read request and there is no data block matching the memory access address in the system-level cache, obtaining a target data block matching the memory access address from the memory; Perform bypass prediction on the target data block to obtain a prediction result.
16. A data cache device, applied to a system-level cache, wherein the system-level cache includes a streaming buffer and a basic cache; The device comprises: A request receiving module, used to receive a memory access request sent by a processor, wherein the memory access request carries a memory access address; A bypass prediction module, configured to perform bypass prediction on a target data block corresponding to the memory access address to obtain a prediction result when the memory access request satisfies a bypass prediction condition; A first writing module, configured to write the target data block into the streaming buffer when the prediction result indicates that the target data block corresponding to the memory access address needs to be bypassed; A second writing module, configured to write the target data block into the basic cache if the prediction result indicates that the target data block corresponding to the memory access address does not need to be bypassed; The bypass prediction condition includes any one of the following: The memory access request is a write request; The memory access request is a read request, and there is no data block matching the memory access address in the system-level cache; The memory access request satisfies a pre-fetch condition; the pre-fetch condition is used to indicate a condition that needs to be satisfied in order to pre-fetch a data block in the memory into the system-level cache in advance.
17. The device according to claim 16, wherein: The bypass prediction module comprises: A first determination submodule, configured to determine a historical hit status of a target data block corresponding to the memory access address when the memory access request satisfies a bypass prediction condition; The first prediction submodule is used to perform bypass prediction on the target data block according to the historical hit situation to obtain a prediction result.
18. The device according to claim 17, wherein: The system-level cache also includes a hit history table, which is used to record the historical hit conditions of the data block for N consecutive accesses in the past, where N is a positive integer; The first prediction submodule comprises: A first query unit, used for querying whether there is an entry matching the memory access address in the hit history table; A first determining unit is configured to determine that the target data block needs to be bypassed if there is no table entry matching the memory access address in the hit history table; or The second determining unit is configured to determine that the target data block needs to be bypassed if there is a table entry matching the memory access address in the hit history table and the target data block has not been hit for N consecutive times in the past.
19. An electronic device, comprising a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the data caching method as described in any one of claims 1 to 15. 20 . A readable storage medium, when instructions in the readable storage medium are executed by a processor of an electronic device, the processor is enabled to execute the data caching method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Bypass predictor for an exclusive last-level cache
CN111382089A
Cache replacement system and method based on instruction stream and memory access mode learning
CN113986774A
Cache performance processing method and related equipment thereof
CN115587052A
Data caching method and device, electronic equipment and readable storage medium
CN117389630A
Cited By
Distributed table look-up method and network equipment
CN121327002A
Processor, data processing method, chip, display card and electronic equipment
CN122018821A
Access control method and device for heterogeneous storage system, equipment, medium and product
CN122219856A
Storage system with coexistence of on-chip storage and cache functions and control method
CN122220302A