Off-chip storage method for prefetcher
Through the design of fragmented offset record table and single-hot code counter, the problem of waste of storage resources of the spatial memory stream prefetcher is solved, the accuracy and efficiency of the prefetcher are improved, and hardware consumption is reduced.
Patent Information
- Application Number
- CN202510312647.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing spatial memory stream prefetcher occupies a large amount of off-chip memory resources when storing a large number of trigger addresses and space-related addresses, resulting in increased hardware consumption and limited prefetching for the first-occurring address queues.
Using fragmented off-chip storage method, by designing a fragmented offset record table, only offsets related to the trigger address are recorded, and offsets are dynamically deleted or replaced during learning and use, and prefetch accuracy is recorded using a single hot code counter to reduce waste of storage space.
It realizes that while ensuring data integrity, the hardware consumption of off-chip storage is reduced, the accuracy and efficiency of prefetching is improved, and the entire address link failure caused by individual errors is avoided.
Smart Images

Figure CN120256333A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer architecture and data prefetching of central processing units, and particularly relates to an off-chip storage method for a prefetcher. Background Art
[0002] With the continuous increase in the number of processor cores and the main frequency, the bottleneck of processor performance has gradually shifted to the cache system between the main memory and the processor cores. To address this bottleneck, cache technology has played an increasingly important role in improving processor efficiency. Due to the high cost of high-performance cache materials, high-speed read / write caches close to the processor usually have only a very small cache capacity. The farther away from the processor, the larger the cache capacity, but the corresponding read / write rate will be significantly reduced. Memory main memory usually uses DRAM, which has a large data capacity but a slow read / write rate and is the farthest from the processor, unable to respond correctly and in a timely manner to the read / write requests issued by the processor. The cost of the processor reading data from the main memory is relatively high, which makes it difficult to significantly improve processor performance, forming a key factor "memory wall" that restricts the increase of the processor frequency. In 1978, American computer scientist Jim Smith first systematically proposed the concept of prefetching technology in his pioneering paper "A Study of Pre-fetching Techniques" to address the challenges brought by the "memory wall". The prefetching technology reduces the access latency caused by cache misses by prefetching data from the high-level cache or the main memory to the current cache before the data request arrives, thereby increasing the cache hit rate. In 1991, JWEinarson and his team implemented the first application of prefetching technology in the CPU in the patent "Instruction Prefetcher", marking a substantial progress in prefetcher technology. [1] 。
[0003] Although the prefetcher has shown significant effects in increasing the cache hit rate, its application also poses certain challenges. Incorrect prefetching not only fails to effectively reduce the access latency of normal requests, but may even lead to arbitration conflicts between excessive prefetch requests and other normal requests, increasing the hardware burden and resource consumption, thus wasting system resources. Therefore, designing an efficient prefetcher requires comprehensively considering the following three key factors: ① how to select an effective prefetch request address; ② the timing of issuing prefetch requests; ③ the storage location of prefetch data. In response to these challenges, researchers and engineers have proposed a variety of different prefetcher schemes in order to optimize prefetch performance.
[0004] The Spatial Memory Streaming (SMS) prefetcher significantly improves the performance of scientific and commercial server applications by leveraging data spatial relationships outside cache blocks. This technique prefetches data by identifying and exploiting repetitive access patterns, streaming data blocks into the main cache to prefetch data before demand misses. The SMS prefetcher effectively improves spatial locality and prefetch accuracy, and significantly reduces storage requirements by decoupling the detection structure, reducing storage overhead by half while increasing the prefetch coverage by 20%.
[0005] Although the Spatial Memory Streaming (SMS) prefetcher performs well in specific scenarios, it still has deficiencies. The current SMS prefetcher mainly performs relatively accurate prefetching for existing and recurring data in the internal storage table, and its prefetching effect is limited for the first-occurring address queue. In addition, due to the low correlation of the spatiality of address generation, this technique needs to store a large number of trigger addresses and spatially related addresses, thus occupying a large amount of off-chip memory resources. These problems limit the effectiveness and practicality of the SMS prefetcher in a wide range of applications.
[0006] [1]Einarson J W, Khan S A, Barrow A, et al. Instruction prefetcher: US19890333818[P]. US5062036A[2024-07-28]. Summary of the Invention
[0007] (I) Technical problems to be solved
[0008] The technical problem to be solved by the present invention is: to propose a fragmented off-chip storage method for the address correlation table to store the prefetched data, and to minimize the hardware consumption caused by off-chip storage as much as possible under the condition of ensuring data integrity.
[0009] (II) Technical solutions
[0010] To solve the above technical problems, the present invention provides an off-chip storage method for a prefetcher. In this method, a fragmented off-chip storage method is designed to store the prefetched data. The fragmented off-chip storage method is implemented by designing a fragmented offset record table, which contains fields such as trigger address, start bit, number of addresses, offset, and end bit. The fragmented offset record table is only used to record the offset of the associated address related to the trigger address relative to the trigger address, and the recorded fragmented offsets are deleted and replaced during the learning and usage process.
[0011] Preferably, the update method of the fragmented offset record table is as follows:
[0012] Let the data address A be sent from the processor core to the prefetcher as the trigger address. At this time, it first enters the prefetch judgment link. If the trigger address A has not been saved in the current prefetch list, that is, the fragmented offset record table, it is preferentially stored in the prefetch list, and at the same time, the start bit and the end bit are set, and the counter cnt and the number of addresses are initialized to 0;
[0013] When the data address B is sent from the processor core, if the data address B has a spatial correlation with the saved trigger address A, that is, the data address B is a spatially correlated address, then the offset between the data address B and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 1 of the trigger address A, and the number of addresses is incremented by 1, indicating that there is one prefetch offset after the trigger address A;
[0014] When the data address C is sent from the processor core, if the data address C has a spatial correlation with the saved trigger address A, then the offset between the data address C and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 2 of the trigger address A, and the number of addresses is incremented by 1, indicating that there are two prefetch offsets after the trigger address A;
[0015] The counter cnt is used to record the flag indicating whether the prefetch is correct when the same address is triggered 8 times.
[0016] Preferably, the counter cnt is an 8-bit one-hot code counter in the prefetcher, with an initial value of 8'b00000000, used to record the flag indicating whether the prefetch is correct when the same address 8'b00000000 is triggered 8 times; when the prefetcher receives the saved trigger address again, the prefetch addresses are sent out in sequence according to the offset. If the prefetch address is correct, cnt is shifted one bit to the left from right, and the low bit is filled with 1; if the prefetch address is incorrect, cnt is shifted one bit to the left from right, and the low bit is filled with 0. When 8 consecutive prefetches fail, the offset field is deleted.
[0017] Preferably, when all the offsets of the trigger address A fail, the trigger address is automatically deleted.
[0018] Preferably, in the data prefetch stage, the offset fields in the prefetch list can be dynamically added and deleted, and the storage table is updated in real time according to the correctness of the data prefetch to improve the prefetch accuracy rate.
[0019] Preferably, the design of the start and end bits enables all trigger addresses and spatially related addresses to be stored in a serial manner.
[0020] Preferably, the design of the offset offset enables only the relative difference value from the trigger address to be stored when storing data.
[0021] (3) Beneficial effects
[0022] The present invention designs a fragmented off-chip storage address method for storing prefetch data. In this invention design, start and end valid flags are added, enabling all trigger addresses and space-related addresses to be stored serially, without the waste of storage space caused by traditional methods; without the waste of storage space caused by traditional methods; the introduction of the offset makes it only necessary to store the relative difference value with the trigger address during data storage, without the need to save all address information, reducing the consumption of memory space; storing fragmented offset addresses allows individual offset addresses to be deleted and added to existing trigger addresses, and the entire trigger address link prediction will not fail due to the prediction error of one address. Description of the drawings
[0023] Figure 1 It is an example diagram of the storage method of space correlation addresses;
[0024] Figure 2 It is an example diagram of the change of the counter cnt when the prefetch address is successful;
[0025] Figure 3 It is an example diagram of the scenario when the prefetch address fails multiple times;
[0026] Figure 4 It is an example diagram of the scenario when the trigger address fails multiple times;
[0027] Figure 5 It is an example diagram of the recording method of the traditional history record table;
[0028] Figure 6 It is an example diagram of the recording method of the fragmented history record table of the present invention. Detailed implementation manners
[0029] To make the objectives, contents and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.
[0030] Current prefetching techniques usually use dynamic increase in storage space for learning and recording, so as to perform prefetching according to the address-related queues that have appeared before. Although the prefetching accuracy of the same trigger address can be improved, the large amount of data generated during the program operation not only causes waste of storage space, but also causes a large time delay when accessing the trigger address by looking up the table. The present invention proposes a fragmented off-chip storage method for the address correlation table to implement the storage of prefetch data, and under the condition of ensuring data integrity, the hardware consumption brought by off-chip storage is reduced as much as possible. To solve the above problems, the present invention proposes a design of fragmented off-chip addresses, as shown in Table 1 (only the table header is shown), that is, a fragmented offset record table is designed.
[0031] Trigger Address 1 Start Bit Number of Addresses Offset 1 Offset 2 … Offset n Stop Bit
[0032] Table 1 Fragmented Offset Record Table
[0033] In this solution, the fragmented offset record table is only used to record the offset of the associated address related to the trigger address relative to the trigger address, and the fragmented offsets recorded therein can be deleted and replaced during the learning and use process, without affecting other offset address information, avoiding the memory occupation situation of storing the same trigger address and space-related addresses multiple times due to the existence of individual special cases. The detailed introduction will be discussed in the technical solution in the next chapter.
[0034] I. Prefetch List Update Measures
[0035] When the data address A (as the trigger address) is sent from the processor core to the prefetch unit, it will first enter the prefetch judgment link. If the trigger address A has not been saved in the current prefetch list (i.e., the fragmented offset record table), it will be preferentially stored in the prefetch list, and at the same time, the start bit and the end bit are set, and the counter cnt and the number of addresses are initialized to 0, as shown in Figure 1 a in.
[0036] When the data address B is sent from the processor core, if the data address B has spatial correlation with the saved trigger address A (i.e., the data address B is a spatially correlated address), the offset (difference) between the data address B and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 1 of the trigger address A, and the number of addresses is incremented by 1, indicating that there is one prefetch offset after the trigger address A, as shown in Figure 1 b in.
[0037] Similarly, when the data address C is issued from the processor core, if the data address C has spatial correlation with the saved trigger address A, the offset (difference) between the data address C and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 2 of the trigger address A, and the address count is incremented by 1, indicating that there are two prefetch offsets after the trigger address A, as Figure 1 b.
[0038] The above counter cnt is an 8-bit one-hot code counter in the prefetcher, with an initial value of 8'b00000000, used to record the flag indicating whether the prefetch is correct when the same address (8'b00000000) is triggered 8 times. When the prefetcher receives the saved trigger address again, it will issue prefetch addresses in sequence according to the offset (the trigger address is the first column in Table 1. When the actual address issued by the processor core is the trigger address, the prefetcher will send the prefetch address in the fourth column with an offset of 1 relative to the trigger address). If the prefetch address is correct, cnt is shifted one bit to the left from right, and the low bit is filled with 1; if the prefetch address is incorrect, cnt is shifted one bit to the left from right, and the low bit is filled with 0 (after the prefetcher issues the prefetch address, the processor core will send the next actual address. After the prefetch address is issued, the corresponding information of the prefetch address will be retrieved. If the retrieved corresponding information is the information required by the actual address, the value of the counter in Table 1 is updated to 1. If the retrieved corresponding information is not the information required by the actual address, the value of the counter in Table 1 is updated to 0). When the prefetch fails continuously 8 times (i.e., the counter returns to 00000000 again), the offset field will be deleted.
[0039] When all the offsets of the trigger address A fail, the trigger address will be automatically deleted and the above learning will be performed again when it is received again.
[0040] II. Prefetch List Learning Phase
[0041] Figure 2 Indicates the prefetch list record value with the trigger address being 11'b10000000000 and the spatially correlated addresses being: 11'b10000001000, 11'b10000100000, 11'b10000000001, 11'b1000000100.
[0042] III. Prefetch List Prefetch Phase
[0043] This section illustrates the changes in the data recorded in Table 1 when the address prefetch is correct, a single address prefetch is incorrect, and multiple address prefetches are incorrect, respectively, through examples.
[0044] Figure 2Indicates the change of the prefetch list when the address 11’b10000000000 exists in the record table and is triggered four times;
[0045] When triggered for the first time, the spatial correlation address prefetch accuracy is as follows: 11’b10000001000, 11’b10000100000, 11’b10000000001, and 11’b1000000100 are all successfully prefetched;
[0046] When triggered for the second time, the spatial correlation address prefetch accuracy is as follows: 11’b10000001000 and 11’b10000000001 are successfully prefetched, while 11’b10000100000, 11’b10000000001, and 11’b1000000100 are failed to be prefetched;
[0047] When triggered for the third time, the spatial correlation address prefetch accuracy is as follows: 11’b10000001000 and 11’b10000000001 are failed to be prefetched, while 11’b10000100000, 11’b10000000001, and 11’b1000000100 are successfully prefetched;
[0048] When triggered for the fourth time, the spatial correlation address prefetch accuracy is as follows: 11’b10000001000, 11’b10000000001, and 11’b1000000100 are successfully prefetched, while 11’b10000100000 is failed to be prefetched;
[0049] Figure 3 Indicates the change of the prefetch list when the prefetch address has multiple prefetch errors and the offset needs to be removed when the address 11’b10000000000 exists in the record table. Next, when the address 11’b10000000000 is triggered, there will be 8 triggers. The prefetch addresses 11’b10000001000, 11’b10000100000, and 11’b1000000001 are successfully prefetched, while 11’b10000000100 is failed to be prefetched.
[0050] Figure 4 Indicates the change of the prefetch table when all prefetch addresses have multiple prefetch errors and the trigger address needs to be removed when the address 11’b10000000000 exists in the record table. Next, when the address 11’b10000000000 is triggered, there will be 10 triggers.
[0051] When triggered for the thirteenth time, the spatial correlation address prediction accuracy is as follows: 11’b10000001000 and 11’b10000100000 are successfully prefetched, while 11’b1000000001 is failed to be prefetched.
[0052] When triggered for the fourteenth time, the prediction accuracy of the space-related address is as follows: 11’b10000001000 prefetch successful, 11’b10000100000 and 11’b1000000001 prefetch failed.
[0053] In the last 8 triggers, all prefetch addresses failed to be prefetched.
[0054] The fragmented offset record table designed in the off-chip storage method of the dynamic adjustment fragmented data prefetcher of the present invention has the following three advantages compared with the traditional history record table:
[0055] 1. Due to the addition of the start and end bits, all trigger addresses and space-related addresses can be stored in a serial manner, without the waste of storage space caused by the traditional method.
[0056] 2. The introduction of the offset makes it only necessary to store the relative difference value from the trigger address when storing data, without the need to save all the address information, reducing the consumption of memory space.
[0057] 3. Due to the fragmented offset address storage method, individual offset addresses can be deleted and added to the existing trigger addresses, and the prediction failure of one address will not cause the prediction failure of the entire trigger address link.
[0058] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. An off-chip storage method for a prefetcher, characterized in that, In this method, a fragmented off-chip storage method is designed to store the prefetched data. The fragmented off-chip storage method is implemented by designing a fragmented offset record table, which contains fields such as trigger address, start bit, number of addresses, offset, and end bit. The fragmented offset record table is only used to record the offset of the associated address related to the trigger address relative to the trigger address, and the recorded fragmented offsets are deleted and replaced during the learning and usage process.
2. The method according to claim 1, characterized in that, The update method of the fragmented offset record table is as follows: Suppose the data address A is sent from the processor core core to the prefetcher as the trigger address. At this time, it first enters the prefetch judgment link. If the current prefetch list, that is, the fragmented offset record table, does not yet store the trigger address A, it is preferentially stored in the prefetch list, and at the same time, the start bit and end bit are set, and the counter cnt and the number of addresses are initialized to 0; When the data address B is sent from the processor core core, if the data address B has spatial correlation with the saved trigger address A, that is, the data address B is a spatially correlated address, then the offset between the data address B and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 1 of the trigger address A, and the number of addresses is incremented by 1, indicating that there is one prefetch offset after the trigger address A; When the data address C is sent from the processor core core, if the data address C has spatial correlation with the saved trigger address A, then the offset between the data address C and the trigger address A and the value of the counter cnt are concatenated and stored in the offset field 2 of the trigger address A, and the number of addresses is incremented by 1, indicating that there are two prefetch offsets after the trigger address A; The counter cnt is used to record the flag indicating whether the prefetch is correct when the same address is triggered 8 times.
3. The method according to claim 1, characterized in that The counter cnt is an 8-bit one-hot code counter in the prefetcher, with an initial value of 8'b00000000, and is used to record the flag indicating whether the prefetch is correct when the same address 8'b00000000 is triggered 8 times; When the prefetcher receives the saved trigger address again, the prefetch addresses are sent sequentially according to the offset. If the prefetch address is correct, cnt is shifted one bit to the left from right, and the low bit is filled with 1; if the prefetch address is incorrect, cnt is shifted one bit to the left from right, and the low bit is filled with 0. When the prefetch fails continuously 8 times, the offset field is deleted.
4. The method according to claim 1, characterized in that When all the offsets of the trigger address A become invalid, the trigger address is automatically deleted.
5. The method according to claim 1, wherein In the data prefetch stage, the offset fields in the prefetch list can be dynamically added and deleted, and the storage table is updated in real time according to the correctness of the data prefetch, improving the prefetch accuracy.
6. The method according to claim 1, characterized in that, The design of the start and end bits enables all trigger addresses and spatially related addresses to be stored in a serial manner.
7. The method according to claim 1, characterized in that The design of the offset offset enables only the relative difference value from the trigger address to be stored when storing data.
8. The method according to claim 1, characterized in that, This method is applied in the design of the prefetcher.
9. The method according to claim 1, wherein This method is applied in the design of the computer architecture.
10. The method according to claim 1, characterized in that, This method is applied in the design of the central processing unit.