An implementation method and system for L1 cache load forward

By checking the store queue, store_buffer and load miss queue when pipelined on the load instruction, directly bypass returns the data required for load, solving the problem that the load instruction cannot obtain data in the existing technology, and improving CPU performance and pipeline utilization.

CN113467935BActive Publication Date: 2025-06-20GUANGDONG STARFIVE TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110665971.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-16
Publication Date
2025-06-20
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

The prior art does not provide forward data in the bypass forward network, resulting in the load instruction being unable to obtain data at the first time, increasing the waiting time of the load instruction in the load queue, and reducing the overall performance of the CPU.

Method used

When pipelined on the load command, check whether the data required for load can be directly bypassed to the data required for load in the store queue, store_buffer and load miss queue. If possible, directly bypass returns the data to reduce access to D_cache.

Benefits of technology

By providing forward data early, the waiting time of load instructions is reduced, the overall performance of the CPU is improved, and the utilization and throughput of pipelines is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113467935B_ABST
    Figure CN113467935B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of microprocessors, and particularly relates to a method and system for implementing L1 cache loadforward, including a pipeline module, a load queue module, a load miss queue module, a store queue module, and a store buffer module. In the bypass forward network of the present invention, when a load instruction enters the pipeline, it checks whether the data required by the load can be directly bypassed from the store queue, the store buffer, and the load miss queue. If the data can be directly bypassed, the return of the load instruction can be greatly accelerated, and the occupancy of the pipeline by the load instruction can also be reduced. The utilization rate and throughput of the pipeline are effectively improved, and the time for the load instruction to be blocked in the load queue is reduced. The overall performance of the CPU is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microprocessors, and particularly relates to a method and system for implementing L1 cache load forward. Background Art

[0002] In the prior art in the bypass forward network, no forward data is provided, resulting in the inability to obtain the loaded data in the first place. As a result, the pipeline also reads the D_cache and considers it a miss. At the same time, it occupies the ports for reading and writing the D_cache, and the latency of the load cannot be effectively reduced.

[0003] Due to the failure to obtain forward data, the load instruction is blocked in the load queue, resulting in the load instruction occupying the load queue for too long and failing to effectively utilize the load queue resources.

[0004] Since forward data is not obtained, the wakeup time of the instructions dependent on the load instruction is postponed, reducing the overall performance of the CPU. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention discloses a method and system for implementing L1 cache load forward. When the LOAD instruction is pipelined, it checks the store queue, store buffer, and load miss queue to see if the data required by the load can be directly bypassed. If the data can be directly bypassed, the return of the load instruction can be greatly accelerated, and the occupation of the pipeline by the load instruction can also be reduced.

[0006] The present invention is achieved through the following technical solutions:

[0007] In a first aspect, the present invention discloses a method for implementing L1 cache load forward, including the following steps:

[0008] S1 Save the LOAD instruction sent from the upstream to the load queue, and select a request from the load queue to participate in the arbiter arbitration. After winning the arbitration, it goes on the pipeline;

[0009] S2 Detect whether there is forward data that satisfies the load, and then directly return the load data or return the load data after checking the D_cache;

[0010] S3 detects whether there is overlap in the store queue / store buffer / load miss queue search:

[0011] S4 determines that there is an overlap relationship and that the data requirement of the load cannot be met, so an overlap haz relationship is established in the ldq entry;

[0012] S5 determines that no load data can be obtained in forward and D_cache, and requests an entry item from the miss queue. The miss queue entry item sends a reload request to L2.

[0013] S6 returns the reload data through L2 and wakes up the instructions that missed the load in the load queue;

[0014] The selected load instruction of S7 goes up the pipeline and forwrads the data needed for load in the refill buf;

[0015] S8 outputs the target and deallocated load queue, then refills the data onto the pipeline and writes the reloaded data into D_cache.

[0016] Furthermore, in the method, when checking whether there is forward data that satisfies the load, if so, the load data is directly returned and the load queue entry item is deallocated; if no forward data is detected to satisfy the load, D_cache is checked at this time. If D_cache hits, the load data is directly returned and the load queue entry item is deallocated.

[0017] Furthermore, in the method, an overlap haz relationship is established in the ldq entry, wherein the haz relationship includes the overlap haz of the store queue / store buffer / load miss queue.

[0018] Furthermore, in the method, after the corresponding entry item is deallocated, the overlap haz relationship is released, and at the same time, the entry item corresponding to the wakeup LDQ is woken up, and the LOAD instruction is selected from the re-appearing ldq entry item to be put back on the pipeline.

[0019] Furthermore, the method is applied to the bypass forward network. When the load instruction enters the pipeline, it checks the store queue, store buffer, and load miss queue to see if the data required by the load can be directly bypassed. If so, the data is directly bypassed.

[0020] In a second aspect, the present invention discloses a system for implementing L1 cache load forward. The system is used to implement the method for implementing L1 cache load forward described in the first aspect, and includes a pipeline module, a load queue module, a load miss queue module, a store queue module, and a store buffer module.

[0021] The beneficial effects of the present invention are as follows:

[0022] The present invention can return the load data as early as possible by providing forward data. Due to forwarding, the load data can be obtained earlier, and the instructions dependent on the load instruction can be woken up earlier, effectively improving the overall performance of the CPU. Since the load data has been obtained through forwarding, the load instruction does not need to occupy the read / write port of the D-cache, effectively improving the utilization rate of the pipeline. Since the load instruction has obtained the data through forwarding, the entry can be deallocated earlier, thereby effectively improving the utilization rate of the load queue entries. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a flowchart of a method for implementing L1 cache load forward;

[0025] Figure 2 It is a schematic diagram of the principle structure of a system for implementing L1 cache load forward. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment 1

[0028] Refer to Figure 1 As shown, this embodiment discloses a method for implementing L1 cache load forward, including the following steps:

[0029] S1 Save the LOAD instruction sent from the upstream to the load queue, and select a request from the load queue to participate in the arbiter arbitration. After winning the arbitration, it goes to the pipeline;

[0030] S2 Detect whether there is forward data that satisfies this load, and then directly return the load data or return the load data after querying the D_cache;

[0031] S3 Detect whether there is an overlap relationship in the store queue / store buffer / load miss queue search:

[0032] S4 Determine that there is an overlap relationship and the data requirements of this load cannot be met, and establish an overlap haz relationship in the ldq entry;

[0033] S5 Determine that the load data cannot be obtained from both the forwad and the D_cache, request an entry item to be allocated from the miss queue, and the miss queue entry item sends a reload request to the L2;

[0034] S6 Return the reload data through the L2, and at the same time wake up the instruction with a load miss in the load queue;

[0035] S7 The selected load instruction goes to the Pipeline, and the data required for the load is forwraded in the refill buf;

[0036] S8 Output the target and the deallocated load queue. Later, the refill data goes to the pipeline, and the reload data is written into the D_cache.

[0037] In this embodiment, when checking whether there is forward data that satisfies the load, if so, the load data is directly returned and the load queue entry item is deallocated; if no forward data is detected to satisfy the load, the D_cache is checked at this time. If the D_cache hits, the load data is directly returned and the load queue entry item is deallocated.

[0038] This embodiment establishes an overlap haz relationship in the ldq entry, where the haz relationship includes the overlap haz of the store queue / store buffer / load miss queue.

[0039] In this embodiment, after knowing that the corresponding entry item is deallocated, the overlap haz relationship is released, and at the same time, the entry item corresponding to the wakeup LDQ is woken up, and the LOAD instruction is selected from the re-appearing ldq entry item and put back on the pipeline.

[0040] This embodiment is applied to the bypass forward network. When the load instruction is pipelined, the storequeue, store_buffer, and load miss queue are checked to see whether the data required by the load can be directly bypassed. If so, the data is directly bypassed.

[0041] In this embodiment, L2 returns the reload data and wakes up the load miss instruction in the load queue. Since the load instruction woken up by the reload has a relatively high priority in both the load queue and arbitration, it is likely to be selected.

[0042] In this embodiment, the selected load instruction is put on the pipeline. At this time, since the previously missed data has been reloaded back, when this instruction is put on the pipeline again, it can forwrad to the data required for load in the refill buf, and output the target and deallocated load queue at the same time.

[0043] Example 2

[0044] This embodiment discloses Figure 2An implementation system of L1 cache load forward as shown, the system is used to execute an implementation method of L1 cache load forward, including a pipeline module, a load queue module, a load miss queue module, a store queue module and a store buffer module.

[0045] In the implementation scheme of this embodiment in the bypass forward network, when the load instruction is pipelined, it goes to the store queue, store buffer, and load miss queue to check whether the data required by the load can be directly bypassed. If the data can be directly bypassed, the return of the load instruction can be greatly accelerated, and the occupancy of the pipeline by the load instruction can also be reduced.

[0046] In summary, the present invention tries to provide forward data as much as possible for load instructions that can be forwarded, thereby reducing the cache miss rate and latency.

[0047] Since the present invention can obtain forward data, it does not need to access the D_cache, reducing the access to the D_cache and giving the access port of the D_cache to other pipelines, thereby effectively improving the utilization rate and throughput of the pipeline.

[0048] Since the present invention can obtain forward data, relevant entry resources can be released as early as possible, thereby reducing the time when the load instruction is blocked in the load queue.

[0049] Since the present invention can obtain forward data, instructions that are dependent on the load instruction can be awakened as early as possible, thereby improving the overall performance of the CPU.

[0050] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for implementing L1 cache load forward, characterized in that, The method comprises the following steps: S1 saves the load instructions sent from the upstream to the load queue, and selects a request from the load queue to participate in the arbiter arbitration. After winning the arbitration, it goes to the pipeline; S2 checks whether there is forward data that satisfies the load, and then directly returns the load data or returns the load data after checking D_cache; S3 detects whether there is overlap in the store queue / store buffer / load miss queue search: S4 determines that there is an overlap relationship and that the data requirement of the load cannot be met, so an overlap relationship is established in the ldq entry; S5 determines that no load data can be obtained in forward and D_cache, and requests an entry item from the miss queue. The miss queue entry item sends a reload request to L2. S6 returns the reload data through L2 and wakes up the instructions that missed the load in the load queue; The selected load instruction of S7 goes up the pipeline and forwrads the data needed for load in the refill buf; S8 outputs the target and deallocated load queue, then refills the data onto the pipeline and writes the reloaded data into D_cache.

2. The method for implementing L1 cache load forward according to claim 1, characterized in that, In the method, when checking whether there is forward data that satisfies the load, if so, the load data is directly returned and the loadqueue entry item is deallocated; if no forward data is detected that satisfies the load, the D_cache is checked at this time, and if the D_cache hits, the load data is directly returned and the load queue entry item is deallocated.

3. The method for implementing L1 cache load forward according to claim 1, characterized in that, In the method, an overlap haz relationship is established in the ldq entry, wherein the haz relationship includes the overlap haz of the store queue / storebuffer / load miss queue.

4. The method for implementing L1 cache load forward according to claim 1, characterized in that, In the method, after the corresponding entry item is deallocated, the overlap haz relationship is released, and at the same time, the entry item corresponding to the wakeup LDQ is woken up, and the load instruction is selected from the re-appearing LDQ entry item and put back on the pipeline.

5. The method for implementing L1 cache load forward according to claim 1, characterized in that, The described method is applied to the bypass forward network. When a load instruction is pipelined, it checks in the store queue, store buffer, and load miss queue to see if the data required by the load can be directly bypassed. If so, it directly bypasses to the data.

6. An implementation system for L1 cache load forward, the system is used to implement the method for implementing L1 cache load forward according to any one of claims 1-5, characterized in that, It includes a pipeline module, a load queue module, a load miss queue module, a store queue module, and a store buffer module.

Citation Information

Patent Citations

  • Processor cache with independent assembly line for accelerating prefetching requests

    CN107038125A

  • Pipeline recirculation for data misprediction in a fast-load data cache

    US20050097304A1