Method and system for implementing memory access optimization of processor with multi-channel concurrent instructions

The memory access unit of the RISC-V processor is optimized through a multi-channel pipeline structure and VIPT address generation, which solves the parallelism and data correlation problems of Load and Store instructions, improves memory access throughput and system efficiency, and adapts to multi-core/multi-threaded environments.

CN120780356APending Publication Date: 2025-10-14SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Patent Information

Application Number
CN202510731461.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

The memory access units of existing RISC-V processors have limited parallelism when processing Load and Store instructions, complex data dependency processing, low address generation and TLB query efficiency, and multi-core/multi-thread support and consistency challenges, resulting in performance bottlenecks and reduced throughput.

Method used

It adopts a multi-channel pipeline structure, separates the processing of Load and Store instructions, sets up an independent RAW check queue, adopts VIPT address generation, parallelizes address generation and TLB query, separates the address and data of Store instructions, merges them into one write, and supports data consistency in multi-core/multi-threaded environments.

Benefits of technology

It significantly improves the parallel processing capability of memory access instructions, reduces load-to-use latency, improves memory access throughput and system efficiency, supports multi-core/multi-threaded high concurrency scenarios, and has good scalability and compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780356A_ABST
    Figure CN120780356A_ABST
Patent Text Reader

Abstract

The invention provides a processor memory access optimization implementation method and system with multi-channel concurrent instructions, and relates to the technical field of processor integrated circuits, and the method comprises the steps: obtaining a Load instruction and a Store instruction to be executed; distributing a plurality of memory access instructions to each memory access channel which is independently processed in parallel, wherein each memory access channel comprises a Load memory access channel and a Store memory access channel; an independent Load queue and an independent RAW check queue are arranged in the Load memory access channel, and an independent Store queue, an independent STD queue and an independent Store Buffer queue are arranged in the Store memory access channel; the Load instruction is sequentially subjected to an address generation stage, a TLB parallel query stage, an L1Cache access stage, a RAW check and data forward push stage and a data write-back stage in the Load memory access channel; according to the method, address and data separation processing is carried out on a Store instruction, an address part enters a Store queue, a data part enters an STD queue, then merging is carried out, multiple Store operations on the same Cache Line are merged into one-time writing, the merged data are written into a Store Buffer queue, and then the Store Buffer queue submits and caches in a unified mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of processor integrated circuits, and in particular to a method and system for optimizing processor memory access with concurrent multi-channel instructions. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the rise of high-load applications such as big data, cloud computing, and artificial intelligence, the performance bottleneck of processors has gradually shifted from arithmetic operations to memory access subsystems. Traditional pipeline processors have limited instruction-level parallelism and are unable to meet the high-bandwidth, low-latency access requirements of modern applications. Superscalar and out-of-order execution architectures have become mainstream, improving instruction throughput through multiple issuances and dynamic scheduling. However, as a key module connecting the processor core and the storage layer, the performance of the memory access unit directly limits the operating efficiency of the entire system.

[0004] RISC-V, an emerging open-source instruction set architecture, has become a mainstream ISA widely adopted in academia and industry due to its simplicity, modularity, and scalability. The memory access unit of a RISC-V processor is primarily responsible for processing load / store instructions, enabling access to data caches (L1 / L2 cache), TLBs, and main memory. The key design of the memory access unit in a RISC-V processor is crucial for efficient data access.

[0005] The memory access process currently implemented in the memory access unit of the RISC-V superscalar processor includes address space calculation, address translation, and separate execution of load / store instructions. Load instructions access the L1D cache based on the physical address, while simultaneously accessing the data portion of store instructions. Finally, based on address overlap, the data is written back to the physical register. Store instructions do not directly access the cache but instead write back data through a write queue, merging cache lines on a per-cache basis. For cache misses or writes back to lower-level storage spaces, corresponding miss queues or write-back queues are designed, interacting on a cache line basis.

[0006] However, the above existing design solutions have the following technical problems:

[0007] 1) Limited parallelism of memory access instructions: The memory access instruction processing capabilities of most processors' memory access units (MAUs) fail to match core parameters such as the out-of-order execution unit (OOO) and issue width, creating a system bottleneck. The MUs can only support one or two MAUs per cycle, failing to fully exploit the high concurrency benefits of OOO scheduling. Load and store instructions compete for resource allocation and queue management, leading to pipeline blockages and reduced throughput.

[0008] 2) Complex data dependency processing: Load and Store instructions have RAW (Read After Write) dependencies. Traditional data forwarding and dependency detection mechanisms typically use queue traversal and flags, resulting in complex logic and high latency. For example, serially traversing the Store queue within the Load queue is inefficient and can easily become a performance bottleneck. Imperfect data forwarding mechanisms require Load instructions to wait for Store instructions to be written back to main memory or cache before retrieving data, increasing latency.

[0009] 3) Inefficient address generation and TLB lookup: Address generation (Agen) and TLB lookups for load / store instructions are often performed serially, failing to parallelize them. This increases overall latency in the memory access pipeline. Inadequate application of PIPT (Physical Index Physical Tag) technology means cache accesses must wait for the TLB to return the physical address, reducing memory access efficiency and increasing pipeline voids, impacting throughput.

[0010] 4) Multi-core / multi-thread support and consistency challenges: With the popularization of multi-core and SMT (Simultaneous Multi-Threading) technologies, memory access units need to support higher concurrency and more complex data consistency protocols (such as MESI, MOESI, etc.). Existing designs find it difficult to balance performance and consistency. Summary of the Invention

[0011] To address the above-mentioned issues, the present disclosure proposes a method and system for optimizing processor memory access with multi-channel concurrent instructions. This method establishes a multi-channel pipeline structure within the memory access unit and designs an independent RAW check queue (RAWQ) to detect read-after-write (RAW) dependencies between loads and stores. The address portion and data portion of the store instruction are processed separately, and the VIPT (Virtual Index Physical Tag) method is used for address generation. Address generation is synchronized with TLB queries, eliminating the need for idle pipeline waiting, effectively reducing load-to-use latency and significantly improving memory access efficiency.

[0012] According to some embodiments, the present disclosure adopts the following technical solutions:

[0013] A method for optimizing processor memory access with concurrent multi-channel instructions, comprising:

[0014] Get multiple memory access instructions to be executed in each cycle, including Load instructions and Store instructions;

[0015] Multiple memory access instructions are assigned to independent and parallel memory access channels based on resource occupancy and scheduling priority, including the Load memory access channel and the Store memory access channel. Independent Load queues and RAW check queues are set up in the Load memory access channel, and independent Store queues, STD queues, and Store Buffer queues are set up in the Store memory access channel.

[0016] The load instruction goes through the address generation, TLB parallel query, L1 cache access, RAW check and data forward, and data write back stages in the load memory access channel.

[0017] The address and data of the Store instruction are separated. The address part enters the Store queue and the data part enters the STD queue. They are merged according to the instruction ID. The Store Buffer queue merges multiple Store operations on the same Cache Line into one write. The write cache operations are submitted uniformly by the Store Buffer queue.

[0018] According to some embodiments, the present disclosure adopts the following technical solutions:

[0019] A system for optimizing processor memory access with concurrent multi-channel instructions, comprising:

[0020] An address generation unit is configured to obtain multiple memory access instructions to be executed in each cycle, the memory access instructions including load instructions and store instructions, calculate virtual addresses for the memory access instructions in the same clock cycle, and simultaneously send address query requests to the TLB and cache arbiter;

[0021] The load instruction execution unit is configured to use independent load queues and RAW check queues for processing. The load instruction goes through the address generation, TLB parallel query, L1 cache access, RAW check and data forwarding, and data write back stages in the load memory access channel.

[0022] The Store instruction execution unit is configured to use independent Store queues, STD queues, and Store Buffer queues for processing. The address and data of the Store instruction are separated, with the address portion entering the Store queue and the data portion entering the STD queue. They are then merged based on the instruction ID. The Store Buffer queue merges multiple Store operations on the same cache line into one write, and the write cache operations are submitted uniformly by the Store Buffer queue.

[0023] According to some embodiments, the present disclosure adopts the following technical solutions:

[0024] A computer program product includes a computer program. When the computer program is executed by a processor, the method for optimizing processor memory access with concurrent multi-channel instructions is implemented.

[0025] According to some embodiments, the present disclosure adopts the following technical solutions:

[0026] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, a method for optimizing processor memory access with multi-channel concurrency of instructions is implemented.

[0027] According to some embodiments, the present disclosure adopts the following technical solutions:

[0028] An electronic device comprises: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the processor memory access optimization implementation method of multi-channel concurrent instructions.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention discloses a method for optimizing processor memory access with multi-channel concurrent instructions. The method employs a multi-channel pipeline structure within the memory access unit, specifically supporting up to four memory access instructions (including Load and Store) to enter the execution unit simultaneously per cycle. By dividing the memory access channel into independent Load and Store channels and dynamically balancing the hardware resource allocation between the two types of instructions, resource competition within the memory access unit is significantly reduced, and the parallel processing capability of memory access instructions is enhanced. The parallel processing capability of memory access instructions is significantly improved, and the load-to-use latency is reduced by 20%.

[0031] The present invention discloses a method for optimizing the memory access of a processor with multi-channel concurrent instructions. An independent RAW check queue (RAWQ) is designed to specifically detect the read-after-write (RAW) dependency between Load and Store. The data push-forward mechanism does not need to wait for Store write-back, greatly shortening the critical path of the Load instruction and adapting to high-concurrency scenarios such as AI and image processing. The efficiency of data dependency detection is improved, the probability of pipeline blocking is reduced, and pipeline blocking caused by data dependency is reduced, thereby improving the memory access throughput under out-of-order execution and increasing the overall throughput by 15%. It supports multi-core and multi-threaded high-concurrency scenarios and has good scalability and compatibility.

[0032] The present invention discloses a method for optimizing processor memory access with multi-channel concurrent instructions, which innovatively separates the address part and the data part of the Store instruction for processing, improves the processing efficiency of the Store instruction, manages the address and data of the Store instruction separately, and cooperates with the age vector and address merging mechanism to improve write bandwidth and sequential consistency, thereby avoiding data conflicts caused by out-of-order submission.

[0033] The present invention discloses a method for optimizing processor memory access with concurrent multi-channel instructions. The method adopts a VIPT (Virtual Index Physical Tag) address generation method. To shorten the execution path of memory access instructions, address generation is performed in parallel with TLB / Cache queries, thereby reducing pipeline bubbles and improving system energy efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings, which constitute a part of the present disclosure, are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.

[0035] Figure 1 Optimized structural block diagram of the memory access unit according to the embodiment of the present disclosure;

[0036] Figure 2 A schematic diagram of the Load instruction pipeline according to an embodiment of the present disclosure;

[0037] Figure 3 This is a schematic diagram of the Store instruction pipeline of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0039] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs.

[0040] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, devices, components and / or combinations thereof, but do not preclude the presence or addition of one or more other features, steps, operations, devices, components and / or combinations thereof.

[0041] Embodiment 1

[0042] In one embodiment of the present disclosure, a method for implementing processor memory access optimization for instruction multi-lane concurrency is provided, comprising:

[0043] Step 1: obtaining a plurality of memory access instructions to be executed in each cycle, the memory access instructions including Load instructions and Store instructions;

[0044] Step 2: distributing the plurality of memory access instructions to each independent parallel processing memory access lane according to resource occupation and scheduling priority, including Load memory access lane and Store memory access lane; setting independent Load queue and RAW check queue in the Load memory access lane, and setting independent Store queue, STD queue and Store Buffer queue in the Store memory access lane;

[0045] Load instructions sequentially pass through address generation, TLB parallel query, L1 Cache access, RAW check and data pre-push, and data write-back stages in the Load memory access lane;

[0046] Store instructions are processed using independent Store queue, STD queue and Store Buffer queue; address and data of the Store instructions are processed separately, the address part enters the Store queue, and the data part enters the STD queue; the Store instructions are merged according to instruction ID, the Store Buffer queue merges a plurality of Store operations on the same Cache Line into one write, and Cache write operation is uniformly submitted by the Store Buffer queue.

[0047] As an embodiment, the present invention discloses a method for optimizing processor memory access with multi-channel concurrent instructions. This method features systematic and innovative designs from multiple aspects, including the memory access pipeline structure, data dependency detection and forwarding mechanism, address generation, and TLB / Cache parallel query. This method supports multi-channel parallel memory access instruction processing, improving throughput; optimizes the data dependency detection and forwarding mechanism for Load / Store instructions, reducing pipeline congestion and latency; achieves efficient parallelization of address generation and TLB / Cache query, shortening the execution path of memory access instructions; supports data consistency and high concurrency in multi-core / multi-threaded environments, and possesses good scalability, verifiability, and industrial application value. The specific contents are as follows:

[0048] Step 1: Get multiple memory access instructions to be executed in each cycle, including Load instructions and Store instructions;

[0049] Step 2: Allocate multiple memory access instructions to independent and parallel memory access channels based on resource occupancy and scheduling priority;

[0050] Specifically, the memory access unit of the present disclosure employs a multi-channel pipeline structure with four receiving channels, supporting up to four memory access instructions (including load and store instructions) entering the execution unit simultaneously per cycle. By dividing the memory access channels into independent load and store channels and dynamically balancing the hardware resource allocation between the two types of instructions, resource competition within the memory access unit is significantly reduced, improving the parallel processing capability of memory access instructions.

[0051] The four channels are channel 0, channel 1, channel 2, and channel 3. These channels are divided into independent load and store channels. Channels 0 and 2 can receive both load and store instructions. Channels 1 and 3 can only receive load instructions. This means that channels 0-3 can receive load instructions, while channels 0 and 2 can receive store instructions. Instruction dispatch is scheduled based on the transmit queue's status.

[0052] The memory access channel includes the Load memory access channel and the Store memory access channel. Each Load channel has an independent Load queue (LDQ) and Raw check queue (RAWQ), which can achieve efficient instruction distribution and parallel execution under out-of-order scheduling.

[0053] Each Store channel has an independent Store Queue (STQ), Data Information Queue (STD) and StoreBuffer (STB), which support separate processing of address and data to improve write bandwidth and efficiency.

[0054] Furthermore, the pipeline structure within the channel includes: Load instructions go through the address generation, TLB parallel query, L1 cache access, RAW check and data forward, and data write-back stages in the Load memory access channel; Store instructions first perform address and data separation processing in the Store memory access channel, the address part enters the Store queue (STQ), and the data part enters the STD queue, and then merges them, merging multiple Store operations on the same cache line into one write. The merged data is written to the Store Buffer Queue (STB), and then the Store Buffer Queue (STB) submits it to the cache uniformly.

[0055] As an embodiment, the RAW Check Queue (RAWQ) is a dedicated hardware queue used to specifically detect the read-after-write (RAW) dependency between Load and Store. RAWQ can know the address and data status of the preceding Store instruction in advance during the Load instruction execution phase, thereby achieving efficient parallel detection. Once it is detected that the relevant Store data is already in STQ or STB, the Load instruction can directly obtain the data through the forwarding mechanism without waiting for the Store to be written back to the main memory or Cache, greatly reducing the Load execution latency. This mechanism is particularly effective for out-of-order execution scenarios, and can prevent pipeline blockage and performance degradation caused by data dependencies.

[0056] As an embodiment, the specific process of the multi-channel memory access pipeline and parallel scheduling disclosed in the present invention is as follows:

[0057] (1) Memory access instruction distribution

[0058] The out-of-order scheduling unit distributes pending memory access instructions to various memory access channels based on resource occupancy and scheduling priority. Each channel can independently process load / store instructions, avoiding resource contention.

[0059] Specifically, a dynamic load balancing algorithm is used to monitor the queue depth and resource utilization of each channel in real time, and dynamically adjust the allocation strategy according to the instruction type and priority to ensure load balancing and maximum throughput of each channel.

[0060] Each cycle can accept four memory access instructions. The LDQ queue adopts a bank-based design, with four queues performing parallel arbitration. Due to out-of-order instruction execution, instructions execute at varying speeds, resulting in some queues being full and others being empty. Load instructions are allocated based on availability. Specifically, load instructions are allocated based on queue availability. If one queue is relatively full, it is assigned to another queue.

[0061] (2) Execution of each pipeline

[0062] As an example, for a Load instruction, the following stages are sequentially performed: address generation, TLB parallel query, L1Cache access, RAW check and data forwarding, and data writeback. Each channel can execute concurrently without blocking each other.

[0063] Specifically, the present disclosure adopts the VIPT (Virtual Index Physical Tag) mechanism, which includes virtual indexes and physical tags, to achieve high parallelism in address generation, TLB lookups, and L1 Cache accesses. The specific process is as follows:

[0064] 1. Address generation

[0065] The address generation unit (Agen) calculates the virtual address for memory access instructions within the same clock cycle and simultaneously sends address query requests to the TLB and cache arbiter. The address generation unit separates the virtual address into index and tag components during instruction issuance. The index component is used directly for L1 data cache indexing without waiting for the TLB. The tag component waits for the TLB to return the physical address before performing a tag comparison.

[0066] The address generation unit corresponds to Addr gen. This part mainly calculates the address space of memory access instructions, records instruction information such as load, store, and fence instructions, and determines address characteristics such as non-alignment.

[0067] Index and Tag are the memory access instruction address ranges, index is [13:6] of the address, and tag corresponds to [47:14].

[0068] 2. TLB parallel query

[0069] Address generation and TLB lookups are performed simultaneously, eliminating pipeline wait times and significantly improving memory access efficiency. If a TLB lookup hits, the physical address tag comparison is completed immediately, achieving a zero-latency path for cache access and reducing pipeline bubbles. This mechanism effectively reduces load-to-use latency and improves instruction throughput. It is particularly suitable for processor architectures with high issue width and out-of-order execution, improving overall system energy efficiency.

[0070] 3. L1 Cache Access

[0071] Cache access requires two arbitrations. The first arbitration determines which modules gain access. In this design, MQ, WBQ, STB, SNQ, and LDQ all require cache access. The priority is STB full > SNQ > LDQ > MQ > STB normal. The second arbitration determines which banks are accessed. When multiple requests enter the data cache to read data, if different requests access the same bank, a bank conflict occurs. Therefore, address arbitration is performed in advance for valid requests. The arbiter prevents conflicts and ensures that P0, P1, and P2 are inaccessible after arbitration.

[0072] 4. RAW inspection and data forwarding

[0073] Due to the complex data dependency between Load instructions and Store instructions, the traditional serial traversal mechanism is inefficient. The present disclosure adopts an independent RAW check queue (RAWQ). The RAW check queue set in the Load memory access channel is used to perform read-after-write dependency detection between Load instructions and Store instructions. The address and data status of the preceding Store instruction are known in advance during the Load instruction emission phase and recorded in real time. Each Load channel can access the RAW check queue in parallel to detect whether there are related Store instructions. When it is detected that the related Store data is already in the Store queue or Store Buffer, the Load instruction directly obtains the data through the forward push mechanism without waiting for the Store data to be written back to the main memory or Cache, thereby achieving efficient parallel detection and forward push. The specific process is as follows:

[0074] (1) The parallel detection process includes: when a Load instruction enters the RAWQ, the hardware automatically compares its target address with the address range of all Store instructions in the RAWQ.

[0075] If there is address overlap, further determine whether the Store data is in place.

[0076] If the data is already in place, the Load instruction can be pushed forward directly to obtain the data without waiting for the Store queue to write it back to the main memory or cache, greatly reducing the Load execution latency.

[0077] (2) The forward push process includes: implementing multi-level forward push priority, prioritizing the most recent store data to ensure data consistency. If multiple stores hit the same address, an age vector is used to ensure that the load retrieves the most recently written data. This mechanism significantly reduces load-to-use latency, reduces pipeline blockages caused by data dependencies, and improves memory access throughput under out-of-order execution.

[0078] The age vector refers to the age relationship of Store instructions within the STQ table entry. During forward checking, when multiple Store instructions overlap with the address of the Load instruction, to ensure data correctness, the youngest Store instruction that is older than the Load instruction but meets the requirements must be found.

[0079] As an example, Figure 2 The following figure shows the execution pipeline of the Load instruction. LD0-3 represents the execution process of the load instruction, which is as follows:

[0080] Stage 1: Address generation and access rights checking

[0081] In the first stage of the pipeline, the load instruction enters the address generation unit (AGEN), which calculates the target memory access address range and detects special cases such as page crossing or unaligned access. During the address generation process, virtual address calculation and TLB address translation are performed in parallel. After the calculation is completed, the address information is simultaneously passed to the following three modules:

[0082] Cache Arbiter: Checks whether the current instruction can obtain access to L1 Dcache.

[0083] LDQ: Stores instruction information in the LDQ entry, recording the obtained physical address and address attributes. For instructions whose addresses span cache lines, wait until they become the oldest instruction and then split them into two separate instructions for processing.

[0084] DTLB: Gets the physical address information of the page table, checks the address exception according to the PMP attribute of the page table, and passes the exception to the instruction queue.

[0085] Stage 2: Cache access and address conflict check

[0086] In the second stage of the pipeline, after obtaining access to the L1 Dcache, the Load instruction performs a tag match to determine the cache line hit. If the cache line hits, the data is read directly from the L1 Dcache and sent to the Forward module after mask tag processing; if the cache line partially hits, the instruction is temporarily stored in the Load queue, and the three queues STB, STQ, and MQ are checked in sequence to see if there are any entries with overlapping addresses. In the STB, the recorded address in the entry is checked to see if it overlaps with the Load instruction address; all entries in the STQ are traversed to see if their addresses overlap with the Load instruction address; all entries in the MQ are traversed to see if their addresses overlap with the Load instruction address. In this case, Load outputs a completion signal and releases its temporarily used physical registers in advance, accelerating the scheduling and execution of subsequent instructions. In the case of a cache line miss, the instruction enters the LDQ and enters the Sleep state, waiting for the Wake-up signal after the MQ is backfilled.

[0087] Stage 3: Data forwarding and cache reading

[0088] In the third stage of the pipeline, when the data for the Load instruction can be completely forwarded from the STB, STQ, MQ, and Cache, that is, when the data source of the Load instruction is determined, if it is detected that only the address is ready in the corresponding Store Entry of the STQ, but the data is not, the Load instruction cannot meet the completion condition. In this case, the Cancel signal in the completion signal output by the Load instruction is set to 1, and the ROB marks it as incomplete after receiving the instruction completion signal. It is important to note that this stage only completes the data preparation work; the actual data has not yet been written back to the physical register. After determining the address overlap in Stage 2, the Forward module merges the four data sources (L1 Dcache, MQ, STB, and STQ) based on the mask tags fed back by each module.

[0089] Stage 4: Data write back

[0090] In the fourth stage of the pipeline, the data merged by the Forward module is written to the physical register stack, the LDQ table entry is cleared, and the age_vector clears the valid flag to prepare for recording the next Load instruction.

[0091] As an example, for Store instructions: In traditional designs, Store instructions often need to wait for the address and data to arrive before entering the Store queue, resulting in a large delay. This disclosure innovatively separates the address and data parts of the Store instruction, with the address information entering the STQ and the data entering the STD queue. After all are complete, they are merged and written to the STB, which then submits them to the cache or main memory. The specific process is as follows:

[0092] 1. Separate address and data processing. The specific process involves: After a Store instruction is issued, the address and data are separated. The address portion enters the STQ for independent management, while the data portion enters the STD queue for independent storage. These two pieces of information can arrive asynchronously, avoiding pipeline blockages caused by waiting for data.

[0093] 2. The write-merge process involves: Once both the STQ and STD information are in place, address merging logic is used to merge multiple store operations on the same cache line into a single write, reducing the total number of writebacks. The merged data is written to the STB and then collectively committed to the cache by the Store Buffer queue. The STB uses a ring buffer structure to support high-concurrency writes. Both the STQ and STB are equipped with age vectors that record the order in which each store instruction was issued. The commit phase strictly follows program order to ensure data consistency and avoid data conflicts caused by out-of-order commits. The address merging mechanism supports merging multiple writes to the same cache line, reducing writeback pressure.

[0094] After the STB obtains the access right to the cache, it writes one cache line (512 bits) at a time.

[0095] As an example, Figure 3 The execution process of the Store instruction is as follows:

[0096] Stage 1: Address generation and access permission check

[0097] In the first stage of the pipeline, the Store instruction enters the address generation unit, which calculates the target memory access address range and checks for special cases such as page spans and unaligned accesses. The TLB is also queried to obtain the physical address and page table attribute information to identify instruction anomalies. The STD information sent by the IQ is matched against the instruction ID in the STQ table. If a match is successful, the STD information is recorded in the table entry; otherwise, it is temporarily stored in the STD Queue for subsequent merging.

[0098] Stage 2: Address write merging and RAW correlation check

[0099] In the second stage of the pipeline, the physical address, write data, and address access rights of the Store instruction are stored in the STQ. In order to ensure the correctness of instruction execution, the oldest Store instruction is checked for RAW dependency with the completed Load instruction. If an address conflict is found, a pipeline flush will be triggered and the Load instruction ID that executed the error will be marked. If the address, data, RAW dependency, and exception checks of the Store instruction are all completed, the Store instruction can be marked as completed. It should be pointed out that the completion of the Store instruction does not mean that the data has been written to the main memory or cache, but that the Store instruction has met the conditions for sequential retirement. The write operation is completed by the STB.

[0100] Stage 3: Data writing and cache access

[0101] The write operation of the Store instruction enters the execution phase after being marked as completed, and is written into the cache by the STB in the form of a cache line. The STB first allocates or merges the Store instruction sent by the STQ into the corresponding Entry according to the address and determines the write priority based on the Entry ID before initiating an access request to the Cache Arbiter. If the target cache line does not hit, it will choose to write directly or wait for the write back to be completed before executing the write operation according to the original state of the cache line. When the Store instruction becomes the oldest and meets the retirement conditions, the ROB will mark it as retired and update the oldest instruction ID. The table entry corresponding to the STQ will be released, and the entire life cycle of the Store will end.

[0102] Furthermore, the present disclosure can achieve data consistency and concurrency support in a multi-core / multi-threaded environment. The memory access unit can be seamlessly integrated into a multi-core or multi-threaded processor system, supports mainstream cache consistency protocols (such as MESI and MOESI), and improves concurrency performance through the following mechanisms:

[0103] (1) Coherence protocol interface: The memory access unit interacts with the cache coherence protocol controller through a dedicated interface, supporting cache line state transition, invalidation notification, write-back and other operations.

[0104] (2) Multi-threaded scheduling support: Each memory access channel can be bound to different threads. The upper-level scheduling unit dynamically adjusts the memory access resource allocation based on the thread priority and load conditions, thereby improving the overall throughput of the multi-threaded system.

[0105] As an example, Figure 1 As shown, the specific contents of the overall memory access unit architecture of the present disclosure are as follows:

[0106] The Load Queue (LDQ) uses a 4-Bank parallel architecture to store Load instruction information. This includes the instruction virtual address, physical address, cache access status, Flush status, and Except status. Compared to the C910's solution of using a separate Read Buffer to handle Cache Miss and Cross Cache line requests, this design reduces pipeline pressure by adding a dedicated entry in the LDQ outside of the original table entries to handle special requests, waiting for the oldest instruction in the Cross Cache line to execute before executing. This design effectively alleviates the pressure on the Read Buffer in multiple Miss situations and reduces pipeline congestion.

[0107] The Store Instruction Queue (STQ) uses a 2-Bank parallel architecture to store Store instruction information. This design sets up a Store Data Queue (STDQ) to temporarily store Store data, waiting for the address to be sent to the execution unit before performing a merge operation. When the Store instruction arrives earlier than the address, the C910 memory access unit allocates an entry to the Store data in the STQ. In the absence of an address, the data correlation condition is not met and only the entry is occupied. When the STQ receives the complete Store instruction data and address, it can be marked as completed. When the instruction becomes the oldest, it requests the STB to perform a merge operation. After the merge is successful, it enters the RAWQ to perform an urgent address correlation check. If there is no correlation conflict, it can be retired in sequence. The instruction information recorded in the remaining table entries is basically consistent with the Load instruction.

[0108] The Store Buffer (STB) merges store instructions in 512-bit cache line units. It checks for address matching. If there's no address overlap, an entry is created. Otherwise, the entry is merged into the existing entry and the entry's age is updated. After the merge is complete, the STB signals success to the STQ. For store instructions accessing the None-Cache (NC) address range, a special entry is created to handle these instructions. NC store instruction data is not written to the L1 D-cache but is directly passed to the WBQ for write-back to the L2 cache. After the STB wins cache access, it writes to the L1 D-cache in cache line units.

[0109] The Snoop Queue (SNQ) stores bus consistency requests. To ensure multi-core cache data consistency, the Interconnect sends requests to other cores to access or modify the same cache line. After receiving the request, the SNQ creates an entry and requests L1 DCache access rights, checks or modifies the corresponding cache line CSA status bit. If the cache line status is Dirty, the cache line needs to be written back to the memory and the other cores are notified of the data modification at the address.

[0110] The Cache Arbitration Block (CAB) is responsible for arbitrating data cache access requests. Due to hardware design characteristics, this cache supports up to four read and write requests. These requests include VA address information sent directly from the Agen, as well as all instructions waiting for cache access in the STQ, LDQ, MQ, and SNQ. Arbitration is performed based on hardware-set priorities.

[0111] The data cache is structured into three parts: Tag, Data, and CSA. The address is divided into Tags, which are used to determine cacheline hits, and Data, which stores the data within the corresponding address space. The CSA records the cache line status. Any cacheline modification or snoop request updates the CSA status. When a cache line's CSA is in the Dirty / Evict state, if a write or snoop request is received, the cache line is first replaced and passed to the WBQ for write back to memory.

[0112] The Read After Write Queue (RAWQ) is used to check for read-after-write dependencies between load and store instructions. The C910 checks data dependencies across all LDQ and STQ entries. This can cause a load instruction to occupy the entry if an older store instruction has not yet completed when the load instruction completes, impacting memory access bandwidth. To address this, this design creates a separate RAWQ specifically for dependency checking.

[0113] The Write Back Queue (WBQ) takes at least four clock cycles to complete a write operation from L1 Cache back to L2 Cache. To improve write-back efficiency, this design uses cache lines as units to complete write-back operations. WBQ is used to manage cache lines that are written back due to Dirty / Slience Evict.

[0114] The Cache Miss Queue (MQ) handles cache misses. When a cache line miss occurs, it initiates a read request to the L2 cache and wakes up any dormant instructions in the queue after the data is backfilled into the cache. Special entries are set for instructions accessing the None Cacheable address space. The queue also supports handling bus anomalies such as Bus Errors.

[0115] The forward unit (Forward) is responsible for processing the data forward and merge of the Load instruction. In the out-of-order execution scenario, the Store instruction may enter the sleep state due to reasons such as the DTLB page table missing and the cache line backfill, causing the Store instruction data to "linger" in various levels of the pipeline. When the old Store instruction data has not been written back, the data obtained by the Load instruction is often wrong. To solve this problem, a Forward mechanism is added to check the Store address in each module to determine the address overlap. Store instruction data may exist in STQ, MQ, STB, and WBQ. If the Load instruction is determined to be able to obtain all the data, the Load instruction is allowed to complete, and feedback is fed back to the previous emission unit to mark the corresponding register as Ready to reduce the instruction waiting time.

[0116] When a cache miss occurs in a memory access instruction, the cache line prefetch queue (PQ) is trained based on the instruction PC and the missing address, predicts the address space that may be accessed based on the confidence level, and initiates a read request in advance.

[0117] Example 2

[0118] In one embodiment of the present disclosure, a system for optimizing processor memory access with concurrent multi-channel instructions is provided, comprising:

[0119] An address generation unit is configured to obtain multiple memory access instructions to be executed in each cycle, the memory access instructions including load instructions and store instructions, calculate virtual addresses for the memory access instructions in the same clock cycle, and simultaneously send address query requests to the TLB and cache arbiter;

[0120] The load instruction execution unit is configured to use independent load queues and RAW check queues for processing. The load instruction goes through the address generation, TLB parallel query, L1 cache access, RAW check and data forwarding, and data write back stages in the load memory access channel.

[0121] The Store instruction execution unit is configured to use independent Store queues, STD queues, and Store Buffer queues for processing. The address and data of the Store instruction are separated, with the address portion entering the Store queue and the data portion entering the STD queue. They are then merged, combining multiple Store operations on the same cache line into one write. The merged data is written to the Store Buffer queue, which then submits it to the cache.

[0122] As an embodiment, the processor memory access optimization implementation system with multi-channel concurrent instructions proposed in the present invention can be widely used in: high-performance general-purpose processors (such as servers, desktop-level CPUs); embedded processors (such as mobile terminals, IoT chips); dedicated accelerators (such as AI reasoning, image processing, encryption chips); multi-core heterogeneous systems (such as system-on-chip SoCs, cloud data centers), etc.

[0123] Example 3

[0124] In one embodiment of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method for optimizing processor memory access with concurrent multi-channel instructions is implemented.

[0125] Example 4

[0126] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, which is used to store computer instructions. When the computer instructions are executed by a processor, a method for optimizing processor memory access with multi-channel concurrency of instructions is implemented.

[0127] Example 5

[0128] In one embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for optimizing processor memory access with multi-channel concurrency of instructions.

[0129] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes ​ A step that specifies a function in one or more boxes.

[0131] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that, based on the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A method for optimizing processor memory access with concurrent multi-channel instructions, characterized in that: include: Get multiple memory access instructions to be executed in each cycle, including Load instructions and Store instructions; Multiple memory access instructions are assigned to independent and parallel memory access channels based on resource occupancy and scheduling priority, including the Load memory access channel and the Store memory access channel. Independent Load queues and RAW check queues are set up in the Load memory access channel, and independent Store queues, STD queues, and Store Buffer queues are set up in the Store memory access channel. The load instruction goes through the address generation, TLB parallel query, L1 cache access, RAW check and data forward, and data write back stages in the load memory access channel. The address and data of the Store instruction are separated. The address part enters the Store queue and the data part enters the STD queue. They are then merged to combine multiple Store operations on the same Cache Line into one write. The merged data is written to the Store Buffer queue, which then submits it to the cache.

2. The method for optimizing processor memory access with multi-channel concurrent instructions according to claim 1, wherein: The out-of-order scheduling unit distributes the memory access instructions to be executed to each memory access channel according to resource occupancy and scheduling priority. Each memory access channel independently processes Load instructions and Store instructions. A dynamic load balancing algorithm is used to monitor the queue depth and resource utilization of each memory access channel in real time, and dynamically adjust the allocation strategy according to resource occupancy, instruction type and priority to ensure load balance among each memory access channel.

3. The method for optimizing processor memory access with multi-channel concurrent instructions according to claim 1, wherein: The RAW check queue set up in the Load memory access channel is used to perform read-after-write dependency detection between Load instructions and Store instructions. The address and data status of the preceding Store instruction are known in advance during the Load instruction issuance phase and recorded in real time. Each Load channel can access the RAW check queue in parallel to detect whether there are related Store instructions. When it is detected that the relevant Store data is already in the Store queue or Store Buffer, the Load instruction directly obtains the data through the forward push mechanism without waiting for the Store data to be written back to the main memory or cache.

4. The method for optimizing processor memory access with multi-channel concurrent instructions according to claim 1, wherein: The address portion and data portion of the Store instruction are processed separately. The address portion is independently managed by the Store queue, and the data portion is independently stored through the STD queue. When both parts of the Store instruction information are in place, the merged data is written to the Store Buffer queue through a merge mechanism, and age vector management is used to ensure that Store instructions are submitted for write in sequence. The address merge mechanism supports merging multiple writes to the same cache line to reduce write-back pressure.

5. The method for optimizing processor memory access with multi-channel concurrent instructions according to claim 1, wherein: The address generation unit completes the virtual address calculation and TLB query of the Load instruction and Store instruction in the same clock cycle. The VIPT mechanism is used to achieve high parallelism of address generation, TLB query and L1Cache access. The index and tag of the virtual address are separated when the instruction is issued; the index part is directly used for L1 Data Cache indexing, and the tag part waits for the TLB to return the physical address before performing tag comparison. Among them, address generation and TLB query are carried out synchronously without serial waiting. If the TLB is queried, the physical address tag comparison is completed immediately, realizing a zero-delay path for cache access.

6. The method for optimizing processor memory access with multi-channel concurrent instructions according to claim 1, wherein: The Store Buffer queue adopts a ring buffer structure, supports high-concurrency writing and sequential submission, improves write bandwidth and ensures program sequential consistency.

7. A processor memory access optimization implementation system with multi-channel concurrent instructions, characterized in that: include: An address generation unit is configured to obtain multiple memory access instructions to be executed in each cycle, the memory access instructions including load instructions and store instructions, calculate virtual addresses for the memory access instructions in the same clock cycle, and simultaneously send address query requests to the TLB and cache arbiter; The load instruction execution unit is configured to use independent load queues and RAW check queues for processing. The load instruction goes through the address generation, TLB parallel query, L1 cache access, RAW check and data forwarding, and data write back stages in the load memory access channel. The Store instruction execution unit is configured to use independent Store queues, STD queues, and Store Buffer queues for processing. The address and data of the Store instruction are separated, with the address portion entering the Store queue and the data portion entering the STD queue. They are then merged, combining multiple Store operations on the same cache line into one write. The merged data is written to the Store Buffer queue, which then submits it to the cache.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for optimizing processor memory access with multi-channel concurrency of instructions as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the processor memory access optimization implementation method of multi-channel concurrent instructions according to any one of claims 1 to 6 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a method for optimizing processor memory access with multi-channel concurrency of instructions as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • IO scheduling method and device for multichannel SSD solid-state disks

    CN107092445A

  • Data caching method, data caching device and processor

    CN117348934A

  • Processor and method for implementing instruction support for hash algorithms

    US20100250966A1

Cited By

  • Memory access delay acquisition method and device, storage medium and program product

    CN122019310A

  • Access delay acquisition method, device, storage medium and program product

    CN122019310B

  • Memory access instruction processing method and device, equipment, storage medium and program product

    CN122132087A

  • Memory access unit, memory access instruction execution method and chip

    CN122240187A