Processor based on RISC-V decoupled vector memory unit
Through the RISC-V decoupled vector memory access unit design, the timing and storage consistency issues between the vector memory access unit and the scalar memory access unit are solved, efficient vector memory access processing is achieved, and system complexity is reduced.
Patent Information
- Application Number
- CN202510815768.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing fifth-generation reduced instruction set vector memory access unit implementation method has the problem of the timing path impact of the vector memory access unit and the scalar memory access unit and the difficulty of storage consistency management.
A RISC-V-based decoupled vector memory access unit design is adopted. The decoding unit identifies and splits vector memory access instructions into multiple sub-requests, uses the scalar memory access unit for storage access, and monitors exceptions through the execution management unit to achieve decoupling and exception handling of vector memory access instructions.
It reduces the timing risk of vector memory access units, reduces the coupling degree between vector and scalar memory access units, reduces the difficulty of managing storage consistency, and improves memory access efficiency.
Smart Images

Figure CN120335868B_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of processor technology, and in particular relates to a processor implemented based on a RISC-V decoupled vector memory access unit. Background Art
[0002] The processor is the computing and control core of an information processing system. Modern processor instruction set architectures (ISAs) can be divided into complex instruction set architectures (CISCs) and reduced instruction set architectures (RISCs). The CISCs contain instructions for reading data from memory, performing calculations in registers, and then writing the results back to memory. These instructions are relatively complex but powerful. The RISCs can only access memory through dedicated memory access instructions. While single instructions are simple to use, they can reduce memory access efficiency. The RISCs have now reached their fifth generation (RISC-V) and have introduced vector extensions, which include dedicated vector memory access instructions for moving data between vector registers and memory. Compared to scalar memory access instructions, which can only load or store a single piece of data at a time, vector memory access instructions can access multiple pieces of data at once, significantly improving memory access efficiency.
[0003] The vector memory access instructions in the fifth-generation reduced instruction set (RISC) include two types of instructions: load and store. Load instructions load data from memory into vector registers, while store instructions write data from vector registers to memory. Each type of instruction has three memory access modes: the first, starting from a base address, continuously loads or stores multiple data items, referred to as unit-stride addressing mode; the second, starting from a base address, loads or stores one data item at a fixed address interval, referred to as strided addressing mode; and the third, consisting of a base address and a set of offset addresses, where the actual memory address of each element is the base address plus an offset address within the set, referred to as indexed addressing mode. The vector memory access instructions in the fifth-generation RISC only support exact vector exceptions, meaning that when a memory access exception occurs, the index of the element where the exception occurred must be precisely returned.
[0004] Existing fifth-generation reduced instruction set vector memory unit implementations fall into two main categories: the first is a hybrid design that combines the vector memory unit with the scalar memory unit. By adding the components and control mechanisms required for vector memory access to the original scalar memory unit, the existing scalar memory is reused for vector memory access, thereby reducing hardware overhead; the second is a separate design of the vector memory unit and the scalar execution unit, each of which accesses memory independently. This approach provides a dedicated memory access path for vector memory access instructions, allowing for specialized parallel memory access optimization design to improve vector memory access throughput.
[0005] However, the two main implementation methods of the existing fifth-generation reduced instruction set vector memory access unit have the following problems: First, scalar memory access has high requirements for memory access latency. Therefore, the mixed design of vector memory access unit and scalar memory access unit will affect the original scalar memory access timing path, and thus affect the latency of scalar memory access; second, the independent design of vector memory access unit and scalar execution unit will make the management of memory access requests difficult, increasing the difficulty of system storage consistency management.
[0006] Therefore, it is necessary to optimize the architecture of the vector memory access unit to solve the above technical problems. Summary of the Invention
[0007] The present invention provides a processor implemented based on a RISC-V decoupled vector memory access unit, aiming to solve the timing and storage consistency management problems existing in the processing logic of vector memory access in existing processors.
[0008] To solve the above technical problems, the present invention provides a processor implemented based on a RISC-V decoupled vector memory access unit, the processor comprising a decoding unit, a storage unit, a vector memory access unit, a scalar memory access unit, an execution management unit, and a front-end instruction fetch unit, wherein:
[0009] The decoding unit is used to parse the instruction stream and identify scalar memory access instructions and vector memory access instructions from the instruction stream, wherein when the vector memory access instruction is identified, the number of vector memory access requests that need to be split from the vector memory access instruction is further identified;
[0010] The storage unit is used to perform memory access according to the memory access request and return memory access response information corresponding to the memory access request;
[0011] The vector memory access unit is used to receive the vector memory access instruction identified by the decoding unit, and parse the vector memory access instruction to generate a number of vector sub-requests that are the same as the number of the vector memory access requests; the vector memory access unit is also used to: send the vector sub-requests to the scalar memory access unit;
[0012] The scalar memory access unit is used to receive the scalar memory access instruction recognized by the decoding unit and execute the scalar memory access instruction; the scalar memory access unit is further used to: receive the vector sub-request sent by the vector memory access unit, send the vector sub-request to the storage unit, receive the memory access response information corresponding to the vector sub-request returned by the storage unit, and return the memory access response information to the vector memory access unit;
[0013] The execution management unit is used to monitor the execution status of the vector memory access instruction, wherein when an imprecise exception is detected in the execution of the vector memory access instruction, a stop signal is sent to the scalar memory access unit, and at the same time, a re-execution signal is sent to the front-end instruction fetch unit;
[0014] The front-end instruction fetch unit is used to respond to the re-execution signal issued by the execution management unit, re-acquire the vector memory access instruction, and obtain the instruction stream after restoring the instruction context data according to the vector memory access instruction, and send the instruction stream to the decoding unit.
[0015] Furthermore, the vector memory access unit further includes a vector register file and a vector index cache queue, wherein the vector register file is used to store data read from the storage unit and to write the stored data back to the storage unit;
[0016] Furthermore, when the vector memory access instruction is in index addressing mode, a vector register value storing an index is read from the vector register file, and the vector register value is written into a vector index cache queue.
[0017] Furthermore, the vector memory access unit further includes a vector load execution unit, a vector load instruction cache queue, a vector load instruction splitting unit and a vector memory access request management unit, wherein:
[0018] The vector load execution unit is configured to receive the vector memory access instruction for common vector load, and allocate a corresponding first instruction entry in the vector load instruction cache queue according to the vector memory access instruction;
[0019] The vector load instruction splitting unit is configured to split the vector load request according to the first instruction entry corresponding to the vector load instruction cache queue, to obtain the vector sub-requests, including a first type of memory access request including a unit stride load memory access request and / or a second type of memory access request including a fixed stride load memory access request and an index load memory access request, and a number corresponding to the number of the vector memory access requests;
[0020] The vector memory access request management unit is used to send the vector sub-request to the scalar memory access unit, and receive the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit; the vector memory access request management unit is also used to: corresponding to the first type of memory access request and the second type of memory access request, write the memory access response information into the first instruction entry corresponding to the vector load instruction cache queue.
[0021] Furthermore, the vector load execution unit is further configured to:
[0022] Performing a data circular right shift process on the data corresponding to the first type of memory access request returned by the scalar memory access unit so that the data of the element that needs to be loaded first is located at the lowest bit, and writing the processed data into the vector register file.
[0023] Furthermore, the vector memory access unit further includes a vector storage instruction execution unit, a vector storage instruction cache queue and a vector storage instruction splitting unit, wherein:
[0024] The vector store instruction execution unit is configured to receive the vector memory access instruction for vector storage, allocate a corresponding second instruction entry in the vector store instruction cache queue according to the vector memory access instruction, and write data required for vector storage by the vector memory access instruction into the second instruction entry;
[0025] The vector storage instruction splitting unit is configured to split the vector storage request according to the corresponding second instruction entry in the vector storage instruction cache queue to obtain the vector sub-requests, including a third type of memory access request including unit stride storage access requests and / or a fourth type of memory access request including fixed stride storage access requests and index storage access requests, and a number corresponding to the number of the vector memory access requests, wherein:
[0026] Corresponding to the third type of memory access request, the data read from the second instruction entry is shifted left so that the data of the element that needs to be written into the storage unit first is aligned to the starting address for starting writing.
[0027] Furthermore, the vector memory access unit further includes an exception handling unit, a vector load-store instruction cache queue, and a vector load-store instruction splitting unit:
[0028] The exception handling unit is configured to receive the vector memory access instruction for generating an imprecise vector exception, and allocate a corresponding third instruction entry in the vector load / store instruction cache queue according to the vector memory access instruction, wherein, if the vector memory access instruction is for vector storage, data required for vector storage by the vector memory access instruction is simultaneously written into the third instruction entry;
[0029] The vector load-store instruction splitting unit is used to split the vector memory access request according to the corresponding third instruction entry in the vector load-store instruction cache queue to obtain the vector sub-requests whose number corresponds to the number of the vector memory access requests.
[0030] Furthermore, the vector memory access request management unit is further configured to:
[0031] The vector sub-requests sent by the corresponding vector load / store instruction splitting unit are sent to the scalar memory access unit in sequence according to the splitting order of the vector sub-requests; after receiving the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit, the next vector sub-request is sent to the scalar memory access unit.
[0032] Furthermore, the exception handling unit is further configured to:
[0033] Receive the memory access response information corresponding to the vector sub-request returned by the vector memory access request management unit, wherein: if the memory access response information shows that a precise exception occurs in the vector sub-request, send the element index of the vector sub-request in the vector load-store instruction cache queue to the vector memory access request management unit, and if the vector sub-request is for vector loading, further write the data before the element index in the vector load-store instruction cache queue into the vector register stack.
[0034] Furthermore, the vector memory access request management unit is further configured to:
[0035] After receiving the memory access response information corresponding to all the vector sub-requests, generating a completion signal and sending it to the execution management unit;
[0036] After receiving the element index information corresponding to the precise exception occurring in the vector sub-request, the element index information is forwarded to the execution management unit to support the precise vector exception.
[0037] The beneficial effect achieved by the present invention lies in proposing a processor for the RISC-V instruction set based on a RISC-V decoupled vector memory access unit. In the processor, a decoupled vector memory access unit design is adopted, and a memory access instruction is split into several memory access sub-requests according to the addressing mode of the vector memory access instruction, and the memory access paths of normal and abnormal vector instructions are distinguished, thereby reducing the timing risk of the vector memory access unit. In addition, by sending the memory access sub-request to the scalar memory access unit for accessing the storage unit, the degree of coupling between the vector and scalar memory access units is reduced, and the existing scalar memory access path is reused, thereby reducing the difficulty of managing storage consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 1 is a schematic diagram of the structure of a processor implemented based on a RISC-V decoupled vector memory access unit provided by an embodiment of the present invention;
[0039] Figure 2 1 is a schematic structural diagram of a vector memory access unit in a processor implemented based on a RISC-V decoupled vector memory access unit according to an embodiment of the present invention;
[0040] Figure 3 1 is a schematic diagram of a vector load instruction processing flow provided by an embodiment of the present invention;
[0041] Figure 4 1 is a schematic diagram of a vector storage instruction processing flow provided by an embodiment of the present invention;
[0042] Figure 5 This is a schematic diagram of the processing flow of an imprecise vector memory access instruction provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0044] Please refer to Figure 1 , Figure 1 1 is a schematic diagram of the structure of a processor 100 implemented based on a decoupled vector memory access unit 103 according to an embodiment of the present invention. The processor 100 includes a decoding unit 101, a storage unit 102, a vector memory access unit 103, a scalar memory access unit 104, an execution management unit 105, and a front-end instruction fetch unit 106, wherein:
[0045] The decoding unit 101 is used to parse the instruction stream and identify scalar memory access instructions and vector memory access instructions from the instruction stream, wherein when the vector memory access instruction is identified, the number of vector memory access requests that need to be split from the vector memory access instruction is further identified;
[0046] The storage unit 102 is used to perform memory access according to the memory access request and return memory access response information corresponding to the memory access request;
[0047] The vector memory access unit 103 is used to receive the vector memory access instruction identified by the decoding unit 101, and parse the vector memory access instruction to generate a number of vector sub-requests equal to the number of the vector memory access request; the vector memory access unit 103 is further used to: send the vector sub-requests to the scalar memory access unit 104;
[0048] The scalar memory access unit 104 is configured to receive the scalar memory access instruction recognized by the decoding unit 101 and execute the scalar memory access instruction; the scalar memory access unit 104 is further configured to: receive the vector sub-request sent by the vector memory access unit 103, send the vector sub-request to the storage unit 102, receive the memory access response information corresponding to the vector sub-request returned by the storage unit 102, and return the memory access response information to the vector memory access unit 103;
[0049] The execution management unit 105 is configured to monitor the execution status of the vector memory access instruction, wherein when an imprecise exception is detected in the execution of the vector memory access instruction, a stop signal is sent to the scalar memory access unit 104 and a re-execution signal is sent to the front-end instruction fetch unit 106;
[0050] The front-end instruction fetch unit 106 is used to respond to the re-execution signal issued by the execution management unit 105, re-acquire the vector memory access instruction, and obtain the instruction stream after restoring the instruction context data according to the vector memory access instruction, and send the instruction stream to the decoding unit 101.
[0051] Based on the processor 100 described above, the instruction flow in the embodiment of the present invention is specifically as follows:
[0052] The front-end instruction fetch unit 106 first obtains the original instruction stream from the memory, forms an instruction sequence to be processed, and sends it to the decoding unit 101.
[0053] The decoding unit 101 parses the instruction stream: when a scalar memory access instruction is identified, it is directly marked as a scalar type; when a vector memory access instruction is identified, it further calculates how many vector sub-requests the instruction needs to be split into.
[0054] For scalar memory access instructions, the decoding unit 101 directly passes the scalar memory access instruction to the scalar memory access unit 104;
[0055] The scalar memory access unit 104 executes the scalar memory access instruction: generates a scalar memory access request and sends it to the storage unit 102; receives memory access response information returned by the storage unit 102, and completes the scalar access process.
[0056] For a vector memory access instruction, the decoding unit 101 transmits the vector memory access instruction and the number of split vector sub-requests to the vector memory access unit 103 .
[0057] The vector memory access unit 103 splits the original vector instruction into multiple vector sub-requests according to the number of sub-requests. Then, the vector memory access unit 103 sends each vector sub-request to the scalar memory access unit 104 in turn, using its execution capability to process the sub-tasks after the vector is split:
[0058] After receiving the quantum request, the scalar memory access unit 104 sends the request to the storage unit 102 according to the process of processing scalar instructions;
[0059] The storage unit 102 performs a memory access and returns memory access response information corresponding to the sub-request;
[0060] The scalar memory access unit 104 returns the response information to the vector memory access unit 103 , which then aggregates the results of all sub-requests to complete the overall execution of the vector instruction.
[0061] During the execution of the instruction stream, the execution management unit 105 monitors the execution status of the vector memory access instruction in real time:
[0062] If an imprecise exception is detected (such as a data error or address out-of-bounds in a pointer-to-memory access instruction), two actions are immediately triggered:
[0063] Sending a stop signal to the scalar memory access unit 104 to terminate the execution of all currently unfinished vector quantum requests and scalar instructions;
[0064] A re-execution signal is sent to the front-end instruction fetch unit 106 to request re-acquisition of the original vector memory access instruction.
[0065] In response to the re-execution signal issued by the execution management unit 105, the front-end instruction fetch unit 106 re-reads the instruction stream including the original vector memory access instruction from the memory;
[0066] Restore instruction context data (such as register status, program counter, etc.) to ensure that the environment during re-execution is consistent with that before the exception;
[0067] The restored instruction stream is sent to the decoding unit 101 again to start a new round of parsing and execution process to correct the exception and complete the vector instruction.
[0068] In the embodiment of the present invention, in order to specifically implement the vector register reading and writing of the vector memory access instruction, please refer to Figure 2 , Figure 2 1 is a schematic diagram of the structure of a vector memory access unit 103 in a processor 100 implemented based on a decoupled vector memory access unit 103 according to an embodiment of the present invention. The vector memory access unit 103 further includes a vector register file 1031 and a vector index cache queue 1032. The vector register file 1031 is used to store data read from the memory unit 102 and to write the stored data back to the memory unit 102.
[0069] When the vector memory access instruction is in indexed addressing mode, the vector register value storing the index is read from the vector register file, and the vector register value is written into the vector index cache queue 1032 .
[0070] In this embodiment of the present invention, the vector register file 1031 serves as a data transfer station for executing vector memory access instructions, facilitating the loading and storing of required data. For indexed addressing mode, the vector register values storing the indexes are read from the vector register file 1031 and written to the vector index cache queue 1032 for splitting vector memory access requests in indexed addressing mode.
[0071] In the embodiment of the present invention, in order to implement common vector load instruction processing, the vector memory access unit 103 further includes a vector load execution unit 1033, a vector load instruction cache queue 1034, a vector load instruction splitting unit 1035, and a vector memory access request management unit 1036, wherein:
[0072] The vector load execution unit 1033 is configured to receive the vector memory access instruction for common vector load, and allocate a corresponding first instruction entry in the vector load instruction cache queue 1034 according to the vector memory access instruction;
[0073] The vector load instruction splitting unit 1035 is configured to split the vector load request according to the first instruction entry corresponding to the vector load instruction cache queue 1034, to obtain the vector sub-requests, including a first type of memory access request including a unit stride load memory access request and / or a second type of memory access request including a fixed stride load memory access request and an index load memory access request, and a number corresponding to the number of the vector memory access requests;
[0074] The vector memory access request management unit 1036 is used to send the vector sub-request to the scalar memory access unit 104, and receive the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit 104; the vector memory access request management unit 1036 is also used to: corresponding to the first type of memory access request and the second type of memory access request, write the memory access response information into the first instruction entry corresponding to the vector load instruction cache queue 1034.
[0075] Ordinary vector load instructions mainly read the required data content from the storage unit 102 according to the load request of the instruction. Among them, the first type of memory access request of the unit stride load memory access request is different from the second type of memory access request including the fixed stride load memory access request and the index load memory access request in that the size of the data bandwidth is different. The unit stride memory access request is to access continuous addresses, so it will be split into memory access requests for large bandwidth data, reading a piece of continuous address data at a time, while the fixed stride load memory access request and the index load memory access request will be split into several memory access requests of ordinary data bandwidth according to the number of elements, and each request only reads the data of one element. For vector load requests for large bandwidth data, the vector memory access unit 103 in the embodiment of the present invention supports address non-aligned access, and the scalar memory access unit 104 is also used to:
[0076] Determine whether the vector sub-request sent by the vector memory access unit 103 belongs to the first type of memory access request: if so, perform a data cyclic right shift on the data corresponding to the first type of memory access request returned by the scalar memory access unit 104, so that the data of the element that needs to be loaded first is located at the lowest bit, and write the processed data into the vector register stack 1031, so as to keep it consistent with the order of elements written into the vector register.
[0077] According to the description in the above embodiment, please combine Figure 3 The vector load instruction processing flow diagram shown in FIG. 1 shows a general vector load instruction processing process including:
[0078] S201, the vector load execution unit 1033 loads a vector memory access instruction and allocates a corresponding first instruction entry in the vector load instruction cache queue 1034;
[0079] S202, the vector load instruction splitting unit 1035 splits the vector load request;
[0080] S203, the vector memory access request management unit 1036 sends the vector sub-request to the scalar memory access unit 104;
[0081] S204, the vector memory access unit 103 performs address alignment processing on the data requested by the vector;
[0082] S205 . The vector memory access request management unit 1036 writes the memory access response information into the first instruction entry corresponding to the vector load instruction cache queue 1034 , and writes the data into the vector register file 1031 .
[0083] In the embodiment of the present invention, in order to implement common vector storage instruction processing, the vector memory access unit 103 further includes a vector storage instruction execution unit 1037, a vector storage instruction cache queue 1038, and a vector storage instruction splitting unit 1039, wherein:
[0084] The vector store instruction execution unit 1037 is configured to receive the vector memory access instruction for vector storage, allocate a corresponding second instruction entry in the vector store instruction cache queue 1038 according to the vector memory access instruction, and write data required for vector storage by the vector memory access instruction into the second instruction entry;
[0085] The vector storage instruction splitting unit 1039 is configured to split the vector storage request according to the corresponding second instruction entry in the vector storage instruction cache queue 1038 to obtain the vector sub-requests, including a third type of memory access request including unit stride storage access requests and / or a fourth type of memory access request including fixed stride storage access requests and index storage access requests, and a number corresponding to the number of the vector memory access requests, wherein:
[0086] In response to the third type of memory access request, the data read from the second instruction entry is shifted left so that the data of the element that needs to be written into the storage unit 102 first is aligned to the starting address for writing.
[0087] Ordinary vector storage instructions mainly write the data content to be stored into the storage unit 102 according to the storage request of the instruction. Among them, similar to the way of processing load instructions, the third type of memory access request of unit stride load memory access request and the fourth type of memory access request including fixed stride load memory access request and index load memory access request are different in the size of data bandwidth. The unit stride memory access request will be split into memory access requests for large bandwidth data, writing a piece of continuous address data at a time, while the fixed stride load memory access request and index load memory access request will be split into several memory access requests of normal data bandwidth according to the number of elements, and each request only writes the data of one element. In an embodiment of the present invention, the vector storage request for large bandwidth data supports address non-aligned access. When splitting, the data read from the vector register stack 1031 will be left shifted, so that the element first written to the memory in the vector register is aligned to the starting address for writing.
[0088] According to the description in the above embodiment, please combine Figure 4 The vector store instruction processing flow diagram shown in FIG. 1 is a flow chart showing the processing flow of a common vector store instruction. The processing process of a common vector store instruction includes:
[0089] S301, the vector storage instruction execution unit 1037 receives a vector memory access instruction, allocates a corresponding second instruction entry in the vector storage instruction cache queue 1038, and performs address alignment processing on the data of the vector memory access instruction;
[0090] S302, the vector storage instruction splitting unit 1039 splits the vector storage request;
[0091] S303 , the vector memory access request management unit 1036 sends the vector sub-request to the scalar memory access unit 104 ; subsequent processing is consistent with step S205 .
[0092] In the embodiment of the present invention, in order to implement the processing of inexact exception vector memory access instructions, the vector memory access unit 103 further includes an exception handling unit 10310, a vector load-store instruction cache queue 10311, and a vector load-store instruction splitting unit 10312:
[0093] The exception handling unit 10310 is configured to receive the vector memory access instruction for generating an imprecise vector exception, and allocate a corresponding third instruction entry in the vector load / store instruction cache queue 10311 according to the vector memory access instruction, wherein, if the vector memory access instruction is for vector storage, data required for vector storage by the vector memory access instruction is simultaneously written into the third instruction entry;
[0094] The vector load store instruction splitting unit 10312 is used to split the vector memory access request according to the corresponding third instruction entry in the vector load store instruction cache queue 10311 to obtain the vector sub-requests whose number corresponds to the number of the vector memory access requests.
[0095] Based on the architecture of the vector memory access unit 103 for processing the aforementioned imprecise exception vector memory access instructions, the exception handling unit 10310 sets a vector load and store instruction cache queue 10311 that is independent of the vector load instruction cache queue 1034 and the vector store instruction cache queue 1038 to decouple the data flow from that of ordinary vector memory access instructions and reduce the mutual influence of timing paths. During implementation, vector load instructions and vector store instructions can share a cache queue to save area overhead. The difference from the split design of ordinary vector memory access instructions is that the memory access request during exception handling is split into memory access requests for normal bandwidth data according to the number of sub-vector requests.
[0096] Furthermore, the vector memory access request management unit 1036 is further configured to:
[0097] The vector sub-request sent by the corresponding vector load-store instruction splitting unit 10312 is sent to the scalar memory access unit 104 in sequence according to the splitting order of the vector sub-request; after receiving the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit 104, the next vector sub-request is sent to the scalar memory access unit 104.
[0098] After sending a sub-vector request to the vector memory access request management unit 1036, the vector load / store instruction splitting unit 10312 waits for a response to the memory access request before sending the next split memory access request. If an exception occurs after receiving a response to the memory access request, the exception handling unit 10310 records the element index corresponding to the memory access request and notifies the vector memory access request management unit 1036, thereby supporting precise vector exceptions.
[0099] The exception handling unit 10310 is further configured to:
[0100] Receive the memory access response information corresponding to the vector sub-request returned by the vector memory access request management unit 1036, wherein: if the memory access response information shows that a precise exception occurs in the vector sub-request, then the element index of the vector sub-request in the vector load storage instruction cache queue 10311 is sent to the vector memory access request management unit 1036, and if the vector sub-request is for vector loading, then the data before the element index in the vector load storage instruction cache queue 10311 is further written into the vector register stack 1031.
[0101] In the embodiment of the present invention, a unified vector memory access request management unit 1036 is provided in the vector memory access unit 103 to implement the management and execution state synchronization of vector memory access instructions. The vector memory access request management unit 1036 is further used to:
[0102] After receiving the memory access response information corresponding to all the vector sub-requests, generating a completion signal and sending it to the execution management unit 105;
[0103] After receiving the element index information corresponding to the precise exception occurring in the vector sub-request, the element index information is forwarded to the execution management unit 105 to support the precise vector exception.
[0104] According to the description in the above embodiment, please combine Figure 5 The flowchart of the inexact vector memory access instruction processing is shown in FIG. The inexact vector instruction processing process includes:
[0105] S401, the exception handling unit 10310 receives the vector memory access instruction (including the vector load instruction and the vector store instruction) in which the inexact vector exception occurs, and allocates a corresponding third instruction entry in the vector load / store instruction cache queue 10311;
[0106] S402, the vector load store instruction splitting unit 10312 splits the vector memory access request;
[0107] S403, the vector memory access request management unit 1036 sends the vector sub-request to the scalar memory access unit 104;
[0108] After step S403, depending on whether an actual exception occurs in the execution of the vector request, the following steps are further included:
[0109] S4041: If no exception occurs, the vector load instruction is processed in the same manner as step S205;
[0110] S4042. If an exception occurs, the exception handling unit 10310 sends the element index of the vector sub-request in the vector load storage instruction cache queue 10311 to the vector memory access request management unit 1036, and if the vector sub-request is for vector loading, further writes the data before the element index in the vector load storage instruction cache queue 10311 into the vector register stack 1031.
[0111] Based on the architectural design of the above-mentioned vector memory access unit 103, the vector memory access request management unit 1036 can achieve full sub-request completion detection to avoid invalid waiting and sub-request level re-execution to reduce redundant operations through refined sub-request status management, targeted response to abnormal scenarios and cross-unit signal coordination; by controlling the impact of exceptions to the minimum range, it supports the reuse of partially correct results; as a transit unit between the vector memory access unit 103 and the execution management unit 105, it ensures the seamless connection of the vector memory access instruction execution process.
[0112] The beneficial effect achieved by the present invention lies in proposing a processor for the RISC-V instruction set based on a RISC-V decoupled vector memory access unit. In the processor, a decoupled vector memory access unit design is adopted, and a memory access instruction is split into several memory access sub-requests according to the addressing mode of the vector memory access instruction, and the memory access paths of normal and abnormal vector instructions are distinguished, thereby reducing the timing risk of the vector memory access unit. In addition, by sending the memory access sub-request to the scalar memory access unit for accessing the storage unit, the degree of coupling between the vector and scalar memory access units is reduced, and the existing scalar memory access path is reused, thereby reducing the difficulty of managing storage consistency.
[0113] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a mobile phone, computer, server, air conditioner, or network equipment) through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0114] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0115] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only a preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.
Claims
1. A processor based on a RISC-V decoupled vector memory access unit, characterized in that: The processor includes a decoding unit, a storage unit, a vector memory access unit, a scalar memory access unit, an execution management unit and a front-end instruction fetch unit, wherein: The decoding unit is used to parse the instruction stream and identify scalar memory access instructions and vector memory access instructions from the instruction stream, wherein when the vector memory access instruction is identified, the number of vector memory access requests that need to be split from the vector memory access instruction is further identified; The storage unit is used to perform memory access according to the memory access request and return memory access response information corresponding to the memory access request; The vector memory access unit is used to receive the vector memory access instruction identified by the decoding unit, and parse the vector memory access instruction to generate a number of vector sub-requests that are the same as the number of the vector memory access requests; the vector memory access unit is also used to: send the vector sub-requests to the scalar memory access unit; The scalar memory access unit is used to receive the scalar memory access instruction recognized by the decoding unit and execute the scalar memory access instruction; the scalar memory access unit is further used to: receive the vector sub-request sent by the vector memory access unit, send the vector sub-request to the storage unit, receive the memory access response information corresponding to the vector sub-request returned by the storage unit, and return the memory access response information to the vector memory access unit; The execution management unit is used to monitor the execution status of the vector memory access instruction, wherein when an imprecise exception is detected in the execution of the vector memory access instruction, a stop signal is sent to the scalar memory access unit, and at the same time, a re-execution signal is sent to the front-end instruction fetch unit; The front-end instruction fetch unit is used to respond to the re-execution signal issued by the execution management unit, re-acquire the vector memory access instruction, and obtain the instruction stream after restoring the instruction context data according to the vector memory access instruction, and send the instruction stream to the decoding unit.
2. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 1, characterized in that: The vector memory access unit further includes a vector register file and a vector index cache queue, wherein the vector register file is used to store data read from the storage unit and to write the stored data back to the storage unit; Furthermore, when the vector memory access instruction is in index addressing mode, a vector register value storing an index is read from the vector register file, and the vector register value is written into a vector index cache queue.
3. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 2, characterized in that: The vector memory access unit further includes a vector load execution unit, a vector load instruction cache queue, a vector load instruction splitting unit, and a vector memory access request management unit, wherein: The vector load execution unit is configured to receive the vector memory access instruction for common vector load, and allocate a corresponding first instruction entry in the vector load instruction cache queue according to the vector memory access instruction; The vector load instruction splitting unit is configured to split the vector load request according to the first instruction entry corresponding to the vector load instruction cache queue, to obtain the vector sub-requests, including a first type of memory access request including a unit stride load memory access request and / or a second type of memory access request including a fixed stride load memory access request and an index load memory access request, and a number corresponding to the number of the vector memory access requests; The vector memory access request management unit is used to send the vector sub-request to the scalar memory access unit, and receive the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit; the vector memory access request management unit is also used to: corresponding to the first type of memory access request and the second type of memory access request, write the memory access response information into the first instruction entry corresponding to the vector load instruction cache queue.
4. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 3, characterized in that: The vector load execution unit is further configured to: Performing a data circular right shift process on the data corresponding to the first type of memory access request returned by the scalar memory access unit so that the data of the element that needs to be loaded first is located at the lowest bit, and writing the processed data into the vector register file.
5. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 3, characterized in that: The vector memory access unit further includes a vector storage instruction execution unit, a vector storage instruction cache queue, and a vector storage instruction splitting unit, wherein: The vector store instruction execution unit is configured to receive the vector memory access instruction for vector storage, allocate a corresponding second instruction entry in the vector store instruction cache queue according to the vector memory access instruction, and write data required for vector storage by the vector memory access instruction into the second instruction entry; The vector storage instruction splitting unit is configured to split the vector storage request according to the corresponding second instruction entry in the vector storage instruction cache queue to obtain the vector sub-requests, including a third type of memory access request including unit stride storage access requests and / or a fourth type of memory access request including fixed stride storage access requests and index storage access requests, and a number corresponding to the number of the vector memory access requests, wherein: Corresponding to the third type of memory access request, the data read from the second instruction entry is shifted left so that the data of the element that needs to be written into the storage unit first is aligned to the starting address for starting writing.
6. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 3, characterized in that: The vector memory access unit also includes an exception handling unit, a vector load and store instruction cache queue, and a vector load and store instruction splitting unit: The exception handling unit is configured to receive the vector memory access instruction for generating an imprecise vector exception, and allocate a corresponding third instruction entry in the vector load / store instruction cache queue according to the vector memory access instruction, wherein, if the vector memory access instruction is for vector storage, data required for vector storage by the vector memory access instruction is simultaneously written into the third instruction entry; The vector load-store instruction splitting unit is used to split the vector memory access request according to the corresponding third instruction entry in the vector load-store instruction cache queue to obtain the vector sub-requests whose number corresponds to the number of the vector memory access requests.
7. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 6, characterized in that: The vector memory access request management unit is further configured to: The vector sub-requests sent by the corresponding vector load / store instruction splitting unit are sent to the scalar memory access unit in sequence according to the splitting order of the vector sub-requests; after receiving the memory access response information corresponding to the vector sub-request returned by the scalar memory access unit, the next vector sub-request is sent to the scalar memory access unit.
8. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 6, characterized in that: The exception handling unit is further configured to: Receive the memory access response information corresponding to the vector sub-request returned by the vector memory access request management unit, wherein: if the memory access response information shows that a precise exception occurs in the vector sub-request, send the element index of the vector sub-request in the vector load-store instruction cache queue to the vector memory access request management unit, and if the vector sub-request is for vector loading, further write the data before the element index in the vector load-store instruction cache queue into the vector register stack.
9. The processor implemented based on the RISC-V decoupled vector memory access unit according to claim 8, characterized in that: The vector memory access request management unit is further configured to: After receiving the memory access response information corresponding to all the vector sub-requests, generating a completion signal and sending it to the execution management unit; After receiving the element index information corresponding to the precise exception occurring in the vector sub-request, the element index information is forwarded to the execution management unit to support the precise vector exception.
Citation Information
Patent Citations
Memory access instruction processing method and processor
CN119621149A
Decoupled scalar / vector computer architecture system and method
US7334110B1