Processor architecture of decoupling vector processing unit based on data path splitting

By equiping each data path of the vector processing unit with a separate channel execution state management subunit and a unified memory access scheduling unit, the timing and storage consistency management problems of micro-instruction split processing logic in the existing processor architecture are solved, more accurate scheduling is achieved and timing risks is reduced, and the reliability of the processor system is improved.

CN120353500AActive Publication Date: 2025-07-22RISECORE (CHENGDU) TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510853445.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the existing processor architecture, the micro-instruction split processing logic has difficulty in managing timing and storage consistency, resulting in difficulty in precise scheduling of micro-instruction execution status and increasing timing risks.

Method used

Each data path of the vector processing unit is equipped with a separate path execution state management subunit. The scalar and vector memory access requests are managed through a unified memory access scheduling unit, and the micro-instruction splitting and register read request paths are optimized.

Benefits of technology

It realizes more accurate micro-instruction scheduling reference, reduces timing risks, and reduces the difficulty of memory consistency management, and enhances the reliability of the processor system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353500A_ABST
    Figure CN120353500A_ABST
Patent Text Reader

Abstract

The invention relates to a processor architecture of a decoupling vector processing unit based on data path splitting. The processor architecture comprises an instruction fetching unit, a decoding unit, a vector processing unit, a scalar processing unit, a memory access scheduling unit and a storage unit. In the processor architecture, each data path of a vector processing unit is provided with an independent path execution state management subunit, so that more accurate scheduling reference is provided for a main scheduler to perform splitting processing on microinstructions, the difficulty of tracking and managing the microinstructions is reduced, meanwhile, a register read request path is optimized, and the time sequence risk is reduced; moreover, the architecture manages the sequence of scalar and vector memory access requests through a uniform memory access scheduling unit, so that the management difficulty of memory consistency is reduced, and the reliability of a processor system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applicable to the field of processor technology, and particularly relates to a processor architecture of a decoupled vector processing unit based on data path splitting. Background Art

[0002] A processor is the computing and control core of an information processing system. Modern processor instruction set architectures (ISAs) can be divided into complex instruction set (CISC) and reduced instruction set (RISC). The former has instructions that read data from memory to a register for calculation and then write the calculation result back to memory. The instruction set is relatively complex but powerful. The latter can only access memory through dedicated memory access instructions. The operation of a single instruction is simple, but the memory access efficiency may be reduced. The reduced instruction set has now developed to the fifth generation (RISC-V), and a vector extension instruction set has been introduced, enabling the processor to support vector processing. Vector processing technology effectively reduces the number of instruction fetches and instruction dispatches by processing multiple data with one instruction, improving the parallelism of data processing, and thus has been widely used in high-performance processors.

[0003] In related technologies, there are mainly two designs for the vector processing unit in a processor: a tightly coupled architecture and a decoupled architecture. The tightly coupled vector architecture mixes a traditional scalar processing unit and a vector processing unit together, and they share a decoding, scheduling, and execution pipeline. Although it saves some hardware overhead, it leads to a high design complexity and also restricts the upper limit of the processor's operating frequency. The decoupled vector architecture, on the other hand, designs the scalar processing unit and the vector processing unit separately. After the decoding module recognizes a vector instruction, it sends it to the vector processing unit, which is independently responsible for executing the vector instruction. In addition, by further splitting the processing of large-bitwidth vector data into multiple independent small-bitwidth data processing paths, the timing bottleneck of the vector data path design can be effectively reduced, and the decoupled vector architecture has gradually become the mainstream of the vector architecture in high-performance processors.

[0004] Generally, in the processor pipeline based on the decoupled vector architecture, instructions are sent from the instruction fetch unit to the decoding unit. When the decoding unit recognizes that an instruction is a vector instruction, it sends it to the vector processing unit at the back end. The vector processing unit includes a main scheduler, a data path unit, a cross-path unit, a memory access unit, and an execution status management unit. Among them: The main scheduler is responsible for splitting an instruction into micro-instructions according to the data path bandwidth, and each micro-instruction is responsible for calculating a part of the data in the vector; The data path unit is divided into several small-bandwidth data paths, and each path includes a path scheduler, a path register file, a path execution unit, and a path register write-back status table; After querying the write-back status table of the path register and determining that all the registers to be read are ready, the path scheduler issues a register read request to the path register file and passes the read data to the path execution unit for micro-instruction execution. After the micro-instruction execution is completed, the corresponding entry in the path register write-back status table is cleared and the execution status management unit is informed. Each path can only receive micro-instructions for reading and writing the path register file of this path, and there is no direct data interaction between all paths, thus reducing the timing risks within and between paths. If there are instructions that need to read and write across data paths, the cross-path unit is responsible for issuing register read and write requests to the corresponding paths and for executing the corresponding instructions. The memory access unit is responsible for executing vector memory access instructions. It splits the memory access instructions into several memory access requests and sends them to the storage unit for reading and writing the memory, and issues register read / write requests to the corresponding paths as needed. After the memory access request is executed, the execution status management unit is informed. When all the micro-instructions and memory access requests of a vector instruction are executed, the execution status management unit informs the backend scalar processing unit, which is responsible for the sequential submission of all instructions.

[0005] However, in the above architecture of the vector processing unit, the splitting of micro-instructions makes it difficult to manage and track the execution status of micro-instructions because the execution information of micro-instructions is only visible within the path, resulting in the main scheduler not being able to timely perceive the changes in the execution status of micro-instructions and thus making it difficult to precisely schedule instructions. Secondly, since both the cross-path unit and the memory access unit have register read requests initiated to the data path, the timing between them and the data path unit will affect each other, further increasing the timing risk. And because vector memory access instructions and scalar memory access instructions independently access the storage unit, it makes it difficult to manage memory consistency.

[0006] Therefore, it is necessary to optimize the architecture of the vector processing unit to solve the above technical problems. Summary of the Invention

[0007] The present invention provides a processor architecture of a decoupled vector processing unit based on data path splitting, aiming to solve the timing and memory consistency management problems existing in the processing logic of micro-instruction splitting in the existing processor architecture.

[0008] To solve the above technical problems, the present invention provides a processor architecture of a decoupled vector processing unit based on data path splitting. The processor architecture includes an instruction fetch unit, a decoding unit, a vector processing unit, a scalar processing unit, a memory access scheduling unit, and a storage unit, wherein: The instruction fetch unit is used to fetch instructions from the processor memory and send them to the decoding unit; The decoding unit is used to parse the instruction and identify scalar instructions and / or vector instructions from the instruction, where: when the vector instruction is identified, the vector instruction is sent to the vector processing unit; when the scalar instruction is identified, the scalar instruction is sent to the scalar processing unit; The vector processing unit is used to execute the vector instruction, and when the vector instruction is a vector memory access instruction, generate a vector memory access request for requesting access to the storage unit according to the vector memory access instruction; The scalar processing unit is used to execute the scalar instruction, and when the scalar instruction is a scalar memory access instruction, generate a scalar memory access request for requesting access to the storage unit according to the scalar memory access instruction; The memory access scheduling unit is used to send the vector memory access request and / or the scalar memory access request to the storage unit according to a preset scheduling method; The storage unit is used to perform memory access according to the vector memory access request and / or the scalar memory access request, and return the corresponding memory access response information to the vector processing unit and / or the scalar processing unit.

[0009] Furthermore, the vector processing unit includes a main scheduler, multiple path register write-back status tables, a data path unit, and an execution status management unit. The data path unit includes a path register file, where: The main scheduler is used to determine whether the vector instruction belongs to an arithmetic instruction or a vector storage instruction. If so, split the vector instruction into multiple micro-instructions with the same number as the number of path register write-back status tables according to the data path bandwidth of the data path unit; Each entry in each path register write-back status table is used to record the write-back status of the path register required by a micro-instruction, and the entry of the write-back status is available or working status; The data path unit includes multiple data paths. Each data path includes an independent path scheduler, a path register file, and a path execution unit, and one data path corresponds to one path register write-back status table, where: The path scheduler is used to monitor the write-back status recorded by the path register write-back status table, and when the path register in the path register file is available, send a corresponding register read request to the path register file according to the micro-instruction; The path execution unit is used to execute the micro-instruction according to the path register data read from the path register file by the register read request; The execution status management unit includes a plurality of path execution status management subunits, and one of the path execution status management subunits corresponds to one of the data paths. The path execution status management subunit is used to monitor the execution status of the microinstructions.

[0010] Furthermore, the vector processing unit further includes a memory access unit, and the memory access unit is used for: Receiving the vector instruction from the main scheduler. If the vector instruction is a vector store instruction, receiving the path register data read by the path execution unit, and splitting the vector memory access instruction into a plurality of memory access sub-requests; Sending the memory access sub-request as the vector memory access request to the memory access scheduler unit, and obtaining the memory access response information corresponding to the vector memory access request from the memory access scheduler unit; After all the memory access sub-requests are executed, sending a first confirmation message to the execution status management unit.

[0011] Furthermore, the vector processing unit further includes a cross-path unit, and the cross-path unit is used for: Receiving the microinstructions split by the main scheduler and the path register data read through the data path unit, executing the microinstructions according to the path register data, and writing the execution results of the microinstructions back to the path register bank of the corresponding data path unit.

[0012] Furthermore, the path execution unit is further used for: After the microinstructions are executed, updating the entries of the write-back status in the path register write-back status table, and sending a second confirmation message to the corresponding path execution status management subunit.

[0013] Furthermore, the execution status management unit is further used for: After receiving the first confirmation message sent by the memory access unit and the second confirmation messages sent by all the path execution status management subunits to their corresponding path execution units, sending a notification signal indicating that the vector instruction execution is completed to the scalar processing unit.

[0014] Furthermore, the memory access unit is further used for: Receiving the vector instruction from the main scheduler. If the vector instruction is a vector load instruction, splitting the vector memory access instruction into a plurality of memory access sub-requests; Send the memory access sub-request as the vector memory access request to the memory access scheduling unit, and obtain the memory access response information corresponding to the vector memory access request from the memory access scheduling unit. Meanwhile, send a corresponding register write request to the path register file according to the memory access response information, where the register write request is used to write the data corresponding to the vector memory access request into the path register file; After all the memory access sub-requests are executed, send a first confirmation message to the execution status management unit.

[0015] Furthermore, the execution status management unit is further configured to: After receiving the first confirmation message sent by the memory access unit and all the register write requests are executed, send a signal indicating that the vector instruction has been executed to the scalar processing unit.

[0016] Furthermore, the preset scheduling method is: According to the order of the scalar instructions and / or vector instructions parsed from the instructions by the decoding unit, send the vector memory access requests and / or scalar memory access requests to the storage unit in sequence; or, Send the vector memory access requests and / or scalar memory access requests to the storage unit out of order according to a preset priority.

[0017] The beneficial effects achieved by the present invention are as follows: A processor architecture implemented based on a decoupled vector memory access unit is proposed. In this processor architecture, a separate path execution status management sub-unit is equipped on each data path of the vector processing unit, providing a more accurate scheduling reference for the main scheduler to perform micro-instruction splitting processing, reducing the difficulty of micro-instruction tracking and management, and at the same time optimizing the register read request path and reducing the timing risk; moreover, this architecture manages the order of scalar and vector memory access requests through a unified memory access scheduling unit, reducing the difficulty of memory consistency management and enhancing the reliability of the processor system. Description of the Drawings

[0018] Figure 1 is a schematic structural diagram of a processor architecture of a decoupled vector processing unit based on data path splitting provided by an embodiment of the present invention; Figure 2 is a schematic structural diagram of a vector processing unit in a processor architecture of a decoupled vector processing unit based on data path splitting provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a scheduling process for a vector processing unit to execute vector instructions provided by an embodiment of the present invention; Figure 4It is a schematic diagram of the scheduling process for a vector processing unit provided by an embodiment of the present invention to execute cross-path vector instructions; Figure 5 It is a schematic diagram of the execution process of a memory access instruction provided by an embodiment of the present invention. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a processor architecture of a decoupled vector processing unit based on data path splitting provided by an embodiment of the present invention. The processor architecture 100 includes an instruction fetch unit 101, a decoding unit 102, a vector processing unit 103, a scalar processing unit 104, a memory access scheduling unit 105, and a storage unit 106, where: The instruction fetch unit 101 is used to fetch instructions from the processor memory and send them to the decoding unit 102; The decoding unit 102 is used to parse the instructions and identify scalar instructions and / or vector instructions from the instructions. When the vector instructions are identified, the vector instructions are sent to the vector processing unit 103; when the scalar instructions are identified, the scalar instructions are sent to the scalar processing unit 104; The vector processing unit 103 is used to execute the vector instructions, and when the vector instruction is a vector memory access instruction, a vector memory access request for requesting access to the storage unit 106 is generated according to the vector memory access instruction; The scalar processing unit 104 is used to execute the scalar instructions, and when the scalar instruction is a scalar memory access instruction, a scalar memory access request for requesting access to the storage unit 106 is generated according to the scalar memory access instruction; The memory access scheduling unit 105 is used to send the vector memory access request and / or the scalar memory access request to the storage unit 106 according to a preset scheduling method; The storage unit 106 is used to perform memory access according to the vector memory access request and / or the scalar memory access request and return the corresponding memory access response information to the vector processing unit 103 and / or the scalar processing unit 104.

[0021] Based on the above-mentioned processor architecture 100, the embodiments of the present invention achieve the design goal of the decoupled architecture, that is, vector instructions and ordinary scalar instructions are split during the decoding stage, and the vector processing unit 103 and the scalar processing unit 104 execute the instructions respectively, which simplifies the pipeline design. At the same time, the processor architecture 100 of the embodiments of the present invention uniformly manages vector and scalar memory access requests through the memory access scheduling unit 105, which conforms to the design principle of the memory hierarchy and is conducive to ensuring cache consistency.

[0022] Further, in the embodiments of the present invention, micro-instruction splitting processing of vector instructions is performed by the vector processing unit 103. Specifically, please refer to Figure 2 , Figure 2 FIG. is a schematic structural diagram of the vector processing unit 103 in the processor architecture 100 implemented based on the decoupled vector processing unit 103 provided by the embodiments of the present invention. The vector processing unit 103 includes a main scheduler 1031, multiple path register write-back status tables 1032, a data path unit 1033, and an execution status management unit 1034. The data path unit 1033 includes a path register file, wherein: The main scheduler 1031 is used to determine whether the vector instruction belongs to an arithmetic instruction or a vector store instruction. If so, according to the data path bandwidth of the data path unit 1033, the vector read instruction is split into multiple micro-instructions with the same number as the number of the path register write-back status tables 1032; Each entry in each of the path register write-back status tables 1032 is used to record the write-back status of the path register required by one of the micro-instructions, and the entry of the write-back status is available or in a working state; The data path unit 1033 includes multiple data paths. Each data path includes an independent path scheduler 10331, a path register file 10332, and a path execution unit 10333, and one data path corresponds to one path register write-back status table 1032, wherein: The path scheduler 10331 is used to monitor the write-back status recorded by the path register write-back status table 1032, and when the path register in the path register file 10332 is available, send a corresponding register read request to the path register file 10332 according to the micro-instruction; The path execution unit 10333 is used to execute the micro-instruction according to the path register data read from the path register file 10332 by the register read request; The execution status management unit 1034 includes a plurality of path execution status management subunits 10341, and one of the path execution status management subunits 10341 corresponds to one of the data paths. The path execution status management subunit 10341 is used to monitor the execution status of the microinstructions.

[0023] In the embodiments of the present invention, the arithmetic instructions described include a first type in which the read / write of the path register can be completed only within the data path, and a second type that requires cross-path read / write of the path register bank; the vector store instructions described in the embodiments of the present invention are used to write the data read from the data path into the storage unit 106. In the embodiments of the present invention, when the main scheduler 1031 performs microinstruction splitting, it mainly targets the read / write instructions related to the path registers in the data path. These instructions will be split by the main scheduler 1031 and sent to the data path unit 1033 to generate register read requests. Among them, the microinstructions corresponding to the first type of arithmetic instructions are executed by the path execution unit 10333 in a single data path.

[0024] In the above structure, by decoupling the data path and separating the path register write-back status table 1032 from the data path and directly setting it as a subordinate structure of the main scheduler 1031, it is possible to provide a more accurate scheduling reference for the main scheduler 1031; and, based on the path register write-back status table 1032, the main scheduler 1031 can quickly query the status of the path registers to automatically adjust the splitting strategy according to the data path bandwidth, thereby making full use of the hardware resources. That is, as Figure 3 shown in the schematic diagram of the scheduling process of the vector instruction, through the design of data path decoupling, when the vector read instruction is executed: S201. First, the main scheduler 1031 executes the query of the path register write-back status table 1032 and the splitting of the microinstructions. S202. The main scheduler 1031 issues microinstructions to the path scheduler 10331 to generate register read requests, and sends the data read from the path register bank 10332 to the path execution unit 10333. S203. The path execution unit 10333 executes the microinstructions. S204. The path execution status management subunit 10341 monitors the execution status of the microinstructions, and at the same time clears or updates the relevant entries of the path register write-back status table 1032 through the path execution unit 10333.

[0025] In the embodiments of the present invention, through the method of decoupling the data path, all read requests for path registers are uniformly set to be issued by the main scheduler. Compared with the existing architecture, the read request path from the cross-path unit and the memory access unit to the path register bank is reduced, and the timing risk is reduced.

[0026] Furthermore, the vector processing unit 103 further includes a memory access unit 1035, and the memory access unit 1035 is configured to: Receive the vector instruction from the main scheduler 1031. If the vector instruction is a vector storage instruction, receive the path register data read by the path execution unit 10333, and split the vector memory access instruction into multiple memory access sub-requests; Send the memory access sub-requests as the vector memory access requests to the memory access scheduler unit 105, and obtain the memory access response information corresponding to the vector memory access requests from the memory access scheduler unit 105; After all the memory access sub-requests are executed, send a first confirmation message to the execution status management unit 1034.

[0027] Furthermore, the vector processing unit 103 further includes a cross-path unit 1036, and the cross-path unit 1036 is configured to: Receive the micro-instructions split by the main scheduler 1031 and the path register data read through the data path unit 1033, execute the micro-instructions according to the path register data, and write the execution results of the micro-instructions back to the path register file 10332 of the corresponding data path unit 1033.

[0028] The cross-path unit 1036, as a lower-level structure of the main scheduler 1031, functions according to whether the micro-instruction type split by the main scheduler 1031 requires cross-data-path vector register reading. The design of the cross-path unit 1036 allows different data paths to directly read and write each other's vector registers, avoiding the round-trip transmission of data between the path register file 10332 and another path register file 10332, thereby reducing the additional access overhead during the execution of different micro-instructions. That is, as Figure 4 shown in the schematic diagram of the scheduling process of the cross-path vector instruction, which involves the execution of the cross-path unit 1036, and the difference from the Figure 3 execution process shown is as follows: S301. While the main scheduler 1031 executes micro-instruction splitting, the corresponding micro-instructions will be sent to the cross-path unit 1036: S302. The path scheduler 10331 generates a cross-path register read request; S303. The path register file 10332 reads data from the corresponding path registers according to the register read request; S304. The cross-path unit 1036 executes the micro-instructions and sends the data read in S303 to the path execution unit 10333 of the corresponding other path.

[0029] In the embodiment of the present invention, the microinstructions executed by the cross-path unit 1036 mainly correspond to the second type of arithmetic instructions for reading / writing the path register file in the above embodiments.

[0030] Specifically, the path execution unit 10333 is further configured to: After the execution of the microinstruction is completed, update the entry of the write-back status in the path register write-back status table 1032, and send a second confirmation message to the corresponding path execution status management subunit 10341.

[0031] The execution status management unit 1034 is further configured to: After receiving the first confirmation message sent by the memory access unit 1035 and the second confirmation messages sent by all the path execution status management subunits 10341 to their corresponding path execution units 10333, send a signal indicating that the execution of the vector instruction is completed to the scalar processing unit 104.

[0032] The path execution unit 10333 directly executes the microinstructions. In the normal pipeline process, after the path execution unit 10333 executes a microinstruction, it updates the relevant entries in the path register write-back status table 1032, and changes the status of the currently used vector register from the working state to the available state, so as to quickly provide a reference basis for the main scheduler 1031.

[0033] In the decoupled processor architecture 100, when the vector processing unit 103 finishes executing all the vector instructions, it is necessary to transmit a signal indicating completion to the scalar processing unit 104. In the embodiment of the present invention, this process is implemented by the execution status management unit 1034 that uniformly manages all data paths. Among them, in the embodiment of the present invention, the first confirmation message is mainly used to determine the execution status of the vector memory access instruction (and its corresponding memory access sub-requests), and the second confirmation message is mainly used to determine the execution status of the microinstructions. It is necessary to send a signal indicating completion to the scalar execution unit 104 after the vector execution unit 103 completely executes all parts (memory access sub-requests and microinstructions) of the vector instruction.

[0034] Different from the above embodiment in which the main scheduler 1031 splits the vector instructions related to reading / writing the path register file, for the processing method of the vector load instruction related to loading data from the storage unit to the vector register file data storage, the memory access unit 1035 is further configured to: Receive the vector instruction from the main scheduler 1031. If the vector instruction is a vector load instruction, split the vector memory access instruction into multiple memory access sub-requests; Send the memory access sub-request as the vector memory access request to the memory access scheduling unit 105, and obtain the memory access response information corresponding to the vector memory access request from the memory access scheduling unit 105. Meanwhile, send a corresponding register write request to the path register file 10332 according to the memory access response information, where the register write request is used to write the data corresponding to the vector memory access request into the path register file 10332; After all the memory access sub-requests are executed, send a first confirmation message to the execution status management unit 1034.

[0035] Different from steps S201 - S204, the vector load instruction for loading data from the storage unit to the vector register file does not require the main scheduler 1031 to split the micro-instructions, but directly sends a memory access request to the storage unit 106 by the memory access unit 1035.

[0036] The execution status management unit 1034 is further configured to: After receiving the first confirmation message sent by the memory access unit 1035 and all the register write requests are executed, send a signal indicating that the vector instruction is executed to the scalar processing unit 104.

[0037] Similarly, for the execution of the vector load instruction related to loading data from the storage unit to the vector register file, the micro-instructions are not split by the main scheduler 1031 during the execution process. Therefore, when determining whether the vector processing unit 103 has executed all the vector instructions, it is determined by the execution status of the memory access sub-requests and the corresponding register write requests.

[0038] In the embodiment of the present invention, the memory access scheduling unit 105, which is at the next level of the scalar processing unit 104 and the vector processing unit 103, is set to execute the operation of sending the vector memory access request and / or the scalar memory access request to the storage unit 106, where the preset scheduling method is: Send the vector memory access request and / or the scalar memory access request to the storage unit 106 in sequence according to the sequence of the scalar instruction and / or the vector instruction parsed from the instruction by the decoding unit 102; or, Send the vector memory access request and / or the scalar memory access request to the storage unit 106 out of order according to the preset priority.

[0039] Through the above design, the memory access scheduling unit 105 in the embodiment of the present invention determines how the vector and scalar memory access requests compete for storage resources, thereby determining the characteristics of the processor architecture 100 according to different hardware usage environments. As Figure 5 shown in the schematic diagram of the memory access instruction execution process: After the decoding unit 102 obtains an instruction, it determines vector memory access instructions and scalar memory access instructions, and sends the instructions to the vector or scalar processing unit according to the determination result; S402. The vector processing unit 103 and the scalar processing unit 104 execute vector instructions correspondingly and generate vector and scalar memory access requests; S403. The memory access scheduling unit 105 determines the sending order of vector and scalar memory access requests according to a preset scheduling method, and sends the memory access requests to the storage unit 106 according to the order; S404. The storage unit 106 reads relevant data according to the vector and scalar memory access requests.

[0040] Among them, the sequential execution mode is applicable to usage environments dominated by scalar operations and sensitive to latency (such as real-time control, low-power devices); by maintaining an out-of-order instruction queue through pre-designed priorities, the out-of-order execution mode is applicable to usage environments with a high proportion of vector operations and sensitive to throughput (such as AI, scientific computing). During implementation, it can be set according to actual needs.

[0041] The beneficial effects achieved by the present invention are that it proposes a processor architecture implemented based on a decoupled vector memory access unit. In this processor architecture, a separate path execution status management subunit is equipped on each data path of the vector processing unit, providing a more accurate scheduling reference for the main scheduler to perform micro-instruction splitting processing, reducing the difficulty of micro-instruction tracking and management, and at the same time optimizing the register read request path and reducing the timing risk; moreover, this architecture manages the order of scalar and vector memory access requests through a unified memory access scheduling unit, reducing the management difficulty of memory consistency and enhancing the reliability of the processor system.

[0042] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0043] It should be noted that, in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that includes a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes such element.

[0044] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. What is disclosed is only the preferred embodiments of the present invention. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can still make many equivalent changes in form without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.

Claims

1. A processor architecture of a decoupled vector processing unit based on data path splitting, characterized in that The processor architecture includes an instruction fetch unit, a decoding unit, a vector processing unit, a scalar processing unit, a memory access scheduling unit, and a storage unit, where: The instruction fetch unit is configured to fetch instructions from the processor memory and send them to the decoding unit; The decoding unit is configured to parse the instructions and identify scalar instructions and / or vector instructions from the instructions, where: when the vector instructions are identified, the vector instructions are sent to the vector processing unit; when the scalar instructions are identified, the scalar instructions are sent to the scalar processing unit; The vector processing unit is configured to execute the vector instructions, and when the vector instruction is a vector memory access instruction, generate a vector memory access request for requesting access to the storage unit according to the vector memory access instruction; The scalar processing unit is configured to execute the scalar instructions, and when the scalar instruction is a scalar memory access instruction, generate a scalar memory access request for requesting access to the storage unit according to the scalar memory access instruction; The memory access scheduling unit is configured to send the vector memory access request and / or the scalar memory access request to the storage unit according to a preset scheduling method; The storage unit is configured to perform a memory access according to the vector memory access request and / or the scalar memory access request, and return corresponding memory access response information to the vector processing unit and / or the scalar processing unit.

2. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 1, characterized in that The vector processing unit includes a main scheduler, a plurality of path register write-back status tables, a data path unit, and an execution status management unit. The data path unit includes a path register bank, where: The main scheduler is configured to determine whether the vector instruction belongs to an arithmetic instruction or a vector store instruction. If so, split the vector instruction into multiple micro-instructions with the same number as the number of the path register write-back status tables according to the data path bandwidth of the data path unit; Each entry in each of the path register write-back status tables is used to record the write-back status of the path register required by a micro-instruction, and the entry of the write-back status is available or in a working state; The data path unit includes a plurality of data paths. Each data path includes an independent path scheduler, a path register bank, and a path execution unit, and one data path corresponds to one path register write-back status table, where: The path scheduler is configured to monitor the write-back status recorded in the path register write-back status table, and when the path register in the path register bank is available, send a corresponding register read request to the path register bank according to the micro-instruction; The path execution unit is configured to execute the micro-instruction according to the path register data read from the path register bank by the register read request; The execution status management unit includes a plurality of path execution status management subunits, and one path execution status management subunit corresponds to one data path. The path execution status management subunit is configured to monitor the execution status of the micro-instruction.

3. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 2, wherein The vector processing unit further includes a memory access unit, and the memory access unit is configured to: Receive the vector instruction from the main scheduler. If the vector instruction is a vector store instruction, receive the data in the path register read by the path execution unit, and split the vector memory access instruction into multiple memory access sub-requests; Send the memory access sub-request as the vector memory access request to the memory access scheduler, and obtain the memory access response information corresponding to the vector memory access request from the memory access scheduler; After all the memory access sub-requests are executed, send a first confirmation message to the execution status management unit.

4. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 2, wherein The vector processing unit further includes a cross-path unit, and the cross-path unit is used for: Receive the micro-instructions split by the main scheduler and the data in the path register read through the data path unit, execute the micro-instructions according to the data in the path register, and write the execution results of the micro-instructions back to the path register file of the corresponding data path unit.

5. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 3, characterized in that, The path execution unit is further used for: After the micro-instruction is executed, update the entry of the write-back status in the path register write-back status table, and send a second confirmation message to the corresponding path execution status management sub-unit.

6. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 5, characterized in that The execution status management unit is further used for: After receiving the first confirmation message sent by the memory access unit and the second confirmation messages sent by all the path execution status management sub-units received by their corresponding path execution units, send a notification signal indicating the completion of the execution of the vector instruction to the scalar processing unit.

7. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 3, characterized in that, The memory access unit is further used for: Receive the vector instruction from the main scheduler. If the vector instruction is a vector load instruction, split the vector memory access instruction into multiple memory access sub-requests; Send the memory access sub-request as the vector memory access request to the memory access scheduler, and obtain the memory access response information corresponding to the vector memory access request from the memory access scheduler. At the same time, send a corresponding register write request to the path register file according to the memory access response information, and the register write request is used to write the data corresponding to the vector memory access request into the path register file; After all the memory access sub-requests are executed, send a first confirmation message to the execution status management unit.

8. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 7, characterized in that, The execution status management unit is further used for: After receiving the first confirmation message sent by the memory access unit and after all the register write requests are executed, send a notification signal indicating the completion of the execution of the vector instruction to the scalar processing unit.

9. The processor architecture of the decoupled vector processing unit based on data path splitting according to claim 1, wherein The preset scheduling method is: Send the vector memory access request and / or the scalar memory access request to the storage unit in sequence according to the order of the scalar instruction and / or the vector instruction parsed from the instruction by the decoding unit; or, Send the vector memory access request and / or the scalar memory access request to the storage unit out of order according to a preset priority.

Citation Information

Patent Citations

  • RISC-V vector memory access processing system and processing method

    CN114579188A

  • Low-hardware-overhead vector processor architecture based on RISC-V vector instruction extension

    CN116521229A

  • Memory access method, processor, electronic equipment and readable storage medium

    CN118796272A

  • Memory access control method and system of decoupling vector coprocessor based on RISC-V

    CN118897695A

  • Vector data processor, instruction processing method, system on chip and computing device

    CN119556982A