RISC-V Instruction Acceleration Method, System, Device and Storage Medium
By specifying acceleration dedicated registers and configuring a parallel processing architecture, the problem of improving the execution efficiency of RISC-V in the prior art is solved, efficient and low-complexity instruction execution is achieved, and energy efficiency ratio is improved.
Patent Information
- Application Number
- CN202510026861.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-08
AI Technical Summary
When the prior art improves the execution efficiency of RISC-V instruction, the hardware is complex and difficult to implement, resulting in a prolonged R&D cycle and a reduced energy efficiency ratio.
By specifying acceleration dedicated registers and configuring a parallel processing architecture, including a first instruction execution path and a second instruction execution path, the parallel processing architecture is used to execute target instructions concurrently to improve instruction execution efficiency.
It realizes efficient execution of RISC-V instructions, reduces hardware complexity and implementation difficulty, improves energy efficiency ratio, and simplifies the R&D process.
Smart Images

Figure CN119415155B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a RISC-V instruction acceleration method, system, device, and storage medium. Background Art
[0002] As the fifth-generation reduced instruction set architecture, the RISC-V instruction set has become the best choice for in-depth research by many enterprises and research institutions due to its open-source, modular, and highly scalable design. It can be fundamentally optimized in the directions of low power consumption, high performance, high reliability, etc. according to actual needs.
[0003] Existing technologies for instruction acceleration usually include increasing the number of pipeline stages, increasing the issue width, improving the cache hit rate, improving the branch prediction accuracy, etc. These methods have high hardware complexity and great implementation difficulty, often leading to an extended R & D cycle and possibly reducing the energy efficiency ratio. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the present invention provides a RISC-V instruction acceleration method, system, device, and storage medium to solve the above technical problems.
[0005] In a first aspect, the present invention provides a RISC-V instruction acceleration method, including:
[0006] Designating an acceleration dedicated register from general-purpose registers;
[0007] Configuring a parallel processing architecture, where the parallel processing architecture includes a first instruction execution path and a second instruction execution path;
[0008] Confirming that the source register of the target instruction includes the acceleration dedicated register, and concurrently executing the target instruction using the parallel processing architecture.
[0009] In an optional embodiment, designating an acceleration dedicated register from general-purpose registers includes:
[0010] Setting x31 among general-purpose registers x0 to x31 as the acceleration dedicated register.
[0011] In an optional embodiment, configuring a parallel processing architecture, where the parallel processing architecture includes a first instruction execution path and a second instruction execution path, includes:
[0012] Configuring a first logical operation unit, a second logical operation unit, a first decoding unit, and a second decoding unit;
[0013] Dividing general-purpose registers other than the acceleration dedicated register into a first general-purpose register group and a second general-purpose register group;
[0014] The first decoding unit and the first logical operation unit are denoted as the first instruction execution path, and the first general register bank is for the use of the first logical operation unit; the second decoding unit and the second logical operation unit are denoted as the second instruction execution path, and the second general register bank is for the use of the second logical operation unit.
[0015] In an optional embodiment, the method further includes:
[0016] Set addressing parameters in the status control register, where the addressing parameters include the starting address, stride, and number of addressing times for the addressing;
[0017] Calculate the address of the data required to execute the target instruction according to the addressing parameters, and load the required data to the acceleration dedicated register based on the address.
[0018] In an optional embodiment, when it is confirmed that the source register of the target instruction includes the acceleration dedicated register, concurrently execute the target instruction using the parallel processing architecture, including:
[0019] Cache the target instruction to a pre-allocated instruction buffer;
[0020] Close the instruction buffer, and the instruction buffer no longer caches new instructions in the closed state;
[0021] Enable the dual-issue function to concurrently execute instructions.
[0022] In an optional embodiment, enable the dual-issue function to concurrently execute instructions, including:
[0023] Read instructions from memory or the instruction buffer;
[0024] Based on the pre-set register usage rules, allocate registers to the first logical operation unit and the second logical operation unit for parallel processing instructions;
[0025] Obtain the data required for the instruction from memory and store the obtained data to the source register of the instruction.
[0026] In an optional embodiment, the register usage rules include:
[0027] When the first logical operation unit and the second logical operation unit execute the same instruction, if the first logical operation unit uses any register in the first general register bank as the source register when executing the instruction, then the second logical operation unit uses the register in the second general register bank as the source register when concurrently executing the instruction;
[0028] If the source registers of two consecutive instructions are the same, set the source register as a shared register, and the shared register allows access by a first logical operation unit and a second logical operation unit;
[0029] When the first logical operation unit and the second logical operation unit execute different instructions, the first logical operation unit executes an instruction that needs to access memory, and selects the source register and the destination register of the instruction from the first general register bank; the second logical operation unit executes the instruction in the instruction buffer, and selects the destination register of the instruction from the second general register bank, and the source register of the instruction is the acceleration dedicated register.
[0030] In a second aspect, the present invention provides a RISC-V instruction acceleration system, including:
[0031] An acceleration configuration module for designating an acceleration dedicated register from general registers;
[0032] A parallel configuration module for configuring a parallel processing architecture, and the parallel processing architecture includes a first instruction execution path and a second instruction execution path;
[0033] An execution control module for confirming that the source register of the target instruction includes the acceleration dedicated register, and concurrently executing the target instruction by using the parallel processing architecture.
[0034] In a third aspect, a device is provided, including:
[0035] A memory for storing a RISC-V instruction acceleration program;
[0036] A processor for implementing the steps of the RISC-V instruction acceleration method provided in the first aspect when executing the RISC-V instruction acceleration program.
[0037] In a fourth aspect, a computer-readable storage medium is provided, and a RISC-V instruction acceleration program is stored on the storage medium. When the RISC-V instruction acceleration program is executed by a processor, the steps of the RISC-V instruction acceleration method provided in the first aspect are implemented.
[0038] The beneficial effects of the present invention are that the RISC-V instruction acceleration method, system, device and storage medium provided by the present invention improve the execution efficiency of RISC-V instructions by designating an acceleration dedicated register and configuring a parallel architecture.
[0039] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0042] Figure 2 It is another schematic flowchart of the method according to an embodiment of the present invention.
[0043] Figure 3 It is a schematic architecture diagram of the system according to an embodiment of the present invention.
[0044] Figure 4 It is a schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners
[0045] In order to enable those skilled in the art of this technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0047] The RISC-V instruction acceleration method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the RISC-V instruction acceleration system runs in the computer device.
[0048] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution entity can be a RISC-V instruction acceleration system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0049] Such as Figure 1 shown, the method includes:
[0050] S1. Designate an acceleration dedicated register from the general-purpose register.
[0051] S2. Configure a parallel processing architecture, which includes a first instruction execution path and a second instruction execution path.
[0052] S3. Confirm that the source register of the target instruction includes the acceleration dedicated register, and concurrently execute the target instruction using the parallel processing architecture.
[0053] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.
[0054] Specifically, set x31 among the general - purpose registers x0 to x31 as the acceleration dedicated register (abbreviation: ADR).
[0055] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.
[0056] Configure a first arithmetic logic unit (abbreviation: ALU1), a second arithmetic logic unit (abbreviation: ALU2), a first decoding unit (abbreviation: DU1), and a second decoding unit (abbreviation: DU2); divide the general - purpose registers other than ADR into a first general - purpose register group (abbreviation: RF1) and a second general - purpose register group (abbreviation: RF2); regard DU1 and ALU1 as the first instruction execution path, and RF1 is for ALU1 to use; regard DU2 and ALU2 as the second instruction execution path, and RF2 is for ALU2 to use.
[0057] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation. Please refer to Figure 2 .
[0058] S301. Cache the target instruction into a pre - allocated instruction buffer.
[0059] Confirm that the target instruction is an instruction that needs to be accelerated, that is, the source register of the target instruction contains the acceleration dedicated register, and cache the target instruction into the instruction buffer (abbreviation: IB). Among them, the target instruction can be a single instruction or a series of related instructions.
[0060] The main function of the IB is to pre - fetch and cache the instructions to be executed. In scenarios such as loop calculations, instructions often need to be executed repeatedly. By storing the instructions in the IB, it is possible to avoid re - fetching instructions from memory every time they are executed, thereby reducing the number of memory accesses, reducing the memory bandwidth pressure, and improving the instruction execution efficiency.
[0061] S302. Close the instruction buffer, and the instruction buffer no longer caches new instructions in the closed state.
[0062] Storing the instructions in the IB and closing the IB before double-issue can also ensure the continuity and consistency of instruction execution. Before the instructions in the IB are fully executed, no new instructions will be inserted into the execution stream. This helps to maintain the sequentiality of instruction execution and avoid errors or performance degradation that may be caused by out-of-order instruction execution.
[0063] In addition, by storing the instructions in the IB and closing the IB, processor resources can be saved to a certain extent. During the period when the IB is closed, the processor does not need to continuously read instructions from memory, but can focus on executing the instructions that have been cached in the IB. This helps to reduce the idle time of the processor and improve the overall performance.
[0064] S303. Enable the double-issue function to execute instructions concurrently.
[0065] Read instructions from memory or the instruction buffer;
[0066] Based on the pre-set register usage rules, allocate registers to the first logical operation unit and the second logical operation unit for parallel processing of instructions;
[0067] Obtain the data required by the instructions from memory and store the obtained data in the source registers of the instructions.
[0068] The register usage rules include:
[0069] When the first logical operation unit and the second logical operation unit execute the same instruction, if the first logical operation unit uses any register in the first general register group as the source register when executing the instruction, then the second logical operation unit uses a register in the second general register group as the source register when executing the instruction in parallel;
[0070] If the source registers of two consecutive instructions are the same, set the source register as a shared register, and the shared register allows the first logical operation unit and the second logical operation unit to access;
[0071] When the first logical operation unit and the second logical operation unit execute different instructions, the first logical operation unit executes the instruction that needs to access memory and selects the source register and target register of the instruction from the first general register group; the second logical operation unit executes the instruction in the instruction buffer and selects the target register of the instruction from the second general register group, and the source register of the instruction is the acceleration dedicated register.
[0072] The following are exemplary instruction execution scenarios:
[0073] (1) Dual-issue can be that ALU1 and ALU2 execute the same instruction, but with different operands. This means that although the functions of the two instructions are the same, the data sets they process are different.
[0074] Execute the same instruction in the IB using dual-issue. The specific execution process is as follows:
[0075] Both DU1 and DU2 decode instruction 1 in the IB. The information obtained by DU1 after decoding is source register 1, source register ADR, and destination register 2. The information obtained by DU2 after decoding is source register 3, source register ADR, and destination register 4. Among them, source register 1 and destination register 2 belong to the first general register bank RF1, and source register 3 and destination register 4 belong to the second general register bank RF2.
[0076] ALU1 and ALU2 respectively execute the arithmetic tasks of instruction 1 based on the decoding information.
[0077] To further improve the instruction execution efficiency,
[0078] Set addressing parameters in the status control register (detect CSR). The addressing parameters include the starting address of the addressing, the stride, and the number of addressing times; calculate the address of the data required to execute the target instruction based on the addressing parameters, and load the required data to the acceleration dedicated register based on the address.
[0079] ALU1 reads data from ADR and source register 1, performs arithmetic operations on the read data, and stores the operation result in destination register 2; ALU2 reads data from ADR and source register 3, performs arithmetic operations on the read data, and stores the operation result in destination register 4.
[0080] (2) Shared register scenario.
[0081] Step 1: Instruction fetch and decode
[0082] Instruction fetch: The processor fetches the next instruction from the IB, and this instruction is decoded by the decoding unit DU1.
[0083] Instruction decode: The instruction decoded by DU1 is an instruction that needs to use a certain source register in RF2 (for example, a load instruction for loading data from memory into this register).
[0084] Step 2: Register allocation and sharing preparation
[0085] Register allocation: The processor determines which source register in RF2 this instruction will use (assumed to be RF2[R1]).
[0086] Shared Preparation: Since the following instructions will continue to use the source register, the processor will prepare to allow both ALU1 and ALU2 to use the register. A status flag is generated to indicate that the register is currently shared.
[0087] Step Three: Execution of the First Instruction
[0088] Execute Load Instruction: ALU1 executes the load instruction to load data from memory into RF2[R1].
[0089] Update Status: The processor updates its internal status, marking that RF2[R1] now contains valid data and that the register is shared.
[0090] Step Four: Fetch and Decode of the Second Instruction
[0091] Instruction Fetch: The processor fetches the next instruction from the IB, which is decoded by the decoding unit DU2.
[0092] Instruction Decode: The instruction decoded by DU2 is an instruction that needs to use the data previously loaded into RF2[R1] (e.g., an arithmetic operation instruction).
[0093] Step Five: Execution of the Second Instruction and Register Sharing
[0094] Execute Arithmetic Operation: ALU2 executes the arithmetic operation instruction, using the data in RF2[R1] as one of the operands.
[0095] Shared Access: Since RF2[R1] is shared, ALU2 can directly access the data in the register without any additional data movement or copying operations.
[0096] Step Six: Instruction Completion and Result Processing
[0097] Instruction Completion: When both instructions have been executed, the processor updates its internal status, marking these instructions as completed.
[0098] Result Processing: The processor may store the operation result in another register, or write it back to memory, or perform other subsequent operations, depending on the program's instruction stream and data stream.
[0099] Through this process example, we can see that in the RISC-V instruction acceleration device, when an instruction fetches an instruction from memory and uses one of the source registers in RF2, and the next instruction will also use this source register, the processor can achieve the shared use of this source register through internal control signals or status flags. This method of sharing registers can improve the flexibility of data processing because it allows two instructions to access the same dataset without the need for additional data movement, thereby reducing the overhead of data movement and improving the performance of the processor.
[0100] (3) Process different types of instruction scenarios.
[0101] Step 1: Instruction fetch and decode
[0102] Fetch instruction: The processor fetches the next instruction from memory and places it in the instruction queue for decoding.
[0103] Decode instruction: DU1 and DU2 simultaneously fetch the instruction from the instruction queue for decoding. Assume that the instruction decoded by DU1 is a load instruction (Load) that needs to access memory, and the instruction decoded by DU2 is an arithmetic operation instruction (such as addition).
[0104] Step 2: Prepare operands and register allocation
[0105] ALU1 preparation: Since the Load instruction needs to access memory, ALU1 will wait for the memory access to complete. During this period, it will not perform any operations. However, it has already started to prepare the operands, that is, to determine which source register (one of the registers in RF1) and target register to use.
[0106] ALU2 preparation: Meanwhile, ALU2 does not need to wait for memory access. It has already fetched the pre-stored addition instruction from the instruction buffer IB and determined to use two registers in RF2 as source registers. Since the instruction of ALU2 comes from IB, it can start to prepare for execution while ALU1 is still waiting for memory access.
[0107] Step 3: Execute instruction
[0108] ALU1 execution: Once the memory access is completed, ALU1 will load the data from memory into the specified source register and execute the load instruction. After the load is completed, ALU1 may store the result in the target register or pass the result to the next instruction as an operand.
[0109] ALU2 execution: While ALU1 is waiting for memory access, ALU2 has already started to execute the addition instruction. It reads the values of the two source registers from RF2, performs the addition operation, and stores the result in the specified target register.
[0110] Step Four: Instruction Completion and Result Processing
[0111] Completion of ALU1: When the load instruction of ALU1 is executed and completed, it updates the processor's state (such as register values, memory addresses, etc.) and marks the instruction as completed.
[0112] Completion of ALU2: Similarly, when the addition instruction of ALU2 is executed and completed, it also updates the processor's state and marks the instruction as completed.
[0113] From this process example, we can see that when ALU2 executes instructions of a different operation type from ALU1, the processor can utilize the instruction buffer IB to prepare and execute ALU2's instructions in advance. In this way, even if ALU1 needs to wait for memory access, the processor can still maintain efficient parallel processing capabilities. This design can significantly improve the overall performance of the processor, especially when executing complex programs containing multiple types of instructions.
[0114] When all the currently cached instructions have been executed and new instructions need to be executed continuously, the IB unit will be reopened to prefetch and cache new instructions, and the dual-issue will be turned off. The timing of reopening the IB unit usually depends on the progress of instruction execution and the caching situation of instructions in the IB unit. When the instructions in the IB unit are emptied or about to be emptied, the processor will send a signal to reopen the IB unit. In addition, in some cases, such as when encountering branch instructions or interrupts, it may also be necessary to reopen the IB unit to obtain a new instruction stream.
[0115] In some embodiments, the RISC-V instruction acceleration system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the RISC-V instruction acceleration system can be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1 description) the functions of RISC-V instruction acceleration.
[0116] In this embodiment, according to the functions it performs, the RISC-V instruction acceleration system can be divided into multiple functional modules, such as Figure 3 shown, mainly including:
[0117] Acceleration Dedicated Register (ADR): Set x31 in the general register file as the acceleration dedicated register ADR. When a certain instruction uses ADR as the source register, that instruction and its subsequent two operations will be executed in parallel, which improves the execution efficiency of the instruction.
[0118] Address Generation Unit (AGU): A new AGU is added to calculate the address where the subsequent data to be loaded into the ADR register should be located based on the starting address and stride. This enables the processor to efficiently handle instructions that require consecutive access to memory addresses, such as array operations.
[0119] Control and Status Register (CSR): A new CSR is added to control the working state of the AGU, including setting the starting address, stride, and number of addressings for addressing. By assigning values to the CSR, the parameters of the AGU can be initialized, thereby controlling the data loading method.
[0120] Instruction Buffer (IB): A new IB is added to store instructions. When an instruction uses the ADR register, the instruction is written into the IB.
[0121] Expansion of the Decoding Unit (DU) and the Arithmetic Logic Unit (ALU): The number of existing DUs and ALUs is increased to 2 each (DU1, DU2, ALU1, and ALU2), forming two complete instruction execution paths. This enables the processor to process two instructions simultaneously, further improving the parallel processing ability.
[0122] Partitioning of General Registers: The general registers except the ADR register are divided into two parts, RF1 and RF2, for use by ALU1 and ALU2 respectively. This helps to avoid register conflicts and improve the parallelism of instruction execution.
[0123] Share the Instruction Fetch Unit (IFU) and the Load Store Unit (LSU). The IFU is used to fetch instructions from memory, and the LSU is used to load data from memory.
[0124] By assigning values to the CSR register, set the initialization parameters of the AGU and enable the IB.
[0125] When an instruction uses the ADR register, the instruction is written into the IB.
[0126] Disable the IB and enable double-issue.
[0127] Double-issue can be that ALU1 and ALU2 execute the same instruction but with different operands, or execute different types of instructions (where the instruction of ALU2 comes from the IB).
[0128] All instructions in the IB can also be executed sequentially and in parallel with the instructions fetched from memory.
[0129] In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0130] Acceleration Configuration Module, used to specify acceleration-specific registers from general registers;
[0131] A parallel configuration module for configuring a parallel processing architecture, which includes a first instruction execution path and a second instruction execution path;
[0132] An execution control module for confirming that the source register of the target instruction includes the acceleration dedicated register and concurrently executing the target instruction using the parallel processing architecture.
[0133] Optionally, as an embodiment of the present invention, set x31 among the general-purpose registers x0 to x31 as the ADR.
[0134] Optionally, as an embodiment of the present invention, the parallel configuration module includes:
[0135] Configure a first arithmetic logic unit (referred to as ALU1), a second arithmetic logic unit (referred to as ALU2), a first decoding unit (referred to as DU1), and a second decoding unit (referred to as DU2); divide the general-purpose registers other than the ADR into a first general-purpose register group (referred to as RF1) and a second general-purpose register group (referred to as RF2); regard DU1 and ALU1 as the first instruction execution path, and RF1 is used by ALU1; regard DU2 and ALU2 as the second instruction execution path, and RF2 is used by ALU2.
[0136] Optionally, as an embodiment of the present invention, the system further includes:
[0137] Set addressing parameters in the status control register, where the addressing parameters include the starting address, stride, and number of addressing operations of the addressing;
[0138] The addressing unit calculates the address of the data required to execute the target instruction according to the addressing parameters, and loads the required data to the acceleration dedicated register based on the address.
[0139] Optionally, as an embodiment of the present invention, the execution control module includes:
[0140] Cache the target instruction in a pre-allocated instruction buffer;
[0141] Close the instruction buffer, and the instruction buffer no longer caches new instructions in the closed state;
[0142] Enable the dual-issue function to concurrently execute instructions.
[0143] Optionally, as an embodiment of the present invention, enabling the dual-issue function to concurrently execute instructions includes:
[0144] Read instructions from memory or the instruction buffer;
[0145] Based on the preset register usage rules, allocate registers for the first logical operation unit and the second logical operation unit for parallel processing instructions;
[0146] Obtain the data required by the instruction from the memory and store the obtained data in the source register of the instruction.
[0147] Optionally, as an embodiment of the present invention, the register usage rules include:
[0148] When the first logical operation unit and the second logical operation unit execute the same instruction, if the first logical operation unit uses any register in the first general register group as the source register when executing the instruction, then the second logical operation unit uses the register in the second general register group as the source register when executing the instruction in parallel;
[0149] If the source registers of two consecutive instructions are the same, set the source register as a shared register, and the shared register allows the first logical operation unit and the second logical operation unit to access;
[0150] When the first logical operation unit and the second logical operation unit execute different instructions, the first logical operation unit executes the instruction that needs to access the memory, and selects the source register and the target register of the instruction from the first general register group; the second logical operation unit executes the instruction in the instruction buffer, and selects the target register of the instruction from the second general register group, and the source register of the instruction is the acceleration special register.
[0151] Figure 4 The RISC-V instruction acceleration method provided by the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In the embodiments of the present invention, the device includes but is not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0152] Among them, the device 400 may include: a processor 410, a memory 420, and a communication unit 430. These components communicate via one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0153] Among them, the memory 420 can be used to store the execution instructions of the processor 410. The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 is enabled to execute some or all of the steps in the above method embodiments.
[0154] The processor 410 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 410 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.
[0155] The communication unit 430 is used to establish a communication channel, so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.
[0156] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.
[0157] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., various media that can store program codes, including several instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0158] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.
[0159] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical, or other forms.
[0160] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0162] Although the present invention has been described in detail by referring to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A RISC-V instruction acceleration method, characterized in that: include: Specify acceleration special registers from general registers; configuring a parallel processing architecture, the parallel processing architecture comprising a first instruction execution path and a second instruction execution path; Confirming that the source register of the target instruction includes the acceleration dedicated register, and concurrently executing the target instruction using the parallel processing architecture; Also included are pre-set register usage rules, including: When the first logic operation unit and the second logic operation unit execute the same instruction, if the first logic operation unit uses any register in the first general register group as a source register when executing the instruction, the second logic operation unit uses a register in the second general register group as a source register when executing the instruction in parallel; If the source registers of two consecutive instructions are the same, the source register is set as a shared register, and the shared register allows the first logic operation unit and the second logic operation unit to access; When the first logic operation unit and the second logic operation unit execute different instructions, the first logic operation unit executes the instruction that needs to access the memory, and selects the source register and the target register of the instruction from the first general register group; The second logic operation unit executes the instruction in the instruction buffer and selects the target register of the instruction from the second general register group, and the source register of the instruction is the acceleration special register.
2. The method according to claim 1, characterized in that Specify acceleration-specific registers from general-purpose registers, including: Set x31 of general-purpose registers x0 to x31 to an acceleration special register.
3. The method according to claim 1, characterized in that Configuring a parallel processing architecture, the parallel processing architecture comprising a first instruction execution path and a second instruction execution path, comprising: Configure a first logic operation unit, a second logic operation unit, a first decoding unit, and a second decoding unit; Dividing the general registers other than the acceleration special registers into a first general register group and a second general register group; The first decoding unit and the first logic operation unit are recorded as the first instruction execution path, and the first general register group is used by the first logic operation unit; the second decoding unit and the second logic operation unit are recorded as the second instruction execution path, and the second general register group is used by the second logic operation unit.
4. The method according to claim 1, characterized in that: The method further comprises: Setting addressing parameters in the state control register, the addressing parameters including the starting address, stride and addressing times of the addressing; The address of the data required to execute the target instruction is calculated according to the addressing parameter, so as to load the required data into the acceleration dedicated register based on the address.
5. The method according to claim 2, characterized in that: Confirming that the source register of the target instruction includes the acceleration dedicated register, and concurrently executing the target instruction using the parallel processing architecture, comprises: caching the target instruction into a pre-allocated instruction buffer; Close the instruction buffer, where the instruction buffer no longer caches new instructions; Enables dual-issue capability to execute instructions concurrently.
6. The method according to claim 5, characterized in that Enables dual-issue functionality to execute instructions concurrently, including: Read instructions from memory or the instruction buffer; Allocating registers to the first logic operation unit and the second logic operation unit for parallel processing instructions based on a preset register usage rule; The data required by the instruction is obtained from the memory and stored in the source register of the instruction.
7. A RISC-V instruction acceleration system, characterized in that: include: An acceleration configuration module, used to specify acceleration-specific registers from general registers; A parallel configuration module, configured to configure a parallel processing architecture, wherein the parallel processing architecture includes a first instruction execution path and a second instruction execution path; An execution control module, configured to confirm that the source register of the target instruction includes the acceleration dedicated register, and concurrently execute the target instruction using the parallel processing architecture; Also included are pre-set register usage rules, including: When the first logic operation unit and the second logic operation unit execute the same instruction, if the first logic operation unit uses any register in the first general register group as a source register when executing the instruction, the second logic operation unit uses a register in the second general register group as a source register when executing the instruction in parallel; If the source registers of two consecutive instructions are the same, the source register is set as a shared register, and the shared register allows the first logic operation unit and the second logic operation unit to access; When the first logic operation unit and the second logic operation unit execute different instructions, the first logic operation unit executes the instruction that needs to access the memory, and selects the source register and the target register of the instruction from the first general register group; The second logic operation unit executes the instruction in the instruction buffer and selects the target register of the instruction from the second general register group, and the source register of the instruction is the acceleration special register.
8. A RISC-V instruction acceleration device, characterized in that: include: A memory, used for storing a RISC-V instruction acceleration program; A processor, used to implement the steps of the RISC-V instruction acceleration method as described in any one of claims 1 to 6 when executing the RISC-V instruction acceleration program.
9. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a RISC-V instruction acceleration program, which, when executed by a processor, implements the steps of the RISC-V instruction acceleration method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for parallel processing real time gathering mixed audio blindness separating unit
CN101203061A
RISC processor apparatus and method capable of supporting X86 virtual machine
CN101256504A
Instruction processing method and device, computer equipment and storage medium
CN117193861A