RISC-V-based processor and electronic equipment
By designing instruction fetch, pre-decoding and decoding modules in the RISC-V processor, multiple microcode instructions are executed per clock cycle, which solves the performance bottleneck of the RISC-V processor in AI applications and improves the processor's computing performance and operation efficiency.
Patent Information
- Application Number
- CN202510788601.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing RISC-V processors have low performance in AI applications, and the single transmission sequence structure causes the IPC (number of instructions executed per cycle) to not be greater than 1, which becomes a performance bottleneck.
Using a RISC-V-based processor, through a combined design of the instruction fetch module, pre-decoding module, decoding module and back-end module, multiple microcode instructions are executed per clock cycle, and the pre-decoding information is used to write source data to the register in advance, improving the computing performance and operation efficiency of the processor.
Multiple microcode instructions can be executed each clock cycle to improve the computing performance and operation efficiency of the processor and solve the performance bottleneck problem of RISC-V processors in AI applications.
Smart Images

Figure CN120295673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chips, and more particularly, to a RISC-V-based processor and an electronic device. Background Art
[0002] As an emerging computing architecture, RISC-V is gradually becoming an important driving force for the development of AI computing power. With its open and flexible characteristics, RISC-V provides great freedom for chip designers, allowing them to customize AI accelerators according to specific requirements. Its instruction set is concise and highly extensible, and designers can add custom instruction set extensions to improve the performance and efficiency of AI computing. Currently, most RISC-V processors used for AI control are single-issue sequential structures. Although single-issue sequential RISC-V processors can meet simple program control requirements, in AI applications, due to their low performance, the IPC (instructions per cycle) is usually not greater than 1, which will become a performance bottleneck in AI applications. Summary of the Invention
[0003] The purpose of the present invention is to provide a RISC-V-based processor and an electronic device to improve the above problems.
[0004] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, an embodiment of the present invention provides a RISC-V-based processor, including: An instruction fetch module, configured to extract instructions according to the instruction fetch start address of the current clock cycle in the sequential instruction fetch state, and write the obtained first target instruction start address into the first register; the first target instruction includes at least two microcode instructions; A pre-decoding module, configured to perform pre-decoding processing on a second target instruction, and send the obtained N groups of pre-decoding information and the second target instruction to the second register; If the current clock cycle is a safe cycle, the second target instruction is the most recently written and valid first target instruction in the first register. If the current clock cycle is a non-safe cycle, the second target instruction is the first target instruction read from the first register in the most recent safe cycle. The jth pre-decoding information includes the destination register address and the source register address corresponding to the jth microcode instruction; A decoding module, configured to read the most recently written second target instruction and N groups of pre-decoding information from the second register, determine whether the current clock cycle is a safe cycle based on them, and feedback the judgment result of the safe cycle to the pre-decoding module; send the corresponding source data to the third register according to the source register address in the pre-decoding information, perform decoding processing on the read second target instruction to obtain N groups of decoding information, and send the N groups of decoding information to the third register; The j-th decoded information includes the instruction type, destination register address, operation action, number of execution cycles, and immediate value corresponding to the j-th microcode instruction; The backend module is used to execute corresponding microcode instructions according to the N groups of decoded information and source data recently written in the 3rd register.
[0005] Optionally, the instruction fetching module includes a first selector, a program counter, an adder, an instruction memory, and an identification unit; The program counter is used to send the instruction fetching start address of the current clock cycle to the instruction memory and the 1st register; The instruction memory is used to perform instruction fetching according to the instruction fetching start address of the current clock cycle in the sequential instruction fetching state, and write the fetched first target instruction into the 1st register; The identification unit is used to determine whether the first target instruction is a branch instruction; if not, it sends a sequential execution signal to the first selector; if so, it sends a branch execution signal to the first selector; The adder is used to determine a first address according to the instruction fetching start address of the current clock cycle and a preset number of bytes, and send the first address to the first selector; The first selector is further used to receive a second address fed back by the backend module, and the second address is the execution result of the branch instruction; The first selector is further used to send the first address to the program counter when receiving the sequential execution signal; and send the second address to the program counter when receiving the branch execution signal; The program counter is further used to store the address sent by the first selector in the current clock cycle, and use it as the instruction fetching start address of the next clock cycle.
[0006] Optionally, the pre-decoding module includes an issue queue, a second selector, and N pre-decoding units, where N represents the total number of microcode instructions in the first target instruction; When obtaining the first type of control signal in the current clock cycle, the issue queue and the second selector are used to read the most recently written and valid first target instruction in the 1st register, where the first type of control signal indicates that the current clock cycle is a safe cycle; When obtaining the second type of control signal in the current clock cycle, the issue queue and the second selector stop reading instructions from the 1st register; the second selector is used to read the most recently written first target instruction in the issue queue; where the second type of control signal indicates that the current clock cycle is a non-safe cycle; The second selector is used to send the instruction read by it in the current clock cycle as the second target instruction to the 2nd register and N pre-decoding units; The j-th pre-decoding unit is used to pre-decode the j-th microcode instruction received in the current clock cycle to obtain the j-th pre-decoding information corresponding to the j-th microcode instruction, and send the j-th pre-decoding information to the second register.
[0007] Optionally, the decoding module includes N decoding units, a microcode scoreboard, a register file, and a register hazard check unit; The j-th decoding unit is used to read the j-th microcode instruction in the second target instruction written most recently in the second register in the current clock cycle, and perform decoding to obtain the j-th decoding information, and send the j-th decoding information to the third register and the microcode scoreboard; The microcode scoreboard is used to store hazard-related information, where the hazard-related information includes the instruction type corresponding to the microcode instruction, the destination register address, and the number of execution cycles; The register hazard check unit is used to read the N groups of pre-decoding information written most recently in the second register in the current clock cycle, and query in the microcode scoreboard based on the N groups of pre-decoding information to determine whether there are data hazards and structural hazards; If there are data hazards or structural hazards, the register hazard check unit is used to send a second type of control signal to the pre-decoding module; if there are no data hazards and structural hazards, the register hazard check unit is used to send a first type of control signal to the pre-decoding module; The register file is used to read the N groups of pre-decoding information written most recently in the second register in the current clock cycle, and send the corresponding source data to the third register according to the source register address in the pre-decoding information.
[0008] Optionally, the first register is used to send the fetch start address received in the previous clock cycle to the second register; the second register is used to send the fetch start address received in the previous clock cycle to the third register; the third register is used to send the source data and the first type of decoding information corresponding to the first type of decoding unit received in the previous clock cycle to the fourth register; where the first type of decoding information is the decoding information generated by the first type of decoding unit, and the first type of decoding unit is a decoding unit capable of processing store instructions or load instructions; The backend module includes: an execution module, a fourth register, a memory access module, a fifth register, and a write-back module; The execution module is used to perform corresponding operations according to the fetch start address, N groups of decoding information, and source data written in the third register in the previous clock cycle, send the obtained operation result and the corresponding destination register address to the fourth register, and send the obtained second address to the instruction fetch module; The fourth register is used to send the operation result received in the previous clock cycle and its corresponding destination register address to the fifth register; The memory access module is used to read the first type of decoding information recently written in the fourth register. When the operation action in the first type of decoding information is a store action, store the source data corresponding to the first type of decoding information to the corresponding address. When the operation action in the first type of decoding information is a load action, obtain the load data from the corresponding address, and send the load data and the corresponding destination register address to the fifth register; The write-back module is used to read the operation result and load data written by the fifth register most recently in the current clock cycle, and write the operation result and load data back to the register file according to the corresponding destination register address.
[0009] Optionally, the execution module includes a branch execution unit and N arithmetic logic units; The branch execution unit is used to read the instruction fetch start address, the second type of decoding information, and the source data written by the third register in the previous clock cycle. Among them, the second type of decoding information is the decoding information generated by the second type of decoding unit, and the second type of decoding unit is a decoding unit with the ability to process system instructions; The branch execution unit is used to generate a second address according to the information it reads, and feedback the second address to the instruction fetch module; The j-th arithmetic logic unit is used to read the j-th decoding information and the source data written by the third register in the previous clock cycle, and perform the corresponding logical operation based on the two to obtain the j-th operation result, and send the j-th operation result and its corresponding destination register address to the fourth register.
[0010] Optionally, the branch execution unit is further used to feedback a branch instruction completion signal to the instruction fetch module after determining that the branch instruction execution is completed, so that the instruction fetch module returns from the branch instruction fetch state to the sequential instruction fetch state.
[0011] Optionally, the memory access module includes a data memory; The data memory is used to read the first type of decoding information recently written in the fourth register. When the operation action in the first type of decoding information is a store action, store the source data corresponding to the first type of decoding information to the corresponding address in the data memory. When the operation action in the first type of decoding information is a load action, obtain the load data from the corresponding address in the data memory, and send the load data and the corresponding destination register address to the fifth register.
[0012] In a second aspect, an embodiment of the present invention provides an electronic device, including the above-mentioned RISC-V based processor.
[0013] Compared with the prior art, in the sequential instruction fetch state, the instruction fetch module of the processor and the electronic device based on RISC-V according to the present invention extracts instructions based on the instruction fetch start address of the current clock cycle, and writes the obtained first target instruction start address into the first register; the first target instruction includes at least two microcode instructions; the pre-decoding module performs pre-decoding processing on the second target instruction, and sends the obtained N groups of pre-decoding information and the second target instruction to the second register; the decoding module reads the second target instruction and N groups of pre-decoding information written most recently from the second register, determines whether the current clock cycle is a safe cycle based on the same, and feeds back the judgment result of the safe cycle to the pre-decoding module; sends the corresponding source data to the third register according to the source register address in the pre-decoding information, decodes the read second target instruction to obtain N groups of decoding information, and sends the N groups of decoding information to the third register; the backend module executes the corresponding microcode instructions according to the N groups of decoding information and source data written most recently in the third register. The first target instruction includes at least two microcode instructions, so that multiple microcode instructions can be executed in each clock cycle, which can improve the computing performance of the processor. Through the pre-decoding information, the source data is written into the third register in advance, which is convenient for directly calling the corresponding source data when executing the microcode instructions, and improves the operating efficiency of the processor.
[0014] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a schematic structural diagram of the RISC-V-based processor provided by the embodiment of the present invention.
[0017] Figure 2 It is a schematic structural diagram of the instruction fetch module provided by the embodiment of the present invention.
[0018] Figure 3 It is a schematic structural diagram of the pre-decoding module provided by the embodiment of the present invention.
[0019] Figure 4 It is a schematic structural diagram of the decoding module provided by the embodiment of the present invention.
[0020] Figure 5Schematic diagram of the execution module provided by the embodiment of the present invention.
[0021] Figure 6 Schematic diagram of the connection between the write-back module and the register file provided by the embodiment of the present invention.
[0022] In the figure: 10 - instruction fetch module; 71 - first register; 20 - pre-decoding module; 72 - second register; 30 - decoding module; 73 - third register; 40 - execution module; 74 - fourth register; 50 - memory access module; 75 - fifth register; 60 - write-back module. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0024] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0026] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.
[0027] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "arrangement" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0028] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0029] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a processor based on RISC-V provided by an embodiment of the present invention. The processor includes a fetch instruction module 10, a first register 71, a pre-decoding module 20, a second register 72, a decoding module 30, a third register 73, and a backend module connected in sequence. Optionally, the backend module includes an execution module 40, a fourth register 74, a memory access module 50, a fifth register 75, and a write-back module 60. Figure 1 What is not shown in
[0030] The fetch instruction module 10 is used to extract instructions according to the fetch start address of the current clock cycle in the sequential fetch state, and write the obtained first target instruction start address into the first register 71.
[0031] Among them, the first target instruction includes at least two microcode instructions. Thus, multiple microcode instructions can be executed in each clock cycle, which can improve the computing performance of the processor. And there is no dependency relationship between any two microcode instructions in the first target instruction, so as to reduce the hardware overhead of the processor.
[0032] When the first target instruction written by the fetch instruction module 10 into the first register 71 is a branch instruction, after the writing is completed and the branch instruction completion signal is not received, the fetch instruction module 10 is in the branch fetch state; otherwise, the fetch instruction module 10 is in the sequential fetch state. It should be noted that when the fetch instruction module 10 is in the branch fetch state, it will not write a valid first target instruction into the first register 71.
[0033] The pre-decoding module 20 is used to perform pre-decoding processing on the second target instruction, and send the obtained N groups of pre-decoding information and the second target instruction to the second register 72.
[0034] If the current clock cycle is a safe cycle, the second target instruction is the first target instruction that was most recently written and is valid in the first register 71. If the current clock cycle is a non-safe cycle, the second target instruction is the first target instruction read from the first register 71 in the most recent safe cycle. The pre-decoding information for the j-th group is the pre-decoding information corresponding to the j-th microcode instruction, where 1 ≤ j ≤ N and N represents the total number of microcode instructions in the first target instruction. The j-th pre-decoding information includes the destination register address and source register address corresponding to the j-th microcode instruction. A non-safe cycle means that one or several microcodes have not been executed completely, so it is necessary to continue waiting for execution. At this time, the instruction needs to remain unchanged and cannot be overwritten by a new instruction.
[0035] It should be understood that the pre-decoding module 20 can read the first target instruction stored in the first register 71 and cache it.
[0036] A safe cycle means that there are no data hazards and structural hazards for the second target instruction written to the second register in the previous clock cycle; a non-safe cycle means that there are data hazards or structural hazards for the second target instruction written to the second register in the previous clock cycle. If the current clock cycle is a safe cycle, the second target instruction written to the second register in the previous clock cycle can be issued for execution.
[0037] The decoding module 30 is used to read the most recently written second target instruction and N groups of pre-decoding information from the second register 72, determine whether the current clock cycle is a safe cycle based on it, and feedback the judgment result of the safe cycle to the pre-decoding module 20; send the corresponding source data to the third register 73 according to the source register address in the pre-decoding information, perform decoding processing on the read second target instruction to obtain N groups of decoding information, and send the N groups of decoding information to the third register 73.
[0038] The j-th decoding information includes the instruction type, destination register address, operation action (such as addition, subtraction, multiplication, division, etc.), number of execution cycles (the number of clock cycles required for the execution of this microcode instruction), and immediate value corresponding to the j-th microcode instruction.
[0039] It should be understood that through pre-decoding, the source register address is obtained in advance and output to the second register, and then the register file can be accessed according to the pre-decoding information while decoding, and the decoding information and source data are synchronously written to the third register, which is convenient for directly calling the corresponding source data when executing the microcode instruction and improving the operation efficiency of the processor.
[0040] The backend module is used to execute the corresponding microcode instructions according to the N groups of decoding information and source data that were most recently written to the third register 73.
[0041] In the RISC-V based processor provided by the embodiments of the present invention, the first target instruction includes at least two microcode instructions, so that multiple microcode instructions can be executed in each clock cycle, which can improve the computing performance of the processor. Through pre-decoding, the source register address is obtained in advance and output to the second register, and then the register file can be accessed according to the pre-decoding information while decoding, and the decoding information and source data are synchronously written into the third register, which is convenient to directly call the corresponding source data when executing the microcode instruction, improving the operation efficiency of the processor.
[0042] Based on Figure 1 this, regarding the structure of the instruction fetch module, the embodiments of the present invention further provide an alternative implementation manner. Please refer to Figure 2 , Figure 2 which is the schematic structural diagram of the instruction fetch module provided by the embodiments of the present invention.
[0043] The instruction fetch module 10 includes a first selector S1, a program counter, an adder, an instruction memory, and an identification unit.
[0044] The first input end of the first selector S1 is connected to the output end of the adder, the second input end of the first selector S1 is connected to the backend module, the output end of the first selector S1 is connected to the program counter, and the control end of the first selector S1 is connected to the output end of the identification unit.
[0045] The output end of the program counter is respectively connected to the adder, the instruction memory, and the first register 71; the output end of the instruction memory is respectively connected to the first register 71 and the identification unit.
[0046] In an alternative implementation manner, the instruction memory (and / or the identification unit) is also connected to the backend module (the branch execution unit in the execution module 40), and can receive the branch instruction completion signal (F0).
[0047] The program counter is used to send the instruction fetch start address of the current clock cycle to the instruction memory and the first register 71.
[0048] The instruction memory is used to perform instruction extraction according to the instruction fetch start address of the current clock cycle in the sequential instruction fetch state, and write the extracted first target instruction into the first register 71.
[0049] When the instruction memory is in the branch instruction fetch state, it will not write a valid first target instruction into the first register 71.
[0050] The identification unit is used to judge whether the first target instruction is a branch instruction; if not, a sequential execution signal is sent to the first selector S1; if so, a branch execution signal is sent to the first selector S1.
[0051] Among them, a branch instruction is an instruction that causes a jump during program execution and belongs to a control hazard. The adder is used to determine a first address according to the instruction fetch start address and a preset number of bytes in the current clock cycle, and send the first address to the first selector S1.
[0052] Among them, the preset number of bytes can be but is not limited to 16 bytes, and the preset number of bytes is related to the total number of microcode instructions in the first target instruction. The adder adds the instruction fetch start address in the current clock cycle and the preset number of bytes to determine the first address.
[0053] The first selector S1 is further used to receive a second address fed back by a backend module (the execution module 40 in the following text), and the second address is the execution result of the branch instruction (i.e., the program jump address after the branch instruction is executed).
[0054] The first selector S1 is further used to send the first address to the program counter when receiving a sequential execution signal; and send the second address to the program counter when receiving a branch execution signal.
[0055] The program counter is further used to store the address sent by the first selector S1 in the current clock cycle and use it as the instruction fetch start address for the next clock cycle.
[0056] Based on Figure 1 above, for the pre-decoding module, the embodiment of the present invention further provides an optional implementation manner. Please refer to Figure 3 , Figure 3 which is the structural schematic diagram of the pre-decoding module provided by the embodiment of the present invention.
[0057] The pre-decoding module 20 includes an issue queue, a second selector S2, and N pre-decoding units (also referred to as microcode slot pre-decoding units).
[0058] The input end of the issue queue and the first input end of the second selector S2 are both connected to the first register 71, the second input end of the second selector S2 is connected to the output end of the issue queue, the control ends of the issue queue and the second selector S2 are connected to the decoding module 30 (register hazard checking unit), the output end of the second selector S2 is connected to the second register and the input ends of the N pre-decoding units, and the output ends of the pre-decoding units are connected to the second register 72.
[0059] When obtaining the first type of control signal in the current clock cycle, the issue queue and the second selector S2 are used to read the most recently written and valid first target instruction in the first register 71, where the first type of control signal indicates that the current clock cycle is a safe cycle.
[0060] When the second type of control signal is obtained in the current clock cycle, the issue queue and the second selector S2 stop reading instructions from the first register 71; the second selector S2 is used to read the first target instruction written most recently in the issue queue; wherein, the second type of control signal indicates that the current clock cycle is a non-safe cycle.
[0061] The second selector S2 is used to send the instruction read by it in the current clock cycle as the second target instruction to the second register 72 and N pre-decoding units.
[0062] It should be understood that the second target instruction also includes N microcode instructions, and the pre-decoding unit can obtain its corresponding microcode instructions. For example, the pre-decoding unit 0 obtains the first microcode instruction as bits 0-31 of the second target instruction, the pre-decoding unit 1 obtains the first microcode instruction as bits 31-63 of the second target instruction, the pre-decoding unit 2 obtains the first microcode instruction as bits 64-95 of the second target instruction, and the pre-decoding unit 3 obtains the first microcode instruction as bits 96-127 of the second target instruction. It should be noted that in the drawings of the embodiments of the present invention, N is taken as 4 as an example, but it is not limited thereto.
[0063] The j-th pre-decoding unit is used to pre-decode the j-th microcode instruction received by it in the current clock cycle to obtain the j-th pre-decoding information corresponding to the j-th microcode instruction, and send the j-th pre-decoding information to the second register 72.
[0064] On the basis of Figure 1 Regarding the decoding module, an optional implementation manner is further provided in the embodiments of the present invention. Please refer to Figure 4 Figure 4 which is the structural schematic diagram of the decoding module provided by the embodiments of the present invention.
[0065] The decoding module 30 includes N decoding units (also referred to as microcode slot decoding units), a microcode scoreboard, a register file, and a register hazard check unit.
[0066] The input ends of the N decoding units are connected to the second register 72, the output ends of the N decoding units are connected to the microcode scoreboard and the third register 73, and the microcode scoreboard is connected to the register hazard check unit.
[0067] The register file is connected to the second register 72 through a microcode read port. The number of microcode read ports corresponding to different pre-decoding units can be different and is configured according to user requirements. The output end of the register file is connected to the third register 73, and the register file is also connected to a backend module (write-back module 60).
[0068] The register hazard check unit is also connected to the control end of the second register, the control end of the issue queue, and the control end of the second selector S2.
[0069] The j-th decoding unit is used to read the j-th microcode instruction in the second target instruction written (and valid) most recently in the second register 72 in the current clock cycle, perform decoding to obtain the j-th decoding information, and send the j-th decoding information to the third register 73 and the microcode scoreboard.
[0070] The microcode scoreboard is used to store hazard-related information, where the hazard-related information includes the instruction type corresponding to the microcode instruction, the destination register address, and the number of execution cycles.
[0071] Among them, the instruction type and the number of execution cycles are used to represent the number of cycles during which the destination register is occupied. After the number of occupied cycles, this hazard-related information can be deleted from the microcode scoreboard. The instruction type is used to represent the execution unit of the instruction, and the number of execution cycles is used to represent the number of clock cycles required for the instruction to complete execution. After the instruction is executed, it means that the data in the destination register is ready, and there will no longer be data and structural hazards in this destination register.
[0072] The register hazard check unit is used to read N groups of pre-decoded information written most recently in the second register 72 in the current clock cycle, and query in the microcode scoreboard based on the N groups of pre-decoded information to determine whether there are data hazards and structural hazards.
[0073] If the source register address or the destination register address in the pre-decoded information is the same as the destination register address in the execution state in the microcode scoreboard, it is considered that there is a hazard dependency. At this time, there is a data hazard or a structural hazard.
[0074] If there is a data hazard or a structural hazard, the register hazard check unit is used to send a second type of control signal to the pre-decoding module 20; if there is no data hazard and no structural hazard, the register hazard check unit is used to send a first type of control signal to the pre-decoding module 20.
[0075] In an alternative embodiment, in a non-safe cycle, the register hazard check unit can also send a second type of control signal to the instruction fetch module 10 (in the instruction memory) to avoid writing a new first target instruction to the first register 71 when the first target instruction in the first register 71 has not been read yet.
[0076] The register file is used to read N groups of pre-decoded information written most recently in the second register 72 in the current clock cycle, and send the corresponding source data to the third register 73 according to the source register address in the pre-decoded information.
[0077] Please continue to refer to Figure 1, after entering the current clock cycle, the first register 71 is used to send the fetch start address it received in the previous clock cycle to the second register 72; the second register 72 is used to send the fetch start address it received in the previous clock cycle to the third register 73; the third register 73 is used to send the source data and the first type of decoding information corresponding to the first type of decoding unit it received in the previous clock cycle to the fourth register 74; wherein, the first type of decoding information is the decoding information generated by the first type of decoding unit, and the first type of decoding unit is a decoding unit capable of processing store instructions or load instructions.
[0078] The execution module 40 is used to perform corresponding operations according to the fetch start address, N groups of decoding information, and source data written by the third register 73 in the previous clock cycle, send the obtained operation result and the corresponding destination register address to the fourth register 74, and send the obtained second address to the instruction fetch module 10.
[0079] The fourth register 74 is used to send the operation result and its corresponding destination register address it received in the previous clock cycle to the fifth register 75.
[0080] The memory access module 50 is used to read the first type of decoding information recently written in the fourth register 74. When the operation in the first type of decoding information is a store operation, store the source data corresponding to the first type of decoding information to the corresponding address. When the operation in the first type of decoding information is a load operation, obtain the load data from the corresponding address, and send the load data and the corresponding destination register address to the fifth register 75.
[0081] Wherein, the corresponding address is the address calculated according to the instruction type, operation, and source data.
[0082] The write-back module 60 is used to read the operation result and load data recently written by the fifth register 75 in the current clock cycle, and write the operation result and load data back to the register file according to the corresponding destination register address.
[0083] On the basis of Figure 1 , regarding the structure of the execution module, the embodiment of the present invention also provides an optional implementation manner. Please refer to Figure 5 , Figure 5 which is the structural schematic diagram of the execution module provided by the present embodiment.
[0084] The execution module 40 includes a branch execution unit and N arithmetic logic units (also called microcode slot arithmetic logic units).
[0085] The branch execution unit is connected to the third register 73, the input ends of the N arithmetic logic units are connected to the third register 73, and the output ends of the N arithmetic logic units are connected to the fourth register 74.
[0086] The branch execution unit is used to read the instruction fetch start address, the second type of decoding information, and the source data written in the third register 73 in the previous clock cycle. Among them, the second type of decoding information is the decoding information generated by the second type of decoding unit, and the second type of decoding unit is a decoding unit with the ability to process system instructions.
[0087] The branch execution unit is used to generate a second address according to the information it reads and feedback the second address to the instruction fetch module 10.
[0088] The jth arithmetic logic unit is used to read the jth decoding information and the source data written in the third register 73 in the previous clock cycle, and perform corresponding logical operations based on the two to obtain the jth operation result, and send the jth operation result and its corresponding destination register address to the fourth register 74.
[0089] It should be noted that the arithmetic logic unit includes any one or more of an addition unit, a logic unit, a multiplication unit, a system instruction & system operation unit, a division unit, and a shift unit. In some scenarios, the branch execution unit can be combined with one arithmetic logic unit to save space.
[0090] Optionally, the branch execution unit is further used to feedback a branch instruction completion signal (F0) to the instruction fetch module 10 (to the instruction memory and the recognition unit) after determining that the branch instruction execution is completed, so that the instruction fetch module 10 returns from the branch instruction fetch state to the sequential instruction fetch state.
[0091] Please continue to refer to Figure 5 , the memory access module 50 includes a data memory.
[0092] The data memory is used to read the first type of decoding information recently written in the fourth register 74. When the operation action in the first type of decoding information is a store action, store the source data corresponding to the first type of decoding information into the corresponding address in the data memory. When the operation action in the first type of decoding information is a load action, obtain the load data from the corresponding address in the data memory, and send the load data and the corresponding destination register address to the fifth register 75.
[0093] Please refer to Figure 6 , Figure 6 is a schematic connection diagram between the write-back module and the register file provided by the embodiment of the present invention. There are N connection channels provided between the write-back module 60 and the register file.
[0094] It should be noted that in the embodiments of the present invention, the module connected to the clock line can receive the clock signal. However, the unit or module that receives the clock signal is not limited to the module connected to the clock line. For the sake of clarity of the drawings, some are not drawn.
[0095] The embodiments of the present invention also provide an electronic device, which includes the above-mentioned RISC-V based processor.
[0096] In summary, for the RISC-V based processor and the electronic device provided by the embodiments of the present invention, when the instruction fetch module is in the sequential instruction fetch state, it extracts instructions according to the instruction fetch start address of the current clock cycle, and writes the obtained first target instruction start address into the first register; the first target instruction includes at least two microcode instructions; the pre-decoding module performs pre-decoding processing on the second target instruction, and sends the obtained N groups of pre-decoding information and the second target instruction to the second register; the decoding module reads the second target instruction and N groups of pre-decoding information written most recently from the second register, determines whether the current clock cycle is a safe cycle based on them, and feeds back the judgment result of the safe cycle to the pre-decoding module; according to the source register address in the pre-decoding information, the corresponding source data is sent to the third register, the second target instruction read is decoded to obtain N groups of decoding information, and the N groups of decoding information are sent to the third register; the backend module executes the corresponding microcode instructions according to the N groups of decoding information and source data written most recently in the third register. The first target instruction includes at least two microcode instructions, so that multiple microcode instructions can be executed in each clock cycle, which can improve the computing performance of the processor. Through pre-decoding, the source register address is obtained in advance and output to the second register, and then the register file can be accessed according to the pre-decoding information while decoding, and the decoding information and source data are synchronously written into the third register, which is convenient for directly calling the corresponding source data when executing the microcode instructions, and improves the operation efficiency of the processor.
[0097] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0098] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any respect, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A RISC-V based processor, characterized in that, Including: An instruction fetching module, which is used to fetch instructions according to the instruction fetching start address of the current clock cycle in the sequential instruction fetching state, and write the obtained first target instruction start address into the first register; the first target instruction includes at least two microcode instructions; A pre-decoding module, which is used to perform pre-decoding processing on the second target instruction, and send the obtained N groups of pre-decoding information and the second target instruction to the second register; If the current clock cycle is a safe cycle, the second target instruction is the first target instruction that was most recently written and is valid in the first register; if the current clock cycle is a non-safe cycle, the second target instruction is the first target instruction read from the first register in the most recent safe cycle. The j-th pre-decoding information includes the destination register address and source register address corresponding to the j-th microcode instruction; A decoding module, which is used to read the second target instruction and N groups of pre-decoding information that were most recently written from the second register, determine whether the current clock cycle is a safe cycle based on them, and feedback the judgment result of the safe cycle to the pre-decoding module; send the corresponding source data to the third register according to the source register address in the pre-decoding information, perform decoding processing on the read second target instruction to obtain N groups of decoding information, and send the N groups of decoding information to the third register; The j-th decoding information includes the instruction type, destination register address, operation action, execution cycle number, and immediate number corresponding to the j-th microcode instruction; A back-end module, which is used to execute the corresponding microcode instructions according to the N groups of decoding information and source data that were most recently written in the third register.
2. The RISC-V based processor according to claim 1, wherein, The instruction fetching module includes a first selector, a program counter, an adder, an instruction memory, and an identification unit; The program counter is used to send the instruction fetching start address of the current clock cycle to the instruction memory and the first register; The instruction memory is used to fetch instructions according to the instruction fetching start address of the current clock cycle in the sequential instruction fetching state, and write the fetched first target instruction into the first register; The identification unit is used to determine whether the first target instruction is a branch instruction; if not, send a sequential execution signal to the first selector; if so, send a branch execution signal to the first selector; The adder is used to determine a first address according to the instruction fetching start address of the current clock cycle and a preset byte number, and send the first address to the first selector; The first selector is further used to receive a second address fed back by the back-end module, and the second address is the execution result of the branch instruction; The first selector is further used to send the first address to the program counter when receiving the sequential execution signal; and send the second address to the program counter when receiving the branch execution signal; The program counter is further used to store the address sent by the first selector in the current clock cycle and use it as the instruction fetching start address of the next clock cycle.
3. The RISC-V based processor according to claim 1, characterized in that, The pre-decoding module includes an issue queue, a second selector, and N pre-decoding units, where N represents the total number of microcode instructions in the first target instruction; When the first type of control signal is obtained in the current clock cycle, the issue queue and the second selector are used to read the most recently written and valid first target instruction in the first register, where the first type of control signal indicates that the current clock cycle is a safe cycle; When the second type of control signal is obtained in the current clock cycle, the issue queue and the second selector stop reading instructions from the first register; the second selector is used to read the most recently written first target instruction in the issue queue; where the second type of control signal indicates that the current clock cycle is a non-safe cycle; The second selector is used to send the instruction read by it in the current clock cycle as the second target instruction to the second register and N pre-decoding units; The j-th pre-decoding unit is used to pre-decode the j-th microcode instruction received by it in the current clock cycle to obtain the j-th pre-decoding information corresponding to the j-th microcode instruction, and send the j-th pre-decoding information to the second register.
4. The RISC-V based processor according to claim 3, wherein The decoding module includes N decoding units, a microcode scoreboard, a register file, and a register hazard check unit; The j-th decoding unit is used to read the j-th microcode instruction in the second target instruction that was most recently written in the second register in the current clock cycle, and perform decoding to obtain the j-th decoding information, and send the j-th decoding information to the third register and the microcode scoreboard; The microcode scoreboard is used to store hazard-related information, where the hazard-related information includes the instruction type corresponding to the microcode instruction, the destination register address, and the execution cycle number; The register hazard check unit is used to read the N groups of pre-decoding information that were most recently written in the second register in the current clock cycle, and query within the microcode scoreboard based on the N groups of pre-decoding information to determine whether there are data hazards and structural hazards; If there is a data hazard or a structural hazard, the register hazard check unit is used to send the second type of control signal to the pre-decoding module; if there are no data hazards and structural hazards, the register hazard check unit is used to send the first type of control signal to the pre-decoding module; The register file is used to read the N groups of pre-decoding information that were most recently written in the second register in the current clock cycle, and send the corresponding source data to the third register according to the source register address in the pre-decoding information.
5. The RISC-V based processor according to claim 4, wherein The first register is used to send the fetch start address received by it in the previous clock cycle to the second register; the second register is used to send the fetch start address received by it in the previous clock cycle to the third register; the third register is used to send the source data and the first type of decoding information corresponding to the first type of decoding unit received by it in the previous clock cycle to the fourth register; where the first type of decoding information is the decoding information generated by the first type of decoding unit, and the first type of decoding unit is a decoding unit capable of processing store instructions or load instructions; The backend module includes: an execution module, a fourth register, a memory access module, a fifth register, and a write-back module; The execution module is used to perform corresponding operations according to the instruction fetch start address, N sets of decoding information, and source data written in the third register in the previous clock cycle, send the obtained operation result and the corresponding destination register address to the fourth register, and send the obtained second address to the instruction fetch module; The fourth register is used to send the operation result and its corresponding destination register address received in the previous clock cycle to the fifth register; The memory access module is used to read the first type of decoding information recently written in the fourth register. When the operation in the first type of decoding information is a store operation, store the source data corresponding to the first type of decoding information at the corresponding address. When the operation in the first type of decoding information is a load operation, obtain the load data from the corresponding address, and send the load data and the corresponding destination register address to the fifth register; The write-back module is used to read the operation result and load data written most recently in the fifth register in the current clock cycle, and write the operation result and load data back to the register file according to the corresponding destination register address.
6. The RISC-V based processor according to claim 5, wherein The execution module includes a branch execution unit and N arithmetic logic units; The branch execution unit is used to read the instruction fetch start address, the second type of decoding information, and the source data written in the third register in the previous clock cycle, where the second type of decoding information is the decoding information generated by the second type of decoding unit, and the second type of decoding unit is a decoding unit with the ability to process system instructions; The branch execution unit is used to generate a second address according to the information it reads, and feedback the second address to the instruction fetch module; The j-th arithmetic logic unit is used to read the j-th decoding information and source data written in the third register in the previous clock cycle, and perform corresponding logical operations based on the two to obtain the j-th operation result, and send the j-th operation result and its corresponding destination register address to the fourth register.
7. The RISC-V based processor according to claim 6, wherein The branch execution unit is further used to feedback a branch instruction completion signal to the instruction fetch module after determining that the branch instruction execution is completed, so that the instruction fetch module returns from the branch instruction fetch state to the sequential instruction fetch state.
8. The RISC-V based processor according to claim 5, wherein The memory access module includes a data memory; The data memory is used to read the first type of decoding information recently written in the fourth register. When the operation in the first type of decoding information is a store operation, store the source data corresponding to the first type of decoding information at the corresponding address in the data memory. When the operation in the first type of decoding information is a load operation, obtain the load data from the corresponding address in the data memory, and send the load data and the corresponding destination register address to the fifth register.
9. An electronic device, characterized in that, Including the RISC-V based processor according to any one of claims 1-8.
Citation Information
Patent Citations
Novel 8-digit RISC microcontroller framework
CN101221494A
Ordered processor using multi-transmit scheme and method of operation thereof
CN118426837A
Processor and method providing instruction support for instructions that utilize multiple register windows
US20110296142A1
Processor and method of executing load instructions out-of-order having reduced hazard penalty
US6868491B1
Cited By
Finger fetching module and related equipment
CN121300857A