A RISC-V based neural network extended execution computer system
By designing a neural network based on RISC-V, combining out-of-order execution and dynamic scheduling, the problem of low computing efficiency of RISC-V processors in low-power design is solved, and efficient and low-power neural network computing acceleration is achieved, suitable for edge computing and AI applications of mobile devices.
Patent Information
- Application Number
- CN202510097175.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing processors based on RISC-V instruction set architectures are less computationally efficient in low-power designs, while high-performance designs are too high in power and cost, making it difficult to meet the needs of edge computing and mobile devices for high computing performance and low power consumption, especially inefficient in neural network inference.
A neural network extended execution computer system based on RISC-V is designed, combining a processor, a neural network hardware accelerator, a command memory and a dynamic scheduler with an out-of-order execution architecture. By dynamically switching the working mode, the RISC-V processor and the neural network hardware accelerator are used to perform general computing and neural network computing tasks respectively, and reduce the clock flip rate at low load to reduce power consumption.
It realizes the real-time and privacy protection of devices while maintaining high computing efficiency while reducing power consumption, adapting to high computing needs in mobile and edge scenarios, and improving the real-time and privacy protection capabilities of devices. It is suitable for AI applications of edge computing devices and mobile devices.
Smart Images

Figure CN119539001B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent computer architectures, and particularly relates to a neural network extended execution computer system based on RISC-V. Background Art
[0002] RISC-V (the fifth generation of Reduced Instruction Set Computing Architecture) is an open-source and modular reduced instruction set architecture with scalability. Its base instruction set includes 32-bit integer instructions (RV32I) and 64-bit integer instructions (RV64I), and its extended instruction sets include multiplication and division (RV32M), floating-point operations (RV32F), etc., and can be combined according to requirements. For example, RV32IM means including the 32-bit integer instruction set and the multiplication and division extension instruction. In addition, users can customize some RISC-V instructions according to the needs of application scenarios to accelerate the speed of certain operations. As a design specification of a reduced instruction set architecture (RISC), the RISC-V instruction set defines the basic instruction set that a computer processor can understand and execute. Therefore, RISC-V computer processors or RISC-V computer systems are actual hardware designed by different researchers according to the RISC-V instruction set architecture. Therefore, different design ideas and hardware architectures will bring different functions and performances to RISC-V computer systems. Moreover, some processors can perform custom extensions and hardware implementations on the RISC-V instruction set for specific applications.
[0003] Currently, the processor design based on the RISC-V instruction set architecture mainly focuses on two major directions: on the one hand, low-power design, mainly for embedded applications such as mobile devices and edge computing; on the other hand, high-performance design, mainly for scenarios that require a large amount of computing resources such as servers and artificial intelligence. Low-power design usually adopts a simplified hardware architecture, such as fewer pipeline stages and sequential execution methods, to reduce the hardware resource overhead and thus meet the power consumption requirements. However, such a design is weak in computing efficiency and is prone to pipeline stalls due to data conflicts, thus wasting computing time. This makes it impossible for some applications with high requirements for timeliness and computing power, such as driverless and edge computing, to meet the real-time requirements. In contrast, existing high-performance designs adopt a complex superscalar architecture and use multi-issue and out-of-order execution technologies to improve instruction parallelism and processing efficiency. Although it can significantly improve computing power, it also leads to an increase in hardware resource consumption and energy consumption. In scenarios such as edge computing that are sensitive to power consumption and cost, how to balance low power consumption and efficient computing and use limited hardware resources to improve computing performance has become an urgent challenge to be solved.
[0004] In addition, as artificial intelligence gradually migrates to edge devices (such as smartphones, drones, autonomous vehicles, intelligent sensors, etc.), these devices have high requirements for the real-time performance of neural network inference and often cannot rely on remote data centers for computing. Therefore, a processor integrated with a neural network accelerator can directly perform neural network inference locally, reducing the dependence on cloud computing. However, the training and inference processes of neural networks usually involve a large number of matrix operations and parallel computing tasks, which are less efficient for traditional processors (such as CPUs); while high-performance graphics processing units (GPUs) can execute neural network training and inference quickly, but their high power consumption and high cost make them difficult to be applied to edge devices. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to propose a neural network extended execution computer system based on RISC-V, aiming to improve the computing efficiency, reduce the power consumption, and have the function of accelerating neural network computing by designing a flexible computer system hardware microarchitecture, so as to improve the versatility and flexibility of the computer system with the function of accelerating neural network computing and adapt to high-computing-demand and power-sensitive scenarios including neural network inference in mobile and edge scenarios, so as to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above object, the present invention provides a neural network extended execution computer system based on RISC-V, including: a RISC-V processor based on an out-of-order execution architecture, a neural network hardware accelerator, an instruction memory, a data memory, and a dynamic scheduler;
[0007] The dynamic scheduler is used to switch different working modes for the computer system according to the current computing load, and allocate the computing tasks to be executed to the RISC-V processor or the neural network hardware accelerator for execution; the working modes include: a standard computing mode, a neural network acceleration mode, and a low-power mode; the computing tasks include general computing tasks and neural network computing tasks;
[0008] The RISC-V processor is used to receive and decode externally input RISC-V instructions in the standard computing mode, form an out-of-order issued instruction stream according to the instruction type and the current computing load, and execute the general computing tasks and the neural network computing tasks by sequentially executing the RISC-V instructions in the out-of-order issued instruction stream;
[0009] The neural network hardware accelerator is used to execute the neural network computing tasks in the neural network acceleration mode;
[0010] The dynamic scheduler is further configured to compile the to-be-executed computing tasks into one or more RISC-V instructions, and write the RISC-V instructions into the instruction memory for storage;
[0011] The instruction memory is used to sequentially store one or more RISC-V instructions corresponding to the to-be-executed computing tasks;
[0012] The data memory is used to store the interaction data transmitted between the RISC-V processor and the neural network hardware accelerator through a dedicated communication protocol;
[0013] The dynamic scheduler is further configured to, when there is no computing task or the computing load is lower than a preset value, adopt a gated clock logic to reduce the clock flip rate of the RISC-V processor and the neural network hardware accelerator, and control the RISC-V processor and the neural network hardware accelerator to operate in the low power consumption mode.
[0014] Preferably, the RISC-V processor sequentially includes: an instruction receiving unit, an instruction decoding unit, a physical register file, an out-of-order execution unit, and a write-back unit;
[0015] The instruction receiving unit is configured to prefetch a plurality of RISC-V instructions from the instruction memory for internal caching, and retrieve whether there is a RISC-V instruction corresponding to the program count value among the prefetched plurality of RISC-V instructions; if so, output the retrieved RISC-V instruction to the instruction decoding unit; if not, access the instruction memory, and output the corresponding RISC-V instruction found in the instruction memory to the instruction decoding unit;
[0016] The instruction decoding unit is configured to parse the received RISC-V instruction, identify the register access information and data operation information carried by the RISC-V instruction, and mix-encode them into instruction identification information;
[0017] The physical register file is used to store the logical register data used by the RISC-V processor during execution through a plurality of physical registers;
[0018] The out-of-order execution unit is configured to classify the decoded RISC-V instructions according to the instruction type to form a plurality of issue queues, form an out-of-order issued instruction stream in each issue queue according to the current computing load, and execute the general computing task and the neural network computing task by out-of-order executing the RISC-V instructions in the instruction stream;
[0019] The write-back unit is configured to write the execution result back to the physical register file.
[0020] Preferably, the neural network hardware accelerator includes:
[0021] A model compression unit for performing knowledge distillation on the convolutional neural network in the neural network computing task and compressing the model parameters of the convolutional neural network;
[0022] A computing unit, including a convolution operation module, a matrix multiplication module, a quantization module, an activation function module, and a pooling module, for performing forward propagation calculation operations in the convolutional neural network;
[0023] A storage unit, including independent input feature map cache, weight cache, bias cache, intermediate value cache, pooling cache, output feature map cache, and shared cache, for reducing storage access conflicts and reducing data access latency through a multi-level cache system;
[0024] A data flow control unit for controlling data flow to ensure efficient data transmission between the computing unit and the storage unit;
[0025] A communication interface for data transmission with the RISC-V processor, loading input image data and the model parameters of the convolutional neural network, and outputting the recognition result of the convolutional neural network.
[0026] Preferably, the convolution operation module includes a convolution calculation array and a first adder tree, for sliding a convolution kernel in the input feature map, sequentially performing element-wise multiplication operations in the sliding area by using the convolution calculation array, and accumulating the intermediate calculation results of the element-wise multiplication operations based on the first adder tree to generate an output feature map;
[0027] The matrix multiplication module includes a multiplication calculation array and a second adder tree for performing output calculations of the fully connected layer of the convolutional neural network; wherein, the multiplication calculation array is used for performing multiplication operations on the accessed multi-channel feature maps and the corresponding point convolution kernels, and the second adder tree is used for accumulating the intermediate calculation results of the multiplication calculation array.
[0028] Compared with the prior art, implementing the technical solution provided by the present invention has the following advantages and technical effects:
[0029] The RISC-V-based neural network extended execution computer system provided by the present invention can, on the one hand, improve the computing efficiency and reduce the power consumption of the overall system under the premise of maintaining efficient execution of instructions, so as to achieve a better balance between the performance and energy efficiency of the computer system; on the other hand, it takes into account the execution requirements of general computing tasks and neural network computing tasks, and can switch different working modes for the computer system based on the current computing load. When the computing load is not large and the neural network computing speed is not demanding, the computing resources in the execution unit of the RISC-V processor can be used to execute the neural network computing task; when the neural network computing speed is required to be high, a dedicated neural network hardware accelerator is called to perform hardware acceleration on the neural network computing task to meet the computing requirements of the neural network computing task. Deploying the computer system provided by the present invention on edge devices and mobile devices can not only reduce data transmission delays and meet real-time requirements, but also improve the privacy protection capabilities of the device and reduce dependence on the network. In addition, this integrated design improves the comprehensive performance of the processor when executing traditional general computing tasks and artificial intelligence (AI) tasks, and can also promote the popularization of AI applications on edge computing devices and mobile devices, bringing a wider range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0031] Figure 1 A schematic diagram of the overall micro-architecture of an embodiment of a RISC-V-based neural network extension execution computer system provided by the present invention;
[0032] Figure 2 A functional block diagram of an embodiment of a RISC-V processor based on an out-of-order execution architecture provided by the present invention;
[0033] Figure 3 A principle block diagram of an embodiment of a neural network hardware accelerator provided by the present invention;
[0034] Figure 4 A principle block diagram of an embodiment of a branch prediction unit of a RISC-V processor provided by the present invention;
[0035] Figure 5 A principle block diagram of an embodiment of a register renaming unit provided by the present invention;
[0036] Figure 6 A logical operation structure diagram of an embodiment of the priority arbitration tree provided by the present invention. DETAILED DESCRIPTION
[0037] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0038] It should be noted that the signal transmission process shown in the accompanying drawings refers to the execution process in a computer system such as a set of computer-executable instructions.
[0039] In view of the common characteristics of limited hardware resources and energy consumption but high requirements for computing performance and intelligence in current application scenarios such as embedded, mobile communication, Internet of Things (IoT), edge computing, and intelligent terminals, the present invention proposes a custom processor based on the RISC-V instruction set, which has high performance, low power consumption, single-issue, out-of-order execution, and a neural network computing acceleration function. This processor seeks the best balance among computing performance, intelligence, and power consumption, aiming to provide higher instruction execution efficiency. Its core goal is to solve three key problems in the design of intelligent computer systems: First, how to improve the computing performance and task processing efficiency of the processor, increase the parallelism of instruction execution, and reduce the instruction execution latency to meet the growing high-speed computing requirements; Second, how to reduce the power consumption of the processor so that it can operate efficiently in an environment with limited hardware resources and energy, such as in embedded and mobile devices; Third, efficiently customize a dedicated hardware accelerator for neural network computing tasks to significantly improve the efficiency of neural network operations. The goal of the embodiments of the present invention is to, while improving the instruction parallelism and efficiency of the processor, comprehensively consider low-power design, and significantly increase the computing throughput of the neural network through an innovative microarchitecture solution, so that the computer system can select the optimal computing method and working mode according to the computing load and application needs, and better adapt to the application scenarios where high-performance computing, intelligent computing, and low-power requirements coexist.
[0040] The embodiments of the present invention propose a flexible, intelligent, high-performance and low-power computer system based on the RISC-V instruction set architecture, which can work in multiple working modes, improves the instruction parallelism of the RISC-V processor to obtain higher computing performance, and at the same time takes into account the low-power characteristics of the microarchitecture, achieving a better balance between computing efficiency and power consumption, and having a neural network computing acceleration function, which can be used for the computing acceleration of convolutional neural networks to meet the growing needs of intelligent devices.
[0041] Such as Figure 1As shown in the figure, it is a schematic structural diagram of an implementable manner of the RISC-V based neural network extended execution computer system provided by the embodiment of the present invention. Specifically, the RISC-V based neural network extended execution computer system includes: a RISC-V processor 100 based on an out-of-order execution architecture, a neural network hardware accelerator 200, an instruction memory 300, a data memory 400, and a dynamic scheduler 500.
[0042] The dynamic scheduler 500 is used to switch different working modes for the computer system according to the current computing load, and allocate the computing tasks to be executed to the RISC-V processor 100 or the neural network hardware accelerator 200 for execution; the working modes include: a standard computing mode, a neural network acceleration mode, and a low power consumption mode; the computing tasks include general computing tasks and neural network computing tasks.
[0043] The RISC-V processor 100 is used to receive and decode externally input RISC-V instructions in the standard computing mode, form an out-of-order issued instruction stream according to the instruction type and the current computing load, and execute the general computing tasks and the neural network computing tasks by sequentially executing the RISC-V instructions in the out-of-order issued instruction stream.
[0044] The neural network hardware accelerator 200 is used to execute the neural network computing tasks in the neural network acceleration mode.
[0045] The dynamic scheduler 500 is further used to compile the computing tasks to be executed into one or more RISC-V instructions, and write the RISC-V instructions into the instruction memory 200 for storage.
[0046] The instruction memory 300 is used to sequentially store one or more RISC-V instructions corresponding to the computing tasks to be executed.
[0047] The data memory 400 is used to store the interaction data transmitted between the RISC-V processor 100 and the neural network hardware accelerator 200 through a dedicated communication protocol.
[0048] The dynamic scheduler 500 is further used to adopt gated clock logic to reduce the clock flip rate of the RISC-V processor 100 and the neural network hardware accelerator 200 when there are no computing tasks or the computing load is lower than a preset value, and control the RISC-V processor 100 and the neural network hardware accelerator 200 to work in the low power consumption mode.
[0049] In the RISC-V instruction set architecture, the instruction types of its 32-bit integer instruction set RV32I mainly include register instructions (R-type instructions), immediate instructions (I-type instructions), store instructions (S-type instructions), conditional jump instructions or branch instructions (B-type instructions), etc. Among them, R-type instructions are mainly used to control arithmetic and logical operations between the 32 logical registers defined by RV32I. Therefore, R-type instructions are also known as arithmetic and logical operation instructions; I-type instructions are used to control operations between logical registers and immediate numbers; S-type instructions are mainly used to perform store operations, writing data to memory or loading data from the outside; B-type instructions are also called branch instructions, which are used to perform conditional branch operations and determine whether to jump according to the value of the logical register. In addition, RV32I also includes U-type instructions and J-type instructions, etc., which can be specifically referred to the standard definition of the RISC-V instruction set. It should be noted that the RISC-V instruction set architecture also includes a 64-bit integer instruction set RV64I, a single-precision floating-point instruction set RV32F, a double-precision floating-point instruction set RV64F, and various extended instruction sets. For example, the instruction sets RV32IM and RV64IM for multiplication and division extensions, and the vector extension instruction sets RV32IMV and RV64IMV based on them. Different RISC-V processors may have significant differences in the supported instruction sets and the corresponding hardware implementation architectures and their complexities.
[0050] In an embodiment of the present invention, the designed RISC-V processor 100 can be compatible with the instruction sets RV32IFMV and RV64IFMV for floating-point operations, multiplication and division operation extensions, and vector calculation extensions on the basis of the basic integer instruction set. The instructions that the RISC-V processor 100 can process include but are not limited to arithmetic and logical operation instructions, floating-point operation instructions, branch instructions, load instructions, memory access instructions, and custom extended vector calculation instructions (for example, vector calculation instructions based on RV32V and RV64V extensions). Those skilled in the art can adaptively expand the compatible extended instructions under the inspiration of the embodiments of the present invention according to the actual application scenarios.
[0051] Based on the above instruction architecture, an embodiment of the present invention designs an out-of-order execution RISC-V processor with high computing efficiency, low power consumption, and intelligence, and its hardware architecture schematic diagram is as Figure 2 shown. In a preferred implementation, the RISC-V processor 100 sequentially includes: an instruction receiving unit 110, an instruction decoding unit 120, a physical register file 130, an out-of-order execution unit 140, and a write-back unit 150.
[0052] Among them, the instruction receiving unit 110 is configured to prefetch several RISC-V instructions from the instruction memory 300 for internal caching, and retrieve whether there is a RISC-V instruction corresponding to the program count value among the prefetched several RISC-V instructions; if so, output the retrieved RISC-V instruction to the instruction decoding unit 120; if not, access the instruction memory 300, and output the corresponding RISC-V instruction found in the instruction memory 300 to the instruction decoding unit 120.
[0053] The instruction decoding unit 120 includes a decoder 121, which is configured to parse the received RISC-V instruction, identify the register access information and data operation information carried by the RISC-V instruction, and mix-encode them into instruction identification information.
[0054] The physical register file 130 is configured to store the logical register data used by the RISC-V processor 100 during execution through multiple physical registers. In the RISC-V instruction architecture, R-type instructions and the like involved therein need to operate on the data in the registers. To distinguish from the physical registers implemented by hardware, these registers defined in the instruction architecture are called logical registers or general-purpose registers. Taking the RV32I basic integer instruction set architecture as an example, it defines a total of 32 logical registers, named from x0 to x31. Among them, the x0 logical register is a constant register, and its value is always "0"; the x1 - x31 logical registers can be used for general-purpose data storage (i.e., destination logical registers) and data reading (i.e., source logical registers). In R-type or I-type instructions, etc., the specific operation on a certain logical register has been defined. However, due to the complexity of computer programs, the compiled instructions may repeatedly read and write to some logical registers multiple times, which is likely to cause read-write conflicts. Therefore, some technical solutions are needed to implement the mapping of physical registers to logical registers to avoid the occurrence of register read-write conflicts.
[0055] The out-of-order execution unit 140 is configured to classify the decoded RISC-V instructions according to the instruction type to form multiple issue queues, and form an out-of-order issued instruction stream in each issue queue according to the current computing load, and execute the general computing task and the neural network computing task by out-of-order executing the RISC-V instructions in the instruction stream.
[0056] The write-back unit 150 is configured to write the execution result back to the physical register file 130.
[0057] Different from existing computer systems, the computer system provided by the embodiments of the present invention is suitable for processing a wide range of computing tasks, has strong flexibility, and can operate in three modes: standard computing mode, neural network acceleration mode, and low-power mode. In the standard computing mode, the computer system provided by this embodiment can execute basic arithmetic logic operations, floating-point operations, and even neural network computing tasks relatively slowly based on the computing resources carried by the RISC-V processor. However, the processing efficiency is usually low, and the power consumption is relatively high. Therefore, in order to be compatible with the rapidly developing neural network computing requirements, by integrating a neural network hardware accelerator 200, a dedicated hardware designed specifically for neural network computing, the neural network tasks can be executed with lower power consumption and higher efficiency by switching to the neural network acceleration mode, which is especially suitable for application scenarios with high computing performance requirements and power consumption constraints, such as edge computing devices, smartphones, autonomous driving, Internet of Things devices, etc. In addition, when there is no computing task, the computer system provided by this embodiment will start the low-power mode to save more energy consumption.
[0058] As Figure 3 shown, it is a structural block diagram of an implementation manner of the neural network hardware accelerator provided by the embodiments of the present invention.
[0059] In a preferred implementation manner, the neural network hardware accelerator 200 includes:
[0060] A model compression unit 210, which is used to perform knowledge distillation on the convolutional neural network in the neural network computing task and compress the model parameters of the convolutional neural network. In a convolutional neural network, its model parameters mainly include weight values, bias values, etc. This embodiment introduces the knowledge distillation (abbreviated as KD) technology. By connecting a complex and powerful large model as a teacher model, online training and compression are performed on the hardware-based neural network model deployed on the neural network hardware accelerator 200, and the knowledge and performance of the teacher model can be transferred to the hardware-based neural network model, thereby avoiding online deployment of a complex and computationally expensive large model. The hardware-based model after knowledge distillation has fewer parameters and smaller computational complexity but has the knowledge of the teacher model, which makes it more efficient than the teacher model during inference and can accelerate the inference process in actual applications.
[0061] The computing unit 220 includes a convolution operation module 221, a matrix multiplication module 222, a quantization module 223, an activation function module 224, and a pooling module 225, and is used to perform forward propagation calculation operations in a convolutional neural network. The forward propagation calculation operations mainly refer to various successive operations from the input image to the output layer in the neural network. Among them, preferably, the convolution operation module 221 includes a convolution calculation array 2211 and a first adder tree 2212, and is used to slide a convolution kernel in the input feature map, and sequentially perform element-wise multiplication operations in the sliding area by using the convolution calculation array 2211, and, based on the first adder tree 2212, accumulate the intermediate calculation results of the element-wise multiplication operations to generate an output feature map. In an implementable manner, the matrix multiplication module 222 includes a multiplication calculation array 2221 and a second adder tree 2222, and is used to perform output calculations for the fully connected layer of the convolutional neural network; among them, the multiplication calculation array 2221 is used to perform multiplication operations on the multi-channel feature maps connected thereto and the corresponding point convolution kernels, and the second adder tree 2222 is used to accumulate the intermediate calculation results of the multiplication calculation array 2221. In specific implementation, the convolution operation module 221 is mainly used for image feature extraction, and it performs element-wise multiplication and accumulation operations between the convolution kernel and the input data, focusing on spatial local information processing. The matrix multiplication module 222 is used to process matrix operations of network layers such as the fully connected layer, focusing on efficiently processing large-scale data and complex matrix calculations.
[0062] The storage unit 230 includes an independent input feature map cache 231, a weight cache 232, a bias cache 233, an intermediate value cache 234, a pooling cache 235, an output feature map cache 236, and a shared cache 237, and is used to reduce storage access conflicts and reduce data access latency through a multi-level cache system. Among them, the sizes of the input feature map cache 231, the output feature map cache 236, and the shared cache 237 can be the same, and the resources of the shared cache 237 can be combined with other caches under different model settings to obtain different storage modes. The selection of the storage mode is mainly determined by the convolution mode, accelerator control signal, and feature map characteristics in the network layer configuration parameters.
[0063] The data flow control unit 240 is used to control the data flow, ensure efficient data transmission between the computing unit 220 and the storage unit 230, reduce the transmission latency between various cache data and the computing unit 220, and reduce power consumption. In hardware design, the access process of storage resources is often the most time-consuming and energy-consuming operation. Therefore, in the design process, the distance between the storage space and the computing resources should be shortened as much as possible, and the access operation should be simplified and the access frequency to the storage space should be reduced.
[0064] A communication interface 250 for data transmission with the RISC-V processor 100, loading input image data and model parameters of the convolutional neural network, and outputting recognition results of the convolutional neural network.
[0065] For the RISC-V processor, it acts as the central processing unit in the computer system provided in this embodiment. Specifically, as Figure 2 shown, the low-power single-issue out-of-order execution RISC-V processor provided in this embodiment includes a five-stage pipeline architecture, which are in turn: an instruction reception unit 110, an instruction decoding unit 120, a physical register file 130, an out-of-order execution unit 140, and a write-back unit 150. It should be noted that for the pipeline architectures provided in this embodiment, according to the needs of the application scenario, the internal components of one or more pipeline architectures can be further split into multi-stage pipeline architectures to meet the needs of different instructions or computing efficiencies, shorten the critical path, and increase the operating frequency of the processor.
[0066] The instruction reception unit 110 is configured to read the corresponding RISC-V instruction from an off-chip memory (i.e., the instruction memory 300) according to the output value (abbreviated as the PC value) of the program counter (Program Counter, abbreviated as PC) in the current cycle. Specifically, the program counter is a register, and the PC value in the program counter points to the off-chip memory address of the instruction to be executed. Therefore, the PC value can be used as the read address to fetch the RISC-V instruction from the instruction memory 300. In the pipeline architecture, although instructions are sequentially executed in each stage, when a conditional branch instruction is encountered, the execution direction or jump address of the instruction stream may change. However, in modern computer processors, the execution of a program is relatively complex. The RISC-V instruction stream obtained after compilation and conversion usually does not completely consist of sequentially executed instructions, but also contains some jump instructions or branch instructions. Without the support of a branch predictor, the processor usually has to wait until decoding is completed to know, and at least one clock cycle needs to be waited, which will cause discontinuous execution and jamming of the processor, resulting in performance degradation. Therefore, this embodiment further provides a mechanism for dynamically predicting instruction branches, which pre-judges the instruction type when receiving an instruction, and pre-judges whether the PC value will jump to a non-continuous fetch address. In a preferred implementation manner, in the RISC-V processor 100, its instruction reception unit 110 includes: a dynamic branch prediction unit 111.
[0067] As Figure 4 shown, the dynamic branch prediction unit 111 includes: a branch target buffer unit 1111, a jump history buffer unit 1112, a global history prediction register unit 1113, and a jump arbitration unit 1114.
[0068] Specifically, the branch target cache unit 1111 is used to record the program count value (i.e., the PC value) corresponding to the branch instruction that has undergone a jump and the jump address. The branch target cache unit 1111 serves as a branch target buffer (BTB), which can accelerate the acquisition of the jump target address of the branch instruction. By pre-storing the jump target address of the branch instruction, the BTB allows the instruction receiving unit 110 to obtain the jump target in advance, avoiding the time for waiting for the calculation of the jump target address, thereby accelerating the execution of the jump instruction.
[0069] The jump history buffer unit 1112 is used to record the jump direction and the number of jumps corresponding to each program counter value in the historical process of instruction execution; the jump direction and the number of jumps corresponding to each program counter value are updated by a saturating counter according to the jump results of the executed instructions. A Pattern History Table (PHT) is stored in the jump history buffer unit 1112, which is used to store the historical behavior patterns of branch instructions. Each entry in the PHT stores a saturating count value, which indicates whether to jump. The PHT helps the branch predictor to speculate on future branch behaviors based on "similar history". For example, if a certain branch instruction has jumped multiple times in the past, the branch predictor will predict that it may also jump next time, and vice versa. Among them, each saturating count value in the PHT can be calculated and updated by a saturating counter (SC) with a 2-bit width. The value range of the 2-bit width SC is the 2-bit binary number 00~11, which is used to represent the jump state of the corresponding instruction, corresponding to strongly not jump (00), weakly not jump (01), weakly jump (10), and strongly jump (11) in sequence. Among them, when the branch actually jumps, the counter is incremented by one (for example, 00 is updated to 01, 01 is updated to 10), indicating an increase in the prediction confidence for the jump; when the branch actually does not jump, the counter is decremented by one (for example, 10 is updated to 01, 11 is updated to 10), indicating an increase in the prediction confidence for not jumping. If the saturating count value is weakly jump (10) or strongly jump (11), the counter remains unchanged or continues to increase, but will not exceed 11 (the update direction is 0001→10→11). If the saturating count value is weakly not jump (01) or strongly not jump (00), the counter remains unchanged or continues to decrease, but will not be lower than 00 (the update direction is 11→10→01→00). In this way, the saturating counter can adjust the prediction of branch jumps through historical information. Especially for stable branches, which long follow the jump or non-jump pattern, so the jump behavior of the branch is relatively regular, and the saturating counter can maintain accurate predictions by gradually increasing the confidence (i.e., entering the strongly jump or strongly not jump state). For unstable branches, that is, branches with irregular jump and non-jump patterns, the saturating counter adopted in this embodiment has anti-interference ability because it can avoid overly frequent prediction changes through the weakly jump or weakly not jump states. For example, in the case of a sudden change in branch behavior, the saturating counter will not make a wrong prediction immediately, but adjust gradually, which can avoid the performance loss caused by wrong predictions. The 2-bit width saturating counter provides high accuracy through a moderately complex design, while avoiding the high hardware overhead brought by more complex prediction mechanisms. Compared with more advanced predictors, the 2-bit width saturating counter is a simple but effective prediction tool that can provide good performance with low overhead.The saturation count value maintained by the saturation counter is stored in the PHT, and the saturation count value is obtained by addressing the PHT to perform the prediction of branch jump.
[0070] It should be noted that the access address of the jump history cache unit 1112 in this embodiment is indexed by the PC. However, due to the hardware resource limitations of the processor, it is difficult for the PHT to cover all PC values (for example, for the RV32I instruction architecture, the PC value is 32 bits). Therefore, this embodiment only uses a relatively small storage space to implement the jump history cache unit 1112, and only addresses the PHT in the jump history cache unit 1112 through a partial bit-width of the PC. For example, when the storage depth of the PHT is 1024, only the lower ten bits of the PC (bits 2-11, and the storage minimum unit is 1 byte, so bits 0 and 1 are invalid) are taken as the access address of the PHT. However, this scheme may also cause the problem that multiple branch instructions can address the same entry in the PHT, which may cause the historical data of these branches to interfere with or overwrite each other, resulting in the aliasing problem of branch instructions, thus affecting the accuracy of branch prediction. Therefore, the embodiment of the present invention designs a global history prediction register unit 1113 to increase the diversity of branch history, thereby reducing the sharing and conflict between different branch instructions, improving the accuracy of branch prediction and reducing performance loss.
[0071] As Figure 4As shown, the global history prediction register unit 1113 is used to predict the instruction jump result after the execution of several consecutive instructions, and generate a global jump history prediction value. The global history prediction register unit 1113 stores the execution history of the most recent series of branch instructions in the program. The global history prediction register unit 1113 can be implemented by a shift register, and each bit of its data represents the jump behavior (jump or not jump) of a branch instruction. After the branch instruction is executed and retired, the global history prediction register unit 1113 will record the jump result of the instruction (the updated value of "1" indicates a jump, and "0" indicates no jump). Specifically, when the global history prediction register unit 1113 is updated, it will shift one bit to the left to provide space for a new branch execution, and its lowest bit will be set to the jump result (1 or 0) of this branch. At this time, the historical information stored in the global history prediction register unit 1113 includes the execution results of all previous branches, forming a sliding window. The number of bits of the global history prediction register unit 1113 determines the length of the branch history that can be stored. The longer the history register, the more historical information can be recorded, which may improve the accuracy of branch prediction. However, due to resource constraints, the global history prediction register unit 1113 with a bit width of 6 bits is preferably used in this embodiment. To solve the alias problem of branch instructions, in this embodiment, by improving the indexing mechanism of the PHT, not only the PC is used as the index of the PHT, but a partial PC value (for example, the local 6-bit PC value: PC[12:7]) is XORed with the output value of the global history prediction register unit 1113 to obtain a new index of the PHT.
[0072] The jump arbitration unit 1114 is used to predict and obtain the second program count value of the currently to-be-read RISC-V instruction according to the first program count value of the previous RISC-V instruction, including:
[0073] Performing a logical operation on the first program count value and the global jump history prediction value to obtain an index address, querying the jump history cache unit according to the index address, and determining whether the second program count value jumps;
[0074] If the second program count value jumps, it is determined whether the first program count value hits any branch instruction in the branch target cache unit. If it hits, it is determined whether the hit branch instruction is a function return instruction; if it is the function return instruction, the function call target address corresponding in the return address stack (RAS) is predicted as the jump address of the second program count value; if it is not the function return instruction, the jump address corresponding to the hit branch instruction is predicted as the jump address of the second program count value;
[0075] If the second program count value does not jump, or if it jumps but the first program count value does not hit any branch instruction in the branch target buffer unit, then predict the jump address of the second program count value as PC + k; where PC is the first program count value and k is a user-defined constant.
[0076] The return address stack is used to handle function calls and returns. By storing the return addresses of function calls, it helps the branch predictor to predict the target addresses of function returns, so that when a function returns, the instruction receiving unit 110 can quickly obtain the target address and perform a jump.
[0077] The jump arbitration unit 1114 arbitrates to obtain the jump address for the next cycle from multiple possible jump addresses. Specifically, if the jump direction found in the jump history buffer unit 1112 for the instruction fetched by the current PC is jump, and the target address of the PC corresponding to this instruction exists in the BTB (i.e., the BTB hits this instruction), and the instruction hit in the BTB is not a function return instruction, then the target address hit in the BTB is used as the jump address for the next cycle of this instruction; if it is a function return address, then pop the address in the RAS as the jump address; when the instruction jump direction is predicted as not jump, or the BTB is not hit, then use PC + 4 as the jump address for the next cycle (k = 4), that is, read instructions sequentially. The reason for choosing k = 4 is that the instructions in the instruction memory 300 are stored in bytes with a width of 8 bits, and for the RV32IM instruction set, the width of an instruction is 32 bits, so it takes every 4 bytes to completely read out the information of an RV32IM instruction. Similarly, for instruction sets with other bit widths and different forms of instruction storage methods, the k value needs to be adjusted adaptively.
[0078] Specifically, the method steps for predicting the jump of a branch instruction provided in this embodiment are as follows:
[0079] Step S1: Read the program counter to obtain the PC value output by it, and only take the partial bit width value (for example, the 6-bit PC middle value PC[11:6]);
[0080] Step S2: Read the global jump history prediction value P-GHR of the global history prediction register unit 1113;
[0081] Step S3: Use the PC value with the middle part width, such as the middle 6-bit PC[11:6], to perform an exclusive OR operation with the 6-bit global history prediction register unit in the global history prediction register unit 1113. Then, splice the lower 4-bit PC[5:2] to the obtained intermediate value to get a 10-bit (corresponding to 1024 bits of the PHT) index that integrates global jump information, that is, the PHT read address is {{PC[11:6] (XOR) P-GHR[5:0]}, PC[5:2]}. Using this index to address the PHT can effectively avoid the alias problem caused by the PHT size limitation.
[0082] Step S4: Use the combined 10-bit query address above to perform a read operation on the pattern history table (PHT) in the jump history cache unit 1112 with a depth of 1024. According to the query result of the PHT, determine whether the PC value corresponding to the instruction to be executed in the next cycle jumps: If it jumps, execute Step S5; if it does not jump, execute Step S6.
[0083] Step S5: Retrieve in the BTB whether there is a record information corresponding to the current PC value. If not, that is, if there is no hit, execute Step S6; if any item in the BTB is equal to the PC of the current cycle, it means that the current PC value hits an item in the BTB. However, since the BTB also records the call records of the return functions with function return instructions and their priority is higher, therefore, if the current PC hits an item in the BTB, it is also necessary to execute Step S7 to determine whether it is a function return instruction. In this embodiment, the BTB is preferably formed in a fully associative cache form with 32 items. Any cache block can be stored at any position in the cache without being restricted by the address.
[0084] Step S6: Predict the PC value of the next cycle as PC + k. In this embodiment, if either the PHT or the BTB does not hit the current PC value, it is predicted as "not jump", that is, directly fetch instructions sequentially. For RV32I, the PC is incremented by k = 4 each time.
[0085] Step S7: Determine whether the instruction in the BTB hit by the current PC is a function return instruction; if not, execute Step S8; if so, execute Step S9.
[0086] Step S8: Use the target address recorded in the item hit in the BTB as the jump address of the current PC, that is, obtain the predicted value of the next PC.
[0087] Step S9: Pop the corresponding function return address in the RAS from the stack as the prediction result.
[0088] The processor will select the instruction address to be read in the next cycle according to the prediction result of the dynamic branch prediction unit 111. Branch prediction provides input for speculative execution, determines the branch path in advance and starts execution, which can improve pipeline throughput, reduce pauses caused by branch delays, maintain efficient operation of the pipeline, and greatly improve the parallelism and execution efficiency of the processor, so that the processor of the embodiment of the present invention can reduce the potential performance bottleneck caused by branch instructions while maintaining high performance. This embodiment can complete the access to PHT and BTB within one clock cycle, that is, the prediction result can be obtained within one clock cycle, achieving the effect of fast prediction and preventing the further deepening of the micro-architecture pipeline.
[0089] The dynamic branch prediction unit 111 of this embodiment predicts whether the instruction fetch address will jump based on the historical information of the branch (such as branch history, global history, etc.). For example, based on the previous branch execution situation (jump or not jump), it is predicted whether the current branch will jump. In addition to predicting the jump direction of whether to jump, the dynamic branch prediction unit 111 of this embodiment also predicts the target address of the jump. Based on the dynamic branch prediction unit 111, this embodiment adopts the technical solution of speculative execution to improve the throughput of instructions. When encountering a conditional branch, the processor will determine in advance whether the branch will jump according to the dynamic branch prediction unit 111, and try to execute the instructions on the branch path in advance.
[0090] After fetching the instruction corresponding to the current PC value, the PC value of the next cycle is predicted by the dynamic branch prediction unit 111 as the read address of the next instruction. The RISC-V instructions to be executed are stored outside the RISC-V processor, which helps to reduce the hardware resource consumption and chip area of the RISC-V processor itself; but there is usually an additional off-chip read delay as a cost. Therefore, in order to reduce the time required for instruction reading and realize efficient reading of off-chip stored instructions, the embodiment of the present invention further sets an instruction prefetch unit 112 and an instruction query unit 113 in the instruction receiving unit 110.
[0091] Furthermore, if Figure 2 As shown, the instruction receiving unit 110 further includes an instruction pre-fetching unit 112 and an instruction querying unit 113 .
[0092] The instruction prefetch unit 112 is used to access the instruction memory 300 through a dedicated bus interface according to the first program counter value, and read the RISC-V instruction corresponding to the first program counter value and several adjacent RISC-V instructions from the instruction memory 300 for caching.
[0093] The instruction query unit 113 is configured to query whether there is a RISC-V instruction corresponding to the second program count value output by the dynamic branch prediction unit 111 in the instruction prefetch unit 112; if so, output the corresponding RISC-V instruction in the instruction prefetch unit 112 to the instruction decoding unit 120; if not, find the RISC-V instruction corresponding to the second program count value in the instruction memory 300 and output it to the instruction decoding unit 120.
[0094] By storing frequently used instructions, the instruction prefetch unit 112 can accelerate the instruction fetch speed and reduce the time for the RISC-V processor to access the computer system memory. When the processor executes a software program, instructions need to be loaded from the instruction memory 300, but the access speed of the memory is relatively slow. By caching copies of some frequently used instructions in the instruction prefetch unit 112, the instruction receiving unit 110 can fetch instructions from the cache of the instruction prefetch unit 112 more quickly, avoiding fetching from the main memory of the computer (off-chip storage for the RISC-V processor) every time. Specifically, when the instruction query unit 113 obtains the address of the current instruction from the PC, it will first send a request to fetch the instruction to the instruction prefetch unit 112. The instruction prefetch unit 112 will first check whether the requested instruction is already stored in the cache. If it exists, it is a "cache hit"; if not, it is a "cache miss". When there is a cache hit, the instruction receiving unit 110 directly fetches the instruction from the cache of the instruction prefetch unit 112, skipping the step of accessing the main memory instruction memory 300, greatly improving the instruction fetch speed. When there is a cache miss, the instruction receiving unit 110 will access the instruction memory 300 through the cache bus and the AXI interface to fetch the corresponding RISC-V instruction, and fetch and cache the RISC-V instruction corresponding to the current fetch address of the instruction memory 300 and several adjacent instructions into the instruction prefetch unit 112 for subsequent quick access by the PC and instruction fetching. The energy consumption required for the instruction receiving unit 110 to access the instruction prefetch unit 112 is much less than that for accessing the off-chip instruction memory 300, and this strategy helps to reduce the number of accesses to the instruction memory 300, thus helping to save the overall power consumption of the processor microarchitecture.
[0095] Furthermore, as Figure 2 shown, the instruction decoding unit 120 further includes a register renaming unit 122, which is configured to dynamically map the logical registers included in the RISC-V instruction set to the corresponding physical registers in the physical register file 130, generating register access control information for preventing read-write conflicts in the register data path; the logical registers include a destination logical register and a source logical register.
[0096] Among them, the architecture of a preferred implementation manner of the register renaming unit provided in this embodiment is as Figure 5 shown. The register renaming unit 122 includes: a physical register free list 1221, a speculative renaming mapping table 1222, a physical register ready list 1223, and a register allocation circuit 1224.
[0097] Specifically, the physical register free list 1221 is used to mark the free state of each physical register. In a specific implementation, a flag bit can be set for the physical register free list 1221 to indicate whether the physical register can be used for renaming allocation: when the corresponding physical register flag bit in the physical register free list is "1", it indicates that the physical register is in a free state and can be used for renaming allocation; when the flag bit is "0", it indicates that the physical register is being occupied and cannot be used for renaming allocation.
[0098] The speculative renaming mapping table 1222 is used to record the dynamic mapping relationship between logical registers and physical registers. The speculative renaming mapping table 1222 can query and update the mapping relationship between logical registers and physical registers. When a write register needs to be executed, when the register allocation circuit 1224 allocates a free physical register for the destination logical register of the instruction, the allocated information is written into the speculative renaming mapping table 1222, that is, the mapping relationship between the logical register and the physical register in the speculative renaming mapping table 1222 is updated. When an instruction needs to read a register, the two source logical registers of the instruction entering the renaming stage need to query the corresponding physical register by reading the speculative renaming mapping table 1222, and the instruction uses the queried physical register as the access entry address of the physical register file 130 to read the data in the corresponding physical register.
[0099] The physical register ready list 1223 is used to mark the data ready state of each physical register.
[0100] The register allocation circuit 1224 is used to detect whether the physical register free list 1221 is available when decoding the destination logical register to be written of the RISC-V instruction; if not, it waits for the release of any physical register; if so, it selects a free physical register in the physical register free list 1221 through the priority arbitration tree 1225, allocates it as the destination physical register to the destination logical register to be written, and updates the mapping relationship between the corresponding logical register and the physical register in the speculative renaming mapping table 1222.
[0101] The register allocation circuit 1224 is further configured to, when decoding and obtaining the source logical register to be read of the RISC-V instruction, look up the corresponding physical register as the source physical register in the prediction renaming mapping table 1222, and look up the data ready status of the corresponding source physical register in the physical register ready table 1223.
[0102] The register allocation circuit 1224 is further configured to integrate the free destination physical register, source physical register, and data ready status of the source physical register associated with the current RISC-V instruction into the register access control information.
[0103] The register renaming unit 122 allocates physical registers for the logical registers in the RISC-V instruction, enabling the operations of each instruction to be performed independently, improving the parallelism of instruction execution, and avoiding register data conflicts. The physical register does not directly correspond to the logical register converted into an instruction in the software program, but is a specific hardware resource in the processor for storing data. The register renaming unit 122 is set up considering the following possible data conflict situations: In a program, multiple instructions may read or write the same logical register, including WAR (write after read) data dependencies and WAW (write after write) data dependencies. Specifically, WAR data dependency means that in a program, two consecutive instructions need to read and then write the same register. The write operation of the second instruction depends on the read operation of the first instruction. However, due to the circuit structure design and the limited number of logical registers in the RISC-V instruction set, data in the pipeline may be rewritten first because the second instruction is executed first, that is, the data of the register is written after being read, resulting in incorrect data writing, and thus a WAR data conflict occurs. And WAW means that two instructions both attempt to write to the same register or storage location, but their execution order is improper, resulting in the write operation executed later overwriting the write operation executed earlier, thus causing a conflict. Given the limitations of the number of logical registers in the RISC-V instruction set (for example, the RV32I base instruction set only defines 32 logical registers), the program is prone to relatively frequent WAW and WAR data conflict situations. To solve the data dependency problem caused by register resource conflicts, this embodiment introduces a newly designed register renaming technology. This technology uses more physical registers than the number of logical registers and dynamically maps the logical registers of each instruction to different physical registers, avoiding competition for the same logical register between different instructions.
[0104] Preferably, in this embodiment, the logical registers in the instruction set are renamed by adopting a way of doubling the physical registers. For example, 64 physical registers are used to rename the 32 logical registers defined by RV32I. Specifically, the 32 logical registers defined by RV32I are dynamically mapped to 64 physical registers. The processor can eliminate WAW and WAR conflicts, enhance instruction parallelism, improve the execution efficiency of the processor, and thus improve the overall performance without adding too much hardware overhead. Experiments have proved that doubling the physical registers is sufficient in this embodiment to rename the logical registers defined by the instruction set, avoid possible WAW or WAR conflicts, and do not consume too many register resources.
[0105] When specifically implemented, in order to quickly arbitrate for an idle physical register in the register renaming unit 122 to map the destination logical register that needs to perform a write operation in the instruction, this embodiment provides a preferred implementation structure of a priority arbitration tree.
[0106] Figure 6 It is the design schematic diagram of an embodiment of the priority arbitration tree provided by the present invention.
[0107] In a preferred implementation manner, the priority arbitration tree 1225 is a comparison selection merging arbiter with fixed low priority. The comparison selection merging arbiter includes: a plurality of one-bit comparators (MAX), and a plurality of two-to-one multiplexers (MUX) corresponding to the one-bit comparators (MAX) one by one. A one-bit comparator means that the bit width of the input signal for comparison is 1 bit (bit), and it is a dual-input comparator. The two input signals of the two-to-one multiplexer (MUX) are the numbers of physical registers, and their values can be any value in 0~n-1 (n>1), so the bit width of its input signals can be multiple bits.
[0108] The plurality of one-bit comparators form a binary tree structure for pairwise comparison of the idle states of all physical registers in the physical register free list 1221; if each physical register is idle, its idle state is marked as "1", and if it is not idle, its idle state is marked as "0".
[0109] Each one-bit comparator is a maximum value comparator, and when pairwise comparing the idle states of the input physical registers: if the idle states of the two input physical registers are different or the idle states of the two input physical registers are both idle, the comparison result of the current one-bit comparator is "1"; if the idle states of the two input physical registers are both not idle, the comparison result of the current one-bit comparator is "0".
[0110] The multiple multiplexers form a binary tree structure for screening out the first free physical register in ascending order of numbers from the lowest bit to the highest bit in the physical register file.
[0111] Each of the multiplexers uses the comparison result of the corresponding one-bit comparator as a selection signal to screen out the physical register with the corresponding number, including: if the comparison result of the current one-bit comparator is "1", it is determined whether the free states of the two physical registers connected to the current one-bit comparator are the same; if they are different, the number of the physical register with the free state marked as "1" is selected as the input of the next-level multiplexer, and if they are the same, the number of the physical register at the lower bit is used as the input of the next-level multiplexer; if the comparison result of the current one-bit comparator is "0", the number of the physical register at the higher bit is used as the input of the next-level multiplexer.
[0112] The comparison selection merging arbiter is used to select the corresponding free physical register as the destination physical register of the current RISC-V instruction when the comparison result of the one-bit comparator at the top of the tree is "1".
[0113] The arbitration logic of the priority arbitration tree 1225 is to perform pairwise comparison and result transmission of multiple groups of MAX and MUX with the fixed lower bit as the priority. The input data of the priority arbitration tree 1225 are two groups of information, one of which is the free states of each physical register saved in the physical register free list 1221, which can be represented by the binary sequence R0~R n-1 indicating the free state of the physical register corresponding to the corresponding bit. When it is free, the corresponding bit is "1", and when it is not free, the corresponding bit is "0"; the other group is the corresponding physical register numbers I0~I n-1 in the physical register file. When there are more than two physical registers, the numbers I0~I n-1 are integer fixed-point values.
[0114] The physical register release mechanism in this embodiment is as follows: when an instruction writes to a logical register, before the next write to this logical register, the value of this register may be read. Therefore, the mapping from this logical register to the physical register should be retained (i.e., the free state of the physical register is not free). When all subsequent instructions that ensure reading this value have been completed, the corresponding physical register can be released. Therefore, in addition to overwriting the mapping from the logical register to the physical register, the original physical register (i.e., the previous destination physical register) also needs to be recorded. When the instruction retires in the ROB, it means that all instructions that may have depended on this physical register before have been completed. At this time, the original physical register can be released, that is, set the corresponding position in the physical register free list 1221 of this physical register to "1".
[0115] It should be noted that if the comparison result of a one-bit comparator at the top of the tree is "0", it indicates that there is no idle physical register currently. However, this situation where all registers are occupied can be overcome in this embodiment. Since the number of physical registers is at least twice the number of logical registers, at least one write destination logical register can be mapped to two physical registers. At this time, when writing to the same destination logical register in the third instruction, the physical register mapped in the first instruction can be overwritten because when the second instruction is executed, the first instruction has already been executed. Therefore, overwriting the physical register for writing the destination mapped in the first instruction when the third instruction is executed will not bring incorrect data to the calculation of the processor. Therefore, in this embodiment, only the information of the previous destination physical register corresponding to the previous instruction of the currently executed instruction needs to be recorded in the register access control information, so that when the next instruction is executed, the physical register of the previous instruction can be used to write data. At the same time, it also proves that this solution can avoid the data read-write conflict problem of logical registers by using twice the number of physical registers for renaming.
[0116] In addition, relying solely on register remapping technology can eliminate WAW and WAR conflicts, but it cannot eliminate RAW (read after write) conflicts. A RAW conflict refers to a situation where, during the execution of a program, a subsequent instruction needs to read the value of a register, but the value of that register has not yet been written by a previous instruction. RAW conflicts essentially involve data dependencies between instructions. Register renaming cannot change the data dependencies. Even if each logical register has a different physical register, when there is a RAW conflict, the subsequent instruction still needs to wait for the write-back data of the previous instruction to be completed, and the data dependency still exists. Therefore, the register renaming unit 122 also needs to handle RAW conflicts between instructions. For this purpose, the physical register ready table 1223 proposed in the embodiments of the present invention records the ready state of each physical register, and detects data dependencies by querying the ready state of the physical registers of the instructions. Specifically, when an instruction completes register renaming, it is assigned a new physical register. If the new physical register has not been written back by the out-of-order execution unit 140 and a subsequent instruction needs to read this physical register, a RAW conflict will occur, that is, this physical register is in an unready state. The corresponding flag bit of this physical register in the physical register ready table 1223 is set to "0". When a physical register is executed and written back by an instruction, the corresponding flag bit of this physical register in the physical register ready table 1223 is set to "1". When a subsequent instruction needs to read this physical register again, no RAW conflict will occur, and the correct result can be read. An instruction obtains the ready status of the source physical registers of this instruction by querying the physical register ready table 1223. If there are physical registers in an unready state, it indicates that this instruction has data dependencies and needs to wait for the data dependencies to be resolved, that is, it needs to wait for the instruction to be executed and written back to the physical register until all the physical registers of this instruction become ready states.
[0117] It should be noted that since the dynamic branch prediction unit 111 determines in advance whether the input instruction jumps, its branch judgment result is "speculative". If the speculation is correct, these instructions will continue to be executed; if the speculation is incorrect, all the instructions that are speculatively executed in the previous pipelines before the mis-executed instruction need to be cancelled, that is, branch prediction failure recovery needs to be performed. The RISC-V processor provided in this embodiment has a complete branch prediction and branch prediction failure recovery function.
[0118] As Figure 4 shown, further, the dynamic branch prediction unit 111 further includes: a global history register 1115, which is used to record the actual jump history of several RISC-V instructions after the execution ends.
[0119] As Figure 5As shown, the register renaming unit 122 is further provided with: a physical register free list checkpoint 1226 and an architectural renaming mapping table 1227. The physical register free list checkpoint 1226 is used to assign a branch number to the current RISC-V instruction and back up the physical register free list 1221 when the current RISC-V instruction is a branch instruction; the architectural renaming mapping table 1227 is used to record the mapping relationship between each logical register and physical register when the previous RISC-V instruction is executed. For a branch instruction, it is necessary to first complete the backup of the physical register free list checkpoint 1226. When it enters the register renaming unit 122, the physical register free list checkpoint 1226 generates a branch number and uses the checkpoint content corresponding to the branch number to back up the current physical register free list 1221. For all RISC-V instructions, it is necessary to complete the dynamic mapping from logical registers to physical registers.
[0120] The write-back unit 150 includes: branch prediction failure recovery logic 151.
[0121] The branch prediction failure recovery logic 151 is used to, when the dynamic branch prediction unit 111 makes a wrong prediction on the jump address of the second program count value, overwrite the data of the global history register 1115 to the global history prediction register unit 1113, overwrite the data of the physical register free list checkpoint 1226 with the corresponding number to the physical register free list 1221, and overwrite the data of the architectural renaming mapping table 1227 to the predicted renaming mapping table 1222.
[0122] The global history register can be abbreviated as GHR. The setting of the update node of GHR is particularly important. GHR can be updated according to the result of branch prediction in the pipeline instruction reception stage, or can be updated after calculating the jump direction of the branch instruction in the execution stage, or can be updated in the write-back retirement stage at the end of the pipeline. However, when the pipeline architecture of the RISC-V processor is relatively deep, if the scheme of updating GHR when the instruction is executed or retired is adopted, it may cause many subsequent instructions to enter the pipeline and cannot enjoy the result of updating GHR by this branch instruction; if the scheme of updating in the instruction reception stage is adopted, the instructions after the branch instruction may all be on the speculative path, and if these instructions update GHR before retirement, it may cause data errors.
[0123] Therefore, the present invention provides a global history pairing update scheme, and a global history prediction register unit 1113 and a global history register 1115 are provided. Among them, the stored data of the global history prediction register unit 1113 is updated by the prediction result of the dynamic branch prediction unit 111 in the instruction receiving stage, while the stored data of the global history register 1115 is updated by the branch prediction failure recovery logic 151 in the retirement write-back stage after the instruction execution is completed. Therefore, the global history register 1115 does not participate in branch prediction, and the stored data therein is the real and accurate global history jump information. Therefore, the stored data of the global history register 1115 can be used to restore the stored data of the global history prediction register unit 1113 when a branch prediction failure occurs. Specifically, when a branch prediction failure occurs, the stored data of the global history register 1115 is completely overwritten into the global history prediction register unit 1113, all instructions on the wrong path are cleared, and the global history prediction register unit 1113 is restored to the correct state of the previous cycle. The restoration time is only one clock cycle, realizing the fast restoration of global history prediction.
[0124] Similarly, various mapping information in the register renaming unit 122 can be backed up in real time first. When a branch prediction failure occurs, the data of the physical register free list 1221 and the predicted renaming mapping table 1222 in the register renaming unit 122 is restored to the state after the execution of the previous instruction without prediction failure by using the backup information, realizing the rollback of the data of the physical register free list 1221 and the predicted renaming mapping table 1222. The rollback operation only consumes one clock cycle, thus completing the fast state restoration of the renaming mapping table. Since the state restoration of this embodiment is only performed when the instruction is committed, all instructions before the instruction commitment are on the wrong path. Therefore, the physical register ready list of this embodiment only needs to perform a full flush operation, and all bits are reset to complete the restoration, without additional resource consumption. The flush operation also only takes one clock cycle to complete, reducing the time cost required to restore the processor state after a branch prediction failure, improving the performance of the kernel and reducing the power consumption of the processor state restoration.
[0125] In summary, the dynamic branch prediction unit 111 provided in the embodiment of the present invention adopts a two-level adaptive structure. The first level records the global branch jump information of the program through the global history prediction register unit 1113, while the second level consists of a PHT, a BTB, and a RAS. The current instruction PC value is obtained by fusing the global branch history information of the first level, and the prediction of the branch jump direction and address is completed through the cooperation of the PHT, BTB, and RAS in the second level. The two-level adaptive structure dynamic branch prediction unit 111 provided in this embodiment can complete the prediction of the current instruction within one clock cycle, quickly obtain the jump information and use it for speculative execution. Adopting such a fast branch prediction structure can effectively balance the accuracy and complexity of the branch predictor, achieve relatively accurate branch prediction with less resource consumption, and balance the power consumption and performance of the overall microarchitecture.
[0126] Further, in a preferred implementation manner, the out-of-order execution unit 140 includes: an instruction issue unit 141, an arithmetic logic unit 142, a floating-point operation unit 143, a load memory access unit 144, and a vector processing unit 145.
[0127] Among them, the instruction issue unit 141 is used to determine whether the decoded RISC-V instruction is ready to be issued, and distribute the ready RISC-V instructions to one of the issue queues according to their instruction types; sort the ages of the RISC-V instructions in each issue queue respectively, and arbitrate the oldest RISC-V instruction in each issue queue for parallel issuance. If all the source physical registers associated with the current instruction have been awakened, it is determined that the current RISC-V instruction is ready to be issued. The instruction issue unit 141 receives and temporarily stores the instruction after renaming, and issues the ready instruction to the corresponding computing unit in the instruction out-of-order execution unit 140 for execution. Therefore, according to the instruction type, the instruction issue unit 141 can allocate various instructions to be executed to a distributed parallel issue architecture to form corresponding issue queues. Multiple parallel issue queues receive the instructions that have been decoded and renamed by the instruction decoding unit 120, and monitor the readiness of the instructions. The unready instructions need to wait to be awakened, and the ready instructions wait for the oldest instruction to be arbitrated and issued to the corresponding computing unit in the out-of-order execution unit 140 for execution. Among them, the computing units include an arithmetic logic unit 142, a floating-point operation unit 143, a load memory access unit 144, and a vector processing unit 145.
[0128] During specific implementation, the arithmetic logic unit 142 is used to receive and execute the arithmetic logic operation instructions emitted by the instruction emission unit 141; the arithmetic logic operation instructions are fixed-point operation instructions. The arithmetic logic unit 142 is mainly responsible for performing arithmetic addition and subtraction, logical AND, OR, NOT, etc., shift, and comparison operations. Some branch operation instructions, including calculating the branch jump address of the branch instruction, and calculating whether the branch prediction is correct, as well as the multiplier and divider, can be implemented in the arithmetic logic unit 142 in this embodiment.
[0129] The floating-point operation unit 143 is used to receive and execute the floating-point operation instructions emitted by the instruction emission unit 141.
[0130] The load memory access unit 144 is used to receive and execute the load instructions or memory access instructions emitted by the instruction emission unit 141 to load or store external data. The load memory access unit 144 calculates a valid memory access target address according to the information such as registers, immediate numbers, offsets, etc. in the instructions, and sends this address to the load queue and the store queue to ensure that data is read or written from the correct memory location.
[0131] The vector processing unit 145 is used to receive and execute the custom extended vector calculation instructions emitted by the instruction emission unit 141 to implement vector operations and matrix operations in the neural network calculation task. During specific implementation,
[0132] The RISC-V instruction set includes extended instructions dedicated to vectorized computing (abbreviated as the RVV instruction set) to support large-scale data parallel computing. Therefore, the RISC-V processor provided in this embodiment designs a vector processing unit 145 that can execute RVV instructions in the out-of-order execution unit 140 according to the definition of the RVV instruction set, to call the hardware resources of the RISC-V processor itself to execute intensive computing tasks such as matrix operations, to call its own computing resources to implement computing tasks for applications such as image processing and machine learning, and to optimize the parallel processing of multiple data elements. The vector processing unit 145 can use vector registers connected by a multi-level pipeline architecture to store multiple data elements for parallel operations, improving the instruction throughput and the parallel processing ability of the processor.
[0133] The out-of-order execution unit 140 performs corresponding instruction operations according to the decoding result of the instruction to obtain an execution result. The execution result will be written back to the physical register file 130 by the write-back unit 150. A first-in-first-out reorder buffer (ROB) is set in the write-back unit 150 to implement the detection and recovery processing of precise exceptions and branch prediction failures of the RISC-V processor. In this embodiment, by setting the ROB, the state of the physical register file is ensured to be consistent with the result of sequential execution. Each ROB entry records the state of the instruction, the destination logical register, the destination physical register, and the previous destination physical register. The ROB checks the instruction at the head of the queue. If the instruction has been executed and there is no exception, the result can be made effective and the instruction can be deleted from the head of the queue, and the subsequent instructions can be continued to be checked. When the instruction decoding unit 120 decodes an instruction, the instruction is inserted into the tail of the ROB, and the state is updated along with the execution process of the instruction. For an instruction that encounters an exception, the instructions in the issue queue and the intermediate data during the issue and execution processes are cleared, and the execution is restarted from the exception handling address.
[0134] In addition, to reduce the overall power consumption of the system, this embodiment equips the computer system with a low-power operating mode. By comprehensively applying various technical means such as gated clock, input gating, and glitch optimization, while ensuring the efficient execution of the computer system, the overall core power consumption is successfully reduced, achieving a better balance between computing performance and energy efficiency. For example, an input gating circuit design mechanism can be adopted in the out-of-order execution unit 140, that is, by controlling whether the operands of the instruction can enter each computing unit in the out-of-order execution unit 140, thereby affecting the function and behavior of the circuit. In a specific implementation, when the gating signal is at a high level, the operands of the instruction are valid, and the input signal can enter the circuit through the gating gate; when the gating signal is at a low level, the operands of the instruction are invalid, and the input signal will be blocked or masked. Only when the current operand is valid does the circuit module need to work, reducing unnecessary toggles of the functional units in the out-of-order execution unit 140 and effectively reducing the power consumption of the execution unit. The design of each pipeline stage of the RISC-V processor 100 not only pursues performance improvement but also focuses on achieving low power consumption. By adopting a balanced design strategy in this embodiment, not only the normal implementation of the function is ensured, but also the logical structure is simplified, and unnecessary changes in the data path are reduced, thus achieving a more ideal balance between the performance and power consumption of the overall computer system.
[0135] In addition, the computer system in this embodiment also integrates a neural network hardware accelerator, which is optimized for the implementation of neural network computing operations such as matrix multiplication, convolution operations, and activation functions. For neural network computing tasks, it can achieve higher execution efficiency than a RISC-V processor, greatly improving the computing throughput, significantly reducing the inference time, and making edge real-time applications such as autonomous driving, speech recognition, and video analysis possible. In this embodiment, the neural network computing tasks can be executed either on the RISC-V processor or on a dedicated neural network hardware accelerator, and are specifically set according to the needs of the actual application scenario and the specific operating load conditions.
[0136] The microarchitecture of the computer system kernel of the present invention has completed prototype verification on an FPGA of model XC7Z020CLG400-2, and the Coremark benchmark program has been run on the FPGA platform. In addition, the operating frequency, power consumption, and area of the system kernel have been evaluated on a 40-nanometer CMOS integrated circuit process library. Among them, the RISC-V processor can stably operate on the XC7Z020CLG400-2 chip at a maximum frequency of 95 MHz, and the benchmark program report reaches 3.6 Coremark / MHz, which is better than most sequentially executed computer kernels. Under the 40nm CMOS integrated circuit process, after comprehensive optimization, the maximum operating frequency of the computer kernel is increased to 555.56 MHz, the normalized dynamic power consumption is only 17.19 μW / MHz, and the chip area is 0.081 mm². Compared with other embedded computer kernels, the processor kernel of the present invention shows significant advantages in terms of frequency and power consumption, which is attributed to the low-power dynamic scheduling scheme, optimized circuit configuration, and a series of low-power design methods for the microarchitecture. The experimental results show that the low-power computer system designed by the present invention performs excellently in terms of performance and power consumption, and is particularly suitable for application scenarios with high computing requirements and power consumption sensitivity.
[0137] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A RISC-V based neural network extended execution computer system, characterized in that, Including: A RISC-V processor based on an out-of-order execution architecture, a neural network hardware accelerator, an instruction memory, a data memory, and a dynamic scheduler; The dynamic scheduler is used to switch different working modes for the computer system according to the current computing load, and allocate the computing tasks to be executed to the RISC-V processor or the neural network hardware accelerator for execution; the working modes include: a standard computing mode, a neural network acceleration mode, and a low-power mode; the computing tasks include general computing tasks and neural network computing tasks; The RISC-V processor is used to receive and decode externally input RISC-V instructions in the standard computing mode, form an out-of-order issued instruction stream according to the instruction type and the current computing load, and execute the general computing tasks and the neural network computing tasks by sequentially executing the RISC-V instructions in the out-of-order issued instruction stream; The neural network hardware accelerator is used to execute the neural network computing tasks in the neural network acceleration mode; The dynamic scheduler is further used to compile the computing tasks to be executed into one or more RISC-V instructions, and write the RISC-V instructions into the instruction memory for storage; The instruction memory is used to sequentially store one or more RISC-V instructions corresponding to the computing tasks to be executed; The data memory is used to store the interaction data transmitted between the RISC-V processor and the neural network hardware accelerator through a dedicated communication protocol; The dynamic scheduler is further used to, when there is no computing task or the computing load is lower than a preset value, adopt gated clock logic to reduce the clock flip rate of the RISC-V processor and the neural network hardware accelerator, and control the RISC-V processor and the neural network hardware accelerator to work in the low-power mode; The RISC-V processor includes: an instruction decoding unit, a physical register file; The instruction decoding unit includes a register renaming unit, which performs renaming in a way of doubling the physical registers, and is used to dynamically map the logical registers included in the RISC-V instruction set to the corresponding physical registers in the physical register file, and generate register access control information for preventing read-write conflicts in the register data path; the logical registers include a destination logical register and a source logical register; the register renaming unit includes: a physical register free list, a predicted renaming mapping table, a physical register ready list, and a register allocation circuit; The register allocation circuit is used to detect whether the physical register free list is available when decoding the destination logical register to be written in the RISC-V instruction; if not, it waits for the release of any physical register; if so, it selects an idle physical register from the physical register free list through a priority arbitration tree, assigns it as the destination physical register to the destination logical register to be written, and updates the mapping relationship between the corresponding logical register and physical register in the predicted rename mapping table; and is also used to, when decoding the source logical register to be read in the RISC-V instruction, find the corresponding physical register in the predicted rename mapping table as the source physical register, and find the data ready status of the corresponding source physical register in the physical register ready list; The priority arbitration tree is a comparison selection merging arbiter with fixed low priority; The physical register release mechanism is: when an instruction retires in the ROB, its corresponding former destination physical register is released.
2. The RISC-V based neural network extended execution computer system according to claim 1, wherein, The RISC-V processor further includes: an instruction receiving unit, an out-of-order execution unit, and a write-back unit The instruction receiving unit is used to prefetch a number of RISC-V instructions from the instruction memory for internal caching, and retrieve whether there is a RISC-V instruction corresponding to the program count value among the prefetched number of RISC-V instructions; if there is, it outputs the retrieved RISC-V instruction to the instruction decoding unit; if not, it accesses the instruction memory and outputs the corresponding RISC-V instruction found in the instruction memory to the instruction decoding unit; The instruction decoding unit includes a decoder, which is used to parse the received RISC-V instruction, identify the register access information and data operation information carried by the RISC-V instruction, and mix-encode them into instruction identification information; The physical register file is used to store the logical register data used by the RISC-V processor during execution through a number of physical registers; The out-of-order execution unit is used to classify the decoded RISC-V instructions into multiple issue queues according to the instruction type, form an out-of-order issued instruction stream in each issue queue according to the current computing load, and execute the general computing task and the neural network computing task by out-of-order executing the RISC-V instructions in the instruction stream; The write-back unit is used to write the execution result back to the physical register file.
3. The RISC-V based neural network extended execution computer system according to claim 2, wherein The out-of-order execution unit includes: An instruction issue unit, which is used to determine whether the decoded RISC-V instruction is issue-ready, and distribute the issue-ready RISC-V instruction to one of the issue queues according to its instruction type; sort the RISC-V instructions in each issue queue by age respectively, and arbitrate the oldest RISC-V instruction in each issue queue for parallel issue; An arithmetic logic unit, which is used to receive and execute the arithmetic logic operation instructions issued by the instruction issue unit; the arithmetic logic operation instructions are fixed-point operation instructions; A floating-point arithmetic unit for receiving and executing the floating-point arithmetic instructions issued by the instruction issuing unit; A load memory access unit for receiving and executing the load instructions or memory access instructions issued by the instruction issuing unit to load or store external data; and, A vector processing unit for receiving and executing the custom extended vector calculation instructions issued by the instruction issuing unit to implement vector operations and matrix operations in the neural network calculation task.
4. The RISC-V based neural network extended execution computer system according to claim 2, wherein The instruction receiving unit includes: a dynamic branch prediction unit; The dynamic branch prediction unit includes: A branch target buffer unit for recording the program count value and the jump address corresponding to the branch instruction that has jumped; A jump history buffer unit for recording the jump direction and the number of jumps corresponding to each program count value in the historical process of instruction execution; the jump direction and the number of jumps corresponding to each program count value are updated using a saturating counter according to the jump result of the executed instruction; A global history prediction register unit for predicting the instruction jump result after the execution of a continuous number of instructions and generating a global jump history prediction value; A jump arbitration unit for predicting the second program count value of the currently to-be-read RISC-V instruction according to the first program count value of the previous RISC-V instruction, including: Performing a logical operation on the first program count value and the global jump history prediction value to obtain an index address, querying the jump history buffer unit according to the index address, and determining whether the second program count value jumps; If the second program count value jumps, determining whether the first program count value hits any branch instruction in the branch target buffer unit; if it hits, determining whether the hit branch instruction is a function return instruction; if it is the function return instruction, predicting the function call target address corresponding in the return address stack as the jump address of the second program count value; if it is not the function return instruction, predicting the jump address corresponding to the hit branch instruction as the jump address of the second program count value; If the second program count value does not jump, or jumps but the first program count value does not hit any branch instruction in the branch target buffer unit, predicting the jump address of the second program count value as PC + k; where PC is the first program count value and k is a user-defined constant.
5. The RISC-V based neural network extended execution computer system according to claim 4, characterized in that, The instruction receiving unit further includes an instruction prefetch unit and an instruction query unit; The instruction prefetch unit is used to access the instruction memory through a dedicated bus interface according to the first program count value, and read the RISC-V instruction corresponding to the first program count value and several adjacent RISC-V instructions from the instruction memory for caching; The instruction query unit is configured to query whether there is a RISC-V instruction corresponding to the second program count value output by the dynamic branch prediction unit in the instruction prefetch unit; if so, output the corresponding RISC-V instruction in the instruction prefetch unit to the instruction decoding unit; if not, find out the RISC-V instruction corresponding to the second program count value in the instruction memory and output it to the instruction decoding unit.
6. The RISC-V based neural network extended execution computer system according to claim 4, wherein The physical register free list is used to mark the free status of each physical register; The predicted renaming mapping table is used to record the dynamic mapping relationship between logical registers and physical registers; The physical register ready list is used to mark the data ready status of each physical register; the register allocation circuit is further configured to integrate the free destination physical register, source physical registers, and the data ready status of the source physical registers associated with the current RISC-V instruction in the register access control information.
7. The RISC-V based neural network extended execution computer system according to claim 6, wherein The comparison, selection, merging, and arbitration unit includes: a plurality of one-bit comparators, and a plurality of two-to-one multiplexers corresponding to the one-bit comparators one by one; The plurality of one-bit comparators form a binary tree structure for pairwise comparison of the free status of all physical registers in the physical register free list; if each physical register is free, its free status is marked as "1", and if it is not free, its free status is marked as "0"; Each one-bit comparator is a maximum value comparator, and when pairwise comparing the input physical register free status: if the free status of two input physical registers is different or the free status of both input physical registers is free, the comparison result of the current one-bit comparator is "1"; if the free status of both input physical registers is not free, the comparison result of the current one-bit comparator is "0"; The plurality of two-to-one multiplexers form a binary tree structure for screening out the first free physical register in the physical register file in the order of numbering from low to high; Each two-to-one multiplexer is configured to use the comparison result of the corresponding one-bit comparator as a selection signal to screen out the physical register with the corresponding number, including: if the comparison result of the current one-bit comparator is "1", determine whether the free status of the two physical registers connected by the current one-bit comparator is the same; if different, select the physical register number with the free status marked as "1" as the input of the next-level two-to-one multiplexer, and if the same, use the physical register number at the lower position as the input of the next-level two-to-one multiplexer; if the comparison result of the current one-bit comparator is "0", use the physical register number at the higher position as the input of the next-level two-to-one multiplexer; The comparison selection merging arbiter is used to select the corresponding numbered idle physical register as the destination physical register of the current RISC-V instruction when the comparison result of the one-bit comparator at the top of the tree is "1".
8. The RISC-V based neural network extended execution computer system according to claim 6, wherein The dynamic branch prediction unit further includes: a global history register for recording the actual jump history of several RISC-V instructions after execution. The register renaming unit is further provided with: a physical register free list checkpoint and an architectural renaming mapping table; the physical register free list checkpoint is used to assign a branch number to the current RISC-V instruction and back up the physical register free list when the current RISC-V instruction is a branch instruction; the architectural renaming mapping table is used to record the mapping relationship between each logical register and the physical register when the previous RISC-V instruction is executed. The write-back unit includes: branch prediction failure recovery logic. The branch prediction failure recovery logic is used to overwrite the data of the global history register to the global history prediction register unit, overwrite the data of the physical register free list checkpoint with the corresponding number to the physical register free list, and overwrite the data of the architectural renaming mapping table to the predicted renaming mapping table when the dynamic branch prediction unit makes a wrong prediction on the jump address of the second program counter value.
9. The RISC-V based neural network extended execution computer system according to any one of claims 1 to 8, characterized in that The neural network hardware accelerator includes: A model compression unit for performing knowledge distillation on the convolutional neural network in the neural network computing task and compressing the model parameters of the convolutional neural network. A computing unit including a convolution operation module, a matrix multiplication module, a quantization module, an activation function module, and a pooling module for performing forward propagation calculation operations in the convolutional neural network. A storage unit including independent input feature map cache, weight cache, bias cache, intermediate value cache, pooling cache, output feature map cache, and shared cache for reducing storage access conflicts and reducing data access latency through a multi-level cache system. A data flow control unit for controlling data flow to ensure efficient data transmission between the computing unit and the storage unit. A communication interface for data transmission with the RISC-V processor, loading input image data and the model parameters of the convolutional neural network, and outputting the recognition results of the convolutional neural network.
10. The RISC-V based neural network extended execution computer system according to claim 9, wherein The convolution operation module includes a convolution calculation array and a first adder tree for sliding a convolution kernel in the input feature map, sequentially performing element-wise multiplication operations in the sliding area using the convolution calculation array, and accumulating the intermediate calculation results of the element-wise multiplication operations based on the first adder tree to generate an output feature map. The matrix multiplication module includes a multiplication calculation array and a second adder tree, and is used to perform the output calculation of the fully connected layer of the convolutional neural network. Among them, the multiplication calculation array is used to perform multiplication operations on the input multi-channel feature maps and the corresponding point convolution kernels, and the second adder tree is used to accumulate the intermediate calculation results of the multiplication calculation array.
Citation Information
Patent Citations
Memory recovery structure and method suitable for neural network accelerator
CN115964169A
High-performance internet-of-things linkage processing system based on edge computing
CN117118720A
Rapid reasoning method, device and system for large language model of smart phone
CN118446321A