Processor based on RISC-V architecture, design method of processor, data processing method, chip and system
By introducing switching modules and computing modules into the processor of the RISC-V architecture, data movement and computing are realized in memory, and energy consumption problems caused by frequent data transfer in the CPU and GPU collaboration mode are solved, reducing latency and ensuring data consistency.
Patent Information
- Application Number
- CN202510740528.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
AI Technical Summary
In the existing CPU-GPU collaboration mode, frequent data transfer between CPU, GPU and memory leads to higher energy consumption and delay.
Using a processor based on the RISC-V architecture, a first switching module is set in the CPU, a second switching module and a computing module are set in the memory, and the switching module is used to realize the movement and computing of data within the memory, reducing the data transmission between the CPU and the memory.
It reduces the data transmission delay and energy consumption between the CPU and memory, reduces the frequent transfer of data between the CPU, GPU and memory, and ensures data consistency through reordering the cache unit.
Smart Images

Figure CN120256375A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of processors and their design, and particularly relates to a processor based on the RISC-V architecture, a design method of the processor, a chip including the processor, as well as a data processing method and a computer system based on the processor. Background Art
[0002] With the rapid development of artificial intelligence technology, deep learning neural networks have been widely applied in fields such as image recognition and speech recognition. However, there are certain problems in the existing cooperation mode between CPU and GPU. Specifically, in the cooperation mode between CPU and GPU, the CPU needs to transfer instructions to the GPU and copy data from the memory to the GPU for operation. After the GPU operation, the result is written back to the memory. The transmission between the CPU core and the memory needs to pass through the internal bus, resulting in additional latency and energy consumption. At the same time, the frequent data transfer between the CPU, GPU, and memory leads to high energy consumption.
[0003] Therefore, there is an urgent need for a processor that can solve the above problems.
[0004] The disclosure of the above background art content is only for assisting in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available before the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0005] In view of this, in order to overcome the defects of the prior art, one of the objectives of the present invention is to provide a processor based on the RISC-V architecture to reduce the energy consumption caused by the frequent data transfer between the CPU, GPU, and memory.
[0006] To achieve the above objective, the present invention adopts the following technical solutions: A processor based on the RISC-V architecture includes a CPU based on the RISC-V architecture and a memory. The CPU includes multiple CPU cores and a first switching module, and multiple CPU cores are electrically connected to the first switching module. The memory includes multiple operation modules and a second switching module, and multiple operation modules are electrically connected to the second switching module. The first switching module is electrically connected to the second switching module. The first switching module is used to receive the control data transmitted by the CPU core and transfer it to the second switching module, and at the same time record the information of the source CPU core that sends the control data. The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the operation module for operation. After the operation module completes the operation, the operation result data is transmitted to the source CPU core through the second switching module and the first switching module.
[0007] According to some preferred implementation aspects of the present invention, the CPU core includes an instruction decoding unit and a reorder buffer unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to both the reorder buffer unit and the first switching module simultaneously; the reorder buffer unit is used to store the control data.
[0008] According to some preferred implementation aspects of the present invention, the first switching module is used to receive the control data transmitted by the instruction decoding unit within the CPU core, and transmit it to the second switching module in the memory, while recording the information of the source CPU core from which the data is sent.
[0009] According to some preferred implementation aspects of the present invention, the first switching module is used to receive the operation result data passed from the second switching module, and send it to the source CPU core corresponding to the control data of the recorded operation result data.
[0010] According to some preferred implementation aspects of the present invention, the reorder buffer unit in the source CPU core receives the operation result data, and reorders the operation result data according to the sequence information in the control data.
[0011] According to some preferred implementation aspects of the present invention, the control data contains target operation module information. The second switching module is used to receive the control data transmitted by the first switching module in the CPU, and allocate the control data to a specified operation module for operation according to the target operation module information in the control data.
[0012] According to some preferred implementation aspects of the present invention, the memory includes a temporary storage module. The temporary storage module is used to store intermediate operation result data and final operation result data during the operation process, and transmit the intermediate operation result data to the operation module for further operation and / or transmit the final operation result data to the corresponding source CPU core through the second switching module and the first switching module.
[0013] According to some preferred implementation aspects of the present invention, the operation module is used to perform operations based on the control data transmitted by the second switching module, and after the operation is completed, transmit the operation result data to the reorder buffer unit within the source CPU core through the second switching module and the first switching module. The reorder buffer unit is used to store the operation result data, and reorder the operation result data according to the sequence information in the corresponding control data.
[0014] The second object of the present invention is to provide a data processing method based on the above-mentioned processor, including the following steps: The CPU core transfers control data to the first switching module; The first switching module is configured to receive the control data transmitted by the CPU core, transfer it to the second switching module, and record the information of the source CPU core that sends the control data; The second switching module is configured to receive the control data transmitted by the first switching module and allocate it to the arithmetic module for arithmetic operations; After the arithmetic operation is completed, the arithmetic result data is transmitted to the source CPU core through the second switching module and the first switching module; The source CPU core reorders the arithmetic result data according to the sequence information in the control data and outputs the arithmetic result.
[0015] A third object of the present invention provides a chip and a computer system including the processor based on the RISC-V instruction set architecture as described above. The computer system is based on the above-mentioned processor and executes the above-mentioned data processing method to perform data processing and calculation.
[0016] A fourth object of the present invention is to provide a design method for the above-mentioned processor based on the RISC-V architecture, including the following steps: Set a first switching module in the CPU. The first switching module is electrically connected to multiple CPU cores; the first switching module is configured to receive the control data transmitted by the CPU core and send it to the second switching module, and record the information of the source CPU core that sends the control data; Set a second switching module, multiple arithmetic modules, and a temporary storage module in the memory. Multiple arithmetic modules are electrically connected to the second switching module; the arithmetic module is configured to receive the control data transmitted by the second switching module and perform arithmetic operations. The intermediate arithmetic result data during the arithmetic operation is stored in the temporary storage module, and after the arithmetic operation is completed, the final arithmetic result data is transmitted to the source CPU core through the second switching module and the first switching module.
[0017] According to some preferred implementation aspects of the present invention, the control data contains target arithmetic module information, and the second switching module allocates the control data to a specified arithmetic module according to the target arithmetic module information in the control data.
[0018] According to some preferred implementation aspects of the present invention, the CPU core includes an instruction decoding unit and a reordering buffer unit. The instruction decoding unit is configured to parse and decode instructions to obtain control data, and send the control data to the reordering buffer unit and the first switching module simultaneously; the reordering buffer unit is configured to store the control data and the corresponding arithmetic result data, and reorder the arithmetic result data according to the sequence information in the control data.
[0019] Due to the above technical solution, compared with the prior art, the beneficial effects of the present invention are as follows: The processor based on the RISC-V architecture of the present invention, by setting a first switching module in the CPU, a second switching module and an arithmetic module in the memory, realizes data transmission through the connection between the first switching module and the second switching module. The arithmetic operation of data does not need to be transmitted to the CPU for arithmetic through the internal bus and then store the result data back to the memory after the operation. Instead, the data movement and operation are directly performed inside the memory, reducing the latency and energy consumption of data transmission between the CPU core and the memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic structural diagram of the CPU core in the preferred embodiment of the present invention; Figure 2 It is a schematic structural diagram of the first switching module in the preferred embodiment of the present invention; Figure 3 It is a schematic structural diagram of the memory in the preferred embodiment of the present invention; Figure 4 It is a schematic structural diagram of the processor in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] In the design of existing processors, in the cooperation mode of the CPU and the GPU, data needs to be frequently transferred among the CPU, the GPU, and the memory, resulting in high energy consumption. Moreover, the data needs to be first transferred from the memory to the CPU, and then stored back in the memory after being processed by the CPU. The transfer between the CPU core and the memory needs to pass through the internal bus, causing additional latency and energy consumption. Therefore, in order to solve the problem of high energy consumption caused by the frequent transfer of data between the CPU core and the memory in the prior art, the present invention provides a processor based on the RISC-V architecture, a design method of the processor, a chip including the processor, and a data processing method and a computer system based on the processor. In the design of the processor based on the RISC-V architecture of the present invention, by using the customizable instruction set and the hardware modular characteristics of the RISC-V instruction set architecture, the switching module in the CPU is combined with the switching module and the arithmetic module in the memory, and the arithmetic module in the memory is used to reduce the data transfer between the CPU and the memory required for the CPU to process data, and reduce the energy consumption caused by the frequent transfer of data among the CPU, the GPU, and the memory.
[0024] As Figures 1 to 4 shown, the processor based on the RISC-V architecture of the present invention includes a RISC-V CPU and a memory that cooperate with each other. Among them, the RISC-V CPU includes a plurality of CPU cores and a first switching module, and the plurality of CPU cores are electrically connected to the first switching module. The memory includes a plurality of arithmetic modules and a second switching module, and the plurality of arithmetic modules are electrically connected to the second switching module.
[0025] The first switching module is electrically connected to the second switching module; preferably, there are a plurality of channels between the first switching module and the second switching module to achieve parallel transmission. The first switching module is used to receive the control data transmitted by the CPU core and transfer it to the second switching module, and at the same time record the information of the source CPU core that sends the control data. The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the arithmetic module for operation. After the arithmetic module finishes the operation, the operation result data is transmitted to the source CPU core through the second switching module and the first switching module.
[0026] In the present invention, the CPU core includes an instruction decoding unit and a reorder buffer unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to both the reorder buffer unit and the first switching module simultaneously; the reorder buffer unit is used to store the control data. The first switching module is used to receive the control data transmitted by the instruction decoding unit within the CPU core and transmit it to the second switching module in the memory, while recording the information of the source CPU core from which the data is sent. The first switching module is used to receive the operation result data passed from the second switching module and send it to the source CPU core corresponding to the control data of the operation result data recorded. The reorder buffer unit in the source CPU core receives the operation result data and reorders the operation result data according to the sequence information in the control data, and then outputs the operation result.
[0027] Further, the CPU core includes an arithmetic unit. In the prior art, the arithmetic unit needs to fetch more than two operands from the memory to the arithmetic unit for operation. The process of moving the operands to the arithmetic unit consumes power and increases latency. However, in the present invention, through the switching module and the arithmetic module, the operands only need to move within the memory, and the control data with a smaller data volume is sent to the memory for operation, with lower power consumption and reduced latency.
[0028] In the present invention, the control data contains target arithmetic module information. The second switching module is used to receive the control data transmitted by the first switching module in the CPU and allocate the control data to the specified arithmetic module for operation according to the target arithmetic module information in the control data. Designating the arithmetic module by the CPU can avoid the delay in allocating the arithmetic module in the switching module, and the switching module can quickly send the control data to the target arithmetic module.
[0029] The arithmetic module is used to perform operations based on the control data passed from the second switching module, and after the operation is completed, transmit the operation result data to the reorder buffer unit in the source CPU core through the second switching module and the first switching module. By setting the arithmetic module in the memory, the moving path of the memory operands is shortened, the power consumption is reduced, and the latency is decreased.
[0030] Further, the memory includes a temporary storage module. The temporary storage module is used to store the intermediate operation result data and the final operation result data during the operation process, and transmit the intermediate operation result to the arithmetic module for re - operation and / or transmit the final operation result to the corresponding source CPU core through the second switching module and the first switching module. On the one hand, the data obtained after the operation is very likely to be used immediately again; on the other hand, the data obtained after the operation may not be used ultimately. Therefore, a temporary storage module is set up for temporarily storing data, and the data that needs to be used is transmitted to the source CPU core through the second switching module and the first switching module, avoiding the frequent movement of data.
[0031] A data processing method for a processor based on the above RISC-V architecture includes the following steps: The CPU core transmits control data to the first switching module; The first switching module is used to receive the control data transmitted by the CPU core and transfer it to the second switching module, and at the same time record the information of the source CPU core that sends the control data; The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the arithmetic module for arithmetic operations; After the arithmetic module completes the operation, the operation result data is transmitted to the source CPU core through the second switching module and the first switching module; The source CPU core reorders the operation result data according to the sequence information in the control data and outputs the operation result.
[0032] A chip and a computer system including the above processor based on the RISC-V instruction set architecture, the computer system is based on the above processor and executes the above data processing method to perform data processing and calculation.
[0033] A design method for a processor based on the RISC-V instruction set architecture includes the following steps: Set a first switching module in the CPU, and the first switching module is electrically connected to multiple CPU cores; the first switching module is used to receive the control data transmitted by the CPU core and send it to the second switching module, and at the same time record the information of the source CPU core that sends the control data; Set a second switching module, multiple arithmetic modules and a temporary storage module in the memory, and multiple said arithmetic modules are all electrically connected to the second switching module; the arithmetic module is used to receive the control data transmitted by the second switching module and perform arithmetic operations, and the intermediate operation result data and the final operation result data during the operation are stored in the temporary storage module, and after the operation is completed, the required final operation result data is transmitted to the source CPU core through the second switching module and the first switching module.
[0034] The control data contains target arithmetic module information, and the second switching module allocates the control data to the specified arithmetic module according to the target arithmetic module information in the control data.
[0035] The CPU core includes an instruction decoding unit and a reordering buffer unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to the reordering buffer unit and the first switching module at the same time; the reordering buffer unit is used to store the control data and the corresponding operation result data, and reorder the operation result data according to the sequence information in the control data, and then output the operation result. Embodiment 1
[0036] As Figures 1 to 4 shown, the processor based on the RISC-V architecture of this embodiment includes a RISC-V CPU and a memory that cooperate with each other. Among them, the CPU includes four CPU cores and a first switching module. The four CPU cores are CPU core 0, CPU core 1, CPU core 2, and CPU core 3 respectively, and the four CPU cores are all electrically connected to the first switching module. The memory includes four arithmetic modules, a second switching module, and a temporary storage module. The four arithmetic modules are all electrically connected to the second switching module. The temporary storage module is used to store the intermediate arithmetic result data and the final arithmetic result data during the arithmetic process of the arithmetic module, and transmit the intermediate arithmetic result data to the arithmetic module for re-arithmetic and / or transmit the required final arithmetic result data to the corresponding source CPU core through the second switching module and the first switching module. In other embodiments, the specific number of CPU cores or arithmetic modules can be increased or decreased according to actual situations.
[0037] The first switching module is electrically connected to the second switching module. The first switching module is located in the CPU, and the second switching module is located in the memory. They are electrically connected to each other for data transmission. Preferably, there are multiple channels between the first switching module and the second switching module to achieve parallel transmission.
[0038] The first switching module is used to receive the control data transmitted by the CPU core and transmit it to the second switching module, and at the same time record the information of the source CPU core that sends the control data. The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the arithmetic module for arithmetic. After the arithmetic module finishes arithmetic, it transmits the required arithmetic result data to the source CPU core through the second switching module and the first switching module.
[0039] In this embodiment, the CPU core includes an instruction decoding unit, an arithmetic unit, and a reorder buffer unit. The instruction stream enters the instruction decoding unit, and the instruction decoding unit parses and decodes the instruction to obtain control data. The instruction decoding unit simultaneously sends the control data to the reorder buffer unit and the first switching module. The reorder buffer unit is used to store the control data.
[0040] The first switching module is used to receive the control data transmitted by the instruction decoding unit in the CPU core and transmit it to the second switching module in the memory, and at the same time record the information of the source CPU core that sends the control data.
[0041] In this embodiment, the control data obtained after the instruction decoding unit parses and decodes the instruction contains target operation module information. The second switching module is used to receive the control data transmitted by the first switching module in the CPU and directly allocate the control data to the specified operation module for operation according to the target operation module information in the control data. The operation module is used to perform operations based on the control data transmitted by the second switching module, and after the operation is completed, transmit the required operation result data to the reorder buffer unit in the source CPU core through the second switching module and the first switching module.
[0042] The first switching module is used to receive the operation result data transmitted by the second switching module and send it to the source CPU core that records the control data corresponding to the operation result data. The reorder buffer unit in the source CPU core receives the operation result data and reorders the operation result data according to the order information in the control data.
[0043] The working process of the processor based on the RISC-V architecture in this embodiment is generally as follows: When the instruction stream enters the RISC-V CPU, after being parsed and decoded by the instruction decoding unit, the obtained control data is written into the reorder buffer unit. The reorder buffer unit is used to store the order information in the control data and the operation result data calculated subsequently based on the control data. The control data contains target operation module information to directly allocate the control data to the specified operation module.
[0044] Next, the first switching module of the CPU transmits the control data to the second switching module in the memory and allocates the control data to the specified operation module for calculation based on the target operation module. The first switching module of the CPU is responsible for managing the data exchange between the CPU and the memory and recording the source CPU core of the control data.
[0045] The second switching module in the memory, as a bridge between the memory and the CPU, allocates the control data to the specified operation module for operation according to the target operation module information in the control data.
[0046] The operation module in the memory can perform various arithmetic and logical operations, including addition, subtraction, multiplication, division, bit operations, etc. The intermediate operation result data during the operation process is stored in the temporary storage module for temporarily storing the operation result; and after the operation is completed, the required final operation result data is stored in the temporary storage module.
[0047] Meanwhile, the final operation result data is transmitted to the first switching module in the CPU through the second switching module in the memory. The first switching module transmits the operation result data to the corresponding source CPU core according to the source CPU core information of the control data corresponding to the operation result data. The reorder buffer unit in the source CPU core receives the operation result data and reorders the operation result data according to the order information in the control data.
[0048] Finally, the CPU performs sequential submission according to the order information in the reorder buffer unit to ensure the consistency and correctness of instruction execution. Embodiment 2
[0049] The data processing method of the processor based on the above RISC-V architecture in this embodiment includes the following steps: Step S1: The instruction stream enters the instruction decoding unit. The instruction decoding unit parses and decodes the instruction to obtain control data, and the instruction decoding unit simultaneously sends the control data to the reorder buffer unit and the first switching module. The reorder buffer unit is used to store the control data.
[0050] The control data generated by the instruction decoding unit contains target operation module information.
[0051] Step S2: The first switching module is used to receive the control data transmitted by the CPU core and transfer it to the second switching module, and simultaneously record the information of the source CPU core that sends the control data; the source CPU core corresponds to the target operation module information to facilitate the subsequent transmission of the operation result data back to the corresponding source CPU core.
[0052] Step S3: The second switching module is used to receive the control data transmitted by the first switching module in the CPU and allocate the control data to the specified operation module for operation according to the target operation module information in the control data.
[0053] Step S4: The operation module performs operations based on the control data transmitted by the second switching module, and after the operation is completed, transmits the required operation result data to the source CPU core through the second switching module and the first switching module.
[0054] The temporary storage module in the memory is used to store the intermediate operation result data and the final operation result data, transmit the intermediate operation result data to the operation module for re-operation, store the final operation result data after the calculation is completed, and simultaneously transmit the required final operation result data to the corresponding source CPU core through the second switching module and the first switching module.
[0055] Step S5: The operation result data is received and stored by the reorder buffer unit in the corresponding source CPU core. The reorder buffer unit reorders the operation result data according to the order information in the control data and outputs the operation result. Embodiment 3
[0056] The design method of a processor based on the RISC-V instruction set architecture in this embodiment includes the following steps: Step S1: Set a first switching module in the CPU. The first switching module is electrically connected to multiple CPU cores.
[0057] The first switching module is used to receive the control data transmitted by the CPU core and send it to the second switching module, and at the same time record the information of the source CPU core that sends the control data.
[0058] Step S2: Set a second switching module, multiple operation modules, and a temporary storage module in the memory. The second switching module is electrically connected to multiple operation modules.
[0059] The operation module is used to receive the control data transmitted by the second switching module and perform operations. The intermediate result data and the final result data during the operation process are stored in the temporary storage module, and the required final operation result data is transmitted to the source CPU core through the second switching module and the first switching module after the operation is completed.
[0060] The control data contains the target operation module information. The second switching module directly distributes the control data to the specified operation module according to the target operation module information in the control data.
[0061] The CPU core includes an instruction decoding unit and a reorder buffer unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to the reorder buffer unit and the first switching module at the same time; the reorder buffer unit is used to store the control data and the corresponding operation result data, reorder the operation result data according to the order information in the control data, and output the operation result.
[0062] The existing patent CN119357121A discloses a general parallel computing architecture coprocessor system based on RISC-V. This patent relates to the architectures of a main processor and a coprocessor. The main processor sends instructions to the coprocessor, and after receiving the instructions, the coprocessor decodes and performs subsequent operations. Data still needs to be transferred to the coprocessor for operation. Facing AI inferences, which involve a large number of scalar operations, transferring data from memory to the coprocessor will result in a large amount of latency and power consumption. However, through the exchange module in the present invention, the computing function of the CPU is extended to the memory, reducing the latency and power consumption caused by data transfer. At the same time, the present invention maintains hardware scalability. In the above patent, instructions are sent to the coprocessor, and the coprocessor decodes and performs operations. When one coprocessor is insufficient to meet the demand, the extension of the coprocessor will have a problem of consistency due to different operation times, and additional hardware units are required for processing, increasing costs and processing latency. The present invention uses the existing reorder buffer technology to ensure consistency without increasing latency.
[0063] In the prior art, in the cooperation mode between the CPU and the GPU, the CPU needs to transfer instructions to the GPU and copy data from memory to the GPU for operation. After the GPU finishes the operation, it writes the result back to memory. This process of frequent data transfer between the CPU, GPU, and memory leads to high energy consumption. To solve this problem, the processor based on the RISC-V architecture of the present invention includes a RISC-V CPU and memory. The RISC-V CPU includes a first exchange module and multiple CPU cores connected to the first exchange module. Each CPU core includes an instruction decoding unit and a reorder buffer unit. The CPU cores are used to execute computing instructions, the instruction decoding unit is used to parse and decode instructions, the reorder buffer unit is used to store control data and reorder operation result data, and the first exchange module is used to achieve high-speed data transmission. The memory includes a second exchange module, a temporary storage module, and multiple operation modules connected to the second exchange module. The operation modules are used to perform operations, and the temporary storage module is used to store operation results and intermediate data.
[0064] During the operation process, the instructions are parsed by the instruction decoding unit of the RISC-V CPU and stored in the reorder buffer unit. At the same time, the control data is transmitted to the memory through the first exchange module. The second exchange module of the memory transmits the control data to the specified operation module for operation. The data after the operation passes through the second exchange module and the first exchange module, returns to the corresponding processor core, updates the data in the reorder buffer unit, and finally completes the sequential submission of the data.
[0065] Compared with the prior art, the processor based on the RISC-V architecture of the present invention has the following beneficial effects: 1. Reduces the data transfer between the CPU and memory required for the CPU to process data, and reduces the energy consumption caused by the frequent transfer of data among the CPU, GPU, and memory; 2. By placing the second switching module and the operation module in the memory, the operation of data does not need to be transferred to the CPU for operation via the internal bus, and the result data is not stored back in the memory after the operation. Instead, the data is directly moved and operated inside the memory, reducing the latency and energy consumption of data transfer between the CPU core and the memory; 3. Utilizes the reorder buffer unit of the CPU to handle the consistency issue, ensuring data consistency and improving scalability.
[0066] The above embodiments obtained by the method of the present invention are only for illustrating the technical concept and features of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A processor based on the RISC-V architecture, characterized in that, It includes a CPU and a memory. The CPU includes multiple CPU cores and a first switching module. All the multiple CPU cores are electrically connected to the first switching module. The memory includes multiple arithmetic modules and a second switching module. All the multiple arithmetic modules are electrically connected to the second switching module. The first switching module is electrically connected to the second switching module. The first switching module is used to receive the control data transmitted by the CPU cores and transfer it to the second switching module, and at the same time record the information of the source CPU core that sends the control data. The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the arithmetic modules for arithmetic operations. After the arithmetic modules complete the operations, they transfer the arithmetic result data to the source CPU core through the second switching module and the first switching module.
2. The processor according to claim 1, wherein The CPU core includes an instruction decoding unit and a reorder buffer unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to both the reorder buffer unit and the first switching module at the same time. The reorder buffer unit is used to store the control data.
3. The processor according to claim 1, wherein The first switching module is used to receive the control data transmitted by the instruction decoding unit in the CPU core and transfer it to the second switching module in the memory, and at the same time record the information of the source CPU core that sends the data.
4. The processor according to claim 1, wherein, The first switching module is used to receive the arithmetic result data transferred by the second switching module and send it to the source CPU core corresponding to the control data of the recorded arithmetic result data.
5. The processor according to claim 2, wherein The reorder buffer unit in the source CPU core receives the arithmetic result data and reorders the arithmetic result data according to the sequence information in the control data.
6. The processor according to claim 1, wherein The control data contains target arithmetic module information. The second switching module is used to receive the control data transmitted by the first switching module in the CPU and allocate the control data to the specified arithmetic module for arithmetic operations according to the target arithmetic module information in the control data.
7. The processor according to claim 1, wherein The memory includes a temporary storage module. The temporary storage module is used to store the intermediate arithmetic result data and the final arithmetic result data during the arithmetic process.
8. The processor according to claim 2, wherein The arithmetic module is used to perform arithmetic operations based on the control data transferred by the second switching module, and transfer the arithmetic result data to the reorder buffer unit in the source CPU core through the second switching module and the first switching module after the arithmetic operations are completed.
9. A chip including the RISC-V architecture-based processor according to any one of claims 1-8.
10. A data processing method for a processor according to any one of claims 1-8, characterized in that, It includes the following steps: The CPU core transfers the control data to the first switching module; The first switching module is used to receive the control data transmitted by the CPU core and transfer it to the second switching module, and at the same time record the information of the source CPU core that sends the control data; The second switching module is used to receive the control data transmitted by the first switching module and allocate it to the arithmetic modules for arithmetic operations; After the arithmetic modules complete the operations, they transfer the arithmetic result data to the source CPU core through the second switching module and the first switching module; The source CPU core reorders the operation result data according to the sequence information in the control data and outputs the operation result.
11. A computer system for performing the data processing method as claimed in claim 10.
12. A design method for a processor based on the RISC-V architecture, characterized in that, It includes the following steps: Set a first switching module in the CPU, and the first switching module is electrically connected to multiple CPU cores; the first switching module is used to receive the control data transmitted by the CPU cores and send it to the second switching module, and at the same time record the information of the source CPU core that sends the control data; Set a second switching module, multiple operation modules and a temporary storage module in the memory, and multiple said operation modules are electrically connected to the second switching module; the operation modules are used to receive the control data transmitted by the second switching module and perform operations, the intermediate operation result data during the operation process is stored in the temporary storage module, and after the operation is completed, the final operation result data is transmitted to the source CPU core through the second switching module and the first switching module.
13. The design method according to claim 12, characterized in that, The control data contains target operation module information, and the second switching module distributes the control data to the specified operation module according to the target operation module information in the control data.
14. The design method according to claim 12, wherein The CPU core includes an instruction decoding unit and a reordering cache unit. The instruction decoding unit is used to parse and decode instructions to obtain control data, and send the control data to the reordering cache unit and the first switching module at the same time; the reordering cache unit is used to store the control data and the corresponding operation result data, and reorder the operation result data according to the sequence information in the control data.
Citation Information
Patent Citations
In-memory computing AI accelerator design architecture based on RISC-V architecture and control method
CN119201836A
Processor based on RISC-V instruction set architecture, design method of processor, data processing method, chip and system
CN120029970A