Processor based on RISC-V instruction set architecture, design method of processor, data processing method, chip and system
By removing complex computing modules from the CPU core and implementing multiple CPU cores to share complex computing modules through switching modules, the problem of resource waste in existing CPUs when handling complex computing is solved, and chip utilization and L1 cache hit rate are improved.
Patent Information
- Application Number
- CN202510512637.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
AI Technical Summary
When existing CPUs deal with complex operations, they occupy a large amount of circuit and bandwidth resources, increasing the area of the CPU core, and complex computing units are often idle, resulting in waste of resources.
A processor based on the RISC-V instruction set architecture is designed. By removing complex computing modules from the CPU core and separating them from the CPU core, the switching module is used to realize the sharing of complex computing modules for multiple CPU cores, and the number of complex computing modules is flexibly increased and decreased according to the application scenario.
This reduces the area of the CPU core, improves chip utilization, reduces the circuit and bandwidth resources occupied by complex computing units in the CPU core, increases the L1 cache hit rate, and reduces the data transfer delay caused by insufficient cache when performing complex computing.
Smart Images

Figure CN120029970A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of processors and their design, and specifically relates to a modular processor based on the RISC-V instruction set architecture, a design method for the processor, a chip including the processor, and a data processing method and a computer system based on the processor. Background Art
[0002] In recent years, with the rapid development of technologies such as artificial intelligence and deep learning, the demand for high-performance computing has been growing. In the design of traditional CPUs (central processing units), in addition to basic addition and multiplication units, the CPU core also integrates arithmetic units for complex operations such as vector addition and matrix multiplication. Although these arithmetic units can reduce the time required for complex operations, they require a large amount of circuit and bandwidth resources, increase the area of the CPU core, and occupy most of the L1 cache during operations, resulting in increased latency in subsequent operations due to cache misses.
[0003] Generally speaking, the CPU spends most of its time processing simple operations, but the design of these complex operations increases the area of the CPU core and is often idle, resulting in a waste of resources. On the contrary, if the complex operation design is removed and simple operations are used for processing, it will cause a lot of delays and reduce computing efficiency.
[0004] That is, the CPU in the prior art has the following problems: 1. The complex computing unit occupies a large amount of circuit and bandwidth resources, increasing the area of the CPU core, but the usage frequency is low, resulting in low overall chip utilization; 2. The complex computing unit occupies most of the L1 cache during calculation, resulting in increased delays in subsequent calculations due to cache misses; 3. If the complex calculation design is removed and simple calculations are used for processing, a large amount of delay will occur, reducing computing efficiency; 4. The existing storage and computing integrated chips have low computing efficiency and high power consumption, and their structure and functions still need to be further optimized.
[0005] Existing patents such as CN118916309A data processing device, data processing system and chip, this patent is not for the improvement of a CPU. Its data processing system includes a RISC-V-based CPU, an NPU connected to the CPU, and the NPU includes a vector calculation unit and a cache. In this patent, the CPU and NPU are in a collaborative relationship. After the bit stream reaches the CPU, it will be handed over to the NPU for subsequent calculations. At the same time, if the number of NPUs is increased, multiple NPUs will have consistency problems such as read-after-write and write-after-write when processing data with dependencies.
[0006] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed before the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the invention
[0007] In view of this, in order to overcome the defects of the prior art, one of the objects of the present invention is to provide a processor based on the RISC-V instruction set architecture, which can effectively improve the utilization of the CPU and reduce data movement delay.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions: A processor based on a RISC-V instruction set architecture comprises a CPU core based on the RISC-V instruction set architecture, a complex operation module and a switching module; the switching module is connected between the CPU core and the complex operation module, the CPU core is used to send data to the switching module, the switching module is used to receive the data and transmit it to the complex operation module, the complex operation module performs operations based on the data and returns the results to the switching module after performing the operations, and the switching module receives the results and transmits them to the CPU core.
[0009] According to some preferred implementation aspects of the present invention, the complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of executing the operation; outside the operation state, the complex operation module is in the idle state.
[0010] According to some preferred implementation aspects of the present invention, the processor includes multiple complex computing modules, and the multiple complex computing modules are all connected to the switching module; after the switching module receives the data, it transmits the data to the complex computing module in an idle state according to the states of the multiple complex computing modules.
[0011] According to some preferred implementation aspects of the present invention, the processor includes multiple CPU cores, and the multiple CPU cores are all connected to the switching module; after the switching module receives the result fed back by the complex operation module, it transmits the result to the CPU core that previously sent the corresponding data.
[0012] According to some preferred implementation aspects of the present invention, the complex operation module includes a cache unit, and the cache unit is used to store intermediate data in the complex operation process. The complex operation unit and another cache unit form an independent complex operation module, which can provide additional cache space for the intermediate data in the complex operation process and reduce the data movement delay caused by insufficient cache when executing complex operations.
[0013] Furthermore, the cache unit is a 64KB cache unit and / or a 128KB cache unit.
[0014] According to some preferred implementation aspects of the present invention, the complex operation module includes a complex operation unit, and the complex operation unit is one or more of a vector operation unit and a matrix operation unit.
[0015] According to some preferred implementation aspects of the present invention, the CPU core includes a simple operation module, and the simple operation module is used to perform basic operations of addition, subtraction, multiplication and division.
[0016] Furthermore, the CPU core includes a decoder, and the decoder is connected to the simple operation module. Meanwhile, the decoder and the simple operation module are both connected to the switching module.
[0017] According to some preferred implementation aspects of the present invention, the switching module includes a first-in-first-out data buffer, namely, a FIFO.
[0018] A second object of the present invention is to provide a data processing method based on the above-mentioned processor, comprising the following steps: The CPU core sends data to the switching module, and the switching module receives the data; The switching module distributes the data to the complex operation modules in an idle state in sequence according to the order of the received data; The complex operation module performs the operation after receiving the data, and sends the result to the switching module after the operation is completed; The exchange module sends the results to the corresponding CPU core according to the order of the received results.
[0019] In some embodiments, a processor based on the RISC-V instruction set architecture includes multiple CPU cores and multiple complex computing modules connected to a switching module, and the switching module has a first-in-first-out data buffer (FIFO). One or more CPU cores send data to the switching module, and the switching module distributes the data to the complex computing modules in an idle state according to the status of the multiple complex computing modules and the order of the received data.
[0020] The complex operation module performs the operation after receiving the data, and sends the result to the switching module after the operation is completed. The switching module sends the result to the corresponding CPU core according to the order of the received results. The corresponding here means that the calculation result is obtained after the data previously sent by the CPU core is calculated by the complex operation module, that is, the data and the result are corresponding, the data is sent by the corresponding CPU core, and the structure is received by the corresponding CPU core.
[0021] A third object of the present invention is to provide a chip including a processor based on the RISC-V instruction set architecture as described above and a computer system, wherein the computer system is based on the above processor and executes the above data processing method to perform data processing and calculation.
[0022] A fourth object of the present invention is to provide a design method for the above-mentioned processor based on the RISC-V instruction set architecture, comprising the following steps: A simple calculation module is set inside the CPU core, a complex calculation module is set outside the CPU core, and a switching module is set between the CPU core and the complex calculation module. The CPU core is used to send data to the switching module, and the switching module is used to receive data and transmit it to the complex calculation module. The complex calculation module performs calculations based on the data and returns the results to the switching module after performing the calculations. The switching module receives the results and transmits them to the CPU core.
[0023] According to some preferred implementation aspects of the present invention, the complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of executing the operation; outside the operation state, the complex operation module is in the idle state.
[0024] According to some preferred implementation aspects of the present invention, the processor includes multiple complex computing modules, and the multiple complex computing modules are all connected to the switching module; after the switching module receives the data, it transmits the data to the complex computing module in an idle state according to the states of the multiple complex computing modules.
[0025] According to some preferred implementation aspects of the present invention, the processor includes multiple CPU cores, and the multiple CPU cores are all connected to the switching module; after the switching module receives the result fed back by the complex operation module, it transmits the result to the CPU core that previously sent the corresponding data.
[0026] Due to the adoption of the above technical scheme, compared with the prior art, the benefits of the present invention are as follows: the processor based on the RISC-V instruction set architecture of the present invention is an improvement on the inside of an overall CPU chip, which removes the complex operation module from the CPU core and separates it from the CPU core, thereby reducing the area of the CPU core and reducing the circuit and bandwidth resources occupied by the complex operation unit in the CPU core; the complex operation module is connected to the CPU core through a switching module, so that multiple CPU cores with simple operation modules can share one or more complex operation modules; not only can the number of complex operation modules be flexibly increased or decreased according to the application scenario to improve chip utilization, but also the L1 cache hit rate can be increased, thereby reducing data movement delays caused by insufficient cache when executing complex operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 is a schematic diagram of the structure of the processor in Embodiment 1 of the present invention; Figure 2 is a schematic diagram of the structure of a processor in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of the processor in Example 3 of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0030] The complex operation unit in the prior art occupies a large amount of circuit and bandwidth resources, increases the CPU core area, and leads to low chip utilization; the complex operation unit occupies most of the L1 cache during operation, resulting in increased delays in subsequent operations due to cache misses; the complex operation unit is often in an idle state, resulting in waste of resources and other problems. In order to solve the above problems, the present invention provides a modular processor based on the RISC-V instruction set architecture.
[0031] The RISC-V instruction set architecture has the characteristics of modular design, and the hardware functions can be modularized. The present invention divides the arithmetic unit into a simple operation module and a complex operation module, sets the simple operation module in the CPU core, removes the complex operation module from the CPU core, and forms an independent module with another cache, and then connects it to the CPU core through a switching module, so that multiple CPU cores with simple operation modules can share one or more complex operation modules. Not only can the number of complex operation modules be flexibly increased or decreased according to the application scenario to improve chip utilization, but also the L1 cache hit rate can be increased, and the data movement delay caused by insufficient cache when performing complex operations can be reduced.
[0032] Specifically, the processor based on the RISC-V instruction set architecture of the present invention includes a CPU core, a complex computing module and a switching module; the switching module is connected between the CPU core and the complex computing module, such as Figure 1 The CPU core is used to send data to the switching module, the switching module is used to receive the data and transmit it to the complex computing module, the complex computing module performs operations based on the data, and returns the results to the switching module after performing the operations, and the switching module receives the results and transmits them to the CPU core.
[0033] The complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of performing operation; outside the operation state, the complex operation module is in the idle state. That is, when the complex operation module is not performing operation, it is in the idle state.
[0034] The complex operation module includes a complex operation unit and a cache unit. The complex operation unit is used to perform complex operations such as vector operations and matrix operations. It can be one or more of the vector operation unit and the matrix operation unit. The cache unit is used to store the intermediate data of the complex operation. It can be a 64KB cache unit and / or 128KB. That is, the complex operation unit and the cache unit constitute an independent complex operation module, which can provide additional cache space for the intermediate data of the complex operation. The complex operation module will not occupy the L1 cache in the CPU core during operation, which can increase the hit rate of the L1 cache and reduce the data movement delay caused by insufficient cache when performing complex operations.
[0035] Each CPU core includes a simple operation module and a decoder, and the simple operation module is used to perform simple basic operations such as addition, subtraction, multiplication, and division. The decoder is connected to the simple operation module, and at the same time, the decoder and the simple operation module are both connected to the switching module.
[0036] The switching module in the present invention includes a first-in-first-out data buffer, i.e., FIFO. It is responsible for forwarding data from the CPU core to the idle complex computing module, and arbitrates the use of the complex computing module by multiple CPU cores using a first-come-first-served algorithm based on the FIFO, so as to prevent the complex computing module from being idle for a long time. The clock signal, program status word and other information required for the operation will be transmitted to the complex computing unit through the switching module to control the complex computing unit to perform the operation, and the switching module can be used to perform flexible configurations of one (CPU core) to many (complex computing modules), many (CPU cores) to one (complex computing modules), and many (CPU cores) to many (complex computing modules).
[0037] Preferably, if Figure 2 and Figure 3As shown, a processor based on the RISC-V instruction set architecture includes multiple CPU cores and multiple complex computing modules connected to a switching module, and the switching module has a first-in-first-out data buffer (FIFO). One or more CPU cores send data to the switching module, and the switching module distributes the data to the complex computing modules in an idle state according to the status of the multiple complex computing modules and the order of the received data.
[0038] The complex operation module performs the operation after receiving the data, and sends the result to the switching module after the operation is completed. The switching module sends the result to the corresponding CPU core according to the order of the received results. The corresponding here means that the calculation result is obtained after the data previously sent by the CPU core is calculated by the complex operation module, that is, the data and the result are corresponding, the data is sent by the corresponding CPU core, and the structure is received by the corresponding CPU core.
[0039] The data processing method based on the above processor includes the following steps: The CPU core sends data to the switch module, and the switch module receives the data; The switching module distributes the data to the complex computing modules in the idle state in sequence according to the order of the received data; The complex operation module performs the operation after receiving the data and sends the result to the switching module after the operation is completed; The switching module sends the results to the corresponding CPU core according to the order in which the results are received.
[0040] A chip including a processor based on the RISC-V instruction set architecture as described above and a computer system. The computer system is based on the above processor and executes the above data processing method to process and calculate data, which can avoid complex calculation modules being idle for a long time, improve resource utilization efficiency, and reduce data movement delays caused by insufficient cache when executing complex calculations.
[0041] A design method for a processor based on the RISC-V instruction set architecture includes the following steps: A simple computing module is set in each CPU core, a complex computing module is set outside the CPU core, and a switching module is set between the CPU core and the complex computing module. The CPU core is used to send data to the switching module, and the switching module is used to receive data and transmit it to the complex computing module. The complex computing module performs operations based on the data and returns the results to the switching module after performing the operations. The switching module receives the results and transmits them to the CPU core.
[0042] The processor based on the RISC-V instruction set architecture of the present invention removes the complex operation module from the CPU core and connects it to the CPU core through a switching module, so that multiple CPU cores with simple operation modules can share one or more complex operation modules. It can not only flexibly increase or decrease the number of complex operation modules according to application scenarios to improve chip utilization, but also increase the L1 cache hit rate and reduce data movement delays caused by insufficient cache when executing complex operations. Example 1
[0043] like Figure 1 As shown, the processor based on the RISC-V instruction set architecture of this embodiment includes a CPU core, a complex computing module and a switching module; the switching module is connected between the CPU core and the complex computing module, such as Figure 1 The CPU core is used to send data to the switching module, the switching module is used to receive the data and transmit it to the complex computing module, the complex computing module performs operations based on the data, and returns the results to the switching module after performing the operations, and the switching module receives the results and transmits them to the CPU core.
[0044] The complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of performing operation; outside the operation state, the complex operation module is in the idle state. That is, when the complex operation module is not performing operation, it is in the idle state.
[0045] The complex operation module includes a complex operation unit and a cache unit. The complex operation unit is used to perform complex operations such as vector operations and matrix operations. It can be one or more of the vector operation unit and the matrix operation unit. The cache unit is used to store the intermediate data of the complex operation. It can be a 64KB cache unit and / or 128KB. That is, the complex operation unit and the cache unit constitute an independent complex operation module, which can provide additional cache space for the intermediate data of the complex operation. The complex operation module will not occupy the L1 cache in the CPU core during operation, which can increase the hit rate of the L1 cache and reduce the data movement delay caused by insufficient cache when performing complex operations.
[0046] Each CPU core includes a simple operation module and a decoder, and the simple operation module is used to perform simple basic operations such as addition, subtraction, multiplication, and division. The decoder is connected to the simple operation module, and at the same time, the decoder and the simple operation module are both connected to the switching module.
[0047] The switching module of this embodiment includes a first-in-first-out data buffer, namely FIFO, which is responsible for forwarding data from the CPU core to the idle complex computing module, and adopts a first-come-first-served algorithm based on FIFO to arbitrate the use of the complex computing module by multiple CPU cores, thereby preventing the complex computing module from being idle for a long time. Example 2
[0048] like Figure 2 As shown, the difference between this embodiment and embodiment 1 is that the processor based on the RISC-V instruction set architecture in this embodiment includes two parallel CPU cores connected to the switching module, and the other components and connection relationships are basically the same as those in embodiment 1. Example 3
[0049] like Figure 3 As shown, the difference between this embodiment and embodiment 1 is that the processor based on the RISC-V instruction set architecture of this embodiment includes four parallel CPU cores connected to the switching module and two parallel complex operation modules connected to the switching module. One of the two complex operation modules includes a vector operation unit and a 64KB cache unit for performing vector operations such as vector addition and vector multiplication; the other includes a matrix operation unit and a 128KB cache unit for performing matrix operations such as matrix multiplication and matrix inversion. Other components and connection relationships are basically the same as those in embodiment 1. Example 4
[0050] This embodiment provides a data processing method based on the processor of Embodiment 3, comprising the following steps: Step S1, one or more CPU cores send data to a switching module, and the switching module receives the data in turn.
[0051] Step S2: The switching module distributes the data to the complex operation modules in the idle state in sequence according to the order of the received data.
[0052] Step S3: After receiving the data, the complex operation module performs the operation and sends the result to the switching module after the operation is completed.
[0053] Step S4: The exchange module sends the result to the corresponding CPU core according to the order of the received results to complete the data processing. The corresponding here means that the result is obtained after the data previously sent by the CPU core is calculated by the complex calculation module, that is, the data and the result are corresponding, the data is sent by the corresponding CPU core, and the structure is received by the corresponding CPU core.
[0054] A computer system including a processor based on the RISC-V instruction set architecture as described above, which processes and calculates data based on the above processor and executes the above data processing method, can avoid complex calculation modules being idle for a long time, improve resource utilization efficiency, and reduce data movement delays caused by insufficient cache when performing complex calculations. Example 5
[0055] This embodiment provides a design method for a processor based on the RISC-V instruction set architecture of Embodiment 3, comprising the following steps: A simple operation module is set in each CPU core to perform simple basic operations such as addition, subtraction, multiplication, and division. The complex operation module is separated from the CPU core and set outside the CPU core, and a switching module is set between multiple CPU cores and multiple complex operation modules.
[0056] The CPU core is used to send data to the switching module, the switching module is used to receive data and transmit it to the complex computing module, the complex computing module performs operations based on the data and returns the results to the switching module after performing the operations, and the switching module receives the results and transmits them to the CPU core.
[0057] That is, the design method of the processor of this embodiment separates the complex operation module from the CPU core, and connects it to the CPU core through the switching module. The switching module is responsible for forwarding data from the CPU core to the idle complex operation module, and uses the first-come-first-served method to allocate the use of the complex operation module by multiple CPU cores. When the CPU core needs to perform a complex operation, the data is sent to the switching module. The switching module forwards the data to the idle complex operation module based on the idle status of the complex operation module. After the complex operation module performs the operation, the result is returned to the switching module, and then sent back to the corresponding CPU core by the switching module.
[0058] The processor based on the RISC-V instruction set architecture of the present invention utilizes the modular characteristics of the RISC-V instruction set architecture to control and configure the complex operation module through the instruction set, thereby realizing the flexibility and scalability of the algorithm. According to the application scenario, the number of complex operation modules can be flexibly increased or decreased to improve the utilization rate of the chip. The processor includes multiple CPU cores and multiple complex operation modules connected to a switching module, and the switching module has a first-in-first-out data buffer (FIFO). One or more CPU cores send data to the switching module, and the switching module sequentially distributes the data to the complex operation modules in an idle state according to the status of the multiple complex operation modules and the order of the received data.
[0059] The complex operation module performs the operation after receiving the data, and sends the result to the switching module after the operation is completed. The switching module sends the result to the corresponding CPU core according to the order of the received results.
[0060] The modular processor based on the RISC-V instruction set architecture provided by the present invention is an improvement within an overall CPU chip, and the bit stream is still written back to the memory after being processed by the CPU. Compared with the existing patents, the advantage of the present invention is the scalability of scale. When facing a large amount of data with dependencies in AI scenarios, the CPU supports handing over the operations to each operation unit. Even if the operation time of each unit is different, it can still be submitted in sequence to avoid consistency problems.
[0061] The present invention provides a modular processor based on the RISC-V instruction set architecture, which divides the arithmetic unit in the CPU core into a simple operation module and a complex operation module. The complex operation module includes the complex operation unit and caches other than L1, L2, and L3, and removes the complex operation module from the CPU core, and then connects it to the CPU core through a switching module, forming multiple CPU cores with simple operation modules. One or more complex operation modules can be used through the switching module. The switching module is responsible for forwarding data from the CPU core to the idle complex operation module and arbitrating the use of the complex operation module by multiple CPU cores. Compared with the prior art, the modular processor based on the RISC-V instruction set architecture of the present invention has the following beneficial effects: 1. The arithmetic unit is divided into a simple operation module and a complex operation module. The simple operation module is set in the CPU core to perform simple operations such as basic addition and multiplication, and the complex operation module is used to perform complex operations such as vector addition and matrix multiplication. By separating the complex operation module from the CPU core and removing it from the CPU core, and forming an independent module with another cache, the number of complex operation modules can be flexibly increased or decreased according to the application scenario, thereby improving the utilization rate of the chip; at the same time, the circuit and bandwidth resources occupied by the complex operation unit in the CPU core can be reduced, thereby reducing the area of the CPU core; 2. The complex computing module includes a complex computing unit and caches other than L1, L2, and L3. The complex computing unit is used to perform complex computing, and the cache is used to store intermediate data of complex computing. The complex computing module does not occupy the L1 cache in the CPU core during computing, which can increase the hit rate of the L1 cache and reduce the delay of subsequent computing due to cache misses; 3. Multiple CPU cores can share one or more complex computing modules through the switching module. The switching module is responsible for forwarding data from the CPU core to the idle complex computing module, and arbitrates the use of the complex computing module by multiple CPU cores based on the FIFO first-come-first-served algorithm, which can prevent the complex computing module from being idle for a long time and improve resource utilization efficiency; 4. Utilize the modular characteristics of the RISC-V instruction set architecture to control and configure complex computing modules through the instruction set to achieve flexibility and scalability of the algorithm; according to the application scenario, the number of complex computing modules can be flexibly increased or decreased to improve chip utilization.
[0062] The above embodiments prepared by the method of the present invention are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be included in the protection scope of the present invention.
Claims
1. A processor based on the RISC-V instruction set architecture, characterized in that: It includes a CPU core based on the RISC-V instruction set architecture, a complex computing module and a switching module; the switching module is connected between the CPU core and the complex computing module, the CPU core is used to send data to the switching module, the switching module is used to receive the data and transmit it to the complex computing module, the complex computing module performs operations based on the data and returns the results to the switching module after performing the operations, and the switching module receives the results and transmits them to the CPU core.
2. The processor according to claim 1, characterized in that The complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of executing the operation; outside the operation state, the complex operation module is in the idle state.
3. The processor according to claim 2, characterized in that The processor includes a plurality of complex computing modules, and the plurality of complex computing modules are all connected to the switching module; after receiving the data, the switching module transmits the data to the complex computing module in an idle state according to the states of the plurality of complex computing modules.
4. The processor according to claim 1, characterized in that The processor includes a plurality of the CPU cores, and the plurality of CPU cores are connected to the switching module; after receiving the result fed back by the complex operation module, the switching module transmits the result to the CPU core that previously sent the corresponding data.
5. The processor according to claim 1, wherein: The complex operation module includes a cache unit, and the cache unit is used to store intermediate data in the complex operation process.
6. The processor according to claim 1, wherein: The complex operation module includes a complex operation unit, and the complex operation unit is one or more of a vector operation unit and a matrix operation unit.
7. The processor according to claim 1, characterized in that The CPU core includes a simple operation module, which is used to perform basic operations of addition, subtraction, multiplication and division.
8. The processor according to any one of claims 1 to 7, characterized in that: The switching module includes a first-in-first-out data buffer.
9. A chip comprising a processor based on the RISC-V instruction set architecture as described in any one of claims 1 to 8.
10. A data processing method based on the processor according to any one of claims 1 to 8, characterized in that: The steps include: The CPU core sends data to the switching module, and the switching module receives the data; The switching module distributes the data to the complex operation modules in an idle state in sequence according to the order of the received data; The complex operation module performs the operation after receiving the data, and sends the result to the switching module after the operation is completed; The exchange module sends the results to the corresponding CPU core according to the order of the received results.
11. A computer system for executing the data processing method according to claim 10.
12. A design method for a processor based on a RISC-V instruction set architecture, characterized in that: The steps include: A simple calculation module is set inside the CPU core, a complex calculation module is set outside the CPU core, and a switching module is set between the CPU core and the complex calculation module. The CPU core is used to send data to the switching module, and the switching module is used to receive the data and transmit it to the complex calculation module. The complex calculation module performs calculations based on the data and returns the results to the switching module after performing the calculations. The switching module receives the results and transmits them to the CPU core.
13. The design method according to claim 12, characterized in that: The complex operation module has an operation state and an idle state. In the operation state, the complex operation module is in the process of executing the operation; outside the operation state, the complex operation module is in the idle state.
14. The design method according to claim 13, characterized in that: The processor includes a plurality of complex computing modules, and the plurality of complex computing modules are all connected to the switching module; after receiving the data, the switching module transmits the data to the complex computing module in an idle state according to the states of the plurality of complex computing modules.
15. The design method according to claim 12, characterized in that: The processor includes a plurality of CPU cores, and the plurality of CPU cores are connected to the switching module; after receiving the result fed back by the complex operation module, the switching module transmits the result to the CPU core that previously sent the corresponding data.
Citation Information
Patent Citations
Data processing device, data processing system, and chip
CN118916309A
Application scene data processing method and system, electronic equipment and storage medium
CN117973465A
AI accelerator architecture and AI chip based on AI accelerator architecture
CN118761448A
In-memory computing AI accelerator design architecture based on RISC-V architecture and control method
CN119201836A
RISC-V-based general parallel computing architecture co-processing system
CN119357121A
Cited By
Processor based on RISC-V architecture, design method of processor, data processing method, chip and system
CN120256375A