3D Interconnected Many-Core Processor Architecture Based on RISC-V and Its Working Method
By designing a three-dimensional interconnected multi-core processor architecture based on RISC-V, the disadvantages of existing RISC-V processors in high-performance processor requirements are solved, and multi-main core collaborative work and efficient accelerator data interaction are realized, which significantly improves processing efficiency.
Patent Information
- Application Number
- CN202111435059.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing RISC-V processors have major disadvantages in the demand for high-performance processors, especially in the areas of multi-core collaborative work and accelerator data interaction.
A three-dimensional interconnected multi-core processor architecture based on RISC-V is designed, including the main control layer, the micro-core array layer and the accelerator layer. The main control layer strengthens the connection between the main cores through the synergy of multiple main cores, and efficiently interacts with the micronuclear array and accelerator layer. The micro-core array layer adopts a 3DRouter link controller to realize the compression and fast data channels of the three-dimensional structure, improving processing efficiency.
Through multi-main core synergy and efficient data interaction, the processor's processing efficiency is significantly improved, providing a new design idea to support the application of RISC-V instruction set on high-performance processors.
Smart Images

Figure CN114356836B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated circuit processor hierarchical structure design, and particularly relates to a three-dimensional interconnected many-core processor architecture based on RISC-V and its working method. Background Technique
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Over the years, with the improvement of chip design technology and the wide range of applications, RISC-V has shown more and more advantages such as being completely open source and having a simple architecture that the traditional ARM and x86 architectures do not have. The inventor found that nowadays, RISC-V is widely used in design industries such as AIoT that have no excessive performance requirements for CPUs, but in some traditional high-performance processor requirements, the existing RISC-V still has significant disadvantages. Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a three-dimensional interconnected many-core processor architecture based on RISC-V and its working method. The solution strengthens the connection between each main core by designing a main core control layer with multi-main core cooperation, can efficiently control the micro-core array through multiple main cores simultaneously, and can also perform efficient data interaction with the accelerator layer, improving the processing efficiency of the processor and providing a new design idea for the application of the RISC-V instruction set in high-performance processors.
[0005] According to the first aspect of the embodiments of the present invention, a three-dimensional interconnected many-core processor architecture based on RISC-V is provided, including a main control layer, a micro-core array layer, and an accelerator layer. Among them,
[0006] The main control layer includes several main cores. The main cores are CPU cores based on the five-stage pipeline RISC-V instruction set, and the main cores interact with each other and with the external environment through independent buses;
[0007] The micro-core array layer includes several micro-unit groups. The micro-units include micro-cores, data storage units, instruction storage units, and link controllers. Among them, the micro-cores are CPU cores based on the RISC-V instruction set that execute some functions of the main cores;
[0008] The accelerator layer is used to optimize the space utilization and running speed of accelerators to meet specific requirements;
[0009] Among them, some of the main cores in the main control layer interact with the accelerator layer, and the other main cores interact with the micro-core array layer. For simple external instructions, they are directly calculated in the main cores, and for complex instructions, they are converted into simple instructions and then processed by the micro-cores.
[0010] Furthermore, the main core control layer is used to issue and dispatch the instructions entering from the outside, and process the received simple instructions in the main cores; for complex instructions, they are converted into simple instructions that can be recognized by the micro-cores in the main cores and then sent to the micro-cores in the micro-core array layer for processing in sequence.
[0011] Furthermore, the instructions in the micro-core control layer interact with the micro-cores in a unidirectional transmission mode; the data interacts with the micro-cores in a bidirectional transmission mode. Among them, in the reverse interaction, after data arbitration with the forward data, it enters the micro-core.
[0012] Furthermore, the link controller in the micro-core control layer is used to compress the three-dimensional structure in the micro-core control layer into a two-dimensional structure. At the same time, a fast channel is set in the link controller. When the operation ends, the data is written back to the main core control layer through the fast channel.
[0013] Furthermore, the specific link mode of the link controller is to perform data interaction with the arrays controlled by other main cores. Among them, except for the interaction between adjacent micro-units, the micro-unit groups under the same main core do not perform data links with other micro-units in the same group.
[0014] According to the second aspect of the embodiments of the present invention, a working method of a three-dimensional interconnected many-core processor architecture based on RISC-V is provided. It utilizes the above-mentioned three-dimensional interconnected many-core processor architecture based on RISC-V. The method includes: the main core control layer receives external data and instructions, and combines the collaborative processing of the micro-core array layer and the accelerator layer to achieve fast processing of external instructions.
[0015] Compared with the prior art, the beneficial effects of the present invention are:
[0016] (1) The present invention provides a three-dimensional interconnected many-core processor architecture based on RISC-V. The solution can strengthen the connection between each main core by designing a main core control layer with the collaborative action of multiple main cores, and can also efficiently control the micro-core array through multiple main cores at the same time. Moreover, it can also perform efficient data interaction with the accelerator layer, providing a brand-new design idea for the application of the RISC-V instruction set in high-performance processors;
[0017] (2) The present invention designs a microkernel array layer with 3D Router, enabling each to perform data interaction for up to two additional groups on the basis of the traditional many-core architecture. The 3D Router not only provides a forward data path but also a reverse fast channel. This fast channel only returns to the main core control array through the fast channel after the microkernel completes the instruction, greatly improving the speed during data write-back and achieving the design concept of the microkernel array being both extensive and fast.
[0018] (3) The present invention connects the accelerator to all the microkernel arrays on the outermost side, enabling the results of accelerator training to be sent to the microkernel arrays without passing through the main core. At the same time, the accelerator layer also has data interaction with the main core, and the main core sends the data received from the outside to the accelerator layer for training after processing.
[0019] (4) The microkernels in the microkernel array layer of the present invention adopt a one-stage pipeline. The microkernels with a one-stage pipeline can complete a single instruction at the fastest speed, maximizing the performance of the microkernel array.
[0020] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings
[0021] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0022] Figure 1 is a schematic diagram of a 3D interconnected many-core processor architecture based on RISC-V described in Embodiment 1 of the present invention;
[0023] Figure 2 is a schematic diagram of the connection interaction and data flow direction of the 3D interconnected many-core processor architecture based on RISC-V described in Embodiment 1 of the present invention. Detailed Description of the Embodiment
[0024] The present invention will be further described below in conjunction with the drawings and embodiments.
[0025] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0027] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0028] Embodiment 1:
[0029] The purpose of this embodiment is to provide a three-dimensional interconnected many-core processor architecture based on RISC-V.
[0030] A three-dimensional interconnected many-core processor architecture based on RISC-V includes a main control layer, a micro-core array layer, and an accelerator layer. Among them,
[0031] The main control layer includes several main cores. The main cores are CPU cores based on the five-stage pipeline RISC-V instruction set, and the main cores interact with each other and with the external environment through independent buses;
[0032] The micro-core array layer includes several micro-unit groups. The micro-unit includes a micro-core, a data storage unit, an instruction storage unit, and a link controller. Among them, the micro-core is a CPU core of the RISC-V instruction set that executes some functions of the main core;
[0033] The accelerator layer is used to optimize the space utilization and running speed of the accelerator to meet specific requirements;
[0034] Among them, some of the main cores in the main control layer perform data interaction with the accelerator layer, and the other main cores interact with the micro-core array layer. For simple external instructions, they are directly operated on in the main cores, and for complex instructions, they are converted into simple instructions and then processed by the micro-cores.
[0035] Further, the main core control layer is used to issue and dispatch the instructions entering from the outside and process the received simple instructions in the main cores; for complex instructions, they are converted into simple instructions that can be recognized by the micro-cores in the main cores and are sequentially sent to the micro-cores in the micro-core array layer for processing.
[0036] Further, the main core control layer includes six main cores. Among them, two main cores perform data interaction with the accelerator layer, and the other four main cores interact with the micro-core array layer.
[0037] Further, the instructions in the microkernel control layer interact with the microkernel in a unidirectional transmission manner; the data interacts with the microkernel in a bidirectional transmission manner. Among them, in the reverse interaction, after data arbitration with the forward data, it enters the microkernel.
[0038] Further, the link controller in the microkernel control layer is used to compress the three-dimensional structure in the microkernel control layer into a two-dimensional structure. At the same time, a fast channel is set in the link controller. When the operation ends, the data is written back to the main kernel control layer through the fast channel.
[0039] Further, the specific link method of the link controller is to perform data interaction with the arrays controlled by other main kernels. Among them, except for the interaction between adjacent micro-units, the micro-unit groups under the same main kernel do not perform data links with other micro-units in the same group.
[0040] Further, the microkernel control layer includes sixteen rows of micro-unit groups, and each main kernel that interacts with the microkernel array layer in the main kernel controls and performs data interaction with four rows of micro-unit groups respectively.
[0041] Further, the accelerator layer is used to achieve the optimal space utilization operating speed by adopting different accelerators according to different requirements for data types, network compression, and tensor calculation performance.
[0042] Further, the five-stage pipeline is specifically instruction fetch, decode, memory access, execute, and write-back.
[0043] Specifically, for the convenience of understanding, the following will describe the solution of the present disclosure in detail with reference to the accompanying drawings:
[0044] As Figure 1 shown, a three-dimensional interconnected many-core processor architecture based on RISC-V consists of three layers of architectures: the main kernel control layer, the microkernel array layer, and the accelerator layer;
[0045] Further, in the main kernel control layer:
[0046] It is composed of six five-stage pipeline RISC-V instruction set CPU cores (hereinafter referred to as main kernels, and the main kernel structure will be described later) that can completely execute instruction fetch, decode, memory access, execute, and write-back. Each core has an independent bus to interact with the external environment;
[0047] At the same time, the main kernels can interact with each other through the bus that exchanges data instructions with the external environment, and the connection between the main kernels performs data or instruction interaction through this bus;
[0048] Among them, two main kernels perform data interaction with the accelerator layer, and the remaining four main kernels interact with the microkernel array layer. Each main kernel controls and performs data interaction with four rows of micro-unit groups, and the four main kernels control a total of 16 rows of micro-unit groups;
[0049] The main functions completed by the main core control layer are to issue and dispatch the instructions entering from the outside. When receiving simple instructions, they can be directly operated in the main core. After receiving some complex arithmetic instructions or a series of instructions, the instructions are converted into a set of simple instructions that can be recognized by the micro cores and sent to the micro core array layer in sequence.
[0050] Furthermore, in the micro core array layer:
[0051] It is composed of micro units combined by 16*n RISC-V instruction set CPUs (hereinafter referred to as micro cores) that can execute part of the functions of the main core, data storage units, instruction storage units, and 3D Routers.
[0052] In such a micro unit:
[0053] Instructions interact with the micro core in a unidirectional transmission mode, and data interacts with the micro core in a bidirectional transmission mode. Among them, during the reverse interaction, after a simple data arbitration selection with the forward data, it flows into the micro core.
[0054] The 3D Router is a link controller that compresses the three-dimensional structure into a two-dimensional structure in the micro core array layer. This link controller enables each micro unit to be able to interact with at most six other micro units simultaneously. At the same time, for the convenience of data write-back, a dedicated fast channel is set in the 3D Router. When the operation is completed, the data is written back to the main core control layer through the fast channel of the 3D Router. Furthermore, the 3D Router is a new link controller different from other traditional many-core connection methods, enabling each micro core to have at most two additional connection methods more than the traditional many-core connection methods. At the same time, the 3D Router also provides a data path for quickly returning to the main core.
[0055] The connection method of the 3D Router:
[0056] Since every four rows of micro-units are controlled by one main core, to avoid data conflicts, the micro-unit array under the same main core does not perform data linking with other micro-units in the same array except for the interaction between adjacent micro-units. The way of linking is to perform data interaction with the arrays controlled by other main cores. Take a [16*3] micro-unit matrix as an example: The micro-unit at the [1,1] position is adjacent to the micro-units at the [1,2] and [2,1] positions both physically and in terms of data. At the same time, the micro-unit at the [1,1] position is adjacent to the micro-unit arrays controlled by other main cores in terms of data. The method adopted is that [1,1] is adjacent to [5,2] in terms of data, that is, each micro-unit is adjacent to the micro-unit whose vertical position is increased by four and horizontal position is increased by one. Similarly, the micro-unit at the [5,2] position is adjacent to the micro-units at the [4,2], [6,2], [5,1], and [5,3] positions both physically and in terms of data, and is adjacent to the micro-units at the [9,3] and [1,1] positions in terms of data. The micro-units linked by the 3D Router can perform data interaction with up to six surrounding micro-units at the same time, greatly improving the running speed of a group of data;
[0057] In the RISC-V micro-core, the functions of instruction fetching, decoding, and memory access are weakened, and five modules are integrated into a first-level pipeline. At the same time, the execution function is strengthened, mainly the computing performance is strengthened;
[0058] Further, in the accelerator layer:
[0059] An AI accelerator is mainly used. Specifically, the accelerator layer generally can use an FPGA or an ASIC for neural network training, etc. for acceleration training. According to different requirements for data types, network compression, tensor computing performance, etc., different accelerators are adopted to optimize the space utilization and running speed.
[0060] In the connection between the accelerator layer and the micro-core array layer, both the accelerator and the 16-row micro-unit array are interconnected, and each micro-unit has a dedicated corresponding accelerator interface.
[0061] Further, the structure of the main core is specifically as follows:
[0062] The main core adopts a five-stage pipeline design structure, and the five-stage pipeline structure is respectively instruction fetching, decoding, memory access, execution, and write-back;
[0063] During the instruction fetching process of the main core: After the external sends the instruction to the instruction storage unit, the instruction is read out from the memory, and at the same time, branch prediction and jump are also performed;
[0064] During the decoding process of the main core: The instruction fetched from the memory is "translated" according to the instruction set format, and the operand register index required by the instruction is obtained after translation. The operands are fetched from the general register bank according to the index;
[0065] During the execution of the main core: After obtaining the operand after decoding and translation, the execution command is then carried out according to the operation type. The arithmetic logic unit (ALU) in the main core only requires simple logical operations. The execution process of the main core is simplified as much as possible, and the logical operations in the arithmetic logic unit maintain a single cycle, which can greatly reduce the area of the main core and shorten the operation of the main core;
[0066] During the memory access process of the main core: The memory access instruction is read from or written to the memory;
[0067] During the write-back process of the main core: The result of the execution is written back to the general register file according to the memory access instruction;
[0068] Furthermore, the structure of the micro core is specifically as follows:
[0069] In terms of structure, in order to minimize the operation cycle, the micro core adopts a single pipeline structure, and all functions of the CPU are completed in one pipeline stage. This structure can most closely connect each core in multi-core collaborative work and can greatly improve the operation speed;
[0070] Although the micro core is a single-stage pipeline, it can still completely execute all functions of a single processor. The instructions received by the micro core from the main core or other micro cores are first stored in its own instruction memory, and then through instruction fetch and decoding, the instructions are dispatched and issued to the execution module;
[0071] In the execution module of the micro core, a high-performance fixed-point high-performance multiplier combining the booth and wallace algorithms is adopted to complete each multiplication operation as quickly as possible. Each micro core only performs the operation of one instruction at the same time to avoid data conflicts. After analyzing the instructions and data for operation, it applies for memory access and then writes back the data;
[0072] In the write-back process, there are two data flow directions. One is to flow into the data memory within the core, arbitrate with the original data and then flow back into the micro core again. The other is to write back to the 3D Router. After writing back to the 3D Router, the core can receive new data again through the handshake signal. It should be noted that since this core may interact with up to six other micro cores in terms of data, each core will have up to six groups of handshake signals at most;
[0073] So far, the work of a single micro core has been completed, and the data will be temporarily stored in the 3D Router, and then sent to other micro cores by the 3D Router. At the same time, in order to facilitate writing the data back to the main core, a dedicated fast channel is set in the 3D Router, and the fast channel of the 3D Router is only used when the data needs to be written back at high speed after the operation ends;
[0074] Example 2:
[0075] The purpose of this embodiment is to provide a working method for a three-dimensional interconnected many-core processor architecture based on RISC-V, which utilizes the above-mentioned three-dimensional interconnected many-core processor architecture based on RISC-V. The method includes: the main core control layer receives external data and instructions, and combines the collaborative processing of the micro-core array layer and the accelerator layer to achieve fast processing of external instructions.
[0076] Specifically, for the convenience of understanding, the following describes the solution of the present disclosure in detail with reference to the accompanying drawings:
[0077] Example 1:
[0078] Taking the multiplication of two [3*3] matrices a and b to obtain matrix c as an example, a simple illustration of the operation method is as follows:
[0079] Among them, matrix a is obtained by accelerator training, and matrix b is obtained by bus transmission to the main core. A total of 18 additions and 27 multiplications are used for the multiplication of the two matrices;
[0080] Since the RISC-V instruction set currently does not support instruction set extensions for tensors and vectors, only the possible future tensor vector instructions are simulated in the implementation;
[0081] In the operation mode of simulating the tensor RISC-V instruction set that supports the entire matrix multiplication (when only scalar instruction sets are used for operation, the following steps can be executed sequentially with separate instructions):
[0082] For the convenience of description, the main cores are numbered from top to bottom: [1-6]; specifically, as Figure 2 shown, a working method for a three-dimensional interconnected many-core processor architecture based on RISC-V is as follows:
[0083] 1. First, external data is transmitted to the main core [6] through the bus. The main core {6} then sorts and sends the data to the accelerator for training. After that, the accelerator sequentially sends the obtained matrix a (a total of 9 integer data) to the main core [5] (in the architecture design, it also supports direct transmission from the accelerator to the micro-core array, but in order to demonstrate the collaborative role of the main core control layer and store matrix b, the data flow mode from the main core [6] to the accelerator to the main core [5] is adopted);
[0084] 2. While the training process is in progress, the external divides matrix b into three groups of vertical three-dimensional vectors and sends them to the main cores [1-3] sequentially (it can also be arbitrarily allocated on the main cores [1-4] in other ways. Here, this data allocation method is used to demonstrate the collaborative operation between the main cores). After obtaining matrix a in main core [5], the matrix is divided into horizontal three-dimensional vectors and stored in main core [5].
[0085] 3. Matrix a in main core [5] is multiplied by matrix b, and a total of 9 inner product operations are performed. Each inner product operation is decomposed into three multiplication operations and two addition operations. Now assume that the micro-core array layer is a 16*8 micro-core matrix. First, main core [5] sequentially sends the first horizontal vector of b to main cores [1-3] through the bus (when there are high requirements for computing resources or power consumption, the data can be transmitted only to the first main core. In this way, after each inner product calculation, all the data is shifted one position to the right as a whole, and b transmits the data to the array controlled by the next main core through the 3DRouter, that is Figure 2 the data flow of the dotted line). Main cores [1-3] split the multiplication instructions and two groups of three-dimensional vectors into separate integer data and send them to the corresponding controlled micro-core arrays in sequence. At this time, a total of nine micro-cores receive the multiplication instructions and perform multiplication operations;
[0086] 4. After the first micro-core multiplication operation is completed, the obtained data is transmitted to the micro-core on the right. Vertically, the data is added sequentially downward, and the final addition result is stored in the micro-core that has not received the instruction below. Taking the micro-core array controlled by main core [1] as an example, the micro-units at positions [1,1], [2,1], and [3,1] of the micro-core array calculate a[1,1]*b[1,1], a[1,2]*b[2,1], and a[1,3]*b[3,1] respectively. The operation results are then shifted one position to the right to the micro-units at positions [1,2], [2,2], and [3,2] in sequence. Then, the data of the [1,2] micro-unit is transmitted downward and added to the data of the [2,2] micro-unit. After that, the operation result is shifted downward and added to the data of the [3,2] micro-unit, and the final result, that is, c[1,1], is placed in the memory of the [4,2] micro-unit;
[0087] 5. After the above process ends, the calculation of the second set of inner products begins. The main core [5] sequentially sends b[2,1], b[2,2], b[2,3] to the micro-core array controlled by the main cores [1-3] to perform the exact same multiplication operation as in step 4. After that, an overall right shift is performed. Still taking the micro-core array controlled by core [1] as an example, after the operation ends, the data stored in the micro-units at positions [1,2], [2,2], [3,2] are a[1,1]*b[1,2], a[1,2]*b[2,2], a[1,3]*b[3,2] respectively. Similarly, by moving down and adding, c[1,2] is obtained on the memory of the micro-unit at [4,2]. It should be noted that due to the overall shift of the previous data, the integer data of c[1,1] has been stored in the memory of the micro-unit at position [4,3] (at this time, the c data is stored on the data memory and not in the core. Because there is no transfer instruction, when the array scale is small, that is, when n is small, the data of the micro-unit at [4,2] can also be written into the micro-unit at [4,3], and then the two data are written into the memory in the core. When the memory in the core reaches the upper limit, the data is marked and sent to the adjacent micro-unit through the 3DRouter).
[0088] 6. Then, the last set of vectors of b is sent to the micro-core array, and the above process is repeated. Similarly, when the array scale is small and there is no right shift possible, the rightmost data will disappear, and the new multiplication data is written into the micro-cores in the rightmost column, and the data of the c matrix will be stored.
[0089] 7. During the execution of the above processes 1-6, the main cores [2] and [3] are also performing the same process. Multiply the second row vector of a with the b matrix, and multiply the third row vector of a with the b matrix. The results of the operations are marked and written into one of the micro-cores in the micro-core array as above.
[0090] 8. Finally, each micro-core unit that obtains the data in the c matrix returns to the inside of the corresponding array-controlled main core through the fast channel of the 3DRouter in turn, and then is sent to the main core [4] through the bus in turn. The complete matrix c is sorted out in the main core [4].
[0091] Example 2:
[0092] Now, two [32*32] large matrices are used as an example for illustration.
[0093] According to the block matrix multiplication formula:
[0094]
[0095] Thus, a large 32*32 matrix can be split into 4 [16*16] matrices. Due to the scale limitation of the microkernel array, all square matrices larger than [16*16] need to be split into several small square matrices using the above method. Different from the previous example, a [16*16] square matrix will occupy the entire array resources. Therefore, the first matrix a is split into 16 horizontal vectors, and each 16-dimensional vector is further split into 4 4-dimensional vectors and stored in main cores [1-4] in sequence. The four main cores simultaneously send the first row of data and multiplication instructions to the microkernels;
[0096] At this time, for matrix b obtained by accelerator training, the first column of data is directly sent to the microkernel array (here, the method of sending to the main core and then transmitting to other cores through the bus as exemplified in the previous example is not adopted. One reason is that the data scale is relatively large, and the other is to demonstrate another operation mode of direct interaction between the accelerator and the microkernel array). At this time, the microkernel array calculates the data of c[1,1];
[0097] According to this vector multiplication, the subsequent operation process of Example 1 can be used to obtain matrix c;
[0098] Overall, for the multiplication of two [32*32] matrices, a total of 8 matrix multiplications and 4 matrix additions are performed. It should be noted that each piece of data returned through the 3DRouter needs to be stored in a main core. To avoid data conflicts, when receiving data during the main core transmission, try to transfer the data to main core [6] that does not store the matrix. The scale of matrix addition operation is much smaller than that of matrix multiplication, which can be performed by the microkernel array or directly completed in the main core to output the result. This will not be elaborated here;
[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0100] The above-described three-dimensional interconnected many-core processor architecture based on RISC-V and its working method provided by the above embodiment can be realized and has broad application prospects.
[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A 3D interconnected many-core processor architecture based on RISC-V, characterized in that, it includes a main control layer, a micro-core array layer and an accelerator layer, where, the main control layer includes several main cores, the main cores are CPU cores based on the five-stage pipeline RISC-V instruction set, and the main cores interact with each other and with the external environment through independent buses; the micro-core array layer includes several micro-unit groups, the micro-unit groups include micro-cores, data storage units, instruction storage units and link controllers, where the micro-cores are CPU cores of the RISC-V instruction set that execute some functions of the main cores; the accelerator layer is used to optimize the space utilization and operating speed by accelerators that meet specific requirements; wherein, some of the main cores in the main control layer perform data interaction with the accelerator layer, and the other main cores interact with the micro-core array layer. For simple external instructions, they are directly calculated in the main cores, and for complex instructions, they are converted into simple instructions and then processed by the micro-cores.
2. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the main core control layer is used to issue and dispatch the instructions entering from the outside, and process the received simple instructions in the main cores; for complex instructions, they are converted into simple instructions that can be recognized by the micro-cores in the main cores and sent to the micro-cores in the micro-core array layer for processing in sequence.
3. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the main core control layer includes six main cores, where two main cores perform data interaction with the accelerator layer, and the other four main cores interact with the micro-core array layer.
4. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the instructions in the micro-core control layer interact with the micro-cores in a one-way transmission manner; the data interacts with the micro-cores in a two-way transmission manner, where, in the reverse interaction, after data arbitration with the forward data, it enters the micro-core.
5. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the link controller in the micro-core control layer is used to compress the three-dimensional structure in the micro-core control layer into a two-dimensional structure. At the same time, a fast channel is set in the link controller, and when the operation ends, the data is written back to the main core control layer through the fast channel.
6. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the specific link method of the link controller is to perform data interaction with the arrays controlled by other main cores. Among them, except for the interaction between adjacent micro-units, the micro-unit groups under the same main core do not perform data links with other micro-units in the same group.
7. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, the micro-core control layer includes sixteen rows of micro-unit groups, and each main core in the main core that interacts with the micro-core array layer controls and performs data interaction with four rows of micro-unit groups respectively.
8. The 3D interconnected many-core processor architecture based on RISC-V according to claim 1, characterized in that, The accelerator layer is used to achieve the optimal space utilization and operating speed by adopting different accelerators according to different requirements for data types, network compression, and tensor computing performance.
9. The RISC-V based three-dimensional interconnected many-core processor architecture according to claim 1, characterized in that, the five-stage pipeline is specifically fetch, decode, memory access, execute, and write-back.
10. A working method of a RISC-V based three-dimensional interconnected many-core processor architecture, characterized in that, it utilizes the RISC-V based three-dimensional interconnected many-core processor architecture according to any one of claims 1-9, and the method includes: the main core control layer receives external data and instructions, and combines the cooperative processing of the micro-core array layer and the accelerator layer to achieve the fast processing of external instructions.
Citation Information
Patent Citations
Multi-core processor and multi-core processor set
CN102446158A
Cache system verification method of efficient multi-core RISC-V processor
CN110750957A