A chip architecture based on an asynchronous many-core vector processor
Through the chip architecture based on asynchronous multi-core vector processor, the high-compute and low-power consumption requirements in embedded applications are solved, and the chip design with high parallelism and low-power consumption is realized, while simplifying the programming process and expanding the application scope.
Patent Information
- Application Number
- CN202310201665.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-06
AI Technical Summary
Existing embedded application chips have challenges in high computing demands, low-power design and flexibility. FPGA development is complex and programming is slow, and SoC chip accelerator design flexibility is limited.
Using a chip architecture based on asynchronous multi-core vector processors, the processing unit interconnection is achieved through an on-chip network, the network and data network are configured to manage clocks and power supplies respectively, and the clock isolation is used for asynchronous first-in first-out buffers, and resource management is achieved in combination with a grid synchronizer and a rotary lock.
It achieves high parallelism, high computing power, low power consumption, simple programming and fast compilation, and is suitable for a wide range of computing tasks.
Smart Images

Figure CN116185939B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semiconductor chip design, and particularly to a low-power many-core chip architecture for embedded applications. Background Art
[0002] Many chips for embedded applications have very high computing requirements. For example, artificial intelligence (AI) edge processing chips, wireless communication baseband chips, robot / autonomous driving chips, and so on. Currently, for these embedded applications, field-programmable gate arrays (FPGAs) are widely used to implement computing. However, the development of FPGAs is difficult, requiring a high level of expertise from R & D personnel, with slow programming and compilation speeds, which brings inconvenience to R & D.
[0003] In recent years, many system-on-chip (SoC) chips have some dedicated accelerators hung on the buses of processors (including CPUs and DSPs). Such SoC chips can meet the computing requirements of some embedded applications. However, since accelerators are often designed for a specific application or algorithm, this limits the flexibility of the chips and is not conducive to the R & D of new products / new technologies.
[0004] Similar to FPGAs, embedded many-core vector processors have the characteristics of high parallelism and high computing power, and have a wide range of application scenarios. However, embedded many-core vector processors are a new type of computing platform, and there are many technical difficulties in their structures that need to be solved. These technical difficulties include:
[0005] 1) The on-chip interconnection method between processor cores in the many-core vector processor;
[0006] 2) The low-power control method; due to the large number of cores and being for embedded applications, low-power design becomes a key requirement;
[0007] 3) The high computing power of the chip; FPGAs are mainly applied to real-time application scenarios with high parallelism and high computing volume. Similarly, the chips adopting this architecture also require very high computing power. Therefore, improving the computing power of the chips becomes a key technology for the chips adopting this architecture;
[0008] 4) A wide range of application scenarios; just as FPGAs are a general computing platform, this chip should also be a general computing platform and can be widely applied to many embedded occasions such as communication, radar, image processing, artificial intelligence, etc. Therefore, the computing power of the chip should have wide universality. Summary of the Invention
[0009] In view of this, the present invention proposes a chip architecture based on an asynchronous many-core vector processor. The present invention can not only meet the high computing requirements of some embedded applications but also meet the low-power requirements of embedded systems.
[0010] The technical solution adopted by the present invention is as follows:
[0011] A chip architecture based on an asynchronous many-core vector processor uses a network-on-chip to implement the interconnection of many-core processing units; the network-on-chip includes a configuration network, a data network, network-on-chip routers, and network-on-chip interconnecting lines. Among them, the configuration network is used to configure the switches and adjustments of the clocks and power supplies of the data network and each processing unit. The data network is used for communication between processing units. Asynchronous first-in-first-out buffers are used between the network-on-chip routers and the processing units for asynchronous clock isolation.
[0012] Further, the processing unit includes a scalar processor, a vector operation unit, a crossbar switch, a direct memory access controller, a network-on-chip interface circuit, and multiple SRAM memories. There is no cache in the processing unit; the vector operation unit has multiple parallel memory read / write interfaces and is hung on the extended instruction interface of the scalar processor; the scalar processor, the vector operation unit, and the direct memory access controller access any SRAM memory through the crossbar switch; both the scalar processor and the direct memory access controller realize data exchange with the network-on-chip through the network-on-chip interface circuit.
[0013] Further, the SRAM memories in the processing unit are divided into two groups; for the crossbar switch on the first group of SRAM memories, the access priority of the scalar processor is higher than that of the vector operation unit and the direct memory access controller; for the crossbar switch on the second group of SRAM memories, the access priorities of the vector operation unit and the direct memory access controller are higher than that of the scalar processor; the addresses of the first group of SRAM memories come first, the addresses of the second group of SRAM memories are immediately after the addresses of the first group of SRAM memories, and the addresses of both the first group of SRAM memories and the second group of SRAM memories are block-interleaved.
[0014] Further, the network-on-chip is a mesh structure or a three-dimensional mesh structure.
[0015] Further, the interconnection of many-core processing units is implemented using a network-on-chip. The specific method is as follows: Every 4 processing units form an orthogonal unit together. Multiple orthogonal units are arranged in a two-dimensional grid array form on a plane and are connected through the network-on-chip; each orthogonal unit is provided with a network-on-chip router, and the network-on-chip routers are connected through a first link; inside the orthogonal unit, the network-on-chip router and the processing units are connected through a second link, and the processing units are connected through a third link.
[0016] Furthermore, the network-on-chip router has 8 input interfaces and 8 output interfaces, corresponding to 4 processing units and the output and input interfaces in the four directions of east, west, south, and north respectively. Each input interface and output interface is an asynchronous first-in-first-out buffer to achieve asynchronous clock isolation. In the input direction, the data packets in each input first-in-first-out buffer, after being output, first go through routing selection and then are sent to the appropriate output port. In the output direction, the data packets sent by multiple routing selection modules, after arbitration and multiplexing, are sent into the output first-in-first-out buffer. The data packets output by the output first-in-first-out buffer are then sent to the local processing unit or the next-level router.
[0017] Furthermore, a synchronizer is also provided in the chip. The synchronizer is connected to the network-on-chip, and a on-chip shared memory is also provided adjacent to the position of the synchronizer. At the four corners of the chip, a DDR memory controller is provided respectively.
[0018] Furthermore, the synchronizer includes a grid synchronizer, a spin lock, and a mutex semaphore.
[0019] The mutex semaphore is a group of registers, and each register acts as a semaphore. When a certain resource is used by a certain processing unit, the semaphore register corresponding to this resource is set to 1, so as to prevent other processing units from using it. When the resource is released, this register is cleared to be provided for other processing units to use. There is no fixed mapping relationship between each semaphore and each processing unit.
[0020] The usage method of the spin lock is as follows: when a certain resource is used by a certain processing unit and another processing unit requests this resource, the spin lock sequentially caches the number information of all other processing units. When the resource is released, the spin lock automatically notifies the first processing unit in the cache to use this resource. There is no fixed mapping relationship between each spin lock and each processing unit.
[0021] The raster synchronizer consists of a group of raster synchronizer units. Each raster synchronization unit is composed of a flag register, a mask register, a synchronization achieved operation logic, and a notification distribution circuit. The bit widths of each flag register and mask register are the same as the number of processing units. Each processing unit occupies one bit in the flag register and one bit in the mask register. In the synchronization setting stage, the bits of the mask register corresponding to the processing units that need to participate in the synchronization are set to 1, and the bits of the mask register corresponding to the positions of the processing units that do not need to participate in the synchronization are set to 0. Before working, all bits in the flag register are cleared. After starting to work, after each processing unit runs to the synchronization point, it sets the corresponding bit in the flag register to 1. The synchronization achieved operation logic takes the inverse of each bit of the mask register and performs a bitwise OR operation with the flag register, and then performs an AND operation on the results of all bitwise OR operations. If the operation result is 1, it means that all processing units participating in the synchronization have reached the synchronization point. At this time, the notification distribution circuit starts to work. The notification distribution circuit sends synchronization achieved notifications to the processing units with the bits of the mask register being 1 according to the content of the mask register. After all synchronization achieved notifications are distributed, the flag register is cleared and waits for the next synchronization process.
[0022] The beneficial effects of the present invention are as follows:
[0023] 1. The many-core chip architecture of the present invention has high parallelism and high computing power. This is mainly because the number of processing units (PEs) is large, and each PE is a vector processor with high computing power.
[0024] 2. Low power consumption. This is mainly because the clock and power supply of each PE are independent, and the clock and power supply of the non-working PEs can be independently turned off.
[0025] 3. Simple programming and fast compilation. Each PE can be programmed in C language, and the entire chip can also be programmed in C language. This is much simpler and faster to compile than programming an FPGA using a hardware description language.
[0026] 4. Wide range of applications. This is mainly because the basic component of the present invention is a vector processor, and vector calculation is the most basic operation in computational mathematics. Therefore, the chip can adapt to a very wide range of computing tasks, and the flexibility of the chip is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of a chip architecture based on an asynchronous many-core vector processor;
[0028] Figure 2 is a schematic diagram of the PE architecture;
[0029] Figure 3 is a schematic diagram of the overall architecture of the network on chip;
[0030] Figure 4 is a schematic diagram of two power rails;
[0031] Figure 5 is a circuit implementation block diagram of a network-on-chip router;
[0032] Figure 6 is a block diagram of a synchronizer and a shared memory;
[0033] Figure 7 is an implementation block diagram of a grid synchronization unit. Embodiment
[0034] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0035] A chip architecture based on an asynchronous many-core vector processor, Figure 1 in which every 4 PEs together form a QPE (Quadrature PE); multiple QPEs are arranged in a 6×6 two-dimensional mesh array on a plane and are connected through a network-on-chip (NoC). Inside each QPE, there is a NoC router (R), and the Rs are connected through NoC L1 links; the R and the PEs are connected through L2 links; and the PEs are connected through L3 links. At the four corners of the left and right sides of the chip, there is a DDR memory controller each; in the center of the chip, there are a synchronizer and a shared memory. The synchronizer and the shared memory are connected to the NoC and occupy the positions of 2×2 QPE units on the NoC.
[0036] Figure 1 Each PE in is a vector processor and thus has high computing power; multiple PEs can work simultaneously and thus have a very high degree of parallelism. Figure 2 shows a structural block diagram of the PE. Figure 2In it, each PE includes a scalar processor, a vector operation unit, a crossbar switch, an SRAM memory bank, a direct memory access controller (DMA), and a NoC interface circuit. The vector operation unit is connected to the extended instruction interface of the scalar processor and has multiple parallel memory access interfaces. The scalar processor, the vector operation unit, and the DMA can access any SRAM in the SRAM memory bank through the crossbar switch. The SRAM memory bank is divided into multiple SRAMs, which helps to improve the parallelism of accessing each SRAM, thereby increasing the memory access bandwidth. For example, the scalar processor and the vector operation unit can access different SRAMs simultaneously. Similarly, the vector operation unit has multiple parallel memory interfaces to adapt to different memory access bit widths, thus leaving more parallel access opportunities for the scalar processor and the DMA. The NoC interface is responsible for data exchange between the scalar processor, the DMA, and the NoC. The scalar processor can perform bus signal format conversion through the NoC interface to directly access remote registers, or it can achieve data exchange between the remote memory and the local SRAM by configuring the DMA. Adopting this form of SRAM organization in the PE instead of a cache helps to realize a real-time system, that is, to control the operation time.
[0037] A typical implementation scheme of the memory bank is 8 SRAMs, divided into 2 groups, with 4 SRAMs in each group. The first group of SRAMs is mainly used for access by the scalar processor, and the second group of SRAMs is mainly used for access by the vector operation unit and the DMA. Therefore, for the crossbar switch on the first group of SRAMs, the access priority of the scalar processor is higher than that of the vector operation unit and the DMA; for the crossbar switch on the second group of SRAMs, the access priority of the vector operation unit and the DMA is higher than that of the scalar processor. In terms of address arrangement, first are the addresses of the first group of SRAMs, interleaved by block; then are the addresses of the second group of SRAMs, interleaved by block. The addresses of the second group of SRAMs are immediately after the last address of the first group of SRAMs, that is, the addresses between groups are arranged in sequence.
[0038] Figure 1 The NoC and its router R in Figure 3 are shown as follows. Figure 3Among them, the NoC is divided into a configuration NoC (C_NoC) and a data NoC (D_NoC). The C_NoC operates at a fixed clock (with a slower rate) and voltage (lower). There is a register configuration file Quad PE Reg connected to it, which can configure the D_NoC router, as well as the clocks and voltages of each PE connected to the D_NoC router. In this way, when a certain PE does not need to work, its clock and power supply can be turned off; or when its workload is not full, its clock frequency can be reduced and its voltage can be lowered (this requires the entire chip to have multiple power rails, such as Figure 4 the two power rails shown). By adopting this method, the lower power consumption of the entire chip can be ensured.
[0039] Figure 5 The implementation block diagram of the router R is given. This block diagram is applicable to both the D_NoC router and the C_NoC router. Figure 5 Among them, each router R has 8 input interfaces and 8 output interfaces, corresponding to the output interfaces and input interfaces of 4 PEs and the four directions of east, west, south, and north of the C_NoC / D_NoC respectively. Each input interface and output interface is an asynchronous FIFO to achieve asynchronous clock isolation (this helps to achieve low power consumption). In the input direction, the data packets in each input FIFO, after being output, first go through routing and are sent to the appropriate output port. At the output port, the data packets sent by multiple routing modules, after arbitration and multiplexing, are sent into the output FIFO. The data packets output by the output FIFO are then sent to the local PE or the next-level C_NoC / D_NoC router.
[0040] Figure 6 The block diagrams of the synchronizer and the shared memory are given. The synchronizer includes three components: a mutex semaphore, a spin lock, and a barrier synchronizer. The mutex semaphore is a set of registers, and each register in it can be used as a semaphore. When a certain resource is used by a certain PE, the semaphore register corresponding to this resource can be set to 1, so as to prevent other PEs from using it; when the resource is released, this register is cleared to be provided for other PEs to use. There is no fixed mapping relationship between each semaphore and each PE.
[0041] The spin lock is an improvement of the mutex semaphore. When a certain resource is used by a certain PE0, if PE1 requests the same resource, the spin lock caches the (number) information of PE1. At this time, if PE2 requests the resource, the information of PE2 is also cached, and so on; until PE0 releases the resource, at this time, the spin lock automatically notifies PE1 to use the resource; if when PE1 is using the resource, PE3 requests to use the resource, then PE3 will also be cached; and so on, until all cached PE requests are serviced. Therefore, each spin lock has a small piece of cache. Similarly, there can be multiple spin locks in the synchronizer. There is no fixed mapping relationship between each spin lock and each PE.
[0042] The grid synchronizer is also composed of a group of grid synchronizer units. As Figure 7 shown, each grid synchronization unit is composed of a flag register, a mask register, a synchronization achievement operation logic, and a notification distribution circuit. The bit widths of each flag register and mask register are the same as the number of PEs. Each PE occupies one bit in the flag register and one bit in the mask register. In the synchronization setup stage, for the PEs that need to participate in the synchronization, the corresponding bit in the mask register is set to 1, and for the PEs that do not need to participate in the synchronization, the corresponding bit in the mask is set to 0. Before operation, all bits in the flag register are cleared. After the operation starts, when each PE reaches the synchronization point, it sets the corresponding bit in the flag register to 1. Inside the operation logic, the operation logic takes the inverse of each bit of the mask register and performs a bitwise OR operation with the flag register, and then performs an AND operation on the results of all bitwise OR operations. If the operation result is 1, it means that all PEs participating in the synchronization have reached the synchronization point. Then, the notification distribution circuit starts to work. The notification distribution circuit sends a synchronization achievement notification to the PEs whose corresponding bits in the mask register are 1 according to the content of the mask register. After all synchronization achievement notifications are distributed, the flag register is cleared and waits for the next synchronization process.
[0043] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A chip based on an asynchronous many-core vector processor, characterized in that, The interconnection of many-core processing units is implemented by a Network-on-Chip (NoC). The NoC includes a configuration network, a data network, NoC routers, and NoC interconnecting wires. Among them, the configuration network is used to configure the switches and adjustments of the clocks and power supplies of the data network and each processing unit. The data network is used for communication between processing units. Asynchronous first-in-first-out (FIFO) buffers are used between the NoC routers and the processing units for asynchronous clock isolation. The processing unit includes a scalar processor, a vector operation unit, a crossbar switch, a direct memory access (DMA) controller, a NoC interface circuit, and multiple SRAM memories. There is no cache in the processing unit. The vector operation unit has multiple parallel memory read / write interfaces and is connected to the extended instruction interface of the scalar processor. The scalar processor, the vector operation unit, and the DMA controller can access any SRAM memory through the crossbar switch. Both the scalar processor and the DMA controller realize data exchange with the NoC through the NoC interface circuit. The SRAM memories in the processing unit are divided into two groups. For the crossbar switch on the first group of SRAM memories, the access priority of the scalar processor is higher than that of the vector operation unit and the DMA controller. For the crossbar switch on the second group of SRAM memories, the access priorities of the vector operation unit and the DMA controller are higher than that of the scalar processor. The addresses of the first group of SRAM memories come first, and the addresses of the second group of SRAM memories follow immediately after those of the first group. The addresses of both the first group and the second group of SRAM memories are interleaved by block.
2. The chip based on an asynchronous many-core vector processor according to claim 1, wherein The NoC is of a mesh structure or a three-dimensional mesh structure.
3. A chip based on an asynchronous many-core vector processor according to claim 1, characterized in that The interconnection of many-core processing units is implemented by a NoC in the following specific way: Every four processing units form an orthogonal unit together. Multiple orthogonal units are arranged in a two-dimensional mesh array on a plane and are connected through the NoC. There is a NoC router in each orthogonal unit, and the NoC routers are connected through the first link. Inside the orthogonal unit, the NoC router and the processing units are connected through the second link, and the processing units are connected through the third link.
4. A chip based on an asynchronous many-core vector processor according to claim 3, characterized in that, The NoC router has 8 input interfaces and 8 output interfaces, corresponding to 4 processing units and the output and input interfaces in the four directions of east, west, south, and north respectively. Each input interface and output interface is an asynchronous FIFO buffer to achieve asynchronous clock isolation. In the input direction, the data packets in each input FIFO buffer are first routed after being output and then sent to the appropriate output port. In the output direction, the data packets sent by multiple routing modules are sent to the output FIFO buffer after arbitration and multiplexing. The data packets output from the output FIFO buffer are sent to the local processing unit or the next-level router.
5. A chip based on an asynchronous many-core vector processor according to claim 1, characterized in that, A synchronizer is also provided in the chip. The synchronizer is connected to the on-chip network, and an on-chip shared memory is also provided adjacent to the synchronizer. At the four corners of the chip, a DDR memory controller is provided respectively.
6. The chip based on an asynchronous many-core vector processor according to claim 5, wherein The synchronizer includes a grid synchronizer, a spin lock, and a mutex semaphore. The mutex semaphore is a set of registers, and each register acts as a semaphore. When a certain resource is used by a certain processing unit, the semaphore register corresponding to the resource is set to 1, so as to prevent other processing units from using it. When the resource is released, the register is cleared to be used by other processing units. There is no fixed mapping relationship between each semaphore and each processing unit. The spin lock is used as follows: When a certain resource is used by a certain processing unit and another processing unit requests the resource, the spin lock sequentially caches the number information of all other processing units. When the resource is released, the spin lock automatically notifies the first processing unit in the cache to use the resource. There is no fixed mapping relationship between each spin lock and each processing unit. The grid synchronizer is composed of a group of grid synchronizer units. Each grid synchronization unit consists of a flag register, a mask register, a synchronization completion operation logic, and a notification distribution circuit. The bit widths of each flag register and mask register are the same as the number of processing units. Each processing unit occupies one bit in the flag register and one bit in the mask register. In the synchronization setting stage, the bit of the mask register corresponding to the processing unit that needs to participate in synchronization is set to 1, and the bit of the mask register corresponding to the position of the processing unit that does not need to participate in synchronization is set to 0. Before working, all bits in the flag register are cleared. After starting to work, when each processing unit runs to the synchronization point, it sets the bit of its corresponding flag register to 1. The synchronization completion operation logic takes the inverse of each bit of the mask register and performs an OR operation with the flag register bit by bit, and then performs an AND operation on the results of all bitwise OR operations. If the operation result is 1, it means that all processing units participating in synchronization have reached the synchronization point, and at this time the notification distribution circuit starts to work. The notification distribution circuit sends a synchronization completion notification to the processing units with the bit of the mask register being 1 according to the content of the mask register. After all synchronization completion notifications are distributed, the flag register is cleared and waits for the next synchronization process.
Citation Information
Patent Citations
Heterogeneous many-core ASIP architecture based on on-chip bus and shared memory
CN107562549A
Computing In parallel processing environments
CN108804348A