Configurable processor for parallel computing
A configurable processor architecture with modular interconnects and flexible data routing addresses the inflexibility of current microprocessors, improving parallel computation efficiency.
Patent Information
- Application Number
- JP2025178418
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-30
- Filing Date
- 2025-10-23
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2040-12-23
AI Technical Summary
Current microprocessors, such as CPUs and GPUs, lack flexibility in hardware configuration, making it difficult to achieve high levels of parallelism and streamlined data flow for repetitive operations common in applications like signal processing and machine learning.
A processor architecture with configurable and reconfigurable processing units interconnected by a modular interconnect fabric, allowing dynamic grouping and pipelined processing, and featuring configurable arithmetic logic circuits and Benes networks with FIFO registers for flexible data routing.
Enables efficient parallel computation by allowing dynamic reconfiguration of data paths and operations, enhancing performance in applications requiring repetitive operations.
Smart Images

Figure 0007791389000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to processor architectures, and more particularly to processor architectures having multiple processing units and data paths that are configurable and reconfigurable to allow parallel computation and data transfer operations to be performed in the processing units. [Background technology]
[0002] Many applications (e.g., signal processing, navigation, matrix inversion, machine learning, large-dataset search) require a significant number of iterative computational steps that are best performed by multiple processors operating in parallel. Current microprocessors, whether traditional "central processing units" (CPUs) powering desktop or mobile computers or more numerically oriented traditional "graphics processing units" (GPUs), are well-suited for such tasks. Even when offered with multiple cores, CPUs or GPUs lack flexibility in hardware configuration. For example, signal processing applications often require a large set of repetitive floating-point operations (e.g., additions and multiplications). As implemented in a traditional CPU or GPU, the operation of a single neuron is implemented as a series of addition, multiplication, and comparison instructions, each of which requires fetching operands from registers or memory, performing the operation in an arithmetic logic unit (ALU), and writing the result of the operation back to registers or memory. While the nature of such operations is well known, the set of instructions or the order in which they are executed varies depending on the data or application. Therefore, due to the way memory, register files, and ALUs are organized in a traditional CPU or GPU, it is difficult to achieve high levels of parallelism and streamlined data flow without the flexibility to reconfigure the data paths that move operands back and forth between memory, register files, and ALUs. In many applications, these operations may be repeated hundreds of millions of times, allowing for significant efficiency gains in a processor with the right architecture. Summary of the Invention [Means for solving the problem]
[0003] According to one embodiment of the present invention, a processor includes: (i) a plurality of configurable processors interconnected by a modular interconnect fabric circuit that is configurable to divide the configurable processors into one or more groups for parallel execution and to interconnect the configurable processors in any order for pipelined processing.
[0004] According to one embodiment, each configurable processor has (i) control circuitry, (ii) a plurality of configurable arithmetic logic circuits, and (iii) configurable interconnect fabric circuitry for interconnecting the configurable arithmetic logic circuits.
[0005] According to one embodiment of the present invention, each configurable arithmetic logic circuit includes (i) a plurality of arithmetic or logic operation circuits and (ii) a configurable interconnect fabric circuit.
[0006] According to one embodiment of the present invention, each configurable interconnect fabric circuit includes (i) a Benes network and (ii) a plurality of configurable first-in-first-out (FIFO) registers.
[0007] The present invention is better understood from the following detailed description considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] 1 illustrates a processor 100 including a 4x4 array of stream processing units (SPUs) 101-1, 101-2, 101-3, . . . , and 101-16, according to one embodiment of the present invention. [Figure 2] 2 illustrates an SPU 200 in one implementation of the SPU of processor 100 of FIG. 1 according to one embodiment of the present invention. [Figure 3A] 3 illustrates an APC 300 in an implementation of one of APCs 201-1, 201-2, 201-3, and 201-4 of FIG. 2 according to one embodiment of the present invention. [Figure 3B]FIG. 10 illustrates an enable signal generated by each operation to signal that the output data stream is ready for processing by the next operation. [Figure 4] 4 illustrates a generalized representative implementation 400 of any of PLF units 102-1, 102-2, 102-3, and 102-4 and PLF subunit 202, according to one embodiment of the present invention.
[0009] To facilitate cross-referencing between the drawings, like elements in the figures are provided with like reference numerals. DETAILED DESCRIPTION OF THE INVENTION
[0010] FIG. 1 illustrates a processor 100 including, for example, a 4×4 array of stream processing units (SPUs) 101-1, 101-2, 101-3, . . . , and 101-16, according to one embodiment of the present invention. Of course, in this detailed description, a 4×4 array is chosen for illustrative purposes. An actual implementation may have any number of SPUs. The SPUs are interconnected to each other by a configurable pipeline fabric (PLF) 102, which allows computation results from a given SPU to be provided, or “streamed,” to another SPU. In this arrangement, the 4×4 array of SPUs within processor 100 is configured at run time into one or more groups of SPUs, with each group of SPUs configured as a pipeline stage for pipelined computational tasks.
[0011] 1, PLF 102 is shown to include PLF units 102-1, 102-2, 102-3, and 102-4, each configured to provide a data path between four SPUs in one of the four quadrants of a 4x4 array. PLF units 102-1, 102-2, 102-3, and 102-4 are also interconnected by appropriately configuring PLF unit 102-5, thereby allowing computation results from any of SPUs 101-1, 101-2, 101-3, ..., and 101-16 to be forwarded to any other one of SPUs 101-1, 101-2, 101-3, ..., and 101-16. In one embodiment, the PLF units of processor 100 are organized hierarchically (the configuration shown in FIG. 1 may be considered a two-level hierarchy, with PLFs 102-1, 102-2, 102-3, and 102-4 forming the first level and PLF 102-5 being the second level). In this embodiment, a host CPU (not shown) configures and reconfigures processor 100 in real time during processing via global bus 104. Interrupt bus 105 is provided to allow each SPU to generate interrupts to the host CPU to indicate task completion or any of a number of exceptional conditions. Input data buses 106-1 and 106-2 stream input data to processor 100.
[0012] In one satellite positioning application, processor 100 may function as a digital baseband circuit that processes real-time digitized samples from a radio frequency (RF) front-end circuit. In this application, the input data samples received by processor 100 on input data buses 106-1 and 106-2 are the in-phase and quadrature components of the signal received at the antenna after signal processing in the RF front-end circuit. The received signal includes navigation signals transmitted from multiple positioning satellites.
[0013] 2 illustrates SPU 200 in one implementation of an SPU of processor 100, according to one embodiment of the present invention. As shown in FIG. 2, SPU 200 has a 2×4 array of arithmetic and logic units, each unit referred to herein as an “arithmetic pipeline complex” (APC), to emphasize that (i) each APC is reconfigurable via a set of configuration registers for any of a number of arithmetic and logic operations, and (ii) the APCs can be configured in any of a number of ways to stream the results of any APC within SPU 200 to another APC. As shown in FIG. 2, APCs 201-1, 201-2, ..., 201-8 of the 2×4 array of APCs in SPU 200 are provided with an inter-APC data path on PLF subunit 202, which is an extension unit from the corresponding PLF unit 101-1, 101-2, 101-3, or 101-4.
[0014] As shown in FIG. 2, SPU 200 has a control unit 203 that executes a small set of instructions from instruction memory 204 loaded by the host CPU via global bus 104. An internal processor bus 209 is accessible by the host CPU via global bus 104 during the configuration phase and by control unit 203 during the computation phase. Switching between the configuration and computation phases is accomplished by an enable signal asserted from the host CPU. When the enable signal is deasserted, the clock signal to the APC, and therefore the data valid signal to operators using the APC, are gated off to conserve power. Any SPU is disabled by the host CPU by gating off the power signal to the SPU. In some embodiments, the power signal to the APC is also gated. Similarly, the PLF may be gated off to conserve power, if desired.
[0015] The enable signals to the APCs are memory-mapped so that the APCs can be accessed via the internal processor bus 209. With this configuration, when multiple APCs are configured in a pipeline, the host CPU or SPU 200 can control the enabling of the APCs in the appropriate order as needed. For example, by enabling the APCs in the reverse order of the data flow in the pipeline, all APCs are ready to process data when the first APC in the data flow is enabled.
[0016] Multiplexer 205 switches control of internal processor bus 209 between the host CPU and control unit 203. SPU 200 includes memory blocks 207-1, 207-2, 207-3, and 207-4 that are accessible via internal processor bus 209 by APCs 201-1, 201-2, ..., 201-8 or by the host CPU or SPU 200 during a computation phase. Switches 208-1, 208-2, 208-3, and 208-4 each switch access to memory blocks 207-1, 207-2, 207-3, and 207-4 between internal processor bus 209 and a corresponding one of internal data buses 210-1, 210-2, 210-3, and 210-4. During the configuration phase, the host CPU configures any elements within SPU 200 by writing to configuration registers via global bus 104, which is at this point extended into internal processor bus 209 by multiplexer 205. During the computation phase, control unit 203 controls the operation of SPU 200 via internal processor bus 209, which includes one or more clock signals that enable APCs 201-1, 201-2, ..., 201-8 to operate in synchronization with one another. At the appropriate times, one or more APCs 201-1, 201-2, ..., 201-8 generate interrupts on interrupt bus 211, which are received by SPU 200 for processing. The SPU forwards these interrupt signals, as well as its own, to the host CPU via interrupt bus 105. Scratch memory 206 is provided to support instruction execution in control unit 203, for example, to store intermediate results, flags, and interrupts. The switch between the construction and computation phases is controlled by the host CPU.
[0017] In one embodiment, memory blocks 207-1, 207-2, 207-3, and 207-4 are accessed by control unit 203 using a local address space mapped to an assigned portion of processor 100's global address space. Configuration registers for APCs 201-1, 201-2, ..., 201-8 are similarly accessible from both the local and global address spaces. APCs 201-1, 201-2, ..., 201-8 and memory blocks 207-1, 207-2, 207-3, and 207-4 may be accessed directly by the host CPU via global bus 104. By configuring multiplexer 205 via the memory-mapped registers, the host CPU can connect and allocate internal processor bus 209 to become part of global bus 104.
[0018] Control unit 203 may be a type of microprocessor known to those skilled in the art as a minimum instruction set computer (MISC) processor that operates under the supervision of a host CPU. In one embodiment, control unit 203 manages lower-level resources (e.g., APCs 201-1, 201-2, 201-3, and 201-4) by handling certain interrupts and configuring configuration registers locally within the resources, thereby reducing the host CPU's monitoring requirements for these resources. In one embodiment, the resources operate without the involvement of control unit 203; that is, the host CPU may handle the interrupts and configuration registers directly. Furthermore, if the configured data processing pipeline requires the participation of multiple SPUs, the host CPU may directly control the entire data processing pipeline.
[0019] FIG. 3A illustrates an APC 300 in one embodiment of one of APCs 201-1, 201-2, 201-3, and 201-4 of FIG. 2, according to one embodiment of the present invention. As shown in FIG. 3A, for purposes of illustration only, APC 300 includes representative operator units 301-1, 301-2, 301-3, and 301-4. Each operator unit may include one or more arithmetic or logic circuits (e.g., adders, multipliers, shifters, suitable combinational logic circuits, suitable sequential logic circuits, or combinations thereof). APC PLF 302, via internal processor bus 209, enables the host CPU to create data paths 303 between operators in any suitable manner. APC PLF 302 and operator units 301-1, 301-2, 301-3, and 301-4 are each configurable by both the host CPU and control unit 203 via internal processor bus 209, and the operator units may be configured to operate on pipelined data streams.
[0020] Within the configured pipeline, the output data stream of each operator is provided as the input data stream of the next operator. As shown in FIG. 3B, a valid signal 401 is generated by each operator and, when asserted, indicates that its output data stream 402 is valid for processing by the next operator. An operator within the pipeline is configured to generate an interrupt signal upon detecting a falling edge of the valid signal 401 to indicate that processing of its input data stream is complete. The interrupt signal is processed by the control unit 203 or the host CPU. Data entering and leaving the APC 300 is provided via the data path of the PLF subunit 202 of FIG. 2.
[0021] Some operators are configured to access an associated memory block (i.e., memory block 207-1, 207-2, 207-3, or 207-4). For example, some operators read data from an associated memory block and write the data to the pipeline's output data stream. Some operators read data from the pipeline's input data stream and write the data to the associated memory block. In either case, the address of the memory location is provided to the operator in its input data stream.
[0022] An APC is provided with one or more buffer operators. A buffer operator may be configured to read from or write to a local buffer (e.g., a FIFO buffer). When congestion occurs at a buffer operator, the buffer operator asserts a pause signal to pause the current pipeline. The pause signal disables all associated APCs until the congestion level returns. The buffer operator then resets the pause signal and resumes processing the pipeline.
[0023] 4 illustrates a generalized representative implementation 400 of any of PLF units 102-1, 102-2, 102-3, and 102-4 and PLF subunit 202, according to one embodiment of the present invention. As shown in FIG. 4, PLF implementation 400 includes Benes network 401 that receives n M-bit input data streams 403-1, 403-2, . . . , 403-n and provides n M-bit output data streams 404-1, 404-2, . . . , 404-n. Benes network 401 is a non-blocking n×n Benes network that can be configured to map and route input data streams to output data streams in any permutation programmed into its configuration registers. Each of output data streams 404-1, 404-2, ..., 404-n is then provided to a corresponding configurable first-in, first-out (FIFO) register in FIFO register 402, thereby causing FIFO output data streams 405-1, 405-2, ..., 405-n to be properly aligned in time for their respective receiving units according to their respective configuration registers. Control buses 410, 411 represent configuration signals to the configuration registers of Benes network 401 and FIFO register 402, respectively.
[0024] The above detailed description is provided to illustrate particular embodiments of the present invention and is not intended to limit the invention. Numerous modifications and variations are possible within the scope of the present invention, which is set forth in the appended claims.
Claims
1. 1. A processor included in a system having a host processor, the processor receiving a system input data stream, the processor comprising: a data bus configured to receive real-time data, over which at least a portion of the system input data stream is received; and a plurality of stream processors each configurable by the host processor to receive an input data stream and provide an output data stream, wherein (i) the input data stream of a selected one or more of the stream processors of the plurality of stream processors comprises the system input data stream; (ii) each of the stream processors comprises an instruction memory, a plurality of arithmetic and logic circuits, a processor bus, and a processing unit, wherein (a) the instruction memory is configurable by the host processor to hold instruction sequences executable by the processing unit, (b) the processing unit executes the instruction sequences in the instruction memory to control operations within the arithmetic and logic circuits, and (c) each of the arithmetic and logic circuits is individually memory mapped and enabled such that the arithmetic and logic circuits can be enabled in a predetermined order by the host processor or the processing unit via the processor bus; a plurality of configurable interconnect circuits connecting the stream processors, each configurable interconnect circuit being configurable by the host processor and one or more of the processing units to selectively route the output data stream of the stream processor as the input data stream of the stream processor, the processing units configuring the configurable interconnect circuit in the course of executing the sequences of instructions in the instruction memory.
2. 2. The processor of claim 1, further comprising a global bus that provides access to said stream processor and said configurable interconnect circuitry and is accessible from said stream processor and said configurable interconnect circuitry.
3. the plurality of stream processors are divided into a plurality of groups; selected ones of the plurality of configurable interconnect circuits connect different groups of the stream processors to each other; 2. The processor of claim 1, wherein configurable interconnect circuits other than the selected configurable interconnect circuit of the plurality of configurable interconnect circuits connect only the stream processors in the same group to each other.
4. The processor of claim 1 , wherein the host processor provides an enable signal to each of the stream processors to initiate a computation phase in the stream processor.
5. The processor of claim 4 , wherein when the enable signal of the stream processor is deasserted, selected circuits within the stream processor are power gated to conserve power.
6. The processor of claim 4 , wherein the processing unit of each of the stream processors is configured to enable power gating of the arithmetic and logic circuit of that stream processor.
7. 2. The processor of claim 1, wherein each said stream processor further comprises an interrupt bus that enables it to generate interrupts to a host computer.
8. The processor of claim 7 , wherein the processing unit of each stream processor processes a selected interrupt on the interrupt bus.
9. the arithmetic and logic unit of each of the stream processors receives an input data stream and provides an output data stream; 2. The processor of claim 1 , wherein an input data stream of an arithmetic logic circuit of the plurality of arithmetic logic circuits comprises the input data stream of the stream processor, and an output data stream of another arithmetic logic circuit of the plurality of arithmetic logic circuits comprises the output data stream of the stream processor.
10. The processor of claim 1, wherein the processor bus transmits instructions and data to the arithmetic logic circuit of the stream processor during execution of the instruction sequence by the processing unit of the stream processor.
11. The processor of claim 10 , wherein each of the stream processors further comprises a plurality of memory circuits directly accessible from the arithmetic logic circuit of the stream processor via the processor bus.
12. each said stream processor further including a plurality of configuration registers accessible by said host processor on a global bus or said processing unit of said stream processor on said processor bus; 11. The processor of claim 10, wherein each of the configuration registers stores values of control parameters of one or more of the arithmetic logic units.
13. The processor of claim 1 , wherein the host processor can access the instruction memory of each of the stream processors via a global bus to provide the instruction sequences.
14. 14. The processor of claim 13, further comprising a processor bus multiplexer configurable by the host processor to connect a portion of the global bus to a processor bus.
15. 2. The processor of claim 1, wherein each of the arithmetic logic circuits receives an enable signal from the processing unit, and when the enable signal is deasserted, the arithmetic logic circuit ceases operation.
16. each said arithmetic logic circuit including a plurality of arithmetic circuits, each receiving an input data stream and providing an output data stream, and configurable interconnect circuitry; 2. The processor of claim 1, wherein the configurable interconnect circuitry is configurable to route (i) the input data stream of the arithmetic logic circuit as the input data stream of an arithmetic circuit of the plurality of arithmetic circuits, (ii) the output data stream of another arithmetic circuit of the plurality of arithmetic circuits back to its own input data stream or as the input data stream of the other arithmetic circuit, and (iii) the output data stream of one arithmetic circuit of the plurality of arithmetic circuits as the output data stream of one arithmetic logic circuit of the plurality of arithmetic logic circuits.
17. 17. The processor of claim 16, wherein each of the arithmetic circuits includes one or more of an adder, a multiplier, or a divider.
18. 17. The processor of claim 16, wherein each arithmetic circuit comprises one or more of a shifter, a combinational logic circuit, a sequential logic circuit, and any combination thereof.
19. 17. The processor of claim 16, wherein each of the operational circuits provides a valid signal to indicate the validity of its output data stream.
20. The processor of claim 16 , wherein at least one of the operational circuits comprises a memory operator.
21. The processor of claim 16 , wherein at least one of the operational circuits comprises a buffer operator.
22. 10. The processor of claim 1, wherein each of the configurable interconnect circuits comprises a non-blocking network that receives one or more input data streams and provides one or more output data streams.
23. 23. The processor of claim 22, wherein the non-blocking network comprises an NxN Benes network.
24. each said configurable interconnect circuit further comprising a plurality of first-in-first-out memories; 23. The processor of claim 22, wherein each of a plurality of the first-in-first-out memories receives a selected one of the one or more output data streams of the non-blocking network and provides a delayed output data stream corresponding to the selected one of the output data streams of the non-blocking network delayed by a configurable delay value.
25. The real-time data includes real-time digitized samples from an RF front-end circuit; The processor of claim 1 , which functions as a digital baseband circuit for processing the real-time data.
26. The processor of claim 25, wherein the samples digitized in real time include in-phase and quadrature components of a signal received at an antenna after signal processing in the RF front-end circuitry.
27. 27. The processor of claim 26, wherein the received signals include navigation signals transmitted from multiple positioning satellites.
Citation Information
Patent Citations
System for configuring and operating adaptive integrated circuits with fixed, application-specific computing elements
JP2005512186A
Tap and wireless payment methods and devices
US20130097080A1
Massively parallel computer, accelerated computing clusters, and two-dimensional router and interconnection network for field programmable gate arrays, and applications
US20170220499A1
Dynamically reconfigurable pipelined pre-processor
WO2013056198A1
Data processing device and control method therefor
WO2014132669A1