Circuit, corresponding device, system and method
Through memory-based hardware accelerator devices, reconfigurable processing units and configurable interconnection networks are used to solve the problems of inflexible hardware accelerator resource usage and low parallel computing efficiency, achieve flexible hardware resource usage and improved parallel computing performance, and meet the processing time requirements of real-time systems.
Patent Information
- Application Number
- CN202110466426.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-07
- Filing Date
- 2021-04-28
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-04-28
AI Technical Summary
Existing hardware accelerators have problems in data processing, such as inflexible resource usage and low parallel computing efficiency, which makes it difficult to meet the processing time requirements of real-time systems.
The memory-based hardware accelerator device, including reconfigurable processing units and configurable interconnection networks, can reconfigure the processing elements at runtime to support multiple signal processing operations, thereby improving resource utilization and parallel computing performance.
It achieves flexible hardware resource utilization and improved parallel computing performance, meeting the processing time requirements of real-time systems.
Smart Images

Figure CN113568864B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of Italian Application No. 102020000009358, filed on April 29, 2020, the contents of which are incorporated herein by reference. Technical Field
[0003] This description relates to digital signal processing circuits, such as hardware accelerators, and related methods, devices, and systems. Background Art
[0004] Various real-time digital signal processing systems (as increasingly required in the automotive sector, for example, for processing video and / or image data, radar data, wireless communication data) can involve processing a significant amount of data per unit time. In various applications, this processing can become very demanding for purely core-based implementations (i.e., implementations involving general-purpose microprocessors or microcontrollers running processing software).
[0005] Therefore, the use of hardware accelerators is becoming increasingly important in certain areas of data processing because it helps speed up the computation of certain algorithms. Compared to core-based implementations, a properly designed hardware accelerator can reduce the processing time of specific operations.
[0006] Conventional hardware accelerators described in the literature or available as commercial products may include different types of processing elements (also referred to as "math units" or "math operators"), where each processing element is dedicated to the computation of a specific operation. For example, such processing elements may include multiply and accumulate (MAC) circuits and / or circuits configured to compute activation functions such as activation nonlinear functions (ANLFs) (e.g., coordinate rotation digital computer (CORDIC) circuits).
[0007] Each of the above-mentioned processing elements is typically designed to implement a specific function (e.g., radix-2 butterfly arithmetic, multiplication of complex vectors, vector / matrix product, trigonometric or exponential or logarithmic functions, convolution, etc.). Consequently, conventional hardware accelerators typically include a variety of such different processing elements connected together via some kind of interconnect network. In some cases, due to data dependencies and / or architectural limitations, only one different processing element is activated at a time, resulting in inefficient use of silicon area and available hardware resources.
[0008] On the other hand, a purely software-implemented, core-based approach (e.g., utilizing a Single Instruction Multiple Data (SIMD) processor) may involve high clock frequencies to meet the typical bandwidth requirements of real-time systems, since in this case each processing element performs a basic operation. Summary of the Invention
[0009] It is an object of one or more embodiments to provide a hardware accelerator device that addresses one or more of the above-mentioned disadvantages.
[0010] In particular, one or more embodiments are directed to providing a memory-based hardware accelerator device (also referred to in the context of this disclosure by the acronym EDPA, Enhanced Data Processing Architecture) comprising one or more processing elements. The processing elements in the hardware accelerator device can be reconfigured at runtime to provide increased flexibility of use and facilitate efficient computation of various signal processing operations that may be particularly demanding in terms of resources (e.g., fast Fourier transforms, digital filtering, implementation of artificial neural networks, etc.).
[0011] One or more embodiments may find application in real-time processing systems where the acceleration of computationally demanding operations (e.g., vector / matrix products, convolutions, FFTs, radix-2 butterfly arithmetic, complex vector multiplications, trigonometric or exponential or logarithmic functions, etc.) may help meet certain performance requirements (e.g., in terms of processing time). This may be the case, for example, in the automotive field.
[0012] According to one or more embodiments, this object may be achieved by means of a circuit (eg, a runtime reconfigurable processing unit) having the features set forth in the following claims.
[0013] One or more embodiments may be directed to corresponding apparatus (eg, a hardware accelerator circuit including one or more runtime reconfigurable processing units).
[0014] One or more embodiments may be directed to a corresponding system (eg, a system-on-chip integrated circuit including a hardware accelerator circuit).
[0015] One or more embodiments may be directed to a corresponding method.
[0016] The claims are an integral part of the technical teaching provided herein with respect to the embodiments.
[0017] According to one or more embodiments, a circuit is provided that may include a set of input terminals configured to receive input digital signals carrying input data; and a set of output terminals configured to provide output digital signals carrying output data. The circuit may include a computational circuit device configured to generate output data based on the input data. The computational circuit device may include a set of multiplier circuits, a set of adder-subtractor circuits, a set of accumulator circuits, and a configurable interconnect network. The configurable interconnect network may be configured to selectively couple the multiplier circuits, the adder-subtractor circuits, the accumulator circuits, the input terminals, and the output terminals in at least two processing configurations. In a first processing configuration, the computational circuit device is configured to compute output data based on a first set of functions, and in at least one second processing configuration, the computational circuit device is configured to compute output data based on a corresponding second set of functions. The second set of functions is different from the first set of functions.
[0018] Thus, one or more embodiments may provide increased flexibility, improved hardware resource usage, and / or improved parallel computing performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] One or more embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0020] Figure 1 is an exemplary circuit block diagram of an electronic system (such as a system on a chip) according to one or more embodiments;
[0021] Figure 2 is an exemplary circuit block diagram of an electronic device implementing a hardware accelerator according to one or more embodiments;
[0022] Figure 3 is an exemplary circuit block diagram of a processing circuit for an electronic device according to an embodiment according to one or more embodiments;
[0023] Figure 4 is another exemplary circuit block diagram of a processing circuit for an electronic device according to an embodiment according to one or more embodiments; and
[0024] Figure 5 This is an example diagram of a multilayer perceptron network structure. DETAILED DESCRIPTION
[0025] In the following description, one or more specific details are described to provide a deeper understanding of the examples of the embodiments of the present description. The embodiments can be obtained without one or more of the specific details, or with other methods, components, materials, etc. In other cases, well-known structures, materials, or operations are not illustrated or described in detail so as not to obscure certain aspects of the embodiments.
[0026] References to "an embodiment" or "one embodiment" in the context of this description are intended to indicate that a particular configuration, structure, or feature described with respect to that embodiment is included in at least one embodiment. Thus, phrases such as "in an embodiment" or "in one embodiment" that may appear in one or more points of this description are not necessarily referring to the same embodiment or embodiments. Furthermore, particular configurations, structures, or features may be combined in any suitable manner in one or more embodiments.
[0027] In the drawings attached hereto, like parts or elements are denoted by like references / numerals, and the corresponding description will not be repeated for the sake of brevity.
[0028] References used herein are for convenience only and do not define the scope of protection or the scope of the embodiments.
[0029] Figure 1 1 is an example of an electronic system 1, such as a system on a chip (SoC), according to one or more embodiments. The electronic system 1 may include various electronic circuits, such as a central processing unit 10 (CPU, e.g., a microprocessor), a main system memory 12 (e.g., system RAM - random access memory), a direct memory access (DMA) controller 14, and a hardware accelerator circuit 16.
[0030] like Figure 1 As shown in FIG, electronic circuits in the electronic system 1 may be connected via a system interconnect network 18 (eg, a SoC interconnect).
[0031] One or more embodiments aim to provide a (runtime) reconfigurable hardware accelerator circuit 16 designed to support the execution of various (basic) arithmetic functions and having improved flexibility of use. Thus, one or more embodiments can help improve the use of silicon area and provide satisfactory processing performance, for example, to meet the processing time requirements of real-time data processing systems.
[0032] like Figure 1 As shown in FIG, in one or more embodiments, the hardware accelerator circuit 16 may include at least one (runtime) configurable processing element 160, preferably a number P of (runtime) configurable processing elements 1600, 1601, ..., 160 P-1 , and a set of local data memory groups M, preferably a number Q=2*P of local data memory groups M0, ..., M Q-1 .
[0033] In one or more embodiments, the hardware accelerator circuit 16 may further include a local control unit 161, a local interconnect network 162, a local data memory controller 163, a local ROM controller 164, (the local ROM controller 164 is coupled to a set of local read-only memories 165, preferably a number P of local read-only memories 1650, 1651, ..., 165 P-1 ) and a local configuration memory controller 166, (the local configuration memory controller 166 is coupled to a set of local configurable coefficient memories 167, preferably a number P of local configurable coefficient memories 1670, 1671, ..., 167 P-1 ). For example, the memory 167 may include a volatile memory (eg, a RAM memory) and / or a non-volatile memory (eg, a PCM memory).
[0034] Different embodiments may include different numbers P of processing elements 160 and / or different numbers Q of local data memory banks M0, . . . , M Q-1 For example, P may be equal to 8 and Q may be equal to 16.
[0035] In one or more embodiments, processing element 160 may be configured to support different (elementary) processing functions with different levels of computational parallelism. For example, processing element 160 may support (e.g., based on appropriate static configuration) different types of arithmetic (e.g., floating point single precision 32-bit, fixed point / integer 32-bit, or 16 or 8-bit with parallel computation or vectorization modes).
[0036] Processing element 160 may include corresponding internal direct memory access (DMA) controllers 1680, 1681, ..., 168 with low complexity. P-1 In particular, the processing element 160 may be configured to read data from the local data memory banks M0, . . . , M0 via the corresponding direct memory access controller 168. Q-1 and / or retrieve input data from the main system memory 12. Thus, the processing element 160 may refine the retrieved input data to generate processed output data. The processing element 160 may be configured to store the processed output data in the local data memory groups M0, ..., M1 through the corresponding direct memory access controller 168. Q-1 and / or main system memory 12.
[0037] Furthermore, processing element 160 may be configured to retrieve input data from local read-only memory 165 and / or from local configurable coefficient memory 167 to perform such refinement.
[0038] In one or more embodiments, a set of local data memory banks M0, ..., M Q-1This can help in processing data in parallel and reduce memory access conflicts.
[0039] Preferably, the local data memory groups M0, ..., M Q-1 Buffering (e.g., double buffering) can be provided, which can help to reduce memory upload time (write operations) and / or download time (read operations). In particular, each local data memory bank can be replicated so that data can be read (e.g., for processing) from one of the two memory banks while (new) data can be stored (e.g., for later processing) in the other memory bank. Thus, moving data can have no negative impact on computing performance because it can be masked.
[0040] In one or more embodiments, local data memory groups M0, ..., M Q-1 The double buffering scheme may be advantageous in combination with streaming mode or back-to-back data processing (eg, as applicable to an FFT N-point processor configured to elaborate a continuous sequence of N data inputs).
[0041] In one or more embodiments, local data memory groups M0, ..., M Q-1 Memory banks with limited storage capacity (and therefore limited silicon footprint) may be included. In the exemplary case of an FFT processor, each local data memory bank may have a storage capacity of at least (maxN) / Q, where maxN is the longest FFT that the hardware can handle. Typical values in applications involving hardware accelerators may be as follows:
[0042] N = 4096 points, for example, each point is a floating point single precision complex number (real number, imaginary number), and its size is 64 bits (or 8 bytes),
[0043] P=8, resulting in Q=16,
[0044] This allows the storage capacity of each local data memory group to be equal to (4096*8 bytes) / 16=2KB (KB=kilobytes).
[0045] In one or more embodiments, local control unit 161 may include a register file that includes information for setting the configuration of processing element 160. For example, local control unit 161 may configure processing element 160 to execute a particular algorithm as directed by a host application running on central processing unit 10.
[0046] In one or more embodiments, local control unit 161 may thus comprise controller circuitry for hardware accelerator circuitry 16. Such controller circuitry may configure (e.g., dynamically) each processing element 160 for computing a specific (basic) function and may configure a corresponding internal direct memory access controller 168 with a specific memory access scheme and cycle time.
[0047] In one or more embodiments, the local interconnect network 162 may include a low-complexity interconnect system, for example, based on a known type of bus network, such as an AXI4-based interconnect. For example, the data parallelism of the local interconnect network 162 may be 64 bits and the address width may be 32 bits.
[0048] Local interconnect network 162 may be configured to connect processing element 160 to local data memory banks M0, . . . , M Q-1 and / or main system memory 12. In addition, local interconnect network 162 may be configured to connect local control unit 161 and local configuration memory controller 166 to system interconnect network 18.
[0049] In particular, the interconnection network 162 may include a set of P master ports MP0, MP1, ..., MP P-1 Each of these master ports is coupled to a corresponding processing element 160; a set of slave ports SP0, SP1, ..., SP P-1 Each of these slave ports can be coupled to local data memory groups M0, ..., M via a local data memory controller 163. Q-1 ; The other pair of ports includes the system master port MP P and the system slave port SP P , configured to be coupled to the system interconnection network 18 (eg, to receive instructions from the central processing unit 10 and / or access data stored in the system memory 12); and another slave port SP P+1 , coupled to the local control unit 161 and the local configuration memory controller 166.
[0050] In one or more embodiments, the interconnection network 162 may be fixed (ie, not reconfigurable).
[0051] In an exemplary embodiment (e.g., see Table I-1 provided below, where an "X" symbol indicates an existing connection between two ports), interconnect network 162 may implement the following connections: coupled to P master ports MP0, MP1, ..., MP P-1 may be connected to respective slave ports SP0, SP1, . . . , SP coupled to the local data memory controller 163. P-1and is coupled to the system interconnection network 18 of the system master port MP P may be connected to a slave port SP coupled to the local control unit 161 P+1 and local configuration memory controller 166 .
[0052] Table I-1 provided below summarizes such exemplary connections implemented through interconnection network 162.
[0053] Table I-1
[0054] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 ]]> <![CDATA[MP0]]> X <![CDATA[MP1]]> X … … <![CDATA[MP P-1 ]]> X <![CDATA[MP P ]]> X
[0055] In another exemplary embodiment (eg, see Table I-2 provided below), the interconnection network 162 may further implement the following connections: P master ports MP0, MP1, ..., MP P-1 Each P master port in the system can be connected to a system slave port SP coupled to the system interconnect network 18. P In this manner, connectivity may be provided between any processing element 160 and the SoC via system interconnect network 18 .
[0056] Table I-2 provided below summarizes such exemplary connections implemented through interconnection network 162.
[0057] Table I-2
[0058] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 ]]> <![CDATA[MP0]]> X X <![CDATA[MP1]]> X X … … … <![CDATA[MP P-1 ]]> X X <![CDATA[MP P ]]> X
[0059] In another exemplary embodiment (e.g., see Table 1-3 provided below, where an "X" symbol indicates an existing connection between two ports, and an "X" in parentheses indicates an optional connection), the interconnection network 162 may further implement the following connections: a system master port MP coupled to the system interconnection network 18; P Can be connected to slave ports SP0, SP1, ..., SP P-1 At least one slave port (here, the P slave port set SP0, SP1, ..., SP P-1 In this way, the first slave port SP0 can be used at the master port MP P Provides connection between the master port MPP and (any) slave port. According to the specific application of the system 1, the connection of the master port MPP can be extended to multiple (eg all) slave ports SP0, SP1, ..., SP P-1 Master Port MP P To slave ports SP0, SP1, ..., SP P-1 The connection of at least one slave port in the local data memory banks M0, ..., M1 can be used (only) to load input data to be processed into the local data memory banks M0, ..., M2. Q-1This is because all memory banks can be accessed via a single slave port. Loading input data can be done using only one slave port, while processing data with the aid of parallel computing can advantageously use multiple (e.g., all) slave ports SP0, SP1, ..., SP P-1 .
[0060] Table I-3 provided below summarizes such exemplary connections implemented via interconnection network 162 .
[0061] Table I-3
[0062] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 <!-- 5 -->]]> <![CDATA[MP0]]> X X <![CDATA[MP1]]> X X … … … <![CDATA[MP P-1 ]]> X X <![CDATA[MP P ]]> X (X) (X) (X) X
[0063] In one or more embodiments, local data memory controller 163 may be configured to arbitrate (eg, by processing element 160 ) access to local data memory banks M0, . . . , M1, . Q-1 For example, the local data memory controller 163 may use a memory access scheme selectable according to a signal received from the local control unit 161 (eg, for calculation of a specific algorithm).
[0064] In one or more embodiments, the local data memory controller 163 may convert an incoming read / write transaction burst (e.g., an AXI burst) generated by the direct read / write memory access controller 168 into a read / write memory access sequence according to a specified burst type, burst length, and memory access scheme.
[0065] Therefore, if Figure 1 One or more embodiments of the hardware accelerator circuit 16 shown in may be intended to reduce the complexity of the local interconnect network 162 by delegating the implementation of the (reconfigurable) connections between the processing elements and the local data memory banks to a local data memory controller 163 .
[0066] In one or more embodiments, local read-only memories 1650, 1651, ..., 165 accessible by processing element 160 via local ROM controller 164 P-1 It may be configured to store digital factors and / or fixed coefficients used to implement a particular algorithm or operation (eg, twiddle factors or other complex coefficients for FFT calculations). The local ROM controller 164 may implement a particular addressing scheme.
[0067] In one or more embodiments, local configurable coefficient memories 1670, 1671, ..., 167 accessible by processing element 160 via local configuration memory controller 166 P-1Can be configured to store application-dependent digital factors and / or coefficients that can be configured by software (e.g., coefficients for implementing FIR filters or beamforming operations, weights of neural networks, etc.) The local configuration memory controller 166 can implement a specific addressing scheme.
[0068] In one or more embodiments, local read-only memories 1650, 1651, ..., 165 P-1 and / or local configurable coefficient memories 1670, 1671, ..., 167 P-1 Advantageously, the local configurable coefficient memory may be divided into a number P of groups equal to the number of processing elements 160 included in the hardware accelerator circuit 16. This helps avoid conflicts during parallel computations. For example, each local configurable coefficient memory may be configured to provide the complete set of coefficients required by each processing element 160 in parallel.
[0069] Figure 2 is the processing element 160 and to the local ROM controller 164, the local configuration memory controller 166 and the local data memory banks M0, ..., M Q-1 An exemplary circuit block diagram of one or more embodiments of the related connections of the processing element 160 (wherein the dotted lines schematically indicate the connection between the processing element 160 and the local data memory groups M0, ..., M Q-1 reconfigurable connections between them).
[0070] like Figure 2 The processing element 160 shown in FIG. 1 may be configured to receive: a first input signal P (eg, indicating a memory access from a local data memory bank M0, ..., M1) via a corresponding direct read memory access 2000 and a buffer register 2020 (eg, a FIFO register). Q-1 a binary-valued digital signal, possibly a complex data with real and imaginary parts); a second input signal Q (e.g., indicating a data memory from the local data memory groups M0, ..., M1, M2, M3, M4, M5, M6, M7, M8, M9, M10, M11, M12, M13, M14, M15, M16, M17, M18, M20, M19, M21, M22, M23, M24, M25, M26, M37, M19, M27, M28, M29, M38, M39, M40, M41, M50, M42, M51, M43, M44, M52, M45, M46, M47, M53, M48, M49, M54, M55, M56, M57, M58, M59, M60, M6 Q-1 a digital signal of a binary value, which may be complex data having a real part and an imaginary part); a first input coefficient W0 (e.g., a digital signal representing a binary value from the local read-only memory 165); and second, third, fourth and fifth input coefficients W1, W2, W3, W4 (e.g., digital signals indicating corresponding binary values from the local configurable coefficient memory 167).
[0071] In one or more embodiments, processing element 160 may include a number of direct read memory accesses 200 equal to the number of input signals P,Q.
[0072] It should be understood that the number of input signals and / or input coefficients received at processing element 160 may vary in different embodiments.
[0073] The processing element 160 may include a computation circuit 20 that may be configured (possibly at runtime) to process input values P, Q and input coefficients W0, W1, W2, W3, W4 to generate a first output signal X0 (e.g., indicating a value to be stored in local data memory banks M0, ..., M1 via corresponding direct write memory access 2040 and buffer registers 2060 (such as FIFO registers)). Q-1 , M0, ..., M1, M2, M3, M4, M5, M6, M7, M8, M9, M10, M11, M12, M13, M14, M15, M16, M17, M18, M20, M19, M21, M22, M23, M24, M25, M26, M37, M19, M27, M28, M29, M30, M31, M32, M43, M10, M11, M29, M33, M12, M13, M24, M14, M25 Q-1 a digital signal with binary values in it).
[0074] In one or more embodiments, processing element 160 may include a number of write direct memory accesses 204 equal to the number of output signals X0 , X1 .
[0075] In one or more embodiments, programming of read and / or write direct memory access 200 , 204 (included in direct memory access controller 168 ) may be performed via an interface (eg, an AMBA interface) that may allow access to internal control registers located in local control unit 161 .
[0076] Additionally, processing element 160 may include ROM address generator circuitry 208 coupled to local ROM controller 164 and memory address generator circuitry 210 coupled to local configuration memory controller 166 to manage data retrieved therefrom.
[0077] Figure 3 is an exemplary circuit block diagram of computing circuitry 20 that may be included in one or more embodiments of processing element 160 .
[0078] like Figure 3 As shown in FIG, the computing circuit 20 may include a processing resource set, for example, including four complex / real multiplier circuits (30a, 30b, 30c, 30d), two complex adder-subtractor circuits (32a, 32b) and two accumulator circuits (34a, 34b), the processing resource set is as shown in FIG. Figure 3 For example, reconfigurable coupling of processing resources can be achieved by means of multiplexer circuits (e.g., 36a to 36j) to form different data paths, where the different data paths correspond to different mathematical operations, where each multiplexer receives a corresponding control signal (e.g., S0 to S7).
[0079] In one or more embodiments, the multiplier circuits 30a, 30b, 30c, 30d can be configured (e.g., by means of an internal multiplexer circuit not visible in the figure) to operate according to two different configurations, which can be selected based on a control signal S8 provided to the multiplier. In a first configuration (e.g., if S8=0), the multiplier can calculate two real product results on four real operands per clock cycle (i.e., each input signal carries two real values). In a second configuration (e.g., if S8=1), the multiplier can calculate one complex product result on two complex operands per clock cycle (i.e., each input signal carries two values, where the first value is the real part of the operand and the second value is the imaginary part of the operand).
[0080] Table II provided below summarizes exemplary possible configurations of the multiplier circuits 30a, 30b, 30c, 30d.
[0081] Table II
[0082]
[0083] By way of example and reference Figure 3 , processing resources can be arranged as follows.
[0084] The first multiplier 30 a may receive a first input signal W1 and a second input signal P (eg, complex operands).
[0085] The second multiplier 30b can receive a first input signal Q and a second input signal selected from the input signals W2 and W4 via a first multiplexer 36a, and the first multiplexer 36a receives a corresponding control signal S2. For example, if S2=0, the multiplier 30b receives the signal W2 as the second input, and if S2=1, the multiplier 30b receives the signal W4 as the second input.
[0086] The third multiplier 30c may receive a first input signal selected from the output signal from the first multiplier 30a and the input signal P.
[0087] For example, Figure 3 As shown in FIG, the second multiplexer 36b can provide either the output signal from the first multiplier 30a (e.g., if S0=0) or the input signal P (e.g., if S0=1) as an output according to the corresponding control signal S0. The third multiplexer 36c can provide either the output signal from the second multiplexer 36b (e.g., if S3=1) or the input signal P (e.g., if S3=0) as an output to the first input of the third multiplier 30c according to the corresponding control signal S3.
[0088] The third multiplier 30 c may receive a second input signal selected from the input signal W3 , the input signal W4 , and the input signal W0 .
[0089] For example, Figure 3 As shown in FIG, the fourth multiplexer 36 d can provide either the input signal W4 (e.g., if S3=0) or the input signal W0 (e.g., if S3=1) as an output according to the corresponding control signal S3. The fifth multiplexer 36 e can provide either the input signal W3 (e.g., if S3=0) or the output signal from the fourth multiplexer 36 d (e.g., if S3=1) as an output to the second input of the third multiplier 30 c according to the corresponding control signal S3.
[0090] The fourth multiplier 30d may receive a first input signal selected from the input signal Q and the output signal from the second multiplier 30b.
[0091] For example, Figure 3 As shown in , the sixth multiplexer 36f can provide either the input signal Q (e.g., if S1=0) or the output signal from the second multiplier 30b (e.g., if S1=1) as an output to the first input of the fourth multiplier 30d according to the corresponding control signal S1.
[0092] The fourth multiplier 30d may receive a second input signal selected from the input signal W4 and the input signal W0.
[0093] For example, Figure 3 As shown, a second input of the fourth multiplier 30d may be coupled to an output of a fourth multiplexer 36d.
[0094] The first adder-subtractor 32a may receive a first input signal selected from the output signal from the first multiplier 30a, the input signal P, and the output signal from the third multiplier 30c.
[0095] For example, Figure 3 As shown in , the seventh multiplexer 36g can provide either the output signal from the second multiplexer 36b (e.g., if S7=1) or the output signal from the third multiplier 30c (e.g., if S7=0) as an output to the first input of the first adder-subtractor 32a.
[0096] The first adder-subtractor 32a may receive a second input signal selected from the input signal Q, the output from the second multiplier 30b, and a zero signal (ie, a binary signal equal to zero).
[0097] For example, Figure 3As shown in FIG, the eighth multiplexer 36h can provide either the input signal Q (e.g., if S6=0) or the output signal from the second multiplier 30b (e.g., if S6=1) as an output according to the corresponding control signal S6. The first AND gate 38a can receive the output signal from the eighth multiplexer 36h as a first input signal and the control signal G0 as a second input signal. The output of the first AND gate 38a can be coupled to the second input of the first adder-subtractor 32a.
[0098] The second adder-subtractor 32 b may receive a first input signal selected from the output signal of the third multiplier 30 c and the output signal of the fourth multiplier 30 d .
[0099] For example, Figure 3 As shown in , the ninth multiplexer 36i can provide either the output signal from the third multiplier 30c (for example, if S5=0) or the output signal from the fourth multiplier 30d (for example, if S5=1) as an output to the first input of the second adder-subtractor 32b according to the corresponding control signal S5.
[0100] The second adder-subtractor 32b may receive a second input signal selected from the output from the fourth multiplier 30d, the output from the second multiplier 30b, and a zero signal (ie, a binary signal equal to zero).
[0101] For example, Figure 3 As shown in FIG, the tenth multiplexer 36j can provide either the output signal from the fourth multiplier 30d (e.g., if S4=0) or the output signal from the second multiplier 30b (e.g., if S4=1) as an output according to the corresponding control signal S4. The second AND gate 38b can receive the output signal from the tenth multiplexer 36j as a first input signal and the control signal G1 as a second input signal. The output of the second AND gate 38b can be coupled to the second input of the second adder-subtractor 32b.
[0102] The first accumulator 34 a may receive an input signal from the output of the first adder-subtractor 32 a and a control signal EN to provide a first output signal X0 of the calculation circuit 20 .
[0103] The second accumulator 34 b may receive an input signal from the output of the second adder-subtractor 32 b and the control signal EN to provide a second output signal X1 of the calculation circuit 20 .
[0104] One or more embodiments including adder-subtractors 32a, 32b may keep their operation "bypassed" via AND gates 38a, 38b, which may be used to force a zero signal at the second input of the adder-subtractors 32a, 32b.
[0105] Figure 4 is an exemplary circuit block diagram of other embodiments of computing circuitry 20 that may be included in one or more embodiments of processing element 160.
[0106] like Figure 4 One or more embodiments shown in the Figure 3 The same arrangement of processing resources and multiplexer circuits discussed above is supplemented with two circuits configured to compute activation non-linear functions (ANLFs) and corresponding multiplexer circuits.
[0107] By way of example and reference Figure 4 , additional processing resources can be arranged as follows.
[0108] The first ANLF circuit 40a can receive an input signal from the output of the first accumulator 34a. The eleventh multiplexer 36k can provide the first output signal X0 of the calculation circuit 20 by selecting either the output signal from the first accumulator 34a (e.g., if S9=0) or the output signal from the first ANLF circuit 40a (e.g., if S9=1) according to the corresponding control signal S9.
[0109] The second ANLF circuit 40b can receive an input signal from the output of the second accumulator 34b. The twelfth multiplexer 36m can provide the second output signal X1 of the calculation circuit 20 by selecting either the output signal from the second accumulator 34b (e.g., if S9=0) or the output signal from the second ANLF circuit 40b (e.g., if S9=1) according to the corresponding control signal S9.
[0110] Therefore, in Figure 4 In one or more embodiments shown in FIG, ANLF circuits 40a and 40b may be "bypassed" by multiplexer circuits 36k and 36m, thereby providing a similar Figure 3 The operation of the embodiment shown in .
[0111] Therefore, reference Figure 3 and Figure 4 As shown, the data paths in computing circuit 20 can be configured to support parallel computing and can facilitate the execution of different functions. In one or more embodiments, the internal pipeline can be designed to meet timing constraints (e.g., clock frequency) for minimum latency.
[0112] In the following, various non-limiting examples are provided of possible configurations of the computation circuit 20. In each example, the computation circuit 20 is configured to compute an algorithm-dependent (elementary) function.
[0113] In the first example, the configuration of the calculation circuit 20 for executing the fast Fourier transform (FFT) algorithm is described.
[0114] Where the hardware accelerator circuit 16 is required to calculate an FFT algorithm, the single processing element 160 may be programmed to implement a radix-2 DIF (decimation in frequency) butterfly algorithm, performing the following complex operations, for example, using signals from the internal control unit 161:
[0115] X0=P+Q
[0116] X1=P*W0-Q*W0
[0117] W0 may be a rotation factor stored in the local read-only memory 165 .
[0118] In this first example, the input signals (P, Q, W0, W1, W2, W3, W4) and the output signals (X0, X1) may be of complex data type.
[0119] Optionally, in order to reduce the impact of discontinuities at the edges of data blocks used in the FFT algorithm on the spectrum, a window function may be applied to the input data before the FFT algorithm is calculated. For example, the processing element 160 may support such window processing by using four multiplier circuits.
[0120] Alternatively, the magnitude or phase of the spectral components can be used instead of the complex values (e.g., in applications such as radar target detection). In this case, the internal (optional) ANLF circuit can be used during the last FFT stage. For example, the input complex vector can be rotated so that it is aligned with the x-axis to calculate the magnitude.
[0121] Table III provided below summarizes some exemplary configurations of computation circuitry 20 for computing different radix-2 algorithms.
[0122] Table III
[0123]
[0124]
[0125] Therefore, the data flow corresponding to the function "Radix-2 Butterfly Algorithm" exemplified above can be:
[0126] X0=P+Q
[0127] X1=P*W0-Q*W0
[0128] The data flow corresponding to the function "radix-2 butterfly algorithm + window" exemplified above can be:
[0129] X0=W1*P+W2*Q
[0130] X1=(W1*P)*W0-(W2*Q)*W0
[0131] The data flow corresponding to the function "radix-2 butterfly algorithm + module" exemplified above can be:
[0132] X0=abs(P+Q)
[0133] X1=abs(P*W0-Q*W0)
[0134] In a first example considered herein, a configuration corresponding to a "radix-2 butterfly algorithm" may involve using two multiplier circuits, two adder-subtractor circuits, no accumulator, and no ANLF circuit.
[0135] In a first example considered herein, a configuration corresponding to a “radix-2 butterfly algorithm + window” may involve using four multiplier circuits, two adder-subtractor circuits, no accumulator, and no ANLF circuit.
[0136] In a first example considered herein, a configuration corresponding to “radix-2 butterfly algorithm + modulo” may involve the use of two multiplier circuits, two adder-subtractor circuits, two ANLF circuits, and no accumulator.
[0137] In the second example, the configuration of the calculation circuit 20 for performing the scalar product of complex data vectors is described.
[0138] The hardware accelerator circuit 16 may be required to calculate the scalar product of complex data vectors. This may be the case, for example, for applications involving filtering operations, such as phased array radar systems, which involve a processing stage known as beamforming. Beamforming techniques can help radar systems resolve targets in angle (azimuth) based on range and radial velocity.
[0139] In this second example, the input signals (P, Q, W0, W1, W2, W3, W4) and the output signals (X0, X1) may be of complex data type.
[0140] In this second example, two different scalar-vector product operations (eg, beamforming operations) may be performed simultaneously by a single processing element 160 (eg, by utilizing all internal hardware resources).
[0141] During beamforming operations, the local configurable coefficient memory 167 may be used to store the phase shifts for the different array antenna elements.
[0142] Similar to the first example, in this second example, if modulo rather than complex values are to be calculated, then one may choose to use an ANLF circuit.
[0143] Table IV provided below illustrates a possible configuration of computation circuitry 20 for computing the scalar product of two vectors simultaneously.
[0144] Table IV
[0145]
[0146] Therefore, the data flow corresponding to the function "scalar product of vectors" exemplified above can be:
[0147] X0=ACC(P*W1+Q*W2)
[0148] X1=ACC(P*W3+Q*W4)
[0149] The data flow corresponding to the function "scalar product + modulus of vectors" exemplified above can be:
[0150] X0=abs(ACC(P*W1+Q*W2))
[0151] X1=abs(ACC(P*W3+Q*W4))
[0152] In a second example considered herein, a configuration corresponding to a "scalar product of vectors" may involve the use of four multiplier circuits, two adder-subtractor circuits, two accumulators, and no ANLF circuit.
[0153] In the second example considered herein, a configuration corresponding to "scalar product of vectors + modulus" may involve the use of four multiplier circuits, two adder-subtractor circuits, two accumulators, and two ANLF circuits.
[0154] In the third example, the configuration of the calculation circuit 20 for performing the scalar product of real number data vectors is described.
[0155] The hardware accelerator circuit 16 may be needed to compute scalar products of real data vectors on large real data structures, for example, for computing digital filters. For example, in many applications, real-world (e.g., analog) signals may be filtered after being digitized in order to extract (only) relevant information.
[0156] In the digital domain, the convolution operation between the input signal and the filter impulse response (FIR) can take the form of a scalar product of two real data vectors. One of the two vectors can hold the input data, while the other can hold the coefficients that define the filtering operation.
[0157] In this third example, the input signals (P, Q, W0, W1, W2, W3, W4) and the output signals (X0, X1) are of real data type.
[0158] In this third example, two different filtering operations may be performed simultaneously by a single processing element 160 on the same data set, eg, by utilizing all internal hardware resources to process four different input data per clock cycle.
[0159] Table V provided below illustrates a possible configuration of computation circuitry 20 for concurrently computing two filtering operations on a real data vector.
[0160] Table V
[0161]
[0162] Therefore, the data flow corresponding to the function shown above is as follows, where the subscript "h" indicates the MSB part and the subscript "l" indicates the LSB part:
[0163] X0 h =ACC(P h *W1 h +Q h *W2 h )
[0164] X0 l =ACC(P l *W1 l +Q l *W2 l )
[0165] X1 h =ACC(P h *W3 h +Q h *W4 h )
[0166] X1 l =ACC(P l *W3 l +Q l *W4 l )
[0167] In a third example considered herein, a configuration corresponding to "scalar product of real vectors" may involve the use of four multiplier circuits, two adder-subtractor circuits, two accumulators, and no ANLF circuit.
[0168] In the fourth example, the configuration of the calculation circuit 20 for calculating a nonlinear function is described.
[0169] A multilayer perceptron (MLP) is a type of fully connected feedforward artificial neural network that can include at least three layers of nodes / neurons. Except for neurons in the input layer, each neuron computes a weighted sum of all nodes in the previous layer and then applies a nonlinear activation function to the result. The processing element 160, as disclosed herein, can handle such nonlinear functions, for example, using internal ANLF circuits. Typically, neural networks process data from the real world and use real weights and functions to calculate class membership probabilities (the output of the last layer). Therefore, for such artificial networks, the scalar product of real data can be the most computationally demanding and frequently used operation.
[0170] Figure 5 is an example diagram of a general structure of a multilayer perceptron network 50.
[0171] like Figure 5 As shown in FIG, the multilayer perceptron network 50 may include an input layer 50a including N input U 1 ,…,U N (U i , i=1, ..., N), the hidden layer 50b includes M hidden nodes X 1 ,…,X M (X k , k=1, ..., M), the output layer 50c includes P output nodes Y 1 ,…,Y P (Y j , j=1,…,P).
[0172] It should be understood that in one or more embodiments, the multilayer perceptron network may include more than one hidden layer 50b.
[0173] like Figure 5 As shown, the multilayer perceptron network 50 may include an input U 1 ,…,U N With hidden node X 1 ,…,X M The first N*M weight set W between i,k , and at the hidden node X 1 ,…,X M With output node Y 1 ,…,Y P The second M*P weight set W between k,j .
[0174] Stored in input U i , hidden node X k and output node Y j The values in can be calculated, for example, as MAC floating points with single precision.
[0175] Hidden node X k The value of the output node Yj can be calculated according to the following equation:
[0176]
[0177]
[0178] In this fourth example, the trained real weights associated with all edges of the MLP may be stored in the local configurable coefficient memory 167. The real layer inputs may be read from the local data memory (e.g., local data memory banks M0, ..., M1) of the hardware accelerator circuit 16. Q-1 ) is retrieved, and the real layer output can be stored in the local data memory of the hardware accelerator circuit 16.
[0179] Since the MLP model is mapped to the hardware accelerator circuit 16, each processing element 160 (e.g., P processing elements) included therein can be used to calculate the scalar product and activation function output associated with two different neurons in the same layer, for example, processing four edges per clock cycle. Therefore, all processing elements 1600, 1601, ..., 160 can be used simultaneously. P-1 .
[0180] Table VI provided below illustrates a possible configuration of computation circuitry 20 for simultaneously computing two activation function outputs associated with two different neurons.
[0181] Table VI
[0182]
[0183] Therefore, the data flow corresponding to the function shown above is as follows, where the subscript "h" indicates the MSB part and the subscript "l" indicates the LSB part:
[0184] X0 h =f(ACC(P h *W1 h +Q h *W2 h ))
[0185] X0 l =f(ACC(P l *W1 l +Q l *W2 l ))
[0186] X1 h =f(ACC(P h *W3 h +Q h *W4 h ))
[0187] X1 l =f(ACC(P l *W3 l +Q l *W4 l ))
[0188] In the fourth example considered herein, a configuration corresponding to the functionality “MLP computation engine” (which may include computing two scalar products of vectors and applying a nonlinear activation function thereto) may involve the use of four multiplier circuits, two adder-subtractor circuits, two accumulators, and two ANLF circuits.
[0189] Table VII provided below illustrates nonlinear functions that may be implemented in one or more embodiments.Some functions denoted with "algorithm=NN" may be used specifically in the context of neural networks.
[0190] Table VII
[0191]
[0192] Thus, one or more embodiments of the hardware accelerator circuit 16, including at least one computation circuit 20 as described herein and / or in the examples above, may facilitate implementation of a digital signal processing system having one or more of the following advantages: flexibility (e.g., the ability to process different types of algorithms), improved use of hardware resources, improved performance of parallel computations, access to local data memory banks M0, . . . , M0 for each processing element 160, and / or the like. Q-1 and / or extended connectivity and high bandwidth to system memory 12 via a simple local interconnect network 162 and internal direct memory access controllers 1680, 1681, ..., 168 P-1 , and support additional algorithms through a scalable architecture that integrates different processing elements.
[0193] In one or more embodiments, the electronic system 1 may be implemented as an integrated circuit in a single silicon chip or chip (e.g., as a system on a chip). Alternatively, the electronic system 1 may be a distributed system comprising multiple integrated circuits interconnected, for example, by means of a printed circuit board (PCB).
[0194] As shown herein, a circuit (e.g., 160) may include a set of input terminals configured to receive a set of input digital signals (e.g., P, Q, W0, W1, W2, W3, W4) carrying input data, a set of output terminals configured to provide a set of output digital signals (e.g., X0, X1) carrying output data, and a computational circuit device (e.g., 20) configured to generate output data based on the input data. The computational circuit device may include: a set of multiplier circuits (e.g., 30a, 30b, 30c, 30d), a set of adder-subtractor circuits (e.g., 32a, 32b), a set of accumulator circuits (e.g., 34a, 34b), and a configurable interconnect network (e.g., 36a, ..., 36j) configured to selectively couple (e.g., S1, ..., S7) the multiplier circuits, the adder-subtractor circuits, the accumulator circuits, the input terminals, and the output terminals in at least two processing configurations.
[0195] As shown herein, in a first processing configuration, the computing circuitry may be configured to compute output data according to a first set of functions, and in at least one second processing configuration, the computing circuitry may be configured to compute output data according to a corresponding second set of functions, the corresponding second set of functions being different from the first set of functions.
[0196] As shown herein, the circuit may include respective configurable direct read memory access controllers (e.g., 2000, 2001) coupled to a first subset of a set of input terminals to receive (e.g., 162, 163) a respective first subset of input digital signals carrying a first subset of input data (e.g., P, Q). The configurable direct read memory access controllers may be configured to control the access of the memory (e.g., M0, ..., M1) from the memory (e.g., M2, ..., M3). Q-1 ) gets the first subset of the input data.
[0197] As shown herein, the circuit can include corresponding configurable direct write memory access controllers (e.g., 2040, 2041) coupled to the set of output terminals to provide output digital signals carrying output data. The configurable direct write memory access controllers can be configured to control storage of output data into the memory.
[0198] As shown herein, the circuitry may include respective input buffer registers (eg, 2020, 2021) coupled to a configurable direct read memory access controller and respective output buffer registers (eg, 2060, 2061) coupled to a configurable write direct memory access controller.
[0199] As shown herein, the circuit may include a ROM address generator circuit (e.g., 208) configured to control the fetching of a second subset of input data (e.g., W0) from at least one read-only memory (e.g., 164, 165) via a second subset of input digital signals, and / or a memory address generator circuit (e.g., 210) configured to control the fetching of a third subset of input data (e.g., W1, W2, W3, W4) from at least one configurable memory (e.g., 166, 167) via a third subset of input digital signals.
[0200] As shown herein, in a circuit according to an embodiment, the multiplier circuit set may include a first multiplier circuit (e.g., 30a), a second multiplier circuit (e.g., 30b), a third multiplier circuit (e.g., 30c), and a fourth multiplier circuit (e.g., 30d). The adder-subtractor circuit set may include a first adder-subtractor circuit (e.g., 32a) and a second adder-subtractor circuit (32b). The accumulator circuit set may include a first accumulator circuit (e.g., 34a) and a second accumulator circuit (e.g., 34b).
[0201] As shown herein, a first multiplier circuit may receive a first input signal (e.g., W1) of a set of input digital signals as a first operand and may receive a second input signal (e.g., P) of the set of input digital signals as a second operand. A second multiplier circuit may receive a third input signal (e.g., Q) of the set of input digital signals as a first operand and may receive a signal selectable from a fourth input signal (e.g., W2) and a fifth input signal (e.g., W4) of the set of input digital signals as a second operand. A third multiplier circuit may receive a signal selectable from an output signal from the first multiplier circuit and a second input signal as a first operand and may receive a signal selected from a sixth input signal (e.g., W3), a seventh input signal (e.g., W0), and a fifth input signal as a second operand. A fourth multiplier circuit may receive a signal selectable from an output signal from the second multiplier circuit and a third input signal as a first operand and may receive a signal selected from the fifth input signal and the seventh input signal as a second operand. The first adder-subtractor circuit can receive as a first operand a signal selectable from the output signal from the first multiplier circuit, the second input signal, and the output signal from the third multiplier circuit, and can receive as a second operand a signal selectable from the third input signal, the output signal from the second multiplier circuit, and a zero signal. The second adder-subtractor circuit can receive as a first operand a signal selectable from the output signal from the third multiplier circuit and the output signal from the fourth multiplier circuit, the output signal from the second multiplier circuit, and a zero signal, and can receive as a second operand a signal selectable from the output signal from the fourth multiplier circuit, the output signal from the second multiplier circuit, and a zero signal. The first accumulator circuit can receive as an input the output signal from the first adder-subtractor circuit, and the second accumulator circuit can receive as an input the output signal from the second adder-subtractor circuit. The first accumulator circuit can be selectively activated (e.g., EN) to provide a first output signal (e.g., X0), and the second accumulator circuit can be selectively activated to provide a second output signal (e.g., X1).
[0202] As shown herein, computational circuitry may include a set of circuits (eg, 40a, 40b) configured to compute non-linear functions.
[0203] As shown herein, a set of circuits configured to calculate a nonlinear function may include a first circuit configured to calculate a nonlinear function (e.g., 40a) and a second circuit configured to calculate a nonlinear function (e.g., 40b). The first circuit configured to calculate a nonlinear function may receive an output signal from a first accumulator circuit as an input. The second circuit configured to calculate a nonlinear function may receive an output signal from a second accumulator circuit as an input. The first output signal may be selectable between an output signal from the first accumulator circuit and an output signal from the first circuit configured to calculate a nonlinear function (e.g., 36k), and the second output signal may be selectable between an output signal from the second accumulator circuit and an output signal from the second circuit configured to calculate a nonlinear function (e.g., 36m).
[0204] As shown herein, a device (eg, 16) may include a set of circuits according to one or more embodiments, a set of data memory groups (eg, M0, . . . , M1, M2, M3, M4, M5, M6, M7, M8, M9, M10, M11, M12, M13, M14, M15, M16, M17, M18, M20, M19, M21, M22, M Q-1 ) and a control unit (e.g., 161). According to configuration data stored in the control unit, the circuits can be configured (e.g., 161, 168) to read data from and write data to the data memory groups via the interconnection network (e.g., 162, 163).
[0205] As shown herein, the data memory bank may include buffer registers, preferably double buffer registers.
[0206] As shown herein, a system (e.g., 1) may include a device according to one or more embodiments and a processing unit (e.g., 10) coupled to the device via a system interconnect (e.g., 18). Based on control signals received from the processing unit, circuits in a circuit set of the device may be configured in at least two processing configurations.
[0207] As described herein, a method of operating a circuit according to one or more embodiments, an apparatus according to one or more embodiments, or a system according to one or more embodiments may include dividing an operating time of a computing circuit device in at least first and second operating intervals, wherein the computing circuit device operates in a first processing configuration and at least one second processing configuration, respectively.
[0208] Without prejudice to the essential principles, the details and embodiments may vary, even significantly, with respect to what has been described purely by way of example, without departing from the scope of protection.
[0209] The scope of protection is defined by the appended claims.
[0210] Although the present invention has been described with reference to exemplary embodiments, this description is not intended to be construed in a limiting sense. Various modifications and combinations of the illustrative embodiments and other embodiments of the present invention will become apparent to those skilled in the art upon reference to the specification. Accordingly, the appended claims are intended to encompass any such modifications or embodiments.
Claims
1. A circuit comprising: a set of input terminals configured to receive a set of corresponding input digital signals carrying input data; a respective configurable direct read memory access controller coupled to a first subset of the set of input terminals to receive a respective first subset of the input digital signals carrying a first subset of input data, wherein the configurable direct read memory access controller is configured to control retrieval of the first subset of input data from a memory; a set of output terminals configured to provide a set of corresponding output digital signals carrying output data; a corresponding configurable direct-write memory access controller coupled to the set of output terminals to provide the output digital signal carrying output data, wherein the configurable direct-write memory access controller is configured to control storage of the output data into the memory; as well as a computing circuit device configured to generate the output data according to the input data, wherein the computing circuit device comprises: A collection of multiplier circuits; Adder-subtractor circuit assembly; a collection of accumulator circuits; and a configurable interconnect network configured to selectively couple the multiplier circuit, the adder-subtractor circuit, the accumulator circuit, the input terminal, and the output terminal in at least two processing configurations; in: In a first processing configuration, the computation circuitry is configured to compute the output data according to a first set of functions; and In at least one second processing configuration, the computation circuitry is configured to compute the output data according to a respective second set of functions, the respective second set of functions being different from the first set of functions.
2. The circuit according to claim 1, further comprising: A respective input buffer register and a respective output buffer register, the respective input buffer register being coupled to the configurable direct read memory access controller, and the respective output buffer register being coupled to the configurable direct write memory access controller.
3. The circuit of claim 1 , further comprising: a read-only memory (ROM) address generator circuit configured to control the acquisition of a second subset of input data from at least one read-only memory via the second subset of the input digital signals; and / or A memory address generator circuit is configured to control the retrieval of a third subset of input data from at least one locally configurable memory via a third subset of the input digital signals.
4. The circuit of claim 1 , wherein the set of multiplier circuits comprises a first multiplier circuit, a second multiplier circuit, a third multiplier circuit, and a fourth multiplier circuit, the set of adder-subtractor circuits comprises a first adder-subtractor circuit and a second adder-subtractor circuit, the set of accumulator circuits comprises a first accumulator circuit and a second accumulator circuit, and wherein: the first multiplier circuit receiving a first input signal of the set of corresponding input digital signals as a first operand and receiving a second input signal of the set of corresponding input digital signals as a second operand; the second multiplier circuit receiving a third input signal of the set of corresponding input digital signals as a first operand and receiving a signal selectable from a fourth input signal and a fifth input signal of the set of corresponding input digital signals as a second operand; the third multiplier circuit receiving as a first operand a signal selectable from an output signal from the first multiplier circuit and the second input signal, and receiving as a second operand a signal selectable from a sixth input signal, a seventh input signal, and the fifth input signal of the set of corresponding input digital signals; the fourth multiplier circuit receiving as a first operand a signal selectable from the output signal from the second multiplier circuit and the third input signal, and receiving as a second operand a signal selectable from the fifth input signal and the seventh input signal; the first adder-subtractor circuit receiving as a first operand a signal selectable from the output signal from the first multiplier circuit, the second input signal, and the output signal from the third multiplier circuit, and receiving as a second operand a signal selectable from the third input signal, the output signal from the second multiplier circuit, and a zero signal; the second adder-subtractor circuit receiving as a first operand a signal selectable from the output signal from the third multiplier circuit and the output signal from the fourth multiplier circuit, and receiving as a second operand a signal selectable from the output signal from the fourth multiplier circuit, the output signal from the second multiplier circuit, and a zero signal; the first accumulator circuit receiving as input an output signal from the first adder-subtractor circuit; the second accumulator circuit receiving as input an output signal from the second adder-subtractor circuit; as well as The first accumulator circuit is selectively activatable to provide a first output signal, and the second accumulator circuit is selectively activatable to provide a second output signal.
5. The circuit of claim 4, wherein the computational circuitry comprises a collection of functional circuits configured to compute a nonlinear function.
6. The circuit of claim 5 , wherein the set of functional circuits configured to calculate a nonlinear function comprises: A first circuit configured to calculate a nonlinear function, and a second circuit configured to calculate a nonlinear function, and wherein: The first circuit configured to calculate a nonlinear function receives as input the output signal from the first accumulator circuit; The second circuit configured to calculate a nonlinear function receives as input the output signal from the second accumulator circuit; the first output signal being selectable between the output signal from the first accumulator circuit and an output signal from the first circuit configured to compute a nonlinear function; and The second output signal is selectable between the output signal from the second accumulator circuit and an output signal from the second circuit configured to compute a nonlinear function.
7. An electronic device comprising: Data storage group collection; control unit; interconnected networks; as well as A collection of circuits, each circuit consisting of: a set of input terminals configured to receive a set of corresponding input digital signals carrying input data; a set of output terminals configured to provide a set of corresponding output digital signals carrying output data; and a computing circuit device configured to generate the output data according to the input data, wherein the computing circuit device comprises: A collection of multiplier circuits; Adder-subtractor circuit assembly; a collection of accumulator circuits; and a configurable interconnect network configured to selectively couple the multiplier circuit, the adder-subtractor circuit, the accumulator circuit, the input terminal, and the output terminal in at least two processing configurations; At least one of the following: a read-only memory (ROM) address generator circuit configured to control the acquisition of a second subset of input data from at least one read-only memory via the second subset of the input digital signals; and a memory address generator circuit configured to control, via a third subset of the input digital signals, retrieval of a third subset of input data from at least one locally configurable memory; in: In a first processing configuration, the computation circuitry is configured to compute the output data according to a first set of functions; and In at least one second processing configuration, the computation circuitry is configured to compute the output data according to a respective second set of functions, the respective second set of functions being different from the first set of functions; The set of circuits is configurable to read data from and write data to the data memory group via the interconnection network as a function of configuration data stored in the control unit.
8. The electronic device of claim 7, wherein the data memory group comprises a buffer register.
9. The electronic device of claim 8, wherein the buffer register is a double-buffered register.
10. The electronic device according to claim 7, further comprising: a respective configurable direct read memory access controller coupled to a third subset of the set of input terminals to receive a respective third subset of the input digital signals carrying a third subset of input data, wherein the configurable direct read memory access controller is configured to control retrieval of the third subset of input data from a memory; as well as A corresponding configurable direct write memory access controller is coupled to the set of output terminals to provide the output digital signal carrying the output data, wherein the configurable direct write memory access controller is configured to control storage of the output data into the memory.
11. The electronic device according to claim 10, further comprising: A respective input buffer register and a respective output buffer register, the respective input buffer register being coupled to the configurable direct read memory access controller, and the respective output buffer register being coupled to the configurable direct write memory access controller.
12. The electronic device of claim 7 , wherein the set of multiplier circuits includes a first multiplier circuit, a second multiplier circuit, a third multiplier circuit, and a fourth multiplier circuit, the set of adder-subtractor circuits includes a first adder-subtractor circuit and a second adder-subtractor circuit, the set of accumulator circuits includes a first accumulator circuit and a second accumulator circuit, and wherein: the first multiplier circuit receiving a first input signal of the set of corresponding input digital signals as a first operand and receiving a second input signal of the set of corresponding input digital signals as a second operand; the second multiplier circuit receiving a third input signal of the set of corresponding input digital signals as a first operand and receiving a signal selectable from a fourth input signal and a fifth input signal of the set of corresponding input digital signals as a second operand; the third multiplier circuit receiving as a first operand a signal selectable from an output signal from the first multiplier circuit and the second input signal, and receiving as a second operand a signal selectable from a sixth input signal, a seventh input signal, and the fifth input signal of the set of corresponding input digital signals; the fourth multiplier circuit receiving as a first operand a signal selectable from the output signal from the second multiplier circuit and the third input signal, and receiving as a second operand a signal selectable from the fifth input signal and the seventh input signal; the first adder-subtractor circuit receiving as a first operand a signal selectable from the output signal from the first multiplier circuit, the second input signal, and the output signal from the third multiplier circuit, and receiving as a second operand a signal selectable from the third input signal, the output signal from the second multiplier circuit, and a zero signal; the second adder-subtractor circuit receiving as a first operand a signal selectable from the output signal from the third multiplier circuit and the output signal from the fourth multiplier circuit, and receiving as a second operand a signal selectable from the output signal from the fourth multiplier circuit, the output signal from the second multiplier circuit, and a zero signal; the first accumulator circuit receiving as input an output signal from the first adder-subtractor circuit; the second accumulator circuit receiving as input an output signal from the second adder-subtractor circuit; as well as The first accumulator circuit is selectively activatable to provide a first output signal, and the second accumulator circuit is selectively activatable to provide a second output signal.
13. The electronic device of claim 12, wherein the computational circuitry comprises a set of functional circuits configured to compute a nonlinear function.
14. The electronic device according to claim 13, wherein the set of functional circuits configured to calculate a nonlinear function comprises: A first circuit configured to calculate a nonlinear function, and a second circuit configured to calculate a nonlinear function, and wherein: The first circuit configured to calculate a nonlinear function receives as input the output signal from the first accumulator circuit; The second circuit configured to calculate a nonlinear function receives as input the output signal from the second accumulator circuit; the first output signal being selectable between the output signal from the first accumulator circuit and an output signal from the first circuit configured to compute a nonlinear function; and The second output signal is selectable between the output signal from the second accumulator circuit and an output signal from the second circuit configured to compute a nonlinear function.
15. An electronic system comprising: System interconnection; processing unit; a device coupled to a processing unit via the system interconnect, wherein the device comprises: Data storage group collection; control unit; interconnecting networks; and A collection of circuits, each circuit consisting of: a set of input terminals configured to receive a set of corresponding input digital signals carrying input data; a respective configurable direct read memory access controller coupled to a first subset of the set of input terminals to receive a respective first subset of the input digital signals carrying a first subset of input data, wherein the configurable direct read memory access controller is configured to control retrieval of the first subset of input data from a memory; a set of output terminals configured to provide a set of corresponding output digital signals carrying output data; a corresponding configurable direct write memory access controller coupled to the set of output terminals to provide the output digital signal carrying output data, wherein the configurable direct write memory access controller is configured to control storage of the output data into the memory; and a computing circuit device configured to generate the output data according to the input data, wherein the computing circuit device comprises: A collection of multiplier circuits; Adder-subtractor circuit assembly; a collection of accumulator circuits; and a configurable interconnect network configured to selectively couple the multiplier circuit, the adder-subtractor circuit, the accumulator circuit, the input terminal, and the output terminal in at least two processing configurations; in: In a first processing configuration, the computation circuitry is configured to compute the output data according to a first set of functions; and In at least one second processing configuration, the computation circuitry is configured to compute the output data according to a respective second set of functions, the respective second set of functions being different from the first set of functions; wherein the set of circuits is configurable to read data from and write data to the data memory group via the interconnection network as a function of configuration data stored in the control unit; and wherein the set of circuits is configurable in at least two processing configurations according to a control signal received from the processing unit.
16. The electronic system of claim 15, wherein the data memory group comprises a buffer register. The electronic system of claim 16 , wherein the buffer register is a double-buffered register.
18. A method of operating a circuit, the circuit comprising: a set of input terminals configured to receive a set of corresponding input digital signals carrying input data; a set of output terminals configured to provide a set of corresponding output digital signals carrying output data; corresponding configurable direct read memory access controllers coupled to a first subset of the set of input terminals to receive a corresponding first subset of the input digital signals carrying a first subset of input data, wherein the configurable direct read memory access controllers are configured to control retrieval of the first subset of input data from a memory; corresponding configurable direct write memory access controllers coupled to the set of output terminals to provide the output digital signals carrying output data, wherein the configurable direct write memory access controllers are configured to control storage of the output data into the memory; and a computational circuit arrangement configured to generate the output data based on the input data, the computational circuit arrangement comprising a set of multiplier circuits, a set of adder-subtractor circuits, a set of accumulator circuits, and a configurable interconnect network, the configurable interconnect network being configured to selectively couple the multiplier circuits, the adder-subtractor circuits, the accumulator circuits, the input terminals, and the output terminals in at least two processing configurations, the computational circuit arrangement being configured to compute the output data in a first processing configuration according to a first set of functions, and the computational circuit arrangement being configured to compute the output data in at least one second processing configuration according to a corresponding second set of functions, the corresponding second set of functions being different from the first set of functions, the method comprising: dividing the operation time of the computing circuit device into at least a first operation interval and a second operation interval; operating the computing circuitry in the first processing configuration during the first operating interval; and During the second operating interval, the computing circuitry is operated in the at least one second processing configuration.
19. A method of operating a circuit, the method comprising: Receiving a set of corresponding input digital signals carrying input data through the set of input terminals, the receiving comprising at least one of the following: (1) controlling, by a read-only memory (ROM) address generator circuit, the acquisition of a first subset of the input data from at least one read-only memory via a first subset of the input digital signals; and (2) controlling, by a memory address generator circuit, access to a second subset of the input data from at least one locally configurable memory via a second subset of the input digital signals; receiving the input data from the set of input terminals via computational circuitry comprising a set of multiplier circuits, a set of adder-subtractor circuits, and a set of accumulator circuits; dividing the operation time of the computing circuit device into at least a first operation interval and a second operation interval; selectively coupling the multiplier circuit, the adder-subtractor circuit, the accumulator circuit, the input terminals of the set of input terminals, and the output terminals in at least two processing configurations via a configurable interconnect network; computing, by the computing circuitry, output data in the first operating interval using a first set of functions in a first processing configuration of the at least two processing configurations; computing, by the computing circuitry, the output data in the second operating interval with a respective second set of functions in at least one second processing configuration of the at least two processing configurations, the respective second set of functions being different from the first set of functions; and A corresponding set of output digital signals carrying the output data is provided via a set of output terminals.
20. The method according to claim 19, further comprising: receiving, via corresponding configurable direct read memory access controllers, a respective third subset of said input digital signals carrying a third subset of input data from a third subset of said set of input terminals; as well as controlling, by the configurable direct read memory access controller, retrieval of the third subset of the input data from a memory; providing the output digital signal carrying the output data to the set of output terminals via a corresponding configurable direct write memory access controller; as well as Storing of the output data into the memory is controlled by the configurable direct write memory access controller.
21. The method of claim 19, further comprising: receiving, via a first multiplier circuit of the set of multiplier circuits, a first input signal of the set of corresponding input digital signals; receiving, via the first multiplier circuit, a second input signal of the set of corresponding input digital signals; receiving, via a second multiplier circuit of the set of multiplier circuits, a third input signal of the set of corresponding input digital signals; receiving, via the second multiplier circuit, a signal selectable from among a fourth input signal and a fifth input signal of the corresponding input digital signal; receiving, by a third multiplier circuit of the set of multiplier circuits, a signal selectable from an output signal from the first multiplier circuit and the second input signal; receiving, via the third multiplier circuit, a signal selectable from a sixth input signal of the corresponding input digital signal, a seventh input signal of the corresponding input digital signal, and a fifth input signal; receiving, by a fourth multiplier circuit of the set of multiplier circuits, a signal selectable from an output signal from the second multiplier circuit and the third input signal; receiving, via the fourth multiplier circuit, a signal selectable from a fifth input signal and a seventh input signal of the corresponding input digital signal; receiving, by a first adder-subtractor circuit of the set of adder-subtractor circuits, a signal selectable from an output signal from the first multiplier circuit, the second input signal, and an output signal from the third multiplier circuit; receiving, by the first adder-subtractor circuit, a signal selectable from the third input signal, the output signal from the second multiplier circuit, and a zero signal; receiving, by a second adder-subtractor circuit of the set of adder-subtractor circuits, a signal selectable from an output signal from the third multiplier circuit and an output signal from the fourth multiplier circuit; receiving, by the second adder-subtractor circuit, a signal selectable from an output signal from the fourth multiplier circuit, an output signal from the second multiplier circuit, and a zero signal; receiving, via a first accumulator circuit of the set of accumulator circuits, an output signal from the first adder-subtractor circuit; receiving, via a second accumulator circuit of the set of accumulator circuits, an output signal from the second adder-subtractor circuit; selectively activating the first accumulator circuit to provide a first output signal; as well as The second accumulator circuit is selectively activated to provide a second output signal.
Citation Information
Patent Citations
Runtime configurable arithmetic and logic cell
US20110010523A1