Circuit system, circuit device and operating method thereof
By identifying the number of accumulated times associated with input data bits in the CIM circuit and dynamically adjusting the enable status of the circuit components, the problem of low utilization of existing CIM circuits when processing different accumulation times is solved, achieving higher utilization and lower computing resources and power usage.
Patent Information
- Application Number
- CN202510071535.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-22
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
When existing CIM circuits process different accumulation times, the utilization rate is low, resulting in insufficient use of computing resources and power for MAC operations.
A configurable CIM circuit is provided to dynamically enable or disable the circuit components by identifying the number of accumulated times associated with input data bits, and generate corresponding control signals to optimize the configuration of the adder circuit.
It improves the utilization rate of CIM circuits, reduces the computing resources and power usage of MAC operations, and enhances the prevention measures for multipliers.
Smart Images

Figure CN119990210A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of electronic circuits, and more particularly, to circuit systems, circuit devices, and operating methods thereof. Background Art
[0002] Computer artificial intelligence (AI) is built on the foundation of machine learning, such as using deep learning techniques. With machine learning, a computing system organized as a neural network calculates the statistical likelihood that input data matches previously calculated data. A neural network refers to many interconnected processing nodes that are capable of analyzing data to compare inputs to "trained" data. Trained data refers to computational analysis of properties of known data to develop a model to compare input data to. An example application of AI and data training is object recognition, where the system analyzes the properties of many (e.g., thousands or more) images to determine patterns that can be used to perform statistical analysis to identify input objects. Summary of the invention
[0003] One embodiment of the present invention provides a circuit system, comprising: a computing circuit; a memory array operably coupled to the computing circuit; and a controller configured to: input a plurality of input data bits to the computing circuit; identify an accumulated quantity associated with the plurality of input data bits; determine whether to enable or disable at least one component of the computing circuit based on the accumulated quantity; and generate a control signal to enable or disable at least one component of the computing circuit based on the determination of enabling or disabling.
[0004] Another embodiment of the present invention provides a circuit device, comprising: a memory array; a computing circuit operably coupled to the memory array, the computing circuit comprising: a first component configured to receive a plurality of input data bits and provide a first output in response to a control signal; a second component configured to receive the first output from the first component and provide a second output in response to a control signal including a first logic value; and a multiplexer configured to output the first output in response to a control signal including a second logic value, and configured to output the second output in response to a control signal including the first logic value.
[0005] Yet another embodiment of the present invention provides a method for operating a circuit, comprising: receiving a plurality of input data bits through a computing circuit; identifying an accumulated quantity associated with the plurality of input data bits; determining whether to enable or disable at least one component of the computing circuit based on the accumulated quantity; and generating a control signal to enable or disable at least one component of the computing circuit based on the determination of enabling or disabling. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various aspects of the present invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, the various components are not drawn to scale. In fact, the sizes of the various components may be arbitrarily increased or reduced for clarity of discussion.
[0007] Figure 1 A block diagram of a data computing circuit according to some embodiments of the present disclosure is shown.
[0008] Figure 2 A block diagram illustrating a portion of an example data computing circuit (hereinafter referred to as a “configurable circuit”) according to some embodiments of the present disclosure.
[0009] Figure 3 Schematic diagram showing an example configurable circuit according to some embodiments of the present disclosure.
[0010] Figure 4A A schematic diagram showing an example adder circuit according to some embodiments of the present disclosure.
[0011] Figure 4B List some embodiments according to the present disclosure Figure 4A Example states of the components in the adder circuit shown in for different amounts of accumulation.
[0012] Figure 5 Some embodiments of the present disclosure are shown. Figure 4A An example diagram of the signals associated with the adder circuit shown in FIG.
[0013] Fig. 6A It shows that some embodiments of the present disclosure can be used with Figure 2 The configurable circuit shown in FIG. 1 is coupled to an example logic circuit.
[0014] Figure 6B Example modes regarding accumulated quantities and corresponding control signals according to some embodiments of the present disclosure are listed.
[0015] Figure 6C It shows that some embodiments of the present disclosure can be used with Figure 2 The example logic components coupled to the configurable circuit shown in FIG.
[0016] Fig. 7A A block diagram illustrating an example configurable circuit according to some embodiments of the present disclosure is shown.
[0017] Figure 7B List some embodiments according to the present disclosure Fig. 7A Example control signals and corresponding outputs of the MUX shown in .
[0018] Figure 7C Some embodiments of the present disclosure are shown Fig. 7A Block diagram of the adder circuit shown in .
[0019] Fig.7D List some embodiments according to the present disclosure Fig. 7A Example number of bits for different adders shown in .
[0020] Figure 8 It shows that some embodiments of the present disclosure can be used with Figure 2 An example selection circuit is shown in which the configurable circuit is coupled.
[0021] Fig.9A It shows that some embodiments of the present disclosure can be used with Figure 2 An example circuit of the configurable circuit coupling shown in FIG.
[0022] Fig. 9B Some embodiments of the present disclosure are shown. Fig.9A An example diagram of the signals associated with the circuit shown in FIG.
[0023] Fig.10 A flow chart illustrating an example method of operating a configurable circuit according to various embodiments.
[0024] Fig.11 A flow chart illustrating an example method of operating a configurable circuit according to various embodiments. DETAILED DESCRIPTION
[0025] The following disclosure provides many different embodiments or examples for realizing different features of the provided subject matter. Specific examples of components and arrangements are described below to simplify the present invention. Of course, these are only examples and are not intended to be limiting. For example, in the following description, forming a first component above or on a second component may include an embodiment in which the first component and the second component are in direct contact, and may also include an embodiment in which an additional component is formed between the first component and the second component so that the first component and the second component may not be in direct contact. Moreover, the present invention may repeat reference numbers and / or letters in various examples. This repetition is for simplicity and clarity, but does not itself specify the relationship between the various embodiments and / or configurations discussed.
[0026] Additionally, for ease of description, spatially relative terms such as "below," "beneath," "lower," "above," "upper," "top," and the like may be used herein to describe the relationship of one element or component to another element or component as illustrated in the figures. The spatially relative terms are intended to encompass different orientations of the device during use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein likewise interpreted accordingly.
[0027] Neural networks compute "weights" to perform calculations on new data (input data "words"). Neural networks use multiple layers of computational nodes, where deeper layers perform calculations based on the results of calculations performed by higher layers. Machine learning currently relies on the calculation of dot products and vector absolute differences, typically computed by performing multiply-accumulate (MAC) operations on the parameters, input data, and weights. The calculations of large deep neural networks typically involve so many data elements that it is impractical to store them in the processor cache. Therefore, these data elements are typically stored in memory.
[0028] As a result, machine learning is extremely computationally intensive, requiring the calculation and comparison of many different data elements. The computations of operations within the processor are orders of magnitude faster than the transfer of data elements between the processor and main memory resources. Placing all data elements in caches close to the processor is prohibitively expensive for the vast majority of practical systems due to the size of the memory required to store the data elements. As a result, the transfer of data elements becomes a major bottleneck for AI computations. As data sets grow, the time and power / energy used by the computing system to move data elements can ultimately be a multiple of the time and power used to actually perform the computation.
[0029] In this regard, computation-in-memory (CIM) circuits have been proposed to perform such MAC operations. CIM circuits perform data processing in situ within a suitable memory circuit. CIM circuits suppress the latency of data / program extraction and output results upload to the corresponding memory (e.g., memory array), thereby solving the memory (or von Neumann) bottleneck of traditional computers. Another key advantage of CIM circuits is high computational parallelism, thanks to the specific architecture of the memory array, where computations can be performed along multiple current paths simultaneously. CIM circuits also benefit from the high density of multiple memory arrays with computing devices, which typically have excellent scalability and 3D integration capabilities. As a non-limiting example, CIM circuits for various machine learning applications can perform MAC operations locally in memory (i.e., without sending data elements to the main processor) to achieve higher throughput dot products of neuron activations and weight matrices, while still providing higher performance and lower energy compared to the main processor's computation.
[0030] The data elements processed by CIM circuits are of various types or forms, such as integers and floating-point numbers. Floating-point numbers are usually represented by a sign portion, an exponent portion, and a significand (mantissa) portion consisting of the significant digits of the number. For example, the Institute of Electrical and Electronics Engineers The specified floating point format size is thirty-two bits, including twenty-three mantissa bits, eight exponent bits, and one sign bit. Another floating point format size is sixteen bits, including ten mantissa bits, five exponent bits, and one sign bit.
[0031] In machine learning applications, CIM circuits are often configured to process dot-product multiplications based on performing MAC operations on a large number of data elements (e.g., input word vectors and weight matrices), which may each be in the form of floating-point numbers, and then processing the addition (or accumulation) of these dot products.
[0032] With this approach, an adder circuit configured for a fixed number of accumulations may face low utilization when processing different numbers of accumulations in a given neural network layer. For example, when an adder circuit (or accumulator, adder tree) of a CIM circuit is designed for 64 accumulations, the utilization of the CIM circuit will decrease when processing a smaller number of accumulations (e.g., 8, 16, 32, etc.).
[0033] The present disclosure provides various embodiments of CIM circuits. The CIM circuit disclosed herein may include a configurable adder circuit to have a configurable number of accumulations (e.g., configurable between various numbers of accumulations). For example, when the CIM circuit can support up to 64 accumulations, the CIM circuit can support 2 groups of 32 accumulations, 4 groups of 16 accumulations, and 8 groups of 8 accumulations. The disclosed CIM circuit may include a feature or component for detecting the number of accumulations, and then configure the adder circuit according to the detected number of accumulations, thereby improving CIM utilization and taking precautions for the multiplier to reduce the calculation / computational resources / power usage of MAC operations. In one aspect, the disclosed CIM circuit may input multiple input data bits into the calculation circuit, identify the number of accumulations associated with the multiple input data bits based on the number of accumulations, determine whether to enable or disable at least one component of the calculation circuit, and based on the determination of enabling or disabling, generate a control signal to enable or disable at least one component of the calculation circuit. In some embodiments, the disclosed CIM circuit may include: a first component configured to receive a plurality of input data bits and provide a first output in response to a control signal; a second component configured to receive the first output from the first component and provide a second output in response to a control signal including a first logic value; and a multiplexer configured to output the first output in response to the control signal including the second logic value and to output the second output in response to the control signal including the first logic value.
[0034] Figure 1 FIG. 1 is a block diagram of a data computing circuit 100 according to some embodiments of the present disclosure. Figure 1In the illustrated embodiment, the data computation circuit 100 (also referred to as (e.g., CIM) circuit 100 or memory circuit 100) includes various components that are collectively configured to perform in-memory computations (e.g., multiply-accumulate (MAC) operations) on input word vectors and weight matrices. The input word vector may include a complex number (N) of input data elements InDE, and the weight matrix may include a complex number (Nd) of weight data elements WtDE. In various embodiments, each of the input data elements InDE and the weight data elements WtDE may include a floating point number.
[0035] As shown, the circuit 100 includes a memory circuit 102, an input circuit 104, a plurality of multiplier circuits 106, a plurality of summing circuits 108, a difference circuit 110 (e.g., sometimes referred to as a subtractor circuit 110), a shift circuit 112, an adder circuit (or adder tree) 114, a first converter 116, a second converter 118, a control circuit 120, and an output multiplexer (MUX) 122. In some embodiments, the number of multiplier circuits 106 may correspond to the number of summing circuits 108 or the number of control circuits 120. For example, the circuit 100 may include N (weight / number of input data elements WtDE / InDE) multiplier circuits 106, N (weight / number of input data elements WtDE / InDE) summing circuits 108, and N (weight / number of input data elements WtDE / InDE) control circuits 120. It should be understood that Figure 1 The block diagram of the circuit shown in is simplified, and thus, circuit 100 may include any of a variety of other components while still within the scope of the present disclosure.
[0036] Memory circuit 102 may include one or more memory arrays and one or more corresponding circuits. Each memory array is a storage device including a plurality of storage elements 103, each storage element 103 including an electrical, electromechanical, electromagnetic or other device configured to store one or more data elements, each data element including one or more data bits represented by a logical state. In some embodiments, the logical state corresponds to a voltage level of a charge stored in a portion or all of storage element 103. In some embodiments, the logical state corresponds to a physical property of a portion or all of storage element 103, such as resistance or magnetic orientation.
[0037] In some embodiments, the storage element 103 includes one or more static random access memory (SRAM) cells. In various embodiments, the SRAM cell includes multiple transistors, such as a five-transistor (5T) SRAM cell, a six-transistor (6T) SRAM cell, an eight-transistor (8T) SRAM cell, a nine-transistor (9T) SRAM cell, etc. In some embodiments, the SRAM cell includes a multi-rail SRAM cell. In some embodiments, the length of the SRAM cell is at least twice the width.
[0038] In some embodiments, the storage element 103 includes one or more dynamic random access memory (DRAM) cells, resistive random access memory (RRAM) cells, magnetoresistive random access memory (MRAM) cells, ferroelectric random access memory (FeRAM) cells, NOR flash memory cells, NAND flash memory cells, conductive bridging random access memory (CBRAM) cells, data registers, non-volatile memory (NVM) cells, 3D NVM cells, or other memory cell types capable of storing bit data.
[0039] In addition to the memory array, the memory circuit 102 may also include a plurality of circuits to access or otherwise control the memory array. For example, the memory circuit 102 may include a plurality of (e.g., word line) drivers that are operably coupled to the memory array. The drivers may apply signals (e.g., voltages) to the corresponding storage elements 103 to allow access (e.g., programming, reading, etc.) to these storage elements 103. For another example, the memory circuit 102 may include a plurality of programming circuits and / or reading circuits that are operably coupled to the memory array.
[0040] The memory arrays of the memory circuit 102 are each configured to store a plurality of weight data elements WtDE. In some embodiments, the programming circuit may write the weight data elements WtDE into the corresponding storage elements 103 of the memory array, respectively, and the reading circuit may read the bits written into the storage elements 103 to verify or otherwise test whether the written weight data elements WtDE are correct. The driver of the memory circuit 102 may include or be operably coupled to a plurality of input activation latches configured to receive and temporarily store input data elements InDE. In some other embodiments, such input activation latches may be part of the input circuit 104, which may also include a plurality of buffers configured to temporarily store the weight data elements WtDE retrieved from the memory array of the memory circuit 102. Therefore, the input circuit 104 may receive the input data elements InDE and the weight data elements WtDE.
[0041] In various embodiments of the present disclosure, the circuit 100 is configured to perform a MAC operation on an input word vector (including, for example, an input data element InDE) and a weight matrix (including, for example, a weight data element WtDE) each including a plurality of floating point numbers. Therefore, each of the data element InDE and the weight data element WtDE includes a sign bit, a plurality of exponent bits, and a plurality of mantissa bits (sometimes referred to as fraction bits).
[0042] For example, each of the data element InDE and the weight data element WtDE has a BF16 format, also referred to in some embodiments as a bfloat format or a brain floating point format, in which the first bit represents the sign of the floating point number, the following eight bits represent the exponent of the floating point number, and the last seven bits represent the mantissa or fraction of the floating point number. Since the mantissa is configured to start with a non-zero value, the last seven bits of each stored data element represent an eight-bit mantissa with a first most significant bit (MSB) equal to one.
[0043] In some embodiments, each of the data element InDE and the weight data element WtDE has an FP16 format, also referred to as a half-precision format in some embodiments, in which the first bit represents the sign of a floating-point number, the following five bits represent the exponent of the floating-point number, and the last ten bits represent the mantissa or fraction of the floating-point number. In this case, the last ten bits of each stored data element represent an eleven-bit mantissa with the first MSB equal to one. In some other embodiments, each of the data element InDE and the weight data element WtDE has a floating-point format other than BF16 or FP16 format, for example, another 16-bit format, 32-bit, 64-bit, 128-bit or 256-bit format, or a 40-bit or 80-bit extended precision format. The sign and mantissa of the data element representing the floating-point number are collectively referred to as the signed mantissa of the floating-point number. The MSB of the mantissa is called a hidden bit or a hidden MSB.
[0044] Still refer to Figure 1 , input circuit 104 is configured to output all of each of data elements InDE and WtDE to multiplier circuit 106 and addition circuit 108. In some embodiments, input circuit 104 is configured to output a signed mantissa of each data element to multiplier circuit 106 and an exponent of each data element to addition circuit 108, which will be described below.
[0045] The multiplier circuits 106 are each an electronic circuit, such as an integrated circuit (IC), configured to receive, for example, a sign bit InS and a mantissa InM (collectively referred to as a signed mantissa InS / InM) of each of the N data elements InDE, and a sign bit WtS and a mantissa WtM (collectively referred to as a signed mantissa WtS / WtM) of each of the N data elements WtDE from the input circuit 104. The addition circuits 108 are each an electronic circuit, such as an IC, configured to receive, for example, an exponent InE of each of the N data elements InDE, and an exponent WtE of each of the N data elements WtDE from the input circuit 104.
[0046] The multiplier circuits 106 may each include one or more data registers (not shown) configured to receive instances of the signed mantissas InS / InM and WtS / WtM. Figure 1 In the illustrated embodiment, the multiplier circuit 106 is configured to receive instances of signed mantissas InS / InM and WtS / WtM corresponding to the data elements InDE and WtDE. In some other embodiments, the multiplier circuit 106 includes one or more data registers configured to receive instances of signed mantissas InS / InM and / or WtS / WtM including a hidden MSB. In some embodiments, the multiplier circuit 106 includes one or more data registers configured to add the hidden MSB to the received instances of signed mantissas InS / InM and / or WtS / WtM.
[0047] The multiplier circuit 106 may include logic circuitry (not shown) configured to reformat each instance of the signed mantissa InS / InM into a two's complement mantissa InTC, also referred to as a reformatted mantissa InTC, and to reformat each instance of the signed mantissa WtS / WtM into a two's complement mantissa WtTC, also referred to as a reformatted mantissa WtTC, in operation. The reformatted mantissa InTC has the same number of bits as the signed mantissa InS / InM, and the reformatted mantissa WtTC has the same number of bits as the signed mantissa WtS / WtM.
[0048] The multiplier circuit 106 may include one or more logic gates M1 configured to multiply, in operation, a partial or complete instance of the reformatted mantissa InTC with an instance of a partial or complete reformatted mantissa WtTC, thereby generating N products, e.g., P[1] to P[N]. In various embodiments, the one or more logic gates M1 include one or more AND or NOR gates or other circuits suitable for performing a partial or complete multiplication operation. The one or more logic gates M1 are configured to generate, in operation, each of the products P[1] to P[N] as a two's complement data element including a number of bits equal to twice the number of bits of the reformatted mantissas InTC and WtTC minus one. The one or more logic gates M1 may be referred to as a multiplier configured to multiply, in operation, a partial or complete instance of the reformatted mantissa InTC with an instance of a partial or complete reformatted mantissa WtTC. In some cases, the multiplier (e.g., one or more logic gates M1) may receive a signed mantissa InS / InM or a signed mantissa WtS / WtM for multiplication.
[0049] The multiplier circuit 106 is configured to generate a number N of products P[1] to P[N] during operation. For example, the number N of products P[1]-P[N] that the multiplier circuit 106 may generate is equal to 16. In some other embodiments, the number N of products P[1]-P[N] that the multiplier circuit 106 may generate is less than or greater than 16.
[0050] In some embodiments, for example, in embodiments where the data elements InDE and WtDE have a BF16 format, the multiplier circuit 106 is configured to generate each of the products P[1]-P[N] having a total of 17 bits based on each of the signed mantissas InS / InM and WtS / WtM and each of the reformatted mantissas InTC and WtTC having a total of 9 bits. In some embodiments, for example, in embodiments where the data elements InDE and WtDE have a FP16 format, the multiplier circuit 106 is configured to generate each of the products P[1]-P[N] having a total of 23 bits based on each of the signed mantissas InS / InM and WtS / WtM and each of the reformatted mantissas InTC and WtTC having a total of 12 bits. Embodiments in which multiplier circuit 106 is configured to generate each product P[1]-P[N] having other total numbers of bits based on each signed mantissa InS / InM and WtS / WtM and the reformatted mantissa InTC and WtTC having other total numbers of bits are within the scope of the present disclosure.
[0051] Thus, the multiplier circuit 106 is configured to perform a multiplication and reformatting operation on the sign and mantissa bits of the input data element InDE and the weight data element WtDE in operation to generate a two's complement product P[1]-P[N]. The multiplier circuit 106 is configured to output the product P[1]-P[N] to the shift circuit 112 on the data bus (not shown).
[0052] In various implementations, the multiplier circuit 106 may include one or more other components to perform multiplication (or simplify the multiplication process). For example, the multiplier circuit 106 may include one or more multiplexers (MUXs), switches, or other types of logic components. The multiplier circuit 106 may include other types of logic components that are configured to perform functions such as selecting one of multiple inputs to provide as an output based on a control signal.
[0053] In another example, one or more logic gates M1 of the multiplier circuit 106 may be configured to receive a third input in addition to receiving the corresponding reformatted mantissa InTc and the reformatted mantissa WtTC. The third input may include or correspond to a control signal from the corresponding control circuit 120, including a value of 0 or 1. The one or more logic gates M1 may multiply the reformatted mantissa InTc and the reformatted mantissa WtTC by the control signal. In this case, according to the control signal, the one or more logic gates M1 may output 0 (e.g., control signal = 0) as the product P[n], or output the product of the reformatted mantissa InTc and the reformatted mantissa WtTC (e.g., control signal = 1).
[0054] The summing circuits 108 each include one or more data registers (not shown) configured to receive instances of exponents InE and WtE corresponding to the number of data elements InDE and WtDE discussed above with respect to the multiplier circuits 106 .
[0055] Each of the summing circuits 108 includes one or more logic gates A1 configured to add each instance of the exponent InE to each instance of the exponent WtE in operation. In various embodiments, the one or more logic gates A1 include one or more full adder gates, half adder gates, ripple carry adder circuits, carry save adder circuits, carry select adder circuits, carry predict adder circuits, or other circuits suitable for performing part or all of the addition operation. Each of the logic gates A1 of the summing circuits 108 is configured to generate exponent sums S[1]-S[N] as data elements, the total number of bits of which is equal to the number of bits of each of the exponents InE and WtE plus one.
[0056] Summation circuit 108 is configured to generate, in operation, exponent sums S[1]-S[N] having a total number of bits N and an ordering of data elements that corresponds to the total number N and ordering of data elements of products P[1]-P[N] discussed above with respect to multiplier circuit 106. Thus, for a total of N combinations of data elements InDE and WtDE, each nth combination corresponds to an nth exponent sum S[n] of the exponent sums S[1]-S[N] and an nth product P[n] of the products P[1]-P[N].
[0057] In some embodiments, for example, in embodiments where the data elements InDE and WtDE have a BF16 format, the summation circuit 108 is configured to generate each corresponding exponent sum S[1]-S[N] having a total of nine bits based on each of the exponents InE and WtE having a total of eight bits. In some embodiments, for example, in embodiments where the data elements InDE and WtDE have an FP16 format, the summation circuit 108 is configured to generate each of the sums S[0]-S[N] having a total of six bits based on each of the exponents InE and WtE having a total of five bits. It is within the scope of the present disclosure that the summation circuit 108 is configured to generate each of the exponent sums S[1]-S[N] having other total number of bits based on each of the exponents InE and WtE having other total number of bits. The summation circuit 108 is configured to output the exponent sums S[1]-S[N] to the difference circuit 110 via a data bus (not shown).
[0058] Differential circuit 110 is an electronic circuit, such as an IC, including one or more logic gates L1 (e.g., corresponding to or as part of selector circuit 111) and one or more logic gates B1, each of which is configured to receive an exponent and S[1]-S[N] from addition circuit 108. The one or more logic gates L1 may sometimes be referred to as a selector, and the one or more logic gates B1 may sometimes be referred to as a subtractor. The one or more logic gates L1 are configured to generate a maximum exponent and MaxExp as a data element in operation, the data element having a value equal to the maximum value of the data elements of the exponent and S[1]-S[N] and having a number of bits equal to the number of bits of the data elements of the exponent and S[1]-S[N]. The one or more logic gates L1 are configured to output the maximum exponent and MaxExp to the one or more logic gates B1 and converter circuit 124, as described below.
[0059] The one or more logic gates B1 are configured to generate difference values D[1]-D[N] by subtracting each data element of the exponent sum S[1]-S[N] from the maximum exponent and MaxExp in operation. Therefore, the difference values D[1]-D[N] have a total number N and order of data elements corresponding to the above exponent sum S[1]-S[N] and product P[1]-P[N]. Figure 1In the illustrated embodiment, one or more logic gates B1 are configured to output the difference values D[1]-D[N] to the shift circuit 112 and the control circuit 120 on one or more data buses (not shown). In some embodiments, one or more logic gates B1 are not configured to output the difference values D[1]-D[N] to the multiplier circuits 106, and each of the multiplier circuits 106 is configured to generate each instance P[n] of the product P[1]-P[N] by always performing a multiplication operation. In some other embodiments, one or more logic gates B1 are configured to output the difference values D[1]-D[N] to the multiplier circuits 106, respectively, and each of the multiplier circuits 106 is configured to generate each instance P[n] of the product P[1]-P[N] by selectively performing a multiplication operation based on the corresponding instance D[n].
[0060] Each of the comparator circuits 120 is an electronic circuit, such as an IC, configured to receive one of the corresponding difference values D[1]-D[N] from the difference circuit 110, the difference representing the difference between at least one of the exponent InE or the exponent WtE and the maximum exponent MaxExp. The comparator circuit 120 is configured to compare the received difference values D[1]-D[N] with an exponent and a threshold value (e.g., sometimes referred to as an exponent difference threshold) in operation. The exponent and threshold value may be predefined or preconfigured for a particular machine learning application. The exponent and threshold value may be configured based on the desired accuracy of the MAC operation output.
[0061] In some configurations, the circuit 100 may set the exponent and threshold based on the precision of the mantissa InM or the mantissa WtM (e.g., a portion of the input value) or the format of the input value (e.g., a data element from the input circuit 104). For example, the data elements InDE and WtDE may have an FP16 format, including 1 sign bit, 5 exponent bits, and 10 mantissa bits. The output of the MAC operation (e.g., the output from the converter 118) may have the same or a different format (e.g., an FP32 format, including 1 sign bit, 8 exponent bits, and 23 mantissa bits, or other formats). In this case, the precision may be set to the number of bits (e.g., precision) of the mantissa InM or the mantissa WtM (e.g., 10 mantissa bits).
[0062] In some configurations, the circuit 100 may set the exponent and threshold based on a predetermined round-up value starting from the least significant bit (LSB), for example, by configuring the exponent and threshold to be the number of mantissa bits plus the number of extra bits. For example, referring to the aforementioned example, where the data elements InDE and WtDE may have an FP16 format and the MAC operation output may have an FP32 format, the circuit 100 may set the exponent and threshold to the precision of the data element plus one or more extra bits. In some cases, the extra bits may be predefined. In some other cases, the extra bits may be based on a specific architecture or implementation of the circuit 100 or CIM, where 6 extra bits may be set for a 64-bit MAC CIM and 5 extra bits may be set for a 32-bit MAC CIM. Taking 6 extra bits as an example, the circuit 100 may set the exponent and threshold to 16 (e.g., depending on the specific architecture, the number of mantissa bits associated with the data element is 10 and the extra bits are 6).
[0063] Comparator circuit 120 is configured to generate control signals C[1]-C[N] in operation, the total number N of which corresponds to the total number N of at least one of multiplier circuit 106, summing circuit 108, and / or difference values D[1]-D[N]. The generated control signals C[1]-C[N] may be based on or in accordance with a comparison of difference values D[1]-D[N] with an index and a threshold value. Each comparator circuit 120 may generate a corresponding instance C[n] of control signals C[1]-C[N]. For example, comparator circuit 120 may include one or more components capable of or adapted to perform comparison and generation operations.
[0064] For example, the control circuit 120 may generate a control signal C[n] based on whether the corresponding difference value D[n] satisfies the index and the threshold value (e.g., by performing a comparison). For example, satisfying the index and the threshold value may refer to the difference value D[n] being greater than or equal to the index and the threshold value. The control signal C[n] may be 0 or 1, depending on the result of the comparison. If the difference value D[n] is less than the index and the threshold value, the control circuit 120 may generate a control signal C[n] of 1. If the difference value D[n] is greater than or equal to the index and the threshold value, the control circuit 120 may generate a control signal C[n] of 0. In some configurations, for example, if the difference value D[n] is greater than or equal to the index and the threshold value, the control circuit 120 may generate a control signal C[n] of 1, and if the difference value D[n] is less than the index and the threshold value, the control circuit 120 may generate a control signal C[n] of 0. The control circuit 120 may provide the control signal C[n] to the corresponding multiplier circuit 106 or at least one component of the multiplier circuit 106.
[0065] It should be noted that the variables or values (e.g., exponents and thresholds, input values, formats, etc.) are not limited to the examples provided herein, and the circuit 100 or other devices or components thereof may similarly use other variables or values (e.g., different exponents and thresholds, formats, etc.) to perform MAC operations of floating point numbers using reduced computing resources. In addition, it should be noted that more or fewer components and / or different arrangements of one or more components may be implemented to perform the features, operations, or processes discussed herein.
[0066] In various arrangements, operations of at least one of summing circuit 108, differencing circuit 110, and / or comparator circuit 120 may be performed before, after, or in parallel with multiplier circuit 106. In some arrangements, operations of the respective summing circuit 108, differencing circuit 110, or comparator circuit 120 may be performed sequentially or in parallel.
[0067] Shift circuit 112 is an electronic circuit, such as an IC, including one or more registers and / or logic gates, configured to perform a shift operation on each instance P[n] of the product P[1]-P[N] based on the value of the corresponding instance D[n] of the difference value D[1]-D[N].
[0068] Each instance P[n] of the products P[1]-P[N] is based on the sign and mantissa of a corresponding combination of data elements InDE and WtDE, and each instance D[n] of the difference values D[1]-D[N] is based on the sum of the exponents of the same combination. The shift circuit 112 is configured to right shift each instance P[n] of the products P[1]-P[N] by an amount equal to the corresponding difference value D[n] in operation, thereby generating shifted products SP[1]-SP[N], wherein the sign and mantissa bits are aligned according to the total exponent used to generate the difference values D[1]-D[N]. Based on this alignment, the shift circuit 112 is configured to generate each instance SP[n] of the shifted products SP[1]-SP[N] having the same exponent using the maximum exponent and MaxExp as a baseline.
[0069] To compensate for the right shift operation, shift circuit 112 may add a sign bit instance (zero or one) of each product P[n] as the leftmost bit of the corresponding shifted product SP[n]. The number of sign bit instances added is equal to the right shift amount determined by the corresponding difference D[n].
[0070] exist Figure 1In the illustrated embodiment, multiplier circuit 106 may generate corresponding instances P[n] of products P[1]-P[N] by performing a multiplication operation, as described above. Shift circuit 112 may include one or more shifters to receive products P[1]-P[N] from multiplier circuit 106 and selectively output (e.g., shift) one or more of the shifted products SP[1]-SP[N] to adder circuit 114 based on corresponding differences D[1]-D[N]. For example, in Figure 1 , the shift product output to the adder circuit 114 may include SP[w]-SP[z], where "w" to "z" may be one of integers from 1 to N. In one aspect of the present disclosure, the sum of the number of SP[w]-SP[z] may be equal to N. In another aspect of the present disclosure, the sum of the number of SP[w]-SP[z] may be less than N.
[0071] The shift circuit 112 (eg, a shifter) may be configured to shift the corresponding difference values among the difference values D[1]-D[N] to a difference threshold value ( Figure 1 The difference threshold can be configured based on the distribution of the difference values D[1]-D[N]. In an example where the difference values D[1]-D[N] present a normal distribution, the difference threshold can be determined to be one standard deviation below the mean of the normal distribution. In another example where the difference values D[1]-D[N] still present a normal distribution, the difference threshold can be determined to be two standard deviations below the mean of the normal distribution. In another example where the difference values D[1]-D[N] still present a normal distribution, the difference threshold can be determined to be any standard deviation below the mean of the normal distribution.
[0072] When any difference value (e.g., D[n], where n is an integer between 1 and N) is equal to or less than a difference threshold value (sometimes referred to as a “small exponent difference value”), the shift circuit 112 (e.g., a shifter) may be deactivated to prevent the corresponding shifted product SP[n] from being received by the adder circuit 114 (e.g., the corresponding product P[n] is not shifted or separated from the adder circuit 114). Equivalently, when any difference value, such as D[n], is greater than the difference threshold value (sometimes referred to as a “normal exponent difference value”), the shift circuit 112 may be activated to output the corresponding shifted product SP[n] to the adder circuit 114.
[0073] In other words, the shift circuit 112 may shift any of the products P[1]-P[N] and output the shifted products SP[1]-SP[N] to the adder circuit 114 based on comparing the respective difference values D[1]-D[N] to the difference threshold. Thus, the sum of the numbers SP[w]-SP[z] may be equal to N. In some configurations, the shift circuit 112 may detect that at least one of the products P[1]-P[N] from the multiplier circuit 106 is zero. In this case, the shift circuit 112 may not shift the corresponding product to a zero value and / or output the product to the adder circuit 114. Thus, the sum of the numbers SP[w]-SP[z] may be less than N.
[0074] In addition, to generate SP[w]-SP[z], shift circuit 112 may right shift each instance P[n] of the product P[w]-P[z] by an amount equal to the corresponding difference DA[n], thereby aligning the sign and mantissa bits according to the summed exponent. In some embodiments, difference DA[n] may be generated (e.g., by difference circuit 110) based on subtracting each data element of sum S[w]-S[z] from the maximum exponent and MaxExp. The maximum exponent and MaxExp may correspond to the maximum value of the data elements of sum S[w]-S[z]. Based on this alignment, shift circuit 112 may use the maximum exponent and MaxExp as a baseline to generate each instance SP[n] of shifted product SP[w]-SP[z] having the same exponent.
[0075] When any difference value (e.g., D[n], where n is an integer between 1 and N) is equal to or less than a difference threshold (sometimes referred to as a "small exponent difference"), shift circuit 112 may be disabled to prevent the corresponding (e.g., shifted) product SP[n] from being received by adder circuit 114. In some embodiments, products P[n] having such large exponent differences may be ignored.
[0076] In other words, the shift circuit 112 may shift all or part of the products P[1]-P[N], and selectively output corresponding products of the shifted products SP[1]-SP[N] to the adder circuit 114 according to the comparison result of the respective differences D[1]-D[N] with the difference threshold. In this way, the sum of the number of SP[w]-SP[z] output by the shift circuit 112 may be less than or equal to N. When one or more products P[1]-P[N] are ignored (e.g., their respective exponent differences D[n] are equal to or greater than the difference threshold), the sum is less than N; and when no products P[1]-P[N] are ignored, the sum is equal to N.
[0077] In some embodiments, multiplier circuit 106 may receive difference D[1]-D[N] from difference circuit 110 to determine whether difference D[n] is greater than or equal to an exponent and a threshold (eg, sometimes referred to as an exponent difference threshold).
[0078] In some embodiments, for example, embodiments where the data elements InDE and WtDE have a BF16 format, the shift circuit 112 is configured to generate each shift product having a total of 21 bits, for example, SP[0]-SP[N], based on each product P[0]-P[N] having a total of 17 bits. In some embodiments, for example, embodiments where the data elements InDE and WtDE have a FP16 format, the shift circuit 112 is configured to generate each shift product having a total of 27 bits, for example, SP[0]-SP[N], based on each product P[0]-P[N] having a total of 23 bits. It is within the scope of the present disclosure for the shift circuit 112 to generate each shift product SP[0]-SP[N] having other total number of bits, based on each product P[0]-P[N] having other total number of bits.
[0079] Based on the product P[0]-P[N] having a two's complement format, the shift circuit 112 is configured to generate a shift product having a two's complement format, such as SP[0]-SP[N]. As described above, Figure 1 In the example shown, the shift circuit 112 is configured to output the shifted products SP[w]-SP[z] to an adder circuit (tree) 114 on a data bus (not shown).
[0080] The adder tree 114 is an electronic circuit, such as an IC, including multiple layers of one or more logic gates (not shown), such as described above with respect to one or more logic gates A1 (logic gates of the summing circuit 108). For example, the adder tree 114 may include: a first layer, which is configured to receive the shifted product SP[w]-SP[z]; and a last layer, which is configured to generate the sum 115 as a data element corresponding to the sum of the shifted products SP[w]-SP[z]. In some embodiments, each of the one or more consecutive layers between the first layer and the last layer is configured to receive a first number of sum data elements generated by the previous layer, and to generate a second number of sum data elements based on the first number of sum data elements, the second number being half of the first number. Therefore, the total number of layers includes the first layer and the last layer and each consecutive layer (if any).
[0081] In some embodiments, the sum PSTC (e.g., corresponding to the sum 115) is sometimes referred to as a partial sum PSTC or a mantissa sum PSTC, and the total number of bits corresponds to the number of bits and the number of data elements of the shift product SP[w]-SP[z]. In some embodiments, the number of bits of the sum PSTC is equal to the number of bits of the shift product SP[w]-SP[z] plus the number of bits that can represent the number of data elements of the shift product SP[w]-SP[z]. In some embodiments, the number of bits of the sum PSTC is equal to the number of bits of the shift product SP[w]-SP[z] plus 4 bits that can represent the 16 data elements of the shift product SP[w]-SP[z].
[0082] In some embodiments, for example, embodiments where the data elements InDE and WtDE have a BF16 format, the adder tree 114 is configured to generate a sum PSTC having a total of 25 bits based on each shift product SP[w]-SP[z] having a total of 21 bits. In some embodiments, for example, embodiments where the data elements InDE and WtDE have a FP16 format, the adder tree 114 is configured to generate a sum PSTC having a total of 31 bits based on each shift product SP[w]-SP[z] having a total of 27 bits. It is within the scope of the present disclosure for the adder tree 114 to generate a sum PSTC based on each shift product SP[w]-SP[z] having other total number of bits.
[0083] According to various embodiments of the present disclosure, based on the shift product SP[w]-SP[z] having a two's complement format, the adder tree 114 is configured to generate a sum PSTC having a two's complement format. Therefore, the adder tree 114 is configured to output the sum PSTC to the converter 116 on the data bus (not shown). In some other embodiments, the adder tree 114 may output the sum PSTC to a circuit (not shown) outside the circuit 100.
[0084] Converter 116 is an electronic circuit, such as an IC, including logic circuits, configured to receive the sum PSTC from adder tree 114 in operation and convert the sum PSTC from two's complement to a sum PSSM having a sign-plus-mantissa format. Converter 116 is configured to generate a sum PSSM having the same number of bits as the sum PSTC. Figure 1 In the illustrated embodiment, converter 116 is configured to further output the sum PSSM to converter 118 on a data bus (not shown). In some other embodiments, converter 116 may output the sum PSSM to a circuit external to circuit 100 (not shown).
[0085] The converter 118 is an electronic circuit, such as an IC, including logic circuits, configured to receive the sum PSSM from the converter 116 and the maximum exponent and MaxExp from the differential circuit 110 in operation, and convert the sum PSSM from a sign-plus-mantissa format to a sum PS, which has an output format based on the sum PSSM and MaxExp and is different from the sign-plus-mantissa format, such as a floating point format as described above. In various embodiments of the present disclosure, the converter 118 can generate a sum PS that is configured to be compatible with a circuit (not shown) external to the circuit 100. For example, the converter 118 is configured to output the sum PS to a circuit (not shown) external to the circuit 100, such as a memory array as part of a convolutional neural network (CNN) or other instances of the circuit 100. In some arrangements, the converter 116 can be a part of the converter 118, and vice versa. The MUX 122 can be located between the converter 116 and the converter 118, so that the MUX 112 can receive the output from the converter 116 and provide an output to the converter 118.
[0086] Figure 2 A block diagram of a portion (hereinafter referred to as "configurable circuit" 200) of an example data computation circuit (e.g., data computation circuit 100) according to some embodiments of the present disclosure is shown. Configurable circuit 200 may include adder circuit 214, control circuit 220, and MUX 222, which may be substantially similar to or include features of adder circuit 114, control circuit 120, and MUX 122, respectively. In brief overview, adder circuit 214 may receive a partial sum (psum) and provide an output to MUX 222 via an internal output bus. MUX 222 may receive an output from adder circuit 214 and output a result of a MAC operation. Control circuit 220 may provide signal 221 to adder circuit 214 and MUX 222 to configure adder circuit 214 and MUX 222. In some embodiments, adder circuit 214 may be configured to have different configurations for different amounts of accumulation, as described in more detail below. For example, adder circuit 214 can be configured to support accumulation in point-wise convolutional layers (e.g., a high number of accumulations, 16, 32, 64, etc.). Adder circuit 214 can be configured to support accumulation in depth-wise convolutional layers (e.g., a low number of accumulations, 8, etc.). In some embodiments, MUX 222 can be configured to output the results of the MAC operation for different numbers of accumulations (e.g., 8, 16, 32, 64, etc.). Figure 2 The configurable circuit 200 shown in FIG. 2 is a non-limiting example.
[0087] Figure 3 Schematic diagram of an example configurable circuit 300 according to some embodiments of the present disclosure is shown. Figure 3 , configurable circuit 300 is shown to include adder circuit 314 and MUX 322, which may be substantially similar to or include features of adder circuit 214 and MUX 222, respectively. Figure 2 The configurable circuit 200 shown in FIG. 2 is a non-limiting example.
[0088] As shown, adder circuit 314 may receive partial sums psum0–psum63 and perform an addition operation on the received psum. Adder circuit 314 may provide the result of the addition operation to MUX 322, which may output the result of the MAC operation. In some embodiments, adder circuit 314 may receive signal 321 (e.g., from control circuit 220) and may be configured to support different numbers of accumulations. For example, adder circuit 314 may receive signal 321 (e.g., 16A_EN) indicating 16 accumulations (16A) and may then be configured to provide 4 16A (16A×4) results without performing the next addition operation (e.g., 32A). MUX 322 may receive signal 321 (e.g., 16A_EN) indicating 16A and may receive the result of 16A×4 from the corresponding adder (e.g., 16A is performed). MUX 322 may output the result of the MAC based on the received 16A×4 result. Likewise, the adder circuit 314 may receive a signal 321 (e.g., 32A_EN) indicating 32A accumulations (32A), and may then be configured to provide two 32A (32A×2) results without performing the next addition operation (e.g., 64A). The MUX 322 may receive a signal 321 (e.g., 32A_EN) indicating 32A, and may receive a result of 32A×2 from a corresponding adder (e.g., perform 32A). The MUX 322 may output a result of the MAC based on the received result of 32A×2. Likewise, the adder circuit 314 may receive a signal 321 (e.g., 64A_EN) indicating 64A accumulations (64A), and may then be configured to provide a result of 1 64A (64A×1). The MUX 322 may receive a signal 321 (e.g., 64A_EN) indicating 64A, and may receive a result of 64A×1 from a corresponding adder (e.g., perform 64A). MUX 322 may output a MAC result based on the received 64A×1 result. This allows for configurable accumulation times (e.g., configurable between various accumulation times), thereby improving CIM utilization and taking precautions against multipliers to reduce computation / computational resources / power usage of MAC operations.
[0089] In some embodiments, Figure 3As shown, the configurable circuit 300 (e.g., adder circuit 314) can receive a plurality of input data bits (e.g., psum) as input. In response to the receipt of the input data bits, the configurable circuit 300 can identify the number of accumulations associated with the received input data bits. For example, based on the input data bits (e.g., psum), the configurable circuit 300 can determine the number of accumulations to be performed. In some embodiments, based on the number of accumulations, the configurable circuit 300 can determine whether to enable or disable at least one component of the adder circuit 314. For example, when the number of accumulations is determined to be 16A×4, the configurable circuit 300 can determine to disable circuit components that perform 32A×2 and 64A×1 addition operations, thereby allowing the result of the 16A×4 addition operation to be provided to the MUX 322. Similarly, when the number of accumulations is determined to be 32A×2, the configurable circuit 300 can determine to disable circuit components that perform 64A×1 addition operations, thereby allowing the result of the 32A×2 addition operation to be provided to the MUX 322. When the number of accumulations is determined to be 64A×1, the configurable circuit 300 may determine to enable or disable circuit components that perform 16A×4, 32A×2, and 64A×1 addition operations. In some embodiments, the signal 321 may include an indication for enabling or disabling at least one component of the configurable circuit 300. For example, when the number of accumulations is determined to be 16A×4, the configurable circuit 300 may generate a signal 321 indicating 16A_EN and disable components that perform 32A×2 and 64A×1 addition operations. Similarly, when the number of accumulations is determined to be 32A×2, the configurable circuit 300 may generate a signal 321 indicating 32A_EN and disable components that perform 64A×1 addition operations. Similarly, when the number of accumulations is determined to be 64A×1, the configurable circuit 300 may generate a signal 321 indicating 64A_EN and enable components that perform 16A×4, 32A×2, and 64A×1 addition operations.
[0090] In some embodiments, MUX 322 may receive a signal 321 indicating the number of accumulations, and may be configured to output the result of the MAC operation according to the number of accumulations. For example, when MUX 322 receives a signal indicating 16A_EN, MUX 322 may provide four results as outputs of a MAC operation of 16A×4. Similarly, when MUX 322 receives a signal indicating 32A_EN, MUX 322 may provide two results as outputs of a MAC operation of 32A×2. Similarly, when MUX 322 receives a signal indicating 64A_EN, MUX 322 may provide one result as output of a MAC operation of 64A×1.
[0091] Figure 4AA schematic diagram of an example adder circuit 414 is shown according to some embodiments of the present disclosure. The adder circuit 414 may be substantially similar to the adder circuit 214 or include features of the adder circuit 214. Figure 4A The adder circuit 414 shown in FIG. 4 is a non-limiting example. Figure 4B Example states of components (eg, N+4-bit adder, N+3-bit adder, etc.) in adder circuit 414 for different numbers of accumulations (eg, 16A, 32A, 64A, etc.) according to some embodiments of the present disclosure are listed.
[0092] In some embodiments, when the adder circuit 414 receives a signal indicating 16A accumulation, at least the adders of 16A (e.g., up to N+2-bit adder 416C, including N+1-bit adder, N-bit adder, etc.) can be enabled (e.g., set to "1") via 16A_EN, while the adders of 32A and 64A (e.g., N+3-bit adder 416B and N+4-bit adder 416A) can be disabled (e.g., set to "0"). This allows the result of the addition operation (e.g., 16A×4) to be output at the N+2-bit adder 416C (e.g., to the MUX 222). When the adder circuit 414 receives a signal indicating 32A, at least the adders of 16A and 32A (e.g., up to N+3-bit adder 416B, including N+2-bit adder 416C, N+1-bit adder, N-bit adder, etc.) can be enabled (e.g., set to "1") through 32A_EN and 16A_EN, while the adder of 64A (e.g., N+4-bit adder 416A) can be disabled (e.g., set to "0"). This allows the result of the addition operation (e.g., 32A×2) to be output at the N+3-bit adder 416B (e.g., to MUX 222). When the adder circuit 414 receives a signal indicating 64A, the adders of 16A, 32A, and 64A (e.g., up to N+4-bit adder 416C, including N+3-bit adder 416B, N+2-bit adder 416C, N+1-bit adder, N-bit adder, etc.) can be enabled (e.g., set to "1") via 64A_EN, 32A_EN, and 16A_EN. This allows the result of the addition operation (e.g., 64A×1) to be output at the N+4-bit adder 416A (e.g., to MUX 222).
[0093] Adder circuit 414 and according to Figure 4A and Figure 4BThe states of the adding components of the different signals shown in are non-limiting examples, and the number of accumulations and / or the number of MAC outputs are not limited to 16, 32, or 64. That is, in some embodiments, the configurable circuit disclosed herein can be used for any number of accumulations, so that the first output of the adder circuit 414 can include a first group (e.g., 2) of output bits (e.g., 32A) when a first number (e.g., 1) of components are disabled, and the second output can include a second group (e.g., 1) of output bits (e.g., 64A) when a second number (e.g., 0) is disabled, wherein the number of the first group is greater than the number of the second group, and the first number is greater than the second number. Although described as 32A×2 and 64A×1, the configurable circuit disclosed herein can be used for any number (e.g., 128) of accumulations.
[0094] Figure 5 An example diagram showing signals associated with adder circuit 414 according to some embodiments of the present disclosure. In some embodiments, adder circuit 414 may receive a signal indicating different accumulation times during different cycles. For example, the signal may include a first logic value at a first time, and a second logic value at a second time. Figure 5 During a first cycle 551 (eg, a first time), the adder circuit 414 may receive a signal indicating 64A (eg, based on Figure 4B , 16A, 32A, and 64A are set to "1" (enabled). The adder circuit 414 can perform an addition operation on psum from psum0 to psum63 and then output the result. (eg, 64A×1). During a second period 552 (eg, a second time), the adder circuit 414 may receive a signal indicating 32A (based on Figure 4B , for example, 16A and 32A are set to "1" (enabled); 64A is set to "0" (disabled). The adder circuit 414 can perform an addition operation on the first group of psums (from psum0 to psum31) and the second group of psums (from psum32 to psum63), and then output the result (e.g., 2 MAC outputs), and (eg, 32A×2). During a third period 553 (eg, a third time), the adder circuit 414 may receive a signal indicating 16A (eg, based on Figure 4B, 16A is set to "1" (enabled); 32A and 64A are set to "0" (disabled). The adder circuit 414 may perform an addition operation on the first group of psums (from psum0 to psum15), the second group of psums (from psum16 to psum31), the third group of psums (from psum32 to psum47), and the fourth group of psums (from psum48 to psum63), and then output the result (e.g., 4 MAC outputs). and (For example, 16A×4).
[0095] Fig. 6A An example logic circuit 601 is shown that can be coupled to the configurable circuit 200 according to some embodiments of the present disclosure. In some embodiments, the logic circuit 601 can receive a signal (e.g., signal 221) from a control circuit (e.g., control circuit 220). In response to receiving the signal, the logic circuit 601 can control the adder circuit 214. In some embodiments, the logic circuit 601 can include a decoder 602 to read / decode the signal from the control circuit and can provide a signal to enable / disable the adder (e.g., Figure 4A In some embodiments, logic circuit 601 may generate a plurality of logic values and / or logic patterns, each logic value and / or logic pattern indicating a number of adders to be disabled or enabled.
[0096] Figure 6BExample modes of the number of accumulations and corresponding control signals according to some embodiments of the present disclosure are listed. In some embodiments, a control circuit (e.g., control circuit 220) may generate a control signal (e.g., signal 221) that may represent four modes (e.g., mode [1:0]). For example, when the control signal represents mode "11", the control circuit may decode the control signal and identify the number of accumulations as "64". In response to identifying the number of accumulations, the control circuit may generate a control signal to set the adders of 16A, 32A, and 64A to "1" (enabled), thereby performing addition operations up to 64A and outputting a result of 64A×1. Similarly, when the control signal represents mode "10", the control circuit may decode the control signal and identify the number of accumulations as "32". In response to identifying the number of accumulations, the control circuit may generate a control signal to set the addition components of 16A and 32A to "1" (enabled), and set the addition component of 64A to "0" (disabled), thereby performing addition operations up to 32A and outputting a result of 32A×2. Similarly, when the control signal represents the mode "01", the control circuit can decode the control signal and recognize that the number of accumulations is "16". In response to recognizing the number of accumulations, the configurable circuit can generate a control signal to set the addition component of 16A to "1" (enabled) and set the addition components of 32A and 64A to "0" (disabled), thereby performing addition operations up to 16A and outputting a result of 16A×4. Similarly, when the control signal represents the mode "00", the control circuit can decode the control signal and recognize that the number of accumulations is "8". In response to recognizing the number of accumulations, the configurable circuit can generate a control signal to set the addition components of 16A, 32A and 64A to "0" (disabled), thereby performing addition operations up to 8A and outputting a result of 8A×8. The logic circuit 601 may include various logic components to decode signals from the control circuit and provide control signals to enable / disable the adder. Figure 6C It shows that some embodiments of the present disclosure can be used with Figure 2 6. The example logic component 651 of the configurable circuit coupling shown in FIG. For example, the logic circuit 601 may include at least one of an OR gate, an AND gate, a NOR gate, a NAND gate, an XOR gate, a NOT gate, or any combination thereof.
[0097] Fig. 7A A block diagram of an example configurable circuit 700 is shown according to some embodiments of the present disclosure. Figure 7B List some embodiments according to the present disclosure Fig. 7A Example control signals and corresponding outputs of the MUX shown in FIG. Fig. 7A, configurable circuit 700 is shown to include adder circuit 714 and MUX 722, which may be substantially similar to or include features of adder circuit 214 and MUX 222, respectively. Fig. 7A The configurable circuit 700 shown in FIG. 7 is a non-limiting example.
[0098] The adder circuit 714 may include different bit adders, including a 16-bit adder 714A, a 17-bit adder 714B, an 18-bit adder 714C, a 19-bit adder 714D, a 20-bit adder 714E, and a 21-bit adder 714F. Each different bit adder may be configured to provide an accumulation of the input as an output. The adder circuit 714 may receive a partial sum (psum) (e.g., 64 psums) and perform an addition operation through at least one different bit adder.
[0099] Figure 7C 1 shows a block diagram of an adder circuit 714 according to some embodiments of the present disclosure. Fig. 7A 7 is a non-limiting example of an adder circuit 714 including bit adders 714A-714F, but Figure 7B The adder circuit 714 shown in FIG. 7 may include a plurality of adders (e.g., 731A, ..., 731N-1, 731N, etc.), where N may be any number of adders. For example, adder 731A may be Fig. 7A The 16-bit adder 714A and adder 731N can be Fig. 7A Each adder, such as 731N-1, can be configured to receive input A n-1 and B n-1 , and can be configured to output the result of the addition operation according to the control signal. When the adder circuit 714 receives a control signal including a first logic value (e.g., "1" or an enable signal) associated with the next adder (e.g., 731N), the adder 731N-1 can provide the result of the addition operation as an input (carry, "CI") to the next adder 731N. When the adder circuit 714 receives a control signal including a second logic value (e.g., "0" or a disable signal) associated with the next adder (e.g., 731N), the adder 731N-1 can provide the result of the addition operation (S) as an input (carry, "CI") to the next adder 731N. n-1 ) is provided to MUX 722.
[0100] Fig.7D List some embodiments according to the present disclosure Fig. 7A. In some embodiments, as shown, a 16-bit adder 714A may have an input bit width (and number of accumulations) of 16, and an output bit width of 17. Similarly, a 17-bit adder 714B may have an input bit width (and number of accumulations) of 17, and an output bit width of 18; an 18-bit adder 714C may have an input bit width (and number of accumulations) of 18, and an output bit width of 19; a 19-bit adder 714D may have an input bit width (and number of accumulations) of 19, and an output bit width of 20; a 20-bit adder 714E may have an input bit width (and number of accumulations) of 20, and an output bit width of 21; and a 21-bit adder 714F may have an input bit width (and number of accumulations) of 21, and an output bit width of 22.
[0101] refer to Fig. 7A In some embodiments, a first component (e.g., 19-bit adder 714D) can be configured to receive a plurality of input data bits (e.g., psum from 18-bit adder 714C) and provide a first output (e.g., 20b(16A_out1-3)). When the adder circuit 714 receives a control signal associated with a second component or a next adder (e.g., 20-bit adder 714E) including a first logic value (e.g., “1” and / or an enable signal; e.g., 32A_EN or 64A_EN to enable 32A), the second component can receive the first output from the first component and provide a second output (e.g., 21b(32A_out0-1)). When the adder circuit 714 receives a control signal associated with a second component or a next adder (e.g., 20-bit adder) including a second logic value (e.g., “0”; e.g., 16A_EN to disable 32A), the second component can be disabled and the first output from the first component can be provided to MUX 722. Thus, MUX 722 may be configured to output a first output in response to a control signal including a second logic value (associated with 20-bit adder 714E), and configured to output a second output in response to a control signal including a first logic value (associated with 20-bit adder 714E).
[0102] Similarly, a second component (e.g., a 20-bit adder 714E) can be configured to receive a plurality of input data bits (e.g., psum from a 19-bit adder) and provide a second output (e.g., 21b(32A_out0-1)). When the adder circuit 714 receives a control signal associated with a third component or a next adder (e.g., a 21-bit adder) including a first logic value (e.g., “1” and / or an enable signal; e.g., 64A_EN to enable 64A), the third component can receive the second output from the second component and provide a third output (e.g., 22b(64A_out0)). When the adder circuit 714 receives a control signal associated with a third component or a next adder (e.g., a 21-bit adder) including a second logic value (e.g., “0”; e.g., 32A_EN to disable 64A), the third component can be disabled and the second output from the second component can be provided to the MUX 722. Thus, MUX 722 may be configured to output a second output in response to a control signal comprising a second logic value (associated with a 21-bit adder) and to output a third output in response to a control signal comprising a first logic value (associated with a 21-bit adder).
[0103] In some embodiments, MUX 722 may be configured to receive different bit groups (e.g., 20b×4, 21b×2, 22b×1, etc.) from different adder groups (e.g., 19-bit adder 714D, 20-bit adder 714E, 21-bit adder 714F, etc.). In response to receiving bits from the adders, MUX 722 may be configured to output the result of the MAC operation corresponding to the received bits. For example, when MUX 722 receives 20b×4 and a signal indicating a corresponding number of accumulations (e.g., 16A) from 19-bit adder 714D, MUX 722 may provide MAC operation outputs 16A_out0, 16A_out1, 16A_out2, and 16A_out3. When MUX 722 receives 21b×2 and a signal indicating a corresponding number of accumulations (e.g., 32A) from 20-bit adder 714E, MUX 722 may provide outputs 32A_out0 and 32A_out1 of the MAC operation. When MUX 722 receives 22b×1 and a signal indicating a corresponding number of accumulations (e.g., 64A) from 21-bit adder 714F, MUX 722 may provide an output 64A_out0 of the MAC operation. In some embodiments, when the number of bits from the adder is less than the number of MUX output bits, MUX 722 may be configured to set at least one of the output bits to a logic state (e.g., “0”). For example, when MUX 722 is configured to output 80 bits (80b as shown in the figure), and MUX 722 receives bits (e.g., two 21 bits) from 20-bit adder 714E, MUX 722 can provide 80-bit output, including 42 bits (32A_out0, 32A_out1) and 38 bits of "0" from 20-bit adder 714E. Similarly, when MUX 722 is configured to output 80 bits (80b as shown in the figure), and MUX 722 receives bits (e.g., one 22 bit) from 21-bit adder 714F, MUX 722 can provide 80-bit output, including 22 bits (64A_out0) and 58 bits of "0".
[0104] Figure 8 An example selection circuit 800 that can be coupled to the configurable circuit 200 according to some embodiments of the present disclosure is shown. In some embodiments, the selection circuit 800 can be coupled to the MUX 222 to receive the result of the addition operation from the adder circuit 214 and provide the result to the MUX 222. In some embodiments, the selection circuit 800 can include a first circuit 801 to receive and output a first group of bits (e.g., 64A×1) in response to enabling the corresponding adder (e.g., for 16A, 32A, 64A). The selection circuit 800 may include a second circuit 802 to receive and output a second group of bits (e.g., 32A×2) in response to disabling at least one adder (e.g., for 64A).
[0105] As shown, in some examples, selection circuit 800 may include multiple circuit components (e.g., switches, transistors, etc.) to receive a set of bits from adder circuit 214 and output it to MUX 222. For example, selection circuit 800 may receive control signals (e.g., signal 221 from control circuit 220), such as 32A_EN and 32_ENB, and select a first circuit or a second circuit to provide the received bits to MUX 222. Although described with respect to addition operations of 32A and 64A, selection circuit 800 may be used for any number of accumulations (e.g., 8A, 16A, 32A, 64A, etc.).
[0106] Fig.9A An example circuit 901 is shown that can be coupled to the configurable circuit 200 according to some embodiments of the present disclosure. In some embodiments, the circuit 901 may include a D flip-flop (DFF) 905 and an adder 914. The adder 914 may be substantially similar to or include features of the adder circuit 214 or the adder therein. For example, the adder 914 may be any of a 16-bit adder, a 17-bit adder, an 18-bit adder, a 19-bit adder, a 20-bit adder, a 21-bit adder in the adder circuit 714, or a combination thereof. The DFF 905 may be configured to store binary data during an addition operation of the adder 914. In some embodiments, the DFF 905 may synchronize and store intermediate results as data passes through the adder 914 during the addition operation. The output of the adder 914 can be transmitted back to the DFF 905 through the first path 930 as the input of the DFF 905, and can be added by the adder 914 to the previously stored data in the DFF 905. The circuit 901 can be configured to repeat (e.g., N cycles) the addition operation through the first path 930 until a predetermined condition (e.g., a loop requirement) is met. When the loop requirement is met, the adder 914 can provide an output through the second path 940.
[0107] Fig. 9BAn example diagram of signals associated with circuit 901 according to some embodiments of the present disclosure is shown. In some embodiments, N accumulations require N cycles. For example, 8 accumulations require 8 cycles, 16 accumulations require 16 cycles, 32 accumulations require 32 cycles, and 64 accumulations require 64 cycles. That is, for example, when the adder 914 performs an addition operation of 8 accumulations, the adder 914 and the DFF 905 can be configured to perform 8 addition operations through the first path 930, and then output the result of the addition operation through the second path 940, thereby calculating P0+P1+P2+P3+P4+P5+P6+P7. Similarly, when the adder 914 performs an addition operation of 16 accumulations, the adder 914 and the DFF 905 can be configured to perform 16 addition operations through the first path 930, and then output the result of the addition operation through the second path 940, thereby calculating P0+…+P14+P15. Likewise, when the adder 914 performs 64 cumulative addition operations, the adder 914 and the DFF 905 may be configured to perform 64 addition operations through the first path 930 and then output the result of the addition operation through the second path 940, thereby calculating P0+…+P62+P63.
[0108] Fig.10 1 is a flow chart showing an example method 1000 of operating a configurable circuit according to various embodiments. The example method 1000 may be performed by the circuit 200 or one or more components of the circuit 200. Thus, it may be combined with but not limited to Figure 1 -9 to describe the following embodiments of method 1000. The illustrated embodiments of method 1000 are provided as examples and do not limit the scope of the present disclosure. Therefore, it should be understood that any of the various operations of method 1000 may be omitted, reordered, and / or added without departing from the scope of the present disclosure.
[0109] In brief overview, method 1000 may begin at operation 1010, where a plurality of input data bits are received into a computing circuit. Method 1000 may continue at operation 1020, where an accumulation number associated with the plurality of input data bits is identified. Method 1000 may continue at operation 1030, where a determination is made whether to enable or disable at least one component of the computing circuit based on the accumulation number. Method 1000 may continue at operation 1040, where a control signal is generated to enable or disable at least one component of the computing circuit based on the determination to enable or disable.
[0110] At operation 1010, a computing circuit (eg, configurable circuit 200) may receive a plurality of input data bits (eg, Figure 2 For example, a first component of the computation circuit (eg, 19-bit adder 714D) may receive 64 psums.
[0111] At operation 1020, the computation circuitry may identify an accumulation number (e.g., 8A, 16A, 32A, 64A, etc.) associated with the plurality of input data bits. In some embodiments, the computation circuitry may determine an addition operation (e.g., 8A, 16A, 32A, 64A, etc.) to be performed based on the received input data bits (e.g., psum). In some embodiments, the computation circuitry may be configured to determine whether to perform the addition operation on a point-wise convolutional layer (e.g., a high number of accumulations, 16, 32, 64, etc.) or a depth-wise convolutional layer (e.g., a low number of accumulations, 8, etc.).
[0112] At operation 1030, the computation circuitry may determine whether to enable or disable at least one component of the computation circuitry (e.g., at least one of the N+2-bit adder 416C, the N+3-bit adder 416B, the N+4-bit adder 416A, etc.). At operation 1040, based on the determination of enabling or disabling, the computation circuitry may generate a control signal (e.g., signal 221) to enable or disable at least one component of the computation circuitry. For example, when the control signal indicates a first accumulated amount (e.g., 16A), the computation circuitry may disable at least one component (e.g., Figure 4A For example, when the control signal indicates the second accumulated number (e.g., 32A), the computing circuit may disable at least one component (e.g., Figure 4A For example, when the control signal indicates the third accumulated number (e.g., 64A), the computing circuit may enable at least one component (e.g., Figure 4A N+2-bit adder 416C, N+3-bit adder 416B, and N+4-bit adder 416A).
[0113] Fig.11 1 is a flow chart showing an example method 1100 of operating a configurable circuit according to various embodiments. The example method 1100 may be performed by the circuit 200 or one or more components of the circuit 200. Thus, it may be combined with but not limited to Figure 1 -9 to describe the following embodiments of method 1100. The illustrated embodiments of method 1100 are provided as examples and do not limit the scope of the present disclosure. Therefore, it should be understood that any of the various operations of method 1100 may be omitted, reordered, and / or added while remaining within the scope of the present disclosure.
[0114] In brief overview, method 1100 may begin at operation 1110 by receiving 64 psums. Method 1100 may continue to operation 1120 by identifying an accumulation quantity associated with the received psums and determining an accumulation mode. Method 1100 may continue to operation 1130 by summing the received psums. Method 1100 may continue to operation 1140 by generating an output of a MAC operation based on the summed psums.
[0115] At operation 1110, an adder circuit (e.g., adder circuit 214) may receive 64 psums. At operation 1120, a control circuit (e.g., control circuit 220) may identify and determine the number of accumulations associated with the received 64 psums (e.g., 8A, 16A, 32A, 64A, etc.). In some embodiments, the control circuit may generate control signals representing four modes (e.g., Figure 6B As shown in the figure, each mode indicates the adder to be enabled / disabled. For example, when the control signal indicates 64A, the adders of 16A, 32A, and 64A can be enabled (e.g., by 16A_EN, 32A_EN, and 64A_EN being set to "1"). When the control signal indicates 32A, the adders of 16A and 32A can be enabled (e.g., by 16A_EN and 32A_EN being set to "1"), and the adder of 64A can be disabled (e.g., set to "0"). When the control signal indicates 16A, the adder of 16A can be enabled (e.g., by 16A_EN being set to "1"), and the adders of 32A and 64A can be disabled (e.g., set to "0").
[0116] At operations 1130A-C, the psums may be summed according to a control signal indicating the number of accumulations. When the control signal indicates 64A, all 64 psums may be added together to generate one output of a MAC operation at operation 1140A. When the control signal indicates 32A, two groups of 32 psums may be added together to generate two outputs of a MAC operation at operation 1140B. When the control signal indicates 16A, four groups of 16 psums may be added together to generate four outputs of a MAC operation at operation 1140C. Although not shown, when the control signal indicates 8A, eight groups of eight psums may be added together to generate eight outputs of a MAC operation.
[0117] In one aspect of the present disclosure, a circuit is disclosed. The system includes a computing circuit, a memory array operably coupled to the computing circuit, and a controller, the controller being configured to input a plurality of input data bits to the computing circuit, identify an accumulated quantity associated with the plurality of input data bits, determine whether to enable or disable at least one component of the computing circuit based on the accumulated quantity, and generate a control signal to enable or disable at least one component of the computing circuit based on the determination of enabling or disabling.
[0118] In some embodiments, the computation circuit includes an adding component, and wherein the controller is further configured to enable or disable the adding component.
[0119] In some embodiments, the system further comprises a multiplexer configured to: receive a plurality of bits from a plurality of components of a computational circuit including at least one component; and configure output bits based on the accumulated number and the plurality of bits.
[0120] In some embodiments, the controller is configured to set at least one of the output bits to a first logic state when a number of the plurality of bits from the plurality of components is less than a number of the output bits.
[0121] In some embodiments, the output bits include: a first group of output bits when a first number of components of the computing circuit are disabled; and a second group of output bits when a second number of components of the computing circuit are disabled, wherein the number of the first group is greater than the number of the second group, and the first number is greater than the second number.
[0122] In some embodiments, the system further comprises: a first circuit that outputs a first set of bits in response to enabling at least one component of the computational circuit; and a second circuit that outputs a second set of bits in response to disabling at least one component of the computational circuit.
[0123] In some embodiments, the system further comprises at least one logic circuit component associated with a plurality of logic values, each logic value indicating a quantity of at least one component of the computing circuit to be disabled. In another aspect of the present disclosure, a device is disclosed. The device comprises a memory array, a computing circuit operably coupled to the memory array, the computing circuit comprising: a first component configured to receive a plurality of input data bits and provide a first output in response to a control signal; a second component configured to receive a first output from the first component and provide a second output in response to a control signal comprising a first logic value; and a multiplexer configured to output the first output in response to a control signal comprising a second logic value, and configured to output the second output in response to a control signal comprising the first logic value.
[0124] In some embodiments, at least one of the first component or the second component is an adder configured to provide an accumulation as an output.
[0125] In some embodiments, the second component is disabled in response to a control signal including the second logic value.
[0126] In some embodiments, the first digit of the first component is less than the second digit of the second component.
[0127] In some embodiments, the multiplexer includes: a first circuit that outputs a first group of bits in response to a control signal including the first logic value; and a second circuit that outputs a second group of bits in response to a control signal including the second logic value, wherein a first group number within the first group is greater than a second group number within the second group.
[0128] In some embodiments, the control signal includes a first signal including the first logic value at a first time and a second signal including the second logic value at a second time.
[0129] In some embodiments, the device further comprises at least one logic circuit component to generate the control signal. In yet another aspect of the present disclosure, a method is disclosed. The method comprises receiving a plurality of input data bits to a computing circuit, identifying an accumulated quantity associated with the plurality of input data bits, determining whether to enable or disable at least one component of the computing circuit based on the accumulated quantity, and generating a control signal to enable or disable at least one component of the computing circuit based on the determination of enabling or disabling.
[0130] In some embodiments, the method further includes: receiving the plurality of input data bits through a first component of the computing circuit; providing a first output to a second component of the computing circuit through the first component in response to the control signal; receiving the first output from the first component through the second component; providing a second output through the second component in response to a control signal including a first logic value; and outputting the first output through a multiplexer in response to a control signal including a second logic value, or outputting the second output in response to a control signal including the first logic value.
[0131] In some embodiments, the method further includes disabling the second component in response to a control signal including the second logic value.
[0132] In some embodiments, the computation circuit comprises an adding component, and wherein the method further comprises enabling or disabling the adding component.
[0133] In some embodiments, the method further comprises: outputting a first set of bits in response to enabling at least one component of the computational circuit; and outputting a second set of bits in response to disabling at least one component of the computational circuit.
[0134] In some embodiments, the control signal is a first control signal provided at a first time to disable a first number of components of the computing circuit, and the method further includes: generating a second control signal provided at a second time to disable a second number of components of the computing circuit, the second number being different from the first number.
[0135] As used herein, the words "about" and "approximately" generally represent a value of a given quantity that may vary based on a particular technology node associated with the subject semiconductor device. Based on a particular technology node, the word "about" may represent a value of a given quantity that varies within, for example, 10-30% of that value (e.g., +10%, ±20%, or ±30% of that value).
[0136] The components of several embodiments are discussed above so that those skilled in the art can better understand the various embodiments of the present invention. It should be understood by those skilled in the art that the present invention can be easily used as a basis to design or change other processing and structures to achieve the same purpose and / or achieve the same advantages as the embodiments introduced in the present invention. It should also be appreciated by those skilled in the art that these equivalent structures do not deviate from the spirit and scope of the present invention, and that various changes, substitutions and modifications may be made without departing from the spirit and scope of the present invention.
Claims
1. A circuit system, comprising: Computational circuits; a memory array operably coupled to the computing circuit; as well as Controller, configured as: inputting a plurality of input data bits into the computation circuit; identifying an accumulated quantity associated with the plurality of input data bits; determining whether to enable or disable at least one component of the computing circuit based on the accumulated amount; and Based on the determination of enabling or disabling, a control signal is generated to enable or disable at least one component of the computing circuit.
2. The system according to claim 1, wherein: The computation circuit includes an adding component, and wherein the controller is further configured to enable or disable the adding component.
3. The system of claim 1 , further comprising a multiplexer configured to: receiving a plurality of bits from a plurality of components of a computational circuit including at least one component; and Output bits are configured based on the accumulated quantity and the plurality of bits.
4. A circuit device comprising: Memory array; A computing circuit operatively coupled to the memory array, the computing circuit comprising: a first component configured to receive a plurality of input data bits and provide a first output in response to a control signal; a second component configured to receive the first output from the first component and provide a second output in response to a control signal including a first logic value; and A multiplexer configured to output the first output in response to a control signal including a second logic value, and configured to output the second output in response to a control signal including the first logic value.
5. The device according to claim 4, wherein At least one of the first component or the second component is an adder configured to provide an accumulation as an output.
6. The device according to claim 4, wherein The second component is disabled in response to a control signal including the second logic value.
7. The device according to claim 4, wherein The first digit of the first component is smaller than the second digit of the second component.
8. The device according to claim 4, wherein The multiplexer comprises: a first circuit that outputs a first set of bits in response to a control signal including the first logic value; and a second circuit that outputs a second set of bits in response to a control signal including the second logic value, The first group number in the first group is greater than the second group number in the second group.
9. A method for operating a circuit, comprising: receiving a plurality of input data bits via a computation circuit; identifying an accumulated quantity associated with the plurality of input data bits; determining whether to enable or disable at least one component of the computing circuit based on the accumulated amount; as well as Based on the determination of enabling or disabling, a control signal is generated to enable or disable at least one component of the computing circuit.
10. The method according to claim 9, further comprising: receiving, by a first component of the computation circuit, the plurality of input data bits; providing, by the first component, a first output to a second component of the computing circuit in response to the control signal; receiving, by the second component, a first output from the first component; providing, by the second component, a second output in response to a control signal comprising a first logic value; as well as The first output is outputted through the multiplexer in response to a control signal including a second logic value, or the second output is outputted in response to a control signal including the first logic value.