Compute-in-memory devices and methods for operating the same

KR103025636B1Active Publication Date: 2026-09-29TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020250000179
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2025-01-02
Publication Date
2026-09-29
Estimated Expiration
2045-01-02

Smart Images

  • Figure R1020250000179_ABST
    Figure R1020250000179_ABST
Patent Text Reader

Abstract

The memory circuit includes a Booth encoder configured to receive a first data element comprising a first code portion and a first data portion. The memory circuit includes a Booth decoder configured to receive a second data element comprising a second code portion and a second data portion and to provide a product based on the first data element and the second data element. The memory circuit includes a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder. The plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and to change the individual logic states of the plurality of encoded signals based on the first code portion and the second code portion, thereby causing the Booth decoder to provide a product.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application claims priority and interest in U.S. provisional application No. 63 / 616,934 filed January 2, 2024, which is incorporated herein by reference in its entirety for all purposes. Background Technology

[0002] Artificial intelligence (AI) has been based on machine learning, for example, using deep learning techniques. In machine learning, a computing system organized as a neural network calculates the statistical likelihood of matching input data with previously computed data. A neural network refers to a certain number of interconnected processing nodes that enable the analysis of data to compare inputs with "trained" data. Trained data refers to the computational analysis of known data characteristics to develop a model to use for comparing input data. Examples of applications for AI and data training are found in object recognition, where the system analyzes the characteristics of many (e.g., thousands or more) images to determine patterns that can be used to perform statistical analysis to identify input objects. Brief explanation of the drawing

[0003] Aspects of the present disclosure are best understood from the following detailed description when read together with the accompanying drawings. Note that, in accordance with standard practice in the relevant industry, various features are not depicted to scale. In practice, the dimensions of various features may be increased or decreased at will for clarity of discussion. FIG. 1 illustrates an exemplary block diagram of a compute-in-memory (CIM) circuit according to some embodiments. FIG. 2 illustrates a block diagram of one of the calculation blocks of the CIM circuit of FIG. 1 according to some embodiments. FIG. 3 illustrates a component block diagram illustrating Booth encoding of data elements for Booth multiplication according to some embodiments. FIG. 4 illustrates a table summarizing the Booth encoding of data elements for Booth multiplication according to some embodiments. FIG. 5 illustrates a schematic diagram of an exemplary implementation of the calculation block of FIG. 1 according to some embodiments. FIG. 6 illustrates a table summarizing the Booth encoding of data elements for Booth multiplication according to some embodiments. FIG. 7 illustrates a circuit diagram of a code recognition multiplexer of the calculation block of FIG. 5 according to some embodiments. FIG. 8 illustrates a block diagram including a plurality of calculation blocks of FIG. 5 according to some embodiments. FIG. 9 illustrates a flowchart of an exemplary method for operating the calculation block of FIG. 5 according to some embodiments. FIG. 10 illustrates a schematic diagram of an exemplary implementation of the computational circuit of FIG. 1 according to some embodiments. FIGS. 11, FIGS. 12, FIGS. 13, and FIGS. 14 each illustrate various combinations of signed / unsigned data elements processed by the calculation circuit of FIG. 10 according to some embodiments. FIG. 15 illustrates a flowchart of an exemplary method for operating the calculation circuit of FIG. 10 according to some embodiments. FIG. 16 illustrates different combinations of signed / unsigned data elements processed by the calculation circuit of FIG. 10 according to some embodiments. FIG. 17 illustrates an exemplary circuit diagram of a bus encoder according to some embodiments. FIG. 18 illustrates an exemplary circuit diagram of a Buss decoder according to some embodiments. Specific details for implementing the invention

[0004] The following disclosure provides a number of different embodiments or examples to implement the different features of the claimed subject matter. To simplify the disclosure, specific examples of components and configurations are described below. Of course, these are merely examples and are not intended to be limiting. For example, in the following detailed description, the formation of the first feature on or above the second feature may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which additional features may be formed between the first and second features so that the first and second features do not come into direct contact. Additionally, the disclosure may repeat reference numbers and / or letters in various examples. This repetition is for the purpose of simplification and clarification and does not itself determine the relationship between the various embodiments and / or configurations discussed.

[0005] Additionally, spatial relative terms such as “bottom,” “below,” “below,” “above,” “above,” “upper,” and “lower” may be used herein for ease of explanation to describe the relationship of one element or feature to other element(s) or feature(s), as illustrated in the drawings. Spatial terms are intended to include different orientations of the device during use or operation, in addition to the orientations shown in the drawings. The device may be oriented differently (rotated 90 degrees or other orientations), and spatial relative terms used herein may likewise be interpreted accordingly.

[0006] Unless otherwise noted, the terms “processor,” “processor core,” “controller,” and “control unit” are used interchangeably herein to refer to one or all of the following: a software-configured processor, a hardware-configured processor, a general-purpose processor, a dedicated processor, a single-core processor, a homogeneous multi-core processor, a heterogeneous multi-core processor, a core of a multi-core processor, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), etc., a controller, a microcontroller, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), other programmable logic devices, discrete gate logic, transistor logic, etc. A processor may be an integrated circuit, which may be configured so that the components of the integrated circuit reside on a single piece of semiconductor material such as silicon.

[0007] Neural networks calculate “weights” to perform calculations on new data (input data “words”). Neural networks use multiple layers of computation nodes, where deeper layers perform calculations based on the results of calculations performed by upper layers. Machine learning currently relies on the calculation of absolute differences and dot products of vectors, which are typically computed using multiply-accumulate (MAC) operations performed on parameters, input data, and weights. Computations in large and deep neural networks generally involve a very large number of data elements, making it impractical to store them in the processor cache. Therefore, these data elements are usually stored in memory.

[0008] Therefore, machine learning is highly computationally intensive due to the computation and comparison of many different data elements. The computation of operations within the processor is orders of magnitude faster than the transfer of data elements between the processor and main memory resources. Placing all data elements closer to the processor in the cache is enormously costly for most practical systems due to the memory size required to store them. Consequently, the transfer of data elements becomes a major bottleneck for AI computation. As datasets grow, the time and power / energy used by computing systems to move data elements can eventually become multiples of the time and power actually used to perform the computation.

[0009] In this regard, compute-in-memory (CIM) circuits or systems have been proposed to perform such MAC operations. CIM circuits instead perform data processing in situ within a suitable memory circuit. CIM circuits suppress latency for data / program fetching and output result uploading from the corresponding memory (e.g., memory array), thereby resolving the memory (or von Neumann) bottleneck of conventional computers. Another major advantage of CIM circuits is high computing parallelism due to the specific architecture of the memory array, where computation can occur simultaneously along multiple current paths. CIM circuits also benefit from the high density of multi-memory arrays with computing devices, which generally feature excellent scalability and capabilities of 3D integration. As a non-limiting example, CIM circuits for various machine learning applications can perform MAC operations locally in memory (i.e., without having to transmit data elements to the host processor) to enable higher throughput inner products of neuron activations and weight matrices while still providing higher performance and lower energy compared to computation by the host processor.

[0010] Data elements processed in CIM circuits have various data types or forms, such as integer data types and floating-point data types. Integer data types, each representing a range of mathematical integers, can have different sizes. For example, integer data types consist of 4 bits (sometimes referred to as the INT4 data type), 8 bits (sometimes referred to as the INT8 data type), and so on. Floating-point data types are generally represented by a sign part, an exponent part, and a significand part consisting of the significant digits of the number. For example, a floating-point number format specified by the IEEE® (Institute of Electrical and Electronics Engineers) has a size of 16 bits, including 10 mantissa bits, 5 exponent bits, and 1 sign bit (sometimes referred to as the FP16 data type). Other floating-point number formats also have a 16-bit size, including 7 mantissa bits, 8 exponent bits, and 1 sign bit (sometimes referred to as the BF16 data type).

[0011] In machine learning applications, CIM circuits are often configured to process inner product multiplication based on performing MAC operations on a large number of data elements (e.g., input word vectors and weight matrices), each of which may be of a floating-point data type, and then to process the addition (or accumulation) of such inner products. Several CIM circuits have been proposed to process MAC operations on data elements provided as floating-point data types. For example, it has been proposed to integrate a Booth multiplier into a CIM circuit that operates in parallel in multiple stages to produce a final product.

[0012] Booth multipliers generally operate based on the principles of the Booth algorithm. The Booth algorithm multiplies two signed binary numbers. As is typical in binary multiplication, the Booth algorithm generates partial products of the multiplication of the multiplier and the multiplicand, and these partial products are shifted and summed to produce the final product. The Booth algorithm uses rules based on the bit group values ​​of the multiplier to determine the operations for generating partial products using the multiplicand. To calculate the final product, after generating all partial products, the Booth multiplier typically shifts the partial products by individual bit(s) and outputs the shifted partial products to an adder tree for summing.

[0013] When processing signed data elements (sometimes referred to as signed data elements), conventional CIM circuits generally require at least one two's complement circuit operatively coupled between a corresponding Booth multiplier and a corresponding adder tree. For example, in a conventional CIM circuit, the Booth multiplier generates partial products based on the respective unsigned parts of the input data element and the weight data element, and provides these partial products to the two's complement circuit. The two's complement circuit then determines whether to perform a two's complement transformation based on the individual signed parts of the input data element and the weight data element. For example, if the input data element and the weight data element have the same sign, the two's complement circuit is deactivated to change the polarity of the partial products; if the input data element and the weight data element have different signs, the two's complement circuit is activated to change the polarity of the partial products. Such a two's complement circuit typically includes at least one additional half-adder, which significantly complicates the design of the CIM circuit and undesirably increases the size of the CIM circuit. Therefore, existing CIM circuits employing Booth multipliers were not entirely satisfactory in certain aspects.

[0014] The present disclosure provides various embodiments of a compute-in-memory (CIM) circuit configured to process a plurality of input data elements and a plurality of weight data elements. In one aspect, a CIM circuit as disclosed herein may perform in-memory calculations (e.g., Multiply-Accumulate (MAC) operations) on a weight data element and an input data element, each of which may be provided with a sign (e.g., a signed input data element and a signed weight data element), without performing the two's complement conversion mentioned above. The disclosed CIM circuit may multiply an input data element by a weight data element based on a plurality of sign-aware Booth decoded values. For example, the CIM circuit may include a Booth encoder, a Booth decoder (sometimes referred to as a Booth multiplier), and a plurality of sign-aware multiplexers coupled between the Booth encoder and the Booth decoder. A Booth encoder may first generate multiple Booth encoded values ​​based on an input data element (e.g., the mantissa portion of the input data element if a floating-point data type is provided). Code-aware multiplexers may determine, based on the XOR signal between the individual code portions of the input data element and the weight data element, whether to pass the Booth encoded value directly to the Booth decoder (without inversion) or to invert the Booth encoded value and then provide the inverted Booth encoded value to the Booth decoder. Upon receiving such a code-aware decoded signal, the Booth decoder may multiply the decoded signal (representing the input data element) by the weight data element (e.g., the mantissa portion of the weight data element if a floating-point data type is provided) to generate multiple partial products to be summed for the final product.

[0015] In another aspect, the CIM circuit disclosed herein may perform MAC operations on input data elements and weight data elements, which may be provided with or without a sign, respectively. The disclosed CIM circuit may multiply the input data element by the weight data element and may optionally perform sign extension depending on whether the input / weight data elements are provided as signed or unsigned. As a representative example, if the input data element is provided as unsigned, the CIM circuit may decide not to perform sign extension on the input data element. Instead, the CIM circuit may append one or more additional "0" bits to the most significant bit of the input data element. If the input data element is provided as signed, the CIM circuit may decide to perform sign extension on the input data element. For example, the CIM circuit may include a Booth encoder, a Booth decoder (sometimes also referred to as a Booth multiplier), and a plurality of logic gates. The Booth encoder may first generate a plurality of Booth encoded values ​​based on the input data element and provide the Booth encoded values ​​to the Booth decoder. Additionally, some of the logic gates coupled to the Booth decoder can determine whether the input data elements are provided as signed or unsigned. If signed, these logic gates can cause the CIM circuit to perform sign extension on the input data elements by adding an additional bit(s) identical to the most significant bit of the input data elements to the most significant bit of the input data elements. If unsigned, these logic gates can prevent the CIM circuit from performing sign extension on the weighted data elements by adding one or more "0" bits to the most significant bit of the input data elements.

[0016] FIG. 1 illustrates a block diagram of a compute-in-memory (CIM) circuit (100) according to various embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 1, the CIM circuit (100), also referred to as a memory circuit (100), comprises various components collectively configured to perform in-memory calculations (e.g., multiply-accumulate (MAC) operations) on an input word vector and a weight matrix. The input word vector may include a plurality of input data elements (XIN), and the weight matrix may include a plurality of weight data elements (W).

[0017] In some embodiments, the input data elements (XIN) and the weight data elements (W) may each be configured or provided as an INT8 data type. In some embodiments, the input data elements (XIN) and the weight data elements (W) may each be configured or provided as an INT4 data type. In some embodiments, the input data elements (XIN) and the weight data elements (W) may each be configured or provided as an FP16 data type. In some embodiments, the input data elements (XIN) and the weight data elements (W) may each be configured or provided as a BF16 data type.

[0018] As illustrated, the CIM circuit (100) includes a memory circuit (102), an input circuit (104), a calculation circuit (106), and an adder circuit (or adder tree) (108). Each of the components shown in FIG. 1 (e.g., 102 to 108) is an electronic circuit comprising a logic circuit section configured to perform an individual function. In some embodiments, the calculation circuit (106) may provide multiple partial products based on multiplying a multiplicand (e.g., input data elements (XIN)) by a multiplier (e.g., weight data elements (W)) using a Booth algorithm. It should be understood that the block diagram of the circuit shown in FIG. 1 is simplified and therefore the circuit (100) may include any of the various other components while remaining within the scope of the present disclosure.

[0019] The memory circuit (102) may include one or more memory arrays and one or more corresponding circuits. Each memory array is a storage device comprising a certain number of storage elements (103), and each storage element (103) comprises an electric, electromechanical, electromagnetic, or other device configured to store one or more data elements, and each data element comprises one or more data bits represented by a logical state. In some embodiments, the logical state corresponds to a voltage level of charge stored in some or all of the storage elements (103). In some embodiments, the logical state corresponds to a physical characteristic of some or all of the storage elements (103), e.g., resistance or magnetic orientation.

[0020] In some embodiments, the storage element (103) includes one or more static random-access memory (SRAM) cells. In various embodiments, the SRAM cell includes a plurality of transistors, for example, a 5-transistor (5T) SRAM cell, a 6-transistor (6T) SRAM cell, an 8-transistor (8T) SRAM cell, a 9-transistor (9T) SRAM cell, etc. In some embodiments, the storage element (103) includes one or more dynamic random-access memory (DRAM) cells, resistive random-access memory (RRAM) cells, magnetoresistive random-access memory (MRAM) cells, ferroelectric random-access memory (FeRAM) cells, NOR flash cells, NAND flash cells, conductive-bridging random-access memory (CBRAM) cells, data registers, non-volatile memory (NVM) cells, 3D NVM cells, or other types of memory cells capable of storing bit data.

[0021] In addition to the memory array(s), the memory circuit (102) may include a plurality of circuits for accessing the memory array or otherwise controlling the memory array. For example, the memory circuit (102) may include a plurality of (e.g., word line) drivers operatively coupled to the memory array. The drivers may apply a signal (e.g., voltage) to the corresponding storage elements (103) to enable the storage elements (103) to be accessed (e.g., programmed, read, etc.). In another example, the memory circuit (102) may include a plurality of programming circuits and / or reading circuits operatively coupled to the memory array.

[0022] The memory arrays of the memory circuit (102) are each configured to store a plurality of weight data elements (W). In some embodiments, the programming circuits may write the weight data elements (W) into the corresponding storage elements (103) of each memory array, while the writing circuit can read the bits written into the storage elements (103) to verify or otherwise test whether the written weight data elements (W) are accurate. The drivers of the memory circuit (102) may include or be operatively coupled to a plurality of input enable latches configured to receive and temporarily store input data elements (XIN). In some other embodiments, these input enable latches may be part of the input circuit (104), which may further include a plurality of buffers configured to temporarily store weight data elements (W) retrieved from the memory arrays of the memory circuit (102). Thus, the input circuit (104) can receive input data elements (XIN) and weight data elements (W).

[0023] In some embodiments, the input word vector (e.g., including input data elements (XIN)) and weight matrix (e.g., including weight data elements (W)) configured for the CIM circuit (100) to perform MAC operations may be composed of any data type among at least INT8 data type, INT4 data type, FP16 data type, and BF16 data type. However, it should be understood that each of the input data elements (XIN) and weight data elements (W) may have any data type among various other integer or floating-point data types, such as, for example, INT16 data type, UINT16 data type, UINT8 data type, UINT4 data type, FP32 data type, FP64 data type, FP128 data type, etc., while remaining within the scope of the present disclosure.

[0024] When configured as an INT8 data type, each input data element (XIN) and weight data element (W) contains 8 bits, with the leftmost bit being the sign bit. When configured as an INT4 data type, each input data element (XIN) and weight data element (W) contains 4 bits, with the leftmost bit being the sign bit. When configured as a UINT8 data type, each input data element (XIN) and weight data element (W) contains 8 bits, excluding a sign bit. When configured as a UINT4 data type, each input data element (XIN) and weight data element (W) contains 4 bits, excluding a sign bit. When configured as an FP16 data type, each input data element (XIN) and weight data element (W) contains 1 sign bit, 5 exponent bits, and 10 mantissa bits. When configured with BF16 data type, each of the input data elements (XIN) and weight data elements (W) includes 1 sign bit, 8 exponent bits, and 7 mantissa bits.

[0025] Referring still to FIG. 1, the input circuit (104) is configured to output the entire input data elements (XIN) and weight data elements (W) to the calculation circuit (106). In some embodiments of the present disclosure, the calculation circuit (106) may include a plurality of calculation blocks corresponding to a plurality of bits of the input data elements (XIN). Each calculation block may include a Booth encoder, a plurality of sign-aware multiplexers, and a Booth decoder, which are collectively configured to generate at least one partial product, which will be discussed in more detail below with respect to FIG. 5. In some other embodiments of the present disclosure, the calculation circuit (106) may include a plurality of Booth encoders and a plurality of corresponding Booth decoders. In such embodiments, the calculation circuit (106) may further include a plurality of logic gates configured to process the input data elements (XIN) and weight data elements (W), regardless of whether they are provided as signed or unsigned, to determine whether to perform sign extension on the weight data elements (W) and / or input data elements (XIN). Details of these embodiments will be discussed below with reference to FIG. 10. An adder tree (108) can receive partial products from a calculation circuit (106) and sum them to generate a final product (P) of input data elements (XIN) and weight data elements (W).

[0026] FIG. 2 illustrates a block diagram (200) of one of the computation blocks (hereinafter “computation block (200)”) of a computation circuit (106) according to various embodiments of the present disclosure. As described above, the computation block (200) (or the computation block of the computation circuit (106)) may receive input data elements (XIN) and weight data elements (W) from the input circuit (104), generate a plurality of partial products based on the Booth algorithm, and provide the partial products to an adder tree (108) to generate a final product. It should be recognized that the block diagram of the computation circuit (200) illustrated in FIG. 2 is simplified, and accordingly, the computation circuit (200) may include any of various other components (e.g., sign-recognizing multiplexers) while remaining within the scope of the present disclosure.

[0027] As illustrated, the computation circuit (200) includes a Booth encoder (210) and a Booth decoder (220). The Booth encoder (210) may receive a multiplicand (e.g., an input data element (XIN) and / or a subset of the input data element (XIN)). The Booth encoder (210) and the Booth decoder (220) may each be a circuit or combination of logic components (e.g., FIG. 17 and FIG. 18). The Booth encoder (210) may generate and output a plurality of Booth encoded signals (e.g., may include an enable bit, a Booth encoded bit, and a select bit) from the multiplicand. Various combinations of the logic states of the Booth encoded signals may correspond to individual Booth encoded values. The Booth decoder (220) may receive a multiplier (e.g., a weight data element (W) and / or a subset of the weight data element (W)). The Booth decoder (220) further receives a Booth encoded signal from the Booth encoder (210) and can multiply the corresponding Booth encoded value by a multiplier to generate a partial product (PP). In one aspect of the disclosure (e.g., FIG. 5), the Booth encoded value received by the Booth decoder (220) may be forwarded or selected by a plurality of code-aware multiplexers coupled between the Booth encoder (210) and the Booth decoder (220). In another aspect of the disclosure (e.g., FIG. 10), the Booth encoded value may be received directly by the Booth decoder (220) without passing through a code-aware multiplexer, for example.

[0028] FIG. 3 illustrates an example of Booth encoding of an input data element for Booth multiplication in a CIM circuit (e.g., 100 of FIG. 1) according to various embodiments of the present disclosure. As illustrated, a Booth encoder (300) (e.g., an implementation of the Booth encoder (210) of FIG. 2) may encode the data element (310) into a plurality of Booth encoded signals (320) corresponding to one of a plurality of Booth encoded values ​​(e.g., 0, -1, 1, -2, 2) or convert it in another way.

[0029] In some embodiments, the data element (310) may include one or more input data elements (XIN) that serve as the multiplicand of the CIM circuit, while one or more corresponding weight data elements (W) may serve as the multiplier. In other embodiments, the data element (310) may include one or more weight data elements (W) that serve as the multiplicand of the CIM circuit, while one or more corresponding input data elements (XIN) may serve as the multiplier. The following discussion will focus on examples where the input data elements (XIN) are encoded (i.e., the input data elements (XIN) are used as the multiplicand and the weight data elements (W) are used as the multiplier).

[0030] The Booth encoder (300) can encode the input data element (XIN) (310) into various cycles in which the Booth encoder (300) can encode subsets (302, 304) of the input data element (XI) (310). The Booth encoding the input data element (XIN) (310) can simplify the input data element (XIN) (310) by converting the input data element (XIN) (310) into Booth encoded signals (320) associated with a limited number of operations for performing Booth multiplication in a CIM circuit. As further described herein, the Booth encoder (300) can convert each of the subsets (302, 304) into a plurality of Booth encoded signals (320) that collectively correspond to individual Booth encoded values. The Booth encoded signals (320) can be configured to control other parts of the corresponding CIM circuit (Booth decoder such as 220 in FIG. 2) so that the Booth decoder multiplies the Booth encoded value corresponding to the weight data element (W) by the corresponding Booth encoded value to generate a partial product.

[0031] In some embodiments, subsets (302 and 304) of the input data element (XIN) (310) may be superimposed. In some embodiments, the subsets (302, 304) may be centered around a bit position and may include a bit position immediately preceding a bit position and a bit position immediately following a bit position. For a subset (302) centered around the least significant bit of the input data element (XIN) (310), a “0” bit may be added to the input data element (XIN) (310) to fill the bit position immediately preceding the least significant bit.

[0032] FIG. 3 illustrates a non-limiting example of a 3-bit Booth encoding that encodes 3-bit subsets (302, 304) of an input data element (XIN) (310). The multiplication operation for execution by a part of the CIM circuit (e.g., the Booth decoder (220) of FIG. 2) may be the multiplication of the input data element (XIN) and the weight data element (W). The input data element (XIN) (310) may have an arbitrary bit length "p", so that the input data element (XIN) (310) consists of bits X p-1 It can include , ..., X0.

[0033] In the example illustrated in FIG. 3, the input data element (XIN) (310) has 4 bits (i.e., p=4). The Booth encoder (300) can encode subsets (302, 304) of the input data element (XIN) (310) into various cycles, and each of the subsets (302, 304) has 3 bits. Each subset (302, 304) can be used to generate an individual number of Booth encoded signals (320). For example, the input data element (XIN) (310) may contain bits (X3, X2, X1, X0). For example, a "0" bit may be added to the input data element (XIN) (310), for example, attached to the least significant bit (X0), so that the input data element (XIN) (310) may contain bits (X3, X2, X1, X0, 0). "0" bits may be added to fill a subset (302) centered on the least significant bit (X0). In this example, the subsets (302, 304) for 3-bit Booth encoding may each contain bits centered on a bit position including the bit position immediately preceding the bit position and the bit position immediately following the bit position. Each consecutive subset (302, 304) may be centered on a bit position consecutively to the previous subset (302, 304). For example, the subsets (302, 304) contain bits (X 2i+1 , X 2i, and X 2i-1 It can be expressed as ), where "i" may be the number of cycle repetitions. In the case of the first cycle, for example, for i=0, since there may not be a lower bit than the least significant bit (X0), X 2i-1 The bit may not exist, and instead, a "0" bit attached to the least significant bit (X0) may be used. As consecutive subsets (302, 304) center bit positions consecutively with the previous subset (302, 304), the least significant bit of the consecutive subset (302, 304) may overlap with the most significant bit of the previous subset (302, 304). In other words, X of the consecutive subset (302, 304) 2i-1 Bits and X of the previous subset (302, 304) 2i+1 Bits can be superimposed in consecutive repetitions (e.g., bit X when i=1). 2i-1 And if i=0, bit X 2i+1 (All are X1 bits). Accordingly, the Booth encoder (300) in successive iterations 2 bits of the previously unencoded input data element (XIN) (310) (e.g., bit X 2i+1 , X 2i ) and 1 bit of the previously encoded input data element (XIN) (310) (e.g., bit X 2i+1 It can encode ).

[0034] For example, from a subset (302, 304) of bits “111” and / or “000”, the Booth encoder (300) can generate Booth encoded signals (320) representing a “0” Booth encoded value for multiplication with a corresponding weight data element (W) by directing a logic gating operation to achieve, for example, the result of multiplication. The logic gating prevents bits of the weight data element (W) from propagating in the CIM circuit, thereby generating a “low” or “0” signal instead of the weight data element (W) and effectively multiplying the weight data element (W) by a “0” value.

[0035] From a subset (302, 304) of bits "001" and / or "010", the Booth encoder (300) can generate Booth-encoded signals (320) representing a "1" Booth-encoded value for multiplication with the corresponding weight data element (W), for example, by representing a direct mapping operation of the weight data element (W) in the CIM circuit to obtain the result of multiplication. Direct mapping in the CIM circuit can enable the bits of the weight data element (W) to be propagated within the CIM circuit without being altered, and consequently generate signals representing the unaltered weight data and effectively multiply the weight data element (W) by a "1" value.

[0036] From a subset (302, 304) of bits "011", the Booth encoder (300) can generate Booth encoded signals (320) representing a "2" Booth encoded value for multiplication with the corresponding weight data element (W) by directing a direct mapping operation and a left shift operation (e.g., a left shift of 1 bit in an adder) on the weight data element (W) in the CIM circuit to achieve the result of multiplication. Directly mapping the weight data element (W) to the left in the CIM circuit generates signals representing the weight data element (W) multiplied by a value of "2" by shifting the bits of the weight data element (W) by an amount that changes the bits of the weight data element (W).

[0037] From a subset (302, 304) of bits "100", the Booth encoder (300) can generate a Booth encoded signal (320) representing a "-2" Booth encoded value for multiplication with the corresponding weight data element (W) by, for example, performing an inversion operation of the weight data element (W), an addition operation of a "1" value in the least significant bit of the inverted weight data element, and a left shift operation for the sum in the CIM circuit (e.g., a 1-bit left shift in the adder) to obtain a multiplication result. By inverting the bits of the weight data element (W) in the CIM circuit and adding a "1" value to the least significant bit of the inverted bits of the weight data element (W), signals representing a negative sign version of the weight data element (W) can be generated, thereby effectively multiplying the weight data element (W) by a "-1" value. In the CIM circuit, the left shift of the negative sign version of the weight data element (W) shifts the bits of the negative sign version of the weight data element (W) by an amount that changes the bits of the negative sign version of the weight data element (W), thereby generating signals indicating that the negative sign version of the weight data element (W) is multiplied by a value of "2". Together, these operations can result in signals indicating the weight data element (W) multiplied by a value of "-2".

[0038] From a subset (302, 304) of bits "101" and / or "110", the Booth encoder (300) can generate Booth encoded signals (320) representing a "-1" Booth encoded value for multiplication with the corresponding weight data element (W) by, for example, directing an inversion operation of the weight data element (W) and an addition operation of the "1" value of the least significant bit of the inverted weight data element (W) in the CIM circuit to achieve the result of multiplication. By inverting the bits of the weight data element (W) in the CIM circuit and adding a "1" value to the least significant bit of the inverted bits of the weight data element (W), signals representing the negative sign version of the weight data element (W) can be generated, thereby effectively multiplying the weight data element (W) by a "-1" value.

[0039] FIG. 4 shows subsets (302, 304) of input data elements (XIN) (310) (e.g., X) for generating Booth encoded signals (320) according to various embodiments of the present disclosure. 2i+1 , X 2i , and X 2i-1 A non-limiting example of a table (400) of a Booth encoder (300) encoding one of the following is illustrated. As a non-limiting example, Booth encoded signals (320) include an enable bit ("ENB"), a Booth encoded bit ("BE"), and a select bit ("S"). Different combinations of the logical states of these bits (ENB, BE, S) may correspond to individual Booth encoded values. Additionally, the bits (ENB, BE, and S) may be provided to a Booth decoder to function as control bits for the Booth decoder. Upon receiving the control bits, the Booth decoder may multiply the received weight data element (W) by the Booth encoded value.

[0040] As a representative example, a Booth encoder (300) receiving a subset (302, 304) of bits "000" and / or "111" may generate and output Booth encoded signals (320) (e.g., ENB, BE, S) of bit "100", which may be configured to cause a corresponding Booth decoder to multiply a weight data element (W) by a value of "0". The Booth decoder may be configured to be interpreted / controlled by the Booth encoded signals (320) of bit "100" to perform logic gating on the weight data element (W). As another representative example, a Booth encoder (300) receiving a subset (302, 304) of bits "001" and / or "110" may generate and output Booth-encoded signals (320) of bits "000" (e.g., ENB, BE, S), which may be configured to cause a corresponding Booth decoder to multiply a weight data element (W) by a value of "1". The Booth decoder may be configured to be interpreted / controlled by the Booth-encoded signals (320) of bits "000" to perform direct mapping to the weight data element (W). Other combinations of the logical states of ENB, BE, and S are summarized in the table (400), along with the individual Booth-encoded values ​​(or operations performed by the corresponding Booth decoder).

[0041] FIG. 5 illustrates a schematic diagram of an exemplary implementation of the calculation block (200) of FIG. 2 (hereinafter “calculation block (500)”) according to various embodiments of the present disclosure. The calculation block (500) may be configured to process (e.g., encode) one of a plurality of subsets of input data elements (XIN) and to multiply the encoded input data elements (XIN) by weight data elements (W). Generally, the input data elements (XIN) and weight data elements (W) may be provided as signed data elements. It should be understood that the schematic diagram of FIG. 5 is simplified, and accordingly, the calculation block (500) may include any of various other components while remaining within the scope of the present disclosure.

[0042] As illustrated, the computation block (500) includes a Booth encoder (510) (e.g., 210 in FIG. 2) and a Booth decoder (520) (e.g., 220 in FIG. 2), and a plurality of code-aware multiplexers (530, 540, 550, 560). In various embodiments, the code-aware multiplexers (530, 540, 550, and 560) are operatively coupled between the Booth encoder (510) and the Booth decoder (520). In an example where the Booth encoder (510) is implemented as a 3-bit Booth encoder (sometimes referred to as a radix-4 Booth encoder), such as the encoder (300) illustrated in FIG. 3, the number of code-aware multiplexers may be four. These four code-aware multiplexers may each correspond to Booth-encoded values ​​(1, -1, -2, and 2) provided by the Booth encoder (510). In other words, the Booth encoder (510) may operatively (e.g., not physically) have four symbol or operation outputs corresponding to (or otherwise providing) the Booth-encoded values ​​1, -1, -2, and 2, respectively. Furthermore, the Booth encoder (510) may be implemented as any of various other Booth encoders (e.g., base-2 Booth encoder, base-8 Booth encoder) which can change the number of corresponding code-aware multiplexers while remaining within the scope of the present disclosure.

[0043] A Booth encoder (510) is configured to encode one of a subset of received input data elements (XIN) based on a Booth algorithm and to provide Booth encoded signals during each cycle. A Booth decoder (520) is configured to receive a weight data element (W) (or one of a plurality of subsets of the weight data element (W)) and to multiply the weight data element (W) by a Booth encoded value determined based on the Booth encoded signals (provided by the encoder (510)) to provide a plurality of partial products. In various embodiments, code-aware multiplexers (530 to 560) are operatively coupled between the Booth encoder (510) and the Booth decoder (520).

[0044] The input data element (XIN) and the weight data element (W) processed by the calculation block (500) may each be an integer data type or a floating-point data type that may have a sign bit. That is, the input data element (XIN) and the weight data element (W) are each provided as signed data elements. Accordingly, the sign-aware multiplexers (530 to 560) receive Booth-encoded signals and can operationally adjust the Booth-encoded signals based on the logically processed signals of the sign bit of the input data element (XIN) (sometimes also referred to as "XINsign") and the sign bit of the weight data element (W) (sometimes also referred to as "Wsign"). However, in some other embodiments, the calculation block (500) may multiply the unsigned input data element by the unsigned weight data element while remaining within the scope of the present disclosure. For example, if unsigned data elements are provided, the calculation block (500) can disable the code-aware multiplexers (530 to 560); if signed data elements are provided, the calculation block (500) can enable the code-aware multiplexers (530 to 560).

[0045] Each of the code-aware multiplexers (530 to 560) has a first input, a second input, and an output. The first input of the code-aware multiplexer can receive a first combination of individual logic states of Booth-encoded signals, and the second input of the code-aware multiplexer can receive a second combination of individual logic states of Booth-encoded signals. Equally, the first combination of logic states of Booth-encoded signals can correspond to a first Booth-encoded value, and the second combination of logic states of Booth-encoded signals can correspond to a second Booth-encoded value. In various embodiments, the first Booth-encoded value and the second Booth-encoded value received equally by the first and second inputs of each of the code-aware multiplexers (530 to 560) have opposite polarities but the same magnitude. For example, in FIG. 5, the code-aware multiplexer (530) can receive Booth-encoded values ​​1 and -1 at the first input and the second input, respectively; The code recognition multiplexer (540) can receive Booth encoded values ​​-1 and 1 from the first input and the second input, respectively; the code recognition multiplexer (550) can receive Booth encoded values ​​-2 and 2 from the first input and the second input, respectively; and the code recognition multiplexer (560) can receive Booth encoded values ​​2 and -2 from the first input and the second input, respectively.

[0046] In some embodiments, each of the code-aware multiplexers (530 to 560) may be controlled by an XOR signal of XINsign and Wsign (sometimes referred to as "XOR(Wsign,XINsign)"). When XINsign and Wsign are provided as "00" or "11", the XOR signal is equal to logic "0"; when XINsign and Wsign are provided as "01" or "10", the XOR signal is equal to logic "1". That is, when the codes of the input data element (XIN) and the weight data element (W) are the same, the XOR signal is equal to logic "0"; and when the codes of the input data element (XIN) and the weight data element (W) are different, the XOR signal is equal to logic "1".

[0047] When the signal XOR(Wsign, XINsign) is logic "0", the signal recognition multiplexers (530 to 560) can each select the signals (or equivalent Booth encoded values) received at the first input; and when the signal XOR(Wsign, XINsign) is logic "1", the signal recognition multiplexers (530 to 560) can each select the signals (or equivalent Booth encoded values) received at the second input. In other words, the code recognition multiplexers (530 to 560) can each select the first Booth encoded value when the input data element (XIN) and the weight data element (W) have the same code; and select the second Booth encoded value when the input data element (XIN) and the weight data element (W) have different codes. Equally, the code-aware multiplexers (530 to 560) can determine whether to adjust Booth-encoded signals based on whether the codes of the input data element (XIN) and the weight data element (W) are the same (positive product) or different (negative product).

[0048] As a representative example, when the signal XOR(Wsign, XINsign) is "0" and the Booth encoded signals provided by the Booth encoder (510) correspond to the Booth encoded value "1", the code-aware multiplexer (530) can select the Booth encoded value "1" and provide it to the Booth decoder (520). That is, when the signal XOR(Wsign, XINsign) is "0", the code-aware multiplexer (530) can directly forward the Booth encoded value provided by the Booth encoder (510) to the Booth decoder (520). As another representative example, when the signal XOR(Wsign, XINsign) is "1" and the Booth encoded signals provided by the Booth encoder (510) correspond to the Booth encoded value "1", the code-aware multiplexer (530) can select the Booth encoded value "-1" and provide it to the Booth decoder (520). Similarly, if the signal XOR(Wsign, XINsign) is identified as being equal to "1", the code-aware multiplexers (530 to 560) can "adjust" the Booth-encoded value provided by the Booth encoder (510) by selecting a Booth-encoded value having opposite polarity, and provide the adjusted Booth-encoded value to the Booth decoder (520).

[0049] FIG. 6 illustrates a subset of an input data element (XIN) (e.g., X) according to various embodiments of the present disclosure. 2i+1 , X 2i and X 2i-1 A non-limiting example of a table (600) summarizing a calculation block (500) (Fig. 5) is illustrated, which encodes ) and generates Booth encoded values ​​(or Booth encoded signals), optionally adjusts the generated Booth encoded values ​​based on the codes of the input data element (XIN) and the weight data element (W), and optionally multiplies the weight data element (W) by the optionally adjusted Booth encoded values.

[0050] FIG. 7 illustrates an exemplary circuit diagram of code-recognizing multiplexers (530 to 560) (hereinafter “multiplexer (700)”) according to various embodiments of the present disclosure. In the example of FIG. 7, the multiplexer (700) is implemented as a 2-input-1-output multiplexer (sometimes referred to as a 2-to-1 MUX or 2:1 MUX) equipped with AND-OR-INVERT (AOI) logic gates. That is, the multiplexer (700) is configured to select one of two input signals based on a control signal. It should be understood that the multiplexer (700) may be implemented as any of various other configurations (e.g., equipped with OR-AND-INVERT (OAI) logic gates) while remaining within the scope of the present disclosure.

[0051] As illustrated, the multiplexer (700) includes a first AND logic gate (710), a second AND logic gate (720), and an OR logic gate (730). The multiplexer (700) may have: (i) a first input connected to one of the inputs of the AND logic gate (710) (the other input of the AND logic gate (710) is configured to directly receive the signal XOR(Wsign, XINsign); and (ii) a second input connected to one of the inputs of the AND logic gate (720) (the other input of the AND logic gate (720) is configured to receive the signal XOR(Wsign, XINsign) through an inverter). The AND logic gate (710) and the AND logic gate (720) may have their outputs connected to the OR logic gate (730). In an example where the code recognition multiplexer (530) is implemented as the multiplexer (700), the first input and the second input of the multiplexer (700) are configured to receive a first Booth encoded value "1" and a second Booth encoded value "-1". Thus, when the signal XOR (Wsign, XINsign) is equal to "0", the multiplexer (700) (or 530) selects a first combination of logical states of Booth encoded signals corresponding to the Booth encoded value "1"; and when the signal XOR (Wsign, XINsign) is equal to "1", the multiplexer (700) (or 530) selects a second combination of logical states of Booth encoded signals corresponding to the Booth encoded value "-1".

[0052] FIG. 8 illustrates an exemplary block diagram (800) of a computation circuit (106) (hereinafter "computation circuit (800)") according to various embodiments of the present disclosure. In the exemplary example of FIG. 8, the computation circuit (800) is 12 bits (X 12 , X 11 , X 10It can be configured to process (e.g., encode) input data elements (XIN) having X9, X8, X7, X6, X5, X4, X3, X2, X1) and to generate multiple partial products by multiplying the encoded input data elements (XIN) by weight data elements (W).

[0053] As illustrated, the computation circuit (800) may have six computation blocks (810A, 810B, 810C, 810D, 810E, and 810F). Each of the computation blocks (810A through 810F) is composed of the computation block (500) of FIG. 5, which, for example, can generate a Booth-encoded value by encoding a 3-bit subset of an input data element (XIN) and generate a partial product by multiplying a weight data element (W) by the corresponding selected Booth-encoded value. However, it should be understood that the computation circuit (800) can process data elements with any number of bits. Accordingly, the number of computation blocks included in the computation circuit (800) may change accordingly. For example, to process data elements having 8 bits, the computation circuit (800) may have four computation blocks, each configured to generate a partial product. Generally, the number of calculation blocks (N1) of the calculation circuit (800) is equal to half the number of data element bits (N2) received by the calculation circuit (800).

[0054] For example, a calculation block (810A) may encode a subset of (X2, X1, 0) to generate a first Booth-encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the first Booth-encoded value by a weight data element (W) to generate a first partial product; a calculation block (810B) may encode a subset of (X4, X3, and X2) to generate a second Booth-encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the second Booth-encoded value by a weight data element (W) to generate a second partial product; a calculation block (810C) may encode a subset of (X6, X5, and X4) to generate a third Booth-encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the third Booth-encoded value by a weight data element (W) to generate a third partial product; The calculation block (810D) can encode a subset of (X8, X7, and X6) to generate a fourth Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the fourth Booth encoded value by a weight data element (W) to generate a fourth partial product; and the calculation block (810E) can (X 10 A subset of (X9 and X8) can be encoded to generate a fifth Booth encoded value (e.g., 0, 1, -1, -2 or 2) and a weight data element (W) can be multiplied by the fifth Booth encoded value to generate a fifth partial product; and the calculation block (810F) (X 12 , X 11 and X 10 A subset of ) can be encoded to generate a sixth Booth encoded value (e.g., 0, 1, -1, -2, or 2), and a weight data element (W) can be multiplied by the sixth Booth encoded value to generate a sixth partial product. These six partial products can then be summed (by an adder tree such as 108 in FIG. 1) to derive the final product of the input data element (XIN) and the weight data element (W).

[0055] FIG. 9 illustrates a flowchart of an exemplary method (900) for performing a MAC operation on an input data element (XIN) and a weight data element (W) according to various embodiments of the present disclosure. In some embodiments, the input data element (XIN) and the weight data element (W) may each be provided as a signed data element. The operations of the method (900) may be performed, for example, by the components described above in FIG. 5, and accordingly, some of the reference numerals used above may be reused in the subsequent discussion of the method (900). Furthermore, it is understood that the method (900) has been simplified, and accordingly, additional operations may be provided before, during, and after the method (900) of FIG. 9, and some other operations may be described only briefly in this specification.

[0056] The method (900) begins with an operation (910) of receiving a first data element and a second data element. The first data element may be an input data element (XIN), and the second data element may be a weight data element (W). In some embodiments, the input data element (XIN) and the weight data element (W) may each be received as signed data elements that may be an integer data type or a floating-point data type. Accordingly, the input data element (XIN) has a first code bit and a plurality of first data bits, and the weight data element (W) has a second code bit and a plurality of second data bits. As a non-limiting example, using the calculation block (500) of FIG. 5, a Booth encoder (510) may receive the input data element (XIN), and a Booth decoder (520) may receive the weight data element (W).

[0057] The method (900) continues with the operation (920) of encoding a first data bit of a first data element to generate a plurality of encoded values. Continuing with the above example, a Booth encoder (510), implemented as a 3-bit Booth encoder, can encode a 3-bit subset of the first data bits during each cycle. In an example where the number of the first data bits is 4 (e.g., X3, X2, X1, X0), the Booth encoder (510) can generate a first combination of logical states of Booth encoded signals corresponding to a first Booth encoded value (e.g., "1") during the first cycle, and can generate a second combination of logical states of Booth encoded signals corresponding to a second Booth encoded value (e.g., "-1") during the second cycle.

[0058] The method (900) continues with the operation (930) of selecting one from a pair of opposite Booth encoded values ​​based on the logically processed signals of the first code bit of the first data element and the second code bit of the second data element. The pair of opposite Booth encoded values ​​have opposite polarities but are of the same magnitude. Continuing with the example above, after the Booth encoder (510) generates a first Booth encoded value "1" and provides it to a corresponding code-aware multiplexer (e.g., 530), the multiplexer (530) can determine whether to directly transmit the first Booth encoded value "1" to the Booth decoder (520) based on the XOR signal of the first code bit and the second code bit, or to select another Booth encoded value opposite to "1," namely "-1". If the XOR signal is equal to "0" (indicating that the input data element (XIN) and the weight data element (W) have the same sign), the multiplexer (530) can directly forward (select) the first Booth encoded value "1" to the Booth decoder (520); if the XOR signal is equal to "1" (indicating that the input data element (XIN) and the weight data element (W) have different signs), the multiplexer (530) can invert the first Booth encoded value to "-1" and provide (select) it to the Booth decoder (520).

[0059] The method (900) continues with the operation (940) of multiplying the second data bits of the second data element by the selected encoded value. Upon receiving the selected Booth encoded value, the Booth decoder (520) may multiply the weight data element (W) by the selected Booth encoded value to generate a partial product. Using the same example above, if the XOR signal is equal to "0" during the first cycle (if the first Booth encoded value is provided as "1"), the Booth decoder (520) then multiplies the weight data element (W) by 1; if the XOR signal is equal to "1" during the first cycle (if the first Booth encoded value is provided as "1"), the Booth decoder (520) then multiplies the weight data element (W) by -1. After the partial products are generated during each cycle, all partial products may be summed to generate a final product. In the above example where the input data element (XIN) has 4 bits, the two partial products can be summed to produce the final product of the input data element (XIN) and the weight data element (W).

[0060] FIG. 10 illustrates a schematic diagram of an exemplary embodiment of the calculation circuit (106) of FIG. 1 or the plurality of calculation blocks (200) of FIG. 2 (hereinafter, "calculation circuit (1000)") according to various embodiments of the present disclosure. The calculation circuit (1000) may be configured to process (e.g., encode) an input data element (XIN) and to multiply the encoded input data element (XIN) by a weight data element (W). In various embodiments, the input data element (XIN) and the weight data element (W) may be provided as signed or unsigned data elements. Accordingly, the calculation circuit (1000) may have control pins that individually indicate two signals (e.g., two bits) indicating whether the input data element (XIN) is signed or unsigned, one of which (XSIGNED) indicates whether the input data element (XIN) is signed or unsigned, and the other (WSIGNED) indicates whether the weight data element (W) is signed or unsigned. It should be understood that the schematic diagram of FIG. 10 is simplified, and accordingly, the calculation circuit (1000) may include any of various other components while remaining within the scope of the present disclosure.

[0061] As described, the computation circuit (1000) includes a plurality of Booth encoders (e.g., each of which may correspond to 210 in FIG. 2) and a plurality of Booth decoders (e.g., each of which may correspond to 220 in FIG. 2) and a plurality of logic components (1030, 1040, 1050). In the exemplary example of FIG. 10, the data elements (e.g., XIN and W) received by the computation circuit (1000) each have 12 bits (e.g., XIN[11:0] and W[11:0]). In such an example, the computation circuit (1000) may include six Booth encoders (1010A to 1010F) and six corresponding Booth decoders (1020A to 1020F). It should be understood that the data element processed by the calculation circuit (1000) may have any other number of bits while remaining within the scope of the present disclosure. The calculation circuit (1000) may be operatively coupled to an adder tree (1060) (an exemplary implementation of the adder tree (108) of FIG. 1), and the adder tree (1060) may include a plurality of full adders (1061, 1062, 1063, 1064, 1065, and 1066).

[0062] Booth encoders (1010A to 1010F) can each be implemented as a 3-bit Booth encoder (e.g., encoder (300) shown in FIG. 3), and each of the Booth encoders (1010A to 1010F) can be operatively coupled with a corresponding Booth decoder among the Booth decoders (1020A to 1020F). In an example where the input data element (XIN) has 12 bits (e.g., signal (1001) which can be presented as XIN[11:0]), each of the Booth encoders can encode one of a plurality of subsets of signal (1001)(XIN[11:0]) and provide the Booth encoded value to the corresponding Booth decoder.

[0063] For example, a Booth encoder (1010A) may encode a first subset (XIN[11:0]) of a signal (1001) to generate a first Booth encoded value and provide the first Booth encoded value to a Booth decoder (1020A); a Booth encoder (1010B) may encode a second subset (XIN[11:0]) of a signal (1001) to generate a second Booth encoded value and provide the second Booth encoded value to a Booth decoder (1020B); and a Booth encoder (1010C) may encode a third subset (XIN[11:0]) of a signal (1001) to generate a third Booth encoded value and provide the third Booth encoded value to a Booth decoder (1020C). A Booth encoder (1010D) can encode a fourth subset (XIN[11:0]) of a signal (1001) to generate a fourth Booth encoded value and provide the fourth Booth encoded value to a Booth decoder (1020D); a Booth encoder (1010E) can encode a fifth subset (XIN[11:0]) of a signal (1001) to generate a fifth Booth encoded value and provide the fifth Booth encoded value to a Booth decoder (1020E); and a Booth encoder (1010F) can encode a sixth subset (XIN[11:0]) of a signal (1001) to generate a sixth Booth encoded value and provide the sixth Booth encoded value to a Booth decoder (1020F).

[0064] In various embodiments of the present disclosure, the calculation circuit (1000) may use logic components (1030, 1040, 1050) to process the input data element (XIN) and the weight data element (W), regardless of whether the input data element (XIN) and the weight data element (W) are provided as unsigned or signed, respectively. For example, the logic component (1030) may be a NAND2 gate, the logic component (1040) may be a NOR2 gate, and the logic component (1050) may be a half adder. The logic component (1030) may NAND the signals (1003, 1005) to provide the signal (1017); the logic component (1040) may NOR the signals (1011, 1017) to provide the signal (1019); The logic component (1050) can add 1 bit to the signal (1013) to provide the signal (1015). Each of these logic components and signals will be described in detail as follows.

[0065] A signal (1003) received at one of the inputs of the logic component (1030) may represent the most significant bit of the signal (1001), e.g., XIN

[11] . A signal (1005) received at another input of the logic component (1030) may represent a logically inverted version of the signal indicated at one of the control pins, e.g., XSIGNEDB. In some embodiments, the logic component (1030) may provide NAND(XIN

[11] ,XSIGNEDB) as the signal (1017).

[0066] A signal (1011) received at one of the inputs of a logic component (1040) may represent a logically inverted version of a weight data element (WB[11:0]). In some embodiments, when receiving a signal (1017) from a logic component (1030) at another input, the logic component (1040) may provide a NOR (NAND(XIN

[11] ,XSIGNEDB),WB[11:0]) as a signal (1019), where NAND(XIN

[11] ,XSIGNEDB) represents the signal (1017). The signal (1019) may represent a partial product of one of the subsets of the signal (1001)(XIN[11:0]), the partial product comprising the most significant bit and one or more bits attached to the left of the most significant bit.

[0067] The signal (1013) received by the logic component (1050) may represent a weighted data element having -W, which is the opposite polarity. To generate the signal (1015) (e.g., -W), the logic component (1050) may receive the signal (1013) represented as NAND(WSIGNED,W

[11] ),WB[11:0] in various embodiments and add a signal (1013) having a single-bit binary integer (not shown). Specifically, the signal (1013) (NAND(WSIGNED,W

[11] ),WB[11:0]) may represent performing a signature extension on WB[11:0]. For example, if the weight data element (W) is provided as signed (i.e., WSIGNED=1), the signal (1013) becomes NAND(1,W

[11] ),WB[11:0], which ultimately becomes WB

[11] ,WB[11:0]. WB

[11] ,WB[11:0] refers to attaching the most significant bit of WB[11:0] to its left, as disclosed herein. In another example, if the weight data element (W) is provided as unsigned (i.e., WSIGNED=0), the signal (1013) becomes NAND(0,W

[11] ),W[11:0], which ultimately becomes 1,WB[11:0]. As disclosed herein, 1,WB[11:0] refers to attaching "1" to the left of WB[11:0]. Accordingly, the signal (1015)(-W) can be presented as WN[12:0].

[0068] Each of the Booth decoders (1020A to 1020F) can receive two signals (1007, 1009), which represent W and -W with sign extension, respectively. In various embodiments, signal (1007) may be presented as NOR(WSIGNEDB,WB

[11] ),W[11:0], and signal (1009) may be presented as WN

[12] ,WN[12:0]. The Booth decoders (1020A to 1020F) can each generate a partial product by multiplying a Booth encoded value corresponding to a weight data element (W) (e.g., provided by the corresponding Booth encoder among the Booth encoders (1010A to 1010F)). Specifically, each of the Booth decoders (1020A to 1020F) can selectively adjust the received W and -W based on the corresponding Booth encoded value. As a representative example, when using Booth decoder (1020F), when receiving a Booth encoded value "2" from Booth encoder (1010A), Booth decoder (1020A) can perform a left shift operation on W. As another representative example, when using Booth decoder (1020A) when receiving a Booth encoded value "-2" from Booth encoder (1010F), Booth decoder (1020F) can perform a left shift operation on -W.

[0069] In such a configuration, based on whether the signal (1001) (XIN[11:0]) is provided as signed or unsigned, the logic components (1030, 1040) can collectively determine how to process the partial product of the most significant bit of the signal (1001) (e.g., XIN

[11] or signal (1003)). Generally, if the signal (1001) (XIN[11:0]) is provided as unsigned, the logic component (1040) can output the signal (1019) based on a logically inverted version of the most significant bit (e.g., XINB

[11] ) of the signal (1001), such that all the bits are equal to "0" or equal to the weighted data element (W[11:0]). Equivalently, if the input data element (XIN) is provided as unsigned, the partial product corresponding to the most significant bit of the input data element (signal (1001) or XIN[11:0]) and the weight data element (W) is "0" or "W". If the signal (1001) (XIN[11:0]) is provided as signed, the logic component (1040) may output the signal (1019) as having all "0", regardless of whether the most significant bit of the signal (1001) (e.g., XINB

[11] ) is "1" or "0". Equivalently, if the input data element (XIN) is provided as signed, the partial product corresponding to the most significant bit of the input data element (signal (1001) or XIN[11:0]) and the weight data element (W) is always "0". Advantageously, even if the ability to process signed or unsigned data elements is provided, the computational load of the computational circuit (1000) (and corresponding circuit design) does not increase accordingly.

[0070] FIGS. 11, 12, 13, and 14 illustrate examples of a computational circuit (1000) that processes four different combinations of a signed or unsigned input data element (XIN) and a signed or unsigned weight data element (W). In the examples of FIGS. 11 through 14, the input data element (XIN) and the weight data element (W) are each provided to have 12 bits. However, it should be understood that the number of bits of each input data element (XIN) and weight data element (W) processed by the computational circuit (1000) may vary while remaining within the scope of the present disclosure (e.g., FIG. 16).

[0071] In FIG. 11, an example is illustrated where the input data element (XIN) is provided as unsigned and the weight data element (W) is provided as unsigned (i.e., XSIGNED=0 and WSIGNED=0). Accordingly, the signal (1005, XSIGNEDB=1) causes the logic component (1030) to NAND 1 and XIN

[11] to output signal (1017) as XINB

[11] . In response, the logic component (1040) NOR XINB

[11] and WB[11:0] to output signal (1019) where all bits are "0" or W[11:0]. For example, when XINB

[11] =1, the signal (1019) is output as 12 bits of "0", which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. When XINB

[11] =0, the signal (1019) is output as W[11:0], which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is W. Note that in the current example, each of the Booth decoders (1020A to 1020F) receives the signal (1007)(W) and the signal (1009)(-W). The signals (1007, 1009) can be presented as NOR(1,WB

[11] ),W[11:0] and WN

[12] ,WN[12:0], respectively, where NOR(1,WB

[11] ),W[11:0] indicates attaching "0" bit(s) to the left of the most significant bit of the weight data element (W[11:0]).

[0072] In FIG. 12, an example is illustrated where the input data element (XIN) is provided as unsigned and the weight data element (W) is provided as signed (i.e., XSIGNED=0 and WSIGNED=1). Accordingly, the signal (1005, XSIGNEDB=1) causes the logic component (1030) to NAND 1 and XIN

[11] to output signal (1017) as XINB

[11] . In response, the logic component (1040) NOR XINB

[11] and WB[11:0] to output signal (1019) where all bits are "0" or W[11:0]. For example, when XINB

[11] =1, the signal (1019) is output as 12 bits of "0", which refers to the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. When XINB

[11] =0, the signal (1019) is output as W[11:0], which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is W. Note that in the current example, each of the Booth decoders (1020A to 1020F) receives the signal (1007)(W) and the signal (1009)(-W). The signals (1007, 1009) can be represented as NOR(0,WB

[11] ),W[11:0] and WN

[12] ,WN[12:0], respectively, where NOR(1,WB

[11] ),W[11:0] indicates attaching additional most significant bit(s) to the left of the most significant bit of the weight data element (W[11:0]).

[0073] In FIG. 13, an example is illustrated where the input data element (XIN) is provided as signed and the weight data element (W) is provided as unsigned (i.e., XSIGNED=1 and WSIGNED=0). Accordingly, the signal (1005, XSIGNEDB=0) causes the logic component (1030) to NAND 0 and XIN

[11] to output the signal (1017) as a logic 1. In response, the logic component (1040) NORs "1" and WB[11:0], regardless of whether XINB

[11] is a logic 1 or 0, to output the signal (1019) having all "0". For example, when XINB

[11] =1, the signal (1019) is output as 12 bits of "0", which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. When XINB

[11] =0, the signal (1019) is still output as 12 bits of "0", which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. Note that in the current example, each of the Booth decoders (1020A to 1020F) receives the signal (1007)(W) and the signal (1009)(-W). The signals (1007, 1009) can be presented as NOR(1,WB

[11] ),W[11:0] and WN

[12] ,WN[12:0], respectively, where NOR(1,WB

[11] ),W[11:0] indicates attaching "0" bit(s) to the left of the most significant bit of the weight data element (W[11:0]).

[0074] In FIG. 14, an example is illustrated where the input data element (XIN) is provided as signed and the weight data element (W) is provided as signed (i.e., XSIGNED=1 and WSIGNED=1). Accordingly, the signal (1005, XSIGNEDB=0) causes the logic component (1030) to NAND 0 and XIN

[11] to output the signal (1017) as a logic 1. In response, the logic component (1040) NORs "1" and WB[11:0], regardless of whether XINB

[11] is a logic 1 or 0, to output the signal (1019) having all "0". For example, when XINB

[11] =1, the signal (1019) is output as 12 bits of "0", which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. When XINB

[11] =0, the signal (1019) is still output as 12 bits of "0", which represents the partial product of a subset containing the most significant bit of the input data element (XIN

[11] ) and the weight data element (W) which is 0. Note that in the current example, each of the Booth decoders (1020A to 1020F) receives the signal (1007)(W) and the signal (1009)(-W). The signals (1007, 1009) can be represented as NOR(0,WB

[11] ),W[11:0] and WN

[12] ,WN[12:0], respectively, where NOR(1,WB

[11] ),W[11:0] indicates attaching additional most significant bit(s) to the left of the most significant bit of the weight data element (W[11:0]).

[0075] FIG. 15 illustrates a flowchart of an exemplary method (1500) for performing a MAC operation on an input data element (XIN) and a weight data element (W) according to various embodiments of the present disclosure. In some embodiments, the input data element (XIN) and the weight data element (W) may each be provided as a signed or unsigned data element. The operations of the method (1500) may be performed, for example, by the components described above in FIG. 10 through 14, and accordingly, some of the reference numerals used above may be reused in the subsequent discussion of the method (1500). Furthermore, it is understood that the method (1500) has been simplified, and accordingly, additional operations may be provided before, during, and after the method (1500) of FIG. 15, and some other operations may be described only briefly in this specification.

[0076] The method (1500) begins with an operation (1510) of receiving a first data element and a second data element. The first data element may be an input data element (XIN), and the second data element may be a weight data element (W). Using the computational circuit (1000) of FIG. 10 as a non-limiting example in which the input data element (XIN) and the weight data element (W) each have 12 bits, Booth encoders (1010A to 1010F) may each receive a subset of the input data element (XIN) (or signal (1001), e.g., XIN[11:0]), and Booth decoders (1020A to 1020F) may receive the weight data element (W) (or signal (1007), e.g., W[11:0]) and its inverse version (-W) (or signal (1009)).

[0077] The method (1500) proceeds to an operation (1520) of identifying whether the first data element is signed or unsigned and whether the second data element is signed or unsigned. In some embodiments, the input data element (XIN) and the weight data element (W) may be received as one of the following combinations: an unsigned input data element and an unsigned weight data element; an unsigned input data element and a signed weight data element; a signed input data element and an unsigned weight data element; and a signed input data element and a signed weight data element. The signed / unsigned input data element may be indicated as XSIGNED, and the signed / unsigned weight data element may be indicated as WSIGNED. For example, whether the input data element is signed or unsigned may be identified by XSIGNED, and whether the weight data element is signed or unsigned may be identified by WSIGNED.

[0078] When identifying whether each of the input data element (XIN) and the weight data element (W) is signed or unsigned (operation (1520)), the method (1500) may proceed to one of the following operations (1532, 1534, 1536 and 1538). Each of the operations (1532 to 1538) will be discussed in more detail below.

[0079] Operation (1532) includes, in response to identifying that the first data element is unsigned and the second data element is unsigned, optionally generating a partial product of the first data element having only the second data element or the second data element equal to "0" and the most significant bit of the first data element. Continuing with the same example, when identifying that the input data element (XIN) is unsigned and the weight data element (W) is unsigned (e.g., XSIGNED=0 and WSIGNED=0), a logic component (1030) (e.g., a NAND2 gate) having inputs provided as XSIGNEDB and XIN

[11] , respectively, may output a signal (1017) indicating XINB

[11] , which may cause a logic component (1040) (e.g., a NOR2 gate) to output a signal (1019) that all bits are equal to "0" or equal to the weight data element (W[11:0]). In various embodiments, the signal (1019) may represent a partial product of a subset of input data elements including a weight data element and its most significant bit.

[0080] Furthermore, operation (1532) comprises providing each of the Booth decoders (1020A to 1020F) with one input (signal (1007)) which is operationally equivalent to W and another input (signal (1009)) which is operationally equivalent to -W. In some embodiments, the computation circuit (1000) may generate signal (1007) using a different NOR. In operation (1532) (where WSIGNEDB=1), signal (1007) may be generated as NOR (1,WB

[11] ),W[11:0], which is equivalent to 0,W[11:0]. Accordingly, at least one "0" bit is attached to the left of the weight data element (W[11:0]). Signal (1009) may be generated as WN

[12] ,WN[12:0], where WN[12:0] is signal (1015). The calculation circuit (1000) can first generate a signal (1015) using another NAND and logic component (1050) (e.g., a half adder). In operation (1532) (where WSIGNED=0), the signal (1015) (WN[12,0]) can be generated as a single bit added to NAND(0,W

[11] ),WB[11:0], which is equivalent to 1,WB[11:0].

[0081] Operation (1534) includes, in response to identifying that the first data element is unsigned and the second data element is signed, optionally generating a partial product of the second data element or a subset of the first data element having the most significant bit of the second data element and the second data element that is equal to "0". Continuing with the same example, when identifying that the input data element (XIN) is unsigned and the weight data element (W) is signed (e.g., XSIGNED=0 and WSIGNED=1), a logic component (1030) (e.g., a NAND2 gate) having inputs provided as XSIGNEDB and XIN

[11] , respectively, may output a signal (1017) indicating XINB

[11] , which may cause a logic component (1040) (e.g., a NOR2 gate) to output a signal (1019) that all bits are equal to "0" or equal to the weight data element (W[11:0]). In various embodiments, the signal (1019) may represent a partial product of a subset of input data elements including a weight data element and its most significant bit.

[0082] Furthermore, operation (1534) includes providing each of the Booth decoders (1020A to 1020F) with one input (signal (1007)) which is operationally identical to W and another input (signal (1009)) which is operationally equivalent to -W. In some embodiments, the computation circuit (1000) may generate signal (1007) using a different NOR. In operation (1534) (where WSIGNEDB=0), signal (1007) may be generated as NOR(0,WB

[11] ),W[11:0], which is identical to W

[11] ,W[11:0]. Accordingly, at least one most significant bit is attached to the left of the weight data element (W[11:0]). Signal (1009) may be generated as WN

[12] ,WN[12:0], where WN[12:0] is signal (1015). The calculation circuit (1000) can first generate a signal (1015) using another NAND and logic component (1050) (e.g., a half adder). In operation (1534) (WSIGNED=1), the signal (1015) (WN[12,0]) can be generated as a single bit added to NAND (1,W

[11] ),WB[11:0], which is identical to WB

[11] ,WB[11:0].

[0083] Operation (1536) comprises generating a partial product of a subset of the first data element having the most significant bit of the first data element and a second data element having the same as "0" in response to identification that the first data element is signed and the second data element is unsigned. Continuing the same example, when identifying that the input data element (XIN) is signed and the weight data element (W) is unsigned (e.g., XSIGNED=1 and WSIGNED=0), a logic component (1030) (e.g., a NAND2 gate) having inputs provided as XSIGNEDB and XIN

[11] , respectively, may output a signal (1017) as "1", which may cause a logic component (1040) (e.g., a NOR2 gate) to output a signal (1019) that all bits are the same as "0". In various embodiments, the signal (1019) may represent a partial product of a subset of the input data element including the weight data element and its most significant bit.

[0084] Furthermore, operation (1536) comprises providing each of the Booth decoders (1020A to 1020F) with one input (signal (1007)) which is operationally equivalent to W and another input (signal (1009)) which is operationally equivalent to -W. In some embodiments, the computational circuit (1000) may generate signal (1007) using a different NOR. In operation (1532) (where WSIGNEDB=1), signal (1007) may be generated as NOR (1,WB

[11] ),W[11:0], which is equivalent to 0,W[11:0]. Accordingly, at least one "0" bit is attached to the left of the weight data element (W[11:0]). Signal (1009) may be generated as WN

[12] ,WN[12:0], where WN[12:0] is signal (1015). The calculation circuit (1000) can first generate a signal (1015) using another NAND and logic component (1050) (e.g., a half adder). In operation (1532) (WSIGNED=0), the signal (1015) (WN[12,0]) can be generated as a single bit added to NAND(0,W

[11] ),WB[11:0], which is equivalent to 1,WB[11:0].

[0085] Operation (1538) comprises generating a partial product of a subset of the first data element having the most significant bit of the first data element and a second data element having the same as "0" in response to identification that the first data element is signed and the second data element is signed. Continuing with the same example, upon identification that the input data element (XIN) is signed and the weight data element (W) is signed (e.g., XSIGNED=1 and WSIGNED=1), a logic component (1030) (e.g., a NAND2 gate) having inputs provided as XSIGNEDB and XIN

[11] , respectively, may output a signal (1017) as "1", which causes a logic component (1040) (e.g., a NOR2 gate) to output a signal (1019) in which all bits are the same as "0". In various embodiments, the signal (1019) may represent a partial product of a subset of the input data element including the weight data element and its most significant bit.

[0086] Furthermore, operation (1538) comprises providing each of the Booth decoders (1020A to 1020F) with one input (signal (1007)) which is operationally identical to W and another input (signal (1009)) which is operationally equivalent to -W. In some embodiments, the computational circuit (1000) may generate signal (1007) using a different NOR. In operation (1538) (where WSIGNEDB=0), signal (1007) may be generated as NOR(0,WB

[11] ),W[11:0], which is identical to W

[11] ,W[11:0]. Accordingly, at least one most significant bit is attached to the left of the weight data element (W[11:0]). Signal (1009) may be generated as WN

[12] ,WN[12:0], where WN[12:0] is signal (1015). The calculation circuit (1000) can first generate a signal (1015) using another NAND and logic component (1050) (e.g., a half adder). In operation (1538) (WSIGNED=1), the signal (1015) (WN[12,0]) can be generated as a single bit added to NAND (1,W

[11] ),WB[11:0], which is identical to WB

[11] ,WB[11:0].

[0087] At any of the operations (1532 to 1538) simultaneously with or subsequently thereto, the method (1500) may further include one or more operations (not shown in FIG. 15 for brevity) for summing all partial products generated by Booth decoders (e.g., Booth decoders (1020A to 1020F)). Next, the adder tree (1060) of the computation circuit (1000) may sum these partial products to generate the final product of the input data element (XIN) and the weight data element (W).

[0088] FIG. 16 illustrates an example of a computation circuit (1600) that processes a signed or unsigned input data element (XIN) and a signed or unsigned weight data element (W). The computation circuit (1600) is substantially similar to the computation circuit (1000) of FIG. 10. In the example of FIG. 16, k bits are provided for each of the input data element (XIN) and the weight data element (W). As such, the number of Booth encoders and the number of Booth decoders of the computation circuit (1600) may vary accordingly. For example, the computation circuit (1600) may include k-2 Booth encoders (1610) and k-2 Booth decoders (1620). Furthermore, the computation circuit (1600) may include other components substantially similar to the components shown in FIG. 10. For example, the computation circuit (1600) also includes a NAND2 gate (1630), a NOR2 gate (1640), a half adder (1650), and a plurality of full adders (1661, 1662, 1663, 1664, 1665, and 1666). For a data element provided with k bits, the corresponding bits of the signals received or otherwise processed by the computation circuit (1600) may change accordingly. These signals (1601, 1603, 1605, 1607, 1609, 1611, 1613, 1615, 1619) are each represented in the form exemplified in FIG. 16. The signals (1601 to 1619) are substantially similar to the signals (1001 to 1019) ( FIG. 10), and thus the corresponding discussion is not repeated.

[0089] FIG. 17 illustrates an exemplary circuit diagram (1700) of a Booth encoder (e.g., 210 in FIG. 2, 300 in FIG. 3, 510 in FIG. 5, 1010A-F in FIG. 10 to 14) according to various embodiments of the present disclosure. Hereinafter, the circuit diagram of FIG. 17 is referred to as the Booth encoder (1700). It should be understood that the circuit diagram of FIG. 17 is a non-limiting implementation of the Booth encoder and is not intended to limit the scope of the present disclosure.

[0090] In some embodiments, the Booth encoder (1700) is a data element (e.g., X 2i+1 , X 2i , and X 2i-1 3-bit Booth encoding can be implemented for a 3-bit subset of ). As illustrated, a first input bit line (e.g., X) having a first signal representing the first bit of the subset 2i-1 A second input bit line having a second signal representing the second bit of ) and a subset (e.g., X 2i ) can be coupled to the input end of an exclusive OR ("XOR") gate (1702). The XOR gate (1702) can receive a first signal and a second signal as inputs and can generate an output as a first intermediate signal ("1x"). A third bit of a subset of the second bit line (e.g., X 2i+1 A third bit line having a third signal representing ) can be coupled to the input end of an exclusive NOR ("XNOR") gate (1708). The XNOR gate (1708) can receive the second signal and the third signal as inputs and generate an output as a second intermediate signal ("2x").

[0091] The first NOR gate (1704) can be coupled to the output end of the XOR gate (1702) and the output end of the XNOR gate (1708) to receive inputs to the first NOR gate (1704). Thus, the first NOR gate (1704) can receive a first intermediate signal (1x) from the XOR gate (1702) and receive a second intermediate signal (2x) as an input from the XNOR gate (1708). The first NOR gate (1704) can generate an output as a Booth-encoded bit ("BE").

[0092] The second NOR gate (1706) can be coupled to the output end of the first NOR gate (1704) to receive Booth-encoded bits (BE) as inputs to the second NOR gate (1706), as well as to the output end of the XOR gate (1702) to receive a first intermediate signal (1x) as an input. Thus, the second NOR gate (1706) can receive the first intermediate signal (1x) from the XOR gate (1702) and the Booth-encoded bits (BE) from the first NOR gate (1704) as inputs. The second NOR gate (1706) can generate an output as an enable bit ("ENB").

[0093] The third NOR gate (1710) can be coupled to the output end of the second NOR gate (1706) at the input end of the third NOR gate (1710) to receive an ENB as an input. The third NOR gate (1710) can also be coupled to the third bit line at the inverted input end to receive an inversion of the third bit line as an input. For example, an inverter can be coupled between the third bit line and the input end of the third NOR gate (1710). Thus, the third NOR gate (1710) can receive an enable bit (ENB) from the second NOR gate (1706) as inputs, and a third signal representing an inversion of a subset of the third bit from the third bit line. In some embodiments, the third NOR gate (1710) can invert the third signal. In some embodiments, the third NOR gate (1710) can receive an inverted third signal from an inverter. The third NOR gate (1710) can generate an output as a select bit ("S").

[0094] FIG. 18 illustrates an exemplary circuit diagram of a Buss decoder (e.g., 220 in FIG. 2, 520 in FIG. 5, 1020A-F in FIG. 10 to 14) according to various embodiments of the present disclosure. Hereinafter, the circuit diagram of FIG. 18 is referred to as a Buss decoder (1800). It should be understood that the circuit diagram of FIG. 18 is a non-limiting implementation of a Buss decoder and is not intended to limit the scope of the present disclosure.

[0095] In some embodiments, the Booth decoder (1800) may be operatively coupled to a corresponding 3-bit Booth encoder (e.g., Booth encoder (1700)) to receive Booth encoded signals, e.g., Booth encoded bit (BE), enable bit (ENB) and select bit (S). As illustrated, the Booth decoder (1800) includes a multiplexer (1810) and an adder (1850).

[0096] The multiplexer (1810) may be coupled to any number of input lines configured to carry weight data elements at the input. For example, the multiplexer (1810) may be coupled to four input lines configured to carry 4-bit weight data elements (e.g., W[3], W[2], W[1], W[0]). The multiplexer (1810) may include a plurality of inverters (1812, 1814) configured to function as buffers for temporary storage of weight data elements. For example, one of the inverters (1812) may be configured to temporarily store weight data elements, and a corresponding inverter among the inverters (1814) may be configured to temporarily store inversions of weight data elements.

[0097] The multiplexer (1810) may be coupled to a selection signal (e.g., a selection bit "S") output by a corresponding Booth encoder on a selection line. The multiplexer (1810) may include a plurality of transmission gates (1816) coupled between the inverters (1812, 1814) and the outputs of the multiplexer (1810). The transmission gates (1816) may also be coupled to a selected signal at the input. The selection signal may determine whether to output from the multiplexer (1810) an input signal or an inversion of each input signal of the input weight data elements (e.g., W[3], W[2], W[1], W[0]). In some embodiments, pairs of transmission gates (1816) coupled to the same output of the multiplexer (1810) may be configured differently to respond to the selection signal. For example, a transmission gate (1810) may be configured to transmit weight data and / or the inverse of weight data elements stored in an inverter (1812), and another transmission gate (1816) may be configured not to transmit weight data elements and / or the inverse of weight data elements stored in an inverter (1814) for the same selection signal, and vice versa. The multiplexer (1810) may output weight data elements and / or the inverse of weight data elements at the output as controlled by the selection signal.

[0098] The adder (1850) may receive, at input, the inverse of the weight data and / or the weight data element output by the multiplexer (1810) (collectively referred to herein as the weight data element for the adder (1850)). The adder (1850) may be coupled to an enable signal (e.g., an enable bit “ENB”) that may be output from a corresponding Booth encoder. The enable signal may trigger the adder (1850) to add the signal received at the inputs to a value (e.g., a shift register) maintained within the adder component (1870). The adder (1850) may include a plurality of NOR gates (1852A, 1852B, 1852C) configured to receive the weight data element at one input and the enable signal at a second input of the NOR gates (1852A-C). NOR gates (1852A-C) may be configured to NOR the weight data element and the enable signal so that the enable signal can control the logic gating operation of the adder (1850). For example, the enable signal may be configured to enable logic gating (e.g., the enable signal is a "1" value), and the NOR gates (1852A-C) may output only "0" values ​​regardless of the value of the weight data. Otherwise, the NOR gates (1852A to 1852C) may output an enable signal configured not to enable the weight data and logic gating at the input (e.g., the enable signal is a "0" value).

[0099] Control of the adder (1850) may be coupled to a Booth encoded bit (e.g., Booth encoded bit "BE") output by a corresponding Booth encoder. The Booth encoded bit may be configured to control whether the adder (1850) performs a left shift operation (e.g., left shift 1 bit). The output of each NOR gate (1852A-C) may be coupled to a shifter (1856). The shifter (1856) may include a plurality of transfer gates (1858) configured to couple the output of each NOR gate to a plurality of inverters (1860). Additionally, the shifter (1856) may be configured to directly couple the inverter (1862) to the output of the NOR gate (1852A) and may include one of the transfer gates (1858) configured to couple the output of the NOR gate (1852A) to one of the inverters (1860). A NOR gate (1852A) may be associated with the input of the most significant bit of a weight data element. An inverter (1860) coupled to the NOR gate (1852A) may correspond to the most significant bit position of the weight data element, and an inverter (1862) coupled to the NOR gate (1852A) may correspond to a higher bit position than the most significant bit position of the weight data element. A shifter (1856) may include another of the transmission gates (1858) configured to combine the output of the NOR gate (1852C) with one of the inverters (1860), and yet another of the transmission gates (1858) configured to combine the output of the NOR gate (1852C) with the inverter (1864). A NOR gate (1852C) may be associated with the input of the least significant bit of a weight data element. An inverter (1864) coupled to the NOR gate (1852C) may correspond to the least significant bit position of the weight data element. The adder (1850) can also be coupled to the supply voltage (VDD).The shifter (1856) may include a transmission gate (1866) configured to combine the supply voltage (VDD) with the inverter (1864).

[0100] Transmission gates (1858, 1866) may also be coupled to Booth encoded (BE) bits. Transmission gates (1858) may be configured to enable and / or prevent the transmission of output from NOR gates (1852A-C) to inverters (1860, 1864). Transmission gate (1866) may be configured to enable and / or prevent the transmission of supply voltage to inverters (1864). In some embodiments, pairs of transmission gates (1858, 1866) coupled to the same inverters (1860, 1864) may be configured differently to respond to Booth encoded bits.

[0101] In one aspect of the present disclosure, a memory device is disclosed. The memory circuit includes a Booth encoder configured to receive a first data element comprising a first code portion and a first data portion. The memory circuit includes a Booth decoder configured to receive a second data element comprising a second code portion and a second data portion and to provide a product based on the first data element and the second data element. The memory circuit includes a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder. The plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and to change individual logic states of the plurality of encoded signals based on the first code portion and the second code portion, thereby causing the Booth decoder to provide a product.

[0102] In another aspect of the present disclosure, a memory circuit is disclosed. The memory circuit includes a memory array. The memory circuit includes a computation circuit coupled to the memory array. The computation circuit includes: a Booth encoder configured to receive a first data element comprising a first code bit and a plurality of first data bits and configured to provide a plurality of encoded values ​​based on the plurality of first data bits; a Booth decoder configured to retrieve a second data element comprising a second code bit and a plurality of second data bits from the memory array and provide a plurality of partial products based on multiplying the first data element by the second data element; and a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder. Each of the plurality of multiplexers is configured to select a first encoded value or a second encoded value among the encoded values ​​based on a logically processed signal of the first code bit and the second code bit.

[0103] In another aspect of the present disclosure, a method for operating a memory circuit is disclosed. The method comprises receiving a first data element and a second data element, wherein the first data element comprises a first code bit and a plurality of first data bits, and the second data element comprises a second code bit and a plurality of second data bits. The method comprises encoding a plurality of first data bits to generate a plurality of encoded values, each of which corresponds to an individual combination of logical states of a subset of the first data bits. The method comprises selecting between a first encoded value and a second encoded value among a plurality of encoded values ​​that are opposite to each other, based on a logically processed signal of the first code bit and the second code bit. The method comprises multiplying the selected first encoded value or the second encoded value by the second data bit.

[0104] As used herein, the terms “about” and “approximately” generally indicate a given amount value that may vary based on a specific technical node associated with the semiconductor device. Based on a specific technical node, the term “about” may indicate a given amount value that varies within, for example, 10-30% of the value (e.g., ±10%, ±20%, or ±30%).

[0105] The foregoing describes an overview of the features of some embodiments to enable those skilled in the art to better understand the aspects of the present disclosure. Those skilled in the art should understand that they can readily use the present disclosure as a basis for designing or modifying other processes and structures to perform the same purpose or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent configurations may make various changes, substitutions, and modifications to the present disclosure without departing from the spirit and scope of the present disclosure.

[0106] Examples

[0107] Example 1. In a memory circuit,

[0108] A Booth encoder configured to receive a first data element including a first code portion and a first data portion;

[0109] A Booth decoder configured to receive a second data element including a second code portion and a second data portion, and to provide a product based on the first data element and the second data element; and

[0110] A plurality of multiplexers operatively coupled between the above-mentioned Booth encoder and the above-mentioned Booth decoder

[0111] Includes,

[0112] A memory circuit wherein the plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and to change the individual logic states of the plurality of encoded signals based on the first code portion and the second code portion, thereby enabling the Booth decoder to provide the product.

[0113] Example 2. In Example 1,

[0114] A memory circuit in which each of the above multiplexers is controlled by the XOR signal of the first code portion and the second code portion.

[0115] Example 3. In Example 1,

[0116] A memory circuit wherein each of the above multiplexers has a first input and a second input configured to receive a first combination of logic states of the encoded signals and a second combination of logic states of the encoded signals, respectively.

[0117] Example 4. In Example 3,

[0118] A memory circuit in which the first combination corresponds to a first encoded value multiplied by the first data portion, and the second combination of the encoded signals corresponds to a second encoded value multiplied by the second data portion.

[0119] Example 5. In Example 4,

[0120] A memory circuit in which the first encoded value and the second encoded value are opposite to each other.

[0121] Example 6. In Example 3,

[0122] A memory circuit wherein each of the above multiplexers is configured to select the first combination in response to receiving an XOR signal of the first code portion and the second code portion identical to the first logic state.

[0123] Example 7. In Example 6,

[0124] A memory circuit wherein each of the above multiplexers is configured to select the second combination in response to receiving an XOR signal of the first code part and the second code part identical to the second logic state.

[0125] Example 8. In Example 1,

[0126] A memory circuit in which the number of the above multiplexers corresponds to the number of the above first data portions.

[0127] Example 9. In Example 1,

[0128] A memory circuit in which the first data element represents a plurality of input activations received by a memory array, and the second data element represents a plurality of weights stored in the memory array.

[0129] Example 10. In Example 1,

[0130] A memory circuit wherein the first data portion represents a plurality of first mantissa bits of the first data element, and the second data portion represents a plurality of second mantissa bits of the second data element.

[0131] Example 11. In a memory circuit,

[0132] memory array; and

[0133] Computation circuit coupled to the above memory array

[0134] Includes, and the above calculation circuit is:

[0135] A Booth encoder configured to receive a first data element comprising a first code bit and a plurality of first data bits, and configured to provide a plurality of encoded values ​​based on the plurality of first data bits;

[0136] A Booth decoder configured to retrieve a second data element including a second sign bit and a plurality of second data bits from the memory array and to provide a plurality of partial products based on multiplying the first data element by the second data element; and

[0137] A plurality of multiplexers operatively coupled between the above-mentioned Booth encoder and the above-mentioned Booth decoder—each of the plurality of multiplexers is configured to select a first encoded value or a second encoded value among the encoded values ​​based on a logically processed signal of the first code bit and the second code bit.

[0138] A memory circuit that includes

[0139] Example 12. In Example 11,

[0140] The above-described Booth decoder is also configured to multiply the second data element by the first encoded value or the second encoded value selected for the corresponding partial product among the plurality of partial products.

[0141] Example 13. In Example 11,

[0142] The first multiplexer among the above multiplexers is:

[0143] (i) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 0, select a first encoded value among the encoded values ​​corresponding to a first combination of logic states of a subset of the first data bits;

[0144] (ii) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 1, select a second encoded value among the encoded values ​​corresponding to a second combination of logic states of a subset of the first data bits.

[0145] A memory circuit that is configured.

[0146] Example 14. In Example 13,

[0147] The second multiplexer among the above multiplexers is:

[0148] (i) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 0, select the second encoded value;

[0149] (ii) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 1, select the first encoded value.

[0150] A memory circuit that is configured.

[0151] Example 15. In Example 14,

[0152] The third multiplexer among the above multiplexers is:

[0153] (i) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 0, select a third encoded value among the encoded values ​​corresponding to a third combination of the logic states of the subset of the first data bits;

[0154] (ii) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 1, select a fourth encoded value among the encoded values ​​corresponding to a fourth combination of logic states of a subset of the first data bits.

[0155] A memory circuit that is configured.

[0156] Example 16. In Example 15,

[0157] The third multiplexer among the above multiplexers is:

[0158] (i) When identifying that the XOR signal of the first sign bit and the second sign bit is the same as logic 0, select the fourth encoded value;

[0159] (ii) When identifying that the XOR signal of the first code bit and the second code bit is the same as logic 1, select the third encoded value.

[0160] A memory circuit that is configured.

[0161] Example 17. In Example 11,

[0162] A memory circuit in which a plurality of the above-mentioned multiplexers correspond to a plurality of the above-mentioned first data bits.

[0163] Example 18. In Example 11,

[0164] A memory circuit in which the first data bits represent a plurality of first mantissa bits of the first data element, and the second data bits represent a plurality of second mantissa bits of the second data element.

[0165] Example 19. In the method,

[0166] A step of receiving a first data element and a second data element - the first data element includes a first code bit and a plurality of first data bits, and the second data element includes a second code bit and a plurality of second data bits - ;

[0167] A step of encoding the plurality of first data bits to generate a plurality of encoded values ​​- each of the encoded values ​​corresponds to an individual combination of logical states of a subset of the first data bits - ;

[0168] A step of selecting between a first encoded value and a second encoded value among a plurality of encoded values ​​that are opposite to each other, based on the logically processed signals of the first code bit and the second code bit; and

[0169] A step of multiplying the selected first encoded value or second encoded value by the second data bits.

[0170] A method including

[0171] Example 20. In Example 19,

[0172] A method in which the first data bits represent a plurality of first mantissa bits of the first data element, and the second data bits represent a plurality of second mantissa bits of the second data element.

Claims

Claim 1 A memory circuit comprising: a Booth encoder configured to receive a first data element including a first code portion and a first data portion; a Booth decoder configured to receive a second data element including a second code portion and a second data portion and to provide a product based on the first data element and the second data element; and a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder, wherein the plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and to change individual logic states of the plurality of encoded signals based on the first code portion and the second code portion, thereby causing the Booth decoder to provide the product. Claim 2 A memory circuit according to claim 1, wherein each of the multiplexers is controlled by an XOR signal of the first code portion and the second code portion. Claim 3 A memory circuit according to claim 1, wherein each of the multiplexers has a first input and a second input configured to receive a first combination of logic states of the encoded signals and a second combination of logic states of the encoded signals, respectively. Claim 4 A memory circuit according to paragraph 3, wherein the first combination corresponds to a first encoded value multiplied by the first data portion, and the second combination of the encoded signals corresponds to a second encoded value multiplied by the second data portion. Claim 5 A memory circuit according to paragraph 3, wherein each of the multiplexers is configured to select the first combination in response to receiving an XOR signal of the first code part and the second code part identical to the first logic state. Claim 6 A memory circuit according to claim 1, wherein the number of multiplexers corresponds to the number of the first data portion. Claim 7 A memory circuit according to claim 1, wherein the first data element represents a plurality of input activations received by a memory array, and the second data element represents a plurality of weights stored in the memory array. Claim 8 A memory circuit according to claim 1, wherein the first data portion represents a plurality of first mantissa bits of the first data element and the second data portion represents a plurality of second mantissa bits of the second data element. Claim 9 A memory circuit comprising: a memory array; and a computational circuit coupled to the memory array, wherein the computational circuit comprises: a Booth encoder configured to receive a first data element comprising a first code bit and a plurality of first data bits and configured to provide a plurality of encoded values ​​based on the plurality of first data bits; a Booth decoder configured to retrieve a second data element comprising a second code bit and a plurality of second data bits from the memory array and provide a plurality of partial products based on multiplying the first data element by the second data element; and a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder, wherein each of the plurality of multiplexers is configured to select a first encoded value or a second encoded value among the encoded values ​​based on a logically processed signal of the first code bit and the second code bit. Claim 10 A method comprising: receiving a first data element and a second data element, wherein the first data element comprises a first code bit and a plurality of first data bits, and the second data element comprises a second code bit and a plurality of second data bits; encoding the plurality of first data bits to generate a plurality of encoded values, wherein each of the encoded values ​​corresponds to an individual combination of logical states of a subset of the first data bits; selecting between a first encoded value and a second encoded value among the plurality of encoded values, which are opposite to each other, based on a logically processed signal of the first code bit and the second code bit; and multiplying the second data bits by the selected first encoded value or the second encoded value.

Citation Information

Patent Citations

  • Method and a system for performing calculationoperations and a device

    KR1020050065672A