Memory circuit and operating method thereof
By using Booth encoder, decoder and symbol-aware multiplexer in the memory calculation circuit, the multiplication operation of signed data is directly processed, and the complexity and efficiency of existing CIM circuit design is solved, and efficient floating-point data operation is achieved.
Patent Information
- Application Number
- CN202510004581.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-22
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
Existing computing (CIM) circuits in memory increase in design complexity and size when processing multiplication-accumulation operations of floating-point data types, especially due to the introduction of two's complement circuits, resulting in poor efficiency.
The Booth encoder and Booth decoder are combined with the symbol-aware multiplexer, and the multiplication operation of the input data elements and the weighted data elements is directly processed through the Booth algorithm, without the need for two's complement conversion, simplifying the design and improving efficiency.
Efficient multiplication operations on signed data elements are realized, reducing the complexity and size of circuit design, while maintaining high computational parallelism and energy efficiency.
Smart Images

Figure CN120255847A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to memory circuits and methods of operating the same. Background Art
[0002] Computer artificial intelligence (AI) is based on machine learning, such as using deep learning techniques. Through machine learning, a computing system organized as a neural network calculates the statistical likelihood that input data matches previously calculated data. A neural network refers to many interconnected processing nodes that can analyze data and compare the input with "training" data. Training data refers to calculating the attributes of known data to develop a model for comparing input data. An example of an AI and data training application is in object recognition, where the system analyzes the attributes of many (e.g., thousands or more) images to determine patterns that can be used to perform statistical analysis to identify the input object. Summary of the Invention
[0003] According to one aspect of an embodiment of the present application, a memory circuit is provided, including: a Booth encoder configured to receive a first data element including a first symbol part and a first data part; a Booth decoder configured to receive a second data element including a second symbol part and a second data part, and provide a product based on the first data element and the second data element; and a plurality of multiplexers operably coupled between the Booth encoder and the Booth decoder; wherein the plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and change corresponding logic states of the plurality of encoded signals based on the first symbol part and the second symbol part, so that the Booth decoder provides the product.
[0004] According to one aspect of an embodiment of the present application, a memory circuit is provided, including: a memory array; and a computing circuit coupled to the memory array, wherein the computing circuit includes: a Booth encoder configured to receive a first data element including a first symbol bit and a plurality of first data bits, and configured to provide a plurality of encoded values based on the plurality of first data bits; a Booth decoder configured to retrieve a second data element including a second symbol bit and a plurality of second data bits from the memory array, and provide a plurality of partial products based on multiplying the first data element by the second data element; and a plurality of multiplexers operably coupled between the Booth encoder and the Booth decoder, wherein each of the plurality of multiplexers is configured to select a first encoded value among the encoded values or a second encoded value among the encoded values based on a logical processing signal of the first symbol bit and the second symbol bit.
[0005] According to another aspect of the embodiments of the present application, a method of operating a memory circuit is provided, including: receiving a first data element and a second data element, where the first data element includes a first sign bit and a plurality of first data bits, and the second data element includes a second sign bit and a plurality of second data bits; encoding the plurality of first data bits to generate a plurality of encoded values, where each of the encoded values corresponds to a respective combination of the logical states of a subset of the first data bits; selecting between a first encoded value and a second encoded value that are opposite to each other among the plurality of encoded values based on a logical processing signal of the first sign bit and the second sign bit; and multiplying the second data bits by the selected first encoded value or second encoded value. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, according to standard practice in the industry, the various components are not drawn to scale and are for illustrative purposes only. In fact, for clarity of discussion, the dimensions of the various components may be increased or decreased arbitrarily.
[0007] Figure 1 An example block diagram of a compute-in-memory (CIM) circuit according to some embodiments is shown.
[0008] Figure 2 One of the compute blocks of a CIM circuit according to some embodiments is shown Figure 1 in block diagram.
[0009] Figure 3 A component block diagram of Booth encoding of data elements for Booth multiplication according to some embodiments is shown.
[0010] Figure 4 A table summarizing Booth encoding of data elements for Booth multiplication according to some embodiments is shown.
[0011] Figure 5 An example implementation of a compute block according to some embodiments is shown Figure 1 in schematic diagram.
[0012] Figure 6 A table summarizing Booth encoding of data elements for Booth multiplication according to some embodiments is shown.
[0013] Figure 7 A symbol-aware multiplexer circuit diagram of a compute block according to some embodiments is shown Figure 5 in block diagram.
[0014] Figure 8 A block diagram of a plurality of compute blocks including Figure 5 is shown according to some embodiments.
[0015] Figure 9 shows a flow chart of an example method for operating a Figure 5 computing block according to some embodiments.
[0016] Figure 10 shows an example implementation of a Figure 1 computing circuit according to some embodiments.
[0017] Figure 11 , Figure 12 , Figure 13 and Figure 14 respectively show different combinations of signed / unsigned data elements processed by a Figure 10 computing circuit according to some embodiments.
[0018] Figure 15 shows a flow chart of an example method for operating a Figure 10 computing circuit according to some embodiments.
[0019] Figure 16 shows different combinations of signed / unsigned data elements processed by a Figure 10 computing circuit according to some embodiments.
[0020] Figure 17 shows an example circuit diagram of a Booth encoder according to some embodiments.
[0021] Figure 18 shows an example circuit diagram of a Booth decoder according to some embodiments. DETAILED DESCRIPTION
[0022] The following disclosure provides many different embodiments or examples for implementing the present disclosure. Specific embodiments or examples of components and arrangements are described below to simplify the present disclosure. Of course, these are merely examples and are not intended to be limiting. For example, in the following description, forming a first component above or on a second component may include embodiments where the first component and the second component are formed in direct contact, and may also include embodiments where additional components may be formed between the first component and the second component such that the first component and the second component may not be in direct contact. Additionally, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for simplicity and clarity purposes and does not itself indicate a relationship between the various embodiments and / or configurations discussed.
[0023] In addition, for ease of description, in this article, relative positional relationship terms such as "below", "beneath", "lower part", "above", "upper part", etc. may be used to describe the relationship between one element or component and another element or component as shown in the figures. Except for the orientations shown in the figures, relative positional relationship terms are intended to include different orientations of the device during use or operation. The device may be positioned otherwise (rotated 90 degrees or in other orientations), and the relative positional relationship descriptors used in this article may be interpreted accordingly.
[0024] Unless otherwise specified, the terms "processor", "processor core", "controller", and "control unit" are used interchangeably in this article, referring to software-configured processors, hardware-configured processors, general-purpose processors, dedicated processors, single-core processors, homogeneous multi-core processors, cores of heterogeneous multi-core processors, microprocessors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), etc., controllers, microcontrollers, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), other programmable logic devices, discrete gate logic, transistor logic, etc. A processor may be an integrated circuit, which may be configured such that the components of the integrated circuit reside on a single semiconductor material (such as silicon).
[0025] A neural network calculates "weights" for calculations on new data (input data "words"). A neural network uses multiple layers of computing nodes, where deeper layers perform calculations based on the results of calculations performed by higher layers. Machine learning currently relies on the calculations of dot products and vector absolute differences, usually calculated by performing multiply-accumulate (MAC) operations on parameters, input data, and weights. The calculations of large deep neural networks usually involve so many data elements that it is impractical to store them in the processor cache. Therefore, these data elements are usually stored in memory.
[0026] Therefore, machine learning is very intensive in calculating and comparing many different data elements. The operation calculations within the processor are several orders of magnitude faster than the transfer of data elements between the processor and the main memory resources. Due to the memory size required to store data elements, placing all data elements closer to the processor in the cache is very expensive for the vast majority of practical systems. Therefore, the transfer of data elements becomes the main bottleneck in AI calculations. As the data set increases, the time and power / energy used by the computing system to move data elements may ultimately be multiples of the time and power used for actual calculations.
[0027] In this regard, in-memory computing (CIM) circuits or systems have been proposed to perform such MAC operations. Instead, CIM circuits perform in-situ data processing within a suitable memory circuit. CIM circuits suppress the latency of data / program extraction and output result upload to the corresponding memory (e.g., memory array), thus solving the memory (or von Neumann) bottleneck of traditional computers. Another key advantage of CIM circuits is high computational parallelism, which is due to the specific architecture of the memory array, where computations can be performed simultaneously along multiple current paths. CIM circuits also benefit from the high density of multiple memory arrays with computing devices, which typically have excellent scalability and 3D integration capabilities. As a non-limiting example, CIM circuits for various machine learning applications can perform MAC operations locally within the memory (i.e., without sending data elements to the main processor) to achieve higher throughput dot products of neuron activations and weight matrices, while still providing higher performance and lower energy compared to the computations of the main processor.
[0028] The data elements processed by CIM circuits have various data types or forms, such as integer data types and floating-point data types. The sizes of integer data types (each integer data type representing a range of mathematical integers) may vary. For example, integer data types include 4 bits (sometimes called INT4 data type), 8 bits (sometimes called INT8 data type), etc. Floating-point data types are typically represented by a sign part, an exponent part, and a significand (mantissa) part consisting of the significant digits of the number. For example, one floating-point format specified by the Institute of Electrical and Electronics Engineers ( ) has a size of 16 bits (sometimes called FP16 data type), which includes 10 mantissa bits, 5 exponent bits, and 1 sign bit. Another floating-point format also has a size of 16 bits, sometimes called BF16 data type, which includes 7 mantissa bits, 8 exponent bits, and 1 sign bit.
[0029] In machine learning applications, CIM circuits are often configured to process dot product multiplications based on performing MAC operations on a large number of data elements that may be of floating-point data type (e.g., input word vectors and weight matrices), and then process the addition (or accumulation) of these dot products. Few CIM circuits have been proposed to handle MAC operations on data elements provided in floating-point data types. For example, it has been proposed to integrate a Booth multiplier into the CIM circuit, which operates in parallel with multiple stages to produce the final product.
[0030] A Booth multiplier typically operates according to the principle of Booth's algorithm. Booth's algorithm multiplies two signed binary numbers. As is typical in binary multiplication, Booth's algorithm generates partial products of the multiplicand multiplied by the multiplier, and these partial products are shifted and summed to produce the final product. Booth's algorithm uses rules based on the bit group values of the multiplier to determine the operations for generating partial products using the multiplicand. To compute the final product, after generating all partial products, the Booth multiplier typically shifts the partial products by corresponding bits and outputs the shifted partial products to an adder tree to sum the shifted partial products.
[0031] When processing signed data elements (sometimes referred to as signed data elements), existing CIM circuits typically require at least one two's complement circuit that is operatively coupled between a corresponding Booth multiplier and a corresponding adder tree. For example, in existing CIM circuits, the Booth multiplier generates partial products based on the corresponding unsigned parts of the input data element and the weight data element, and provides these partial products to the two's complement circuit. The two's complement circuit then determines whether to perform a two's complement conversion based on the corresponding sign parts of the input data element and the weight data element. For example, if the input data element and the weight data element have the same sign, the two's complement circuit is disabled to change the polarity of the partial products; if the input data element and the weight data element have different signs, the two's complement circuit is activated to change the polarity of the partial products. Such a two's complement circuit typically includes at least one additional half adder, which greatly increases the design complexity of the CIM circuit and unfavorably increases the size of the CIM. Therefore, existing CIM circuits employing Booth multipliers are not entirely satisfactory in some respects.
[0032] The present disclosure provides various embodiments of in-memory computing (CIM) circuits configured to process multiple input data elements and multiple weight data elements. In one aspect, as disclosed herein, a CIM circuit can perform in-memory computations (e.g., multiply-accumulate (MAC) operations) on input data elements and weight data elements, each of which can be assigned a sign (e.g., signed input data elements and signed weight data elements) without performing the binary complement conversion described above. The disclosed CIM circuit can multiply an input data element by a weight data element based on a plurality of sign-aware Booth decoding values. For example, a CIM circuit can include a Booth encoder, a Booth decoder (sometimes referred to as a Booth multiplier), and a plurality of sign-aware multiplexers coupled between the Booth encoder and the Booth decoder. The Booth encoder can first generate a plurality of Booth encoding values based on the input data element (e.g., the mantissa portion of the input data element if a floating-point data type is provided). The sign-aware multiplexer can determine whether to forward the Booth encoding value directly to the Booth decoder (without inversion), or invert the Booth encoding value and then provide the inverted Booth encoding value to the Booth decoder based on the XOR (exclusive OR) signal of the respective sign portions of the input data element and the weight data element. After receiving such sign-aware decoded signals, the Booth decoder can multiply the decoded signals (representing the input data element) by the weight data element (e.g., the mantissa portion of the weight data element if a floating-point data type is provided) to generate a plurality of partial products to be summed for the final product.
[0033] In another aspect, as disclosed herein, a CIM circuit can perform MAC operations on input data elements and weight data elements (each element can be provided as signed or unsigned). The disclosed CIM circuit can multiply the input data elements by the weight data elements and selectively perform sign extension depending on whether the input / weight data elements are provided as signed or unsigned. As a representative example, if the input data element is provided as unsigned, the CIM circuit can determine not to perform sign extension on the input data element. Instead, the CIM circuit can append one or more additional "0" bits to the most significant bit (MSB) of the input data element. If the input data element is provided as signed, the CIM circuit can determine to perform sign extension on the input data element. For example, the CIM circuit can include a Booth encoder, a Booth decoder (sometimes referred to as a Booth multiplier), and a plurality of logic gates. The Booth encoder can first generate a plurality of Booth encoded values based on the input data element and provide the Booth encoded values to the Booth decoder. Additionally, some of the logic gates coupled to the Booth decoder can determine whether the input data element is provided as signed or unsigned. If it is signed, these logic gates can cause the CIM circuit to perform sign extension on the input data element by appending additional bits that are the same as the most significant bit of the input data element to the input data element. If it is unsigned, these logic gates can cause the CIM circuit not to perform sign extension on the weight data element by appending one or more "0" bits to the most significant bit of the input data element.
[0034] Figure 1 FIG. shows a block diagram of a compute-in-memory (CIM) circuit 100 according to various embodiments of the present disclosure. In Figure 1 the illustrated embodiment, the CIM circuit 100, also referred to as the memory circuit 100, includes various components that are collectively configured to perform compute-in-memory operations (e.g., multiply-accumulate (MAC) operations) on an input word vector and a weight matrix. The input word vector can include a plurality of input data elements XIN, and the weight matrix can include a plurality of weight data elements W.
[0035] In some embodiments, each of the input data elements XIN and the weight data elements W can be configured or provided in an INT8 data type. In some embodiments, each of the input data elements XIN and the weight data elements W can be configured or provided in an INT4 data type. In some embodiments, each of the input data elements XIN and the weight data elements W can be configured or provided in an FP16 data type. In some embodiments, each of the input data elements XIN and the weight data elements W can be configured or provided in a BF16 data type.
[0036] As shown, the CIM circuit 100 includes a memory circuit 102, an input circuit 104, a computing circuit 106, and an adder circuit (or adder tree) 108. Figure 1 Each component shown (e.g., 102 to 108) is an electronic circuit including logic circuitry configured to perform a corresponding function. In some embodiments, the computing circuit 106 can use Booth's algorithm to provide multiple partial products based on multiplying a multiplicand (e.g., input data element XIN) by a multiplier (e.g., weight data element W). It should be understood that Figure 1 the block diagram of the circuit depicted is simplified, and thus, the CIM circuit 100 can include any of a variety of other components while still being within the scope of the present disclosure.
[0037] The memory circuit 102 can include one or more memory arrays and one or more corresponding circuits. Each memory array is a storage device including a plurality of storage elements 103, each storage element 103 including an electrical, electromechanical, electromagnetic, or other device configured to store one or more data elements, each data element including one or more data bits represented by a logic state. In some embodiments, the logic state corresponds to the voltage level of the charge stored in part or all of the storage element 103. In some embodiments, the logic state corresponds to the physical property of part or all of the storage element 103, such as resistance or magnetic orientation.
[0038] In some embodiments, the storage element 103 includes one or more static random access memory (SRAM) cells. In various embodiments, the SRAM cell includes a plurality of transistors, such as a five-transistor (5T) SRAM cell, a six-transistor (6T) SRAM cell, an eight-transistor (8T) SRAM cell, a nine-transistor (9T) SRAM cell, etc. In some embodiments, the storage element 103 includes one or more dynamic random access memory (DRAM) cells, resistive random access memory (RRAM) cells, magnetoresistive random access memory (MRAM) cells, ferroelectric random access memory (FeRAM) cells, NOR flash memory cells, NAND flash memory cells, conductive-bridge random access memory (CBRAM) cells, data buffers, non-volatile memory (NVM) cells, 3D NVM cells, or other memory cell types capable of storing bit data.
[0039] In addition to the memory array, the memory circuit 102 may further include a plurality of circuits to access or otherwise control the memory array. For example, the memory circuit 102 may include a plurality of (e.g., word line) drivers operatively coupled to the memory array. The drivers may apply signals (e.g., voltages) to corresponding storage elements 103 to allow access (e.g., programming, reading, etc.) to these storage elements 103. For another example, the memory circuit 102 may include a plurality of programming circuits and / or reading circuits operatively coupled to the memory array.
[0040] The memory arrays of the memory circuit 102 are each configured to store a plurality of weight data elements W. In some embodiments, the programming circuits may write the weight data elements W into the corresponding storage elements 103 of the memory array respectively, while the reading circuits may read the bits written into the storage elements 103 to verify or otherwise test whether the written weight data elements W are correct. The drivers of the memory circuit 102 may include or be operatively coupled to a plurality of input activation latches configured to receive and temporarily store input data elements XIN. In some other embodiments, such input activation latches may be part of the input circuit 104, which may further include a plurality of buffers configured to temporarily store the weight data elements W retrieved from the memory arrays of the memory circuit 102. Thus, the input circuit 104 may receive the input data elements XIN and the weight data elements W.
[0041] In some embodiments, the input word vectors (including, e.g., the input data elements XIN) and the weight matrices (including, e.g., the weight data elements W) for which the CIM circuit 100 is configured to perform MAC operations may be configured as any of at least the following data types: INT8 data type, INT4 data type, FP16 data type, and BF16 data type. However, it should be understood that each of the input data elements XIN and the weight data elements W may have any of a variety of other integer or floating-point data types, such as INT16 data type, UINT16 data type, UINT8 data type, UINT4 data type, FP32 data type, FP64 data type, FP128 data type, etc., while still being within the scope of the present disclosure.
[0042] When configured for the INT8 data type, each of the input data element XIN and the weight data element W includes 8 bits, with the leftmost bit serving as its sign bit. When configured for the INT4 data type, each of the input data element XIN and the weight data element W includes 4 bits, with the leftmost bit serving as its sign bit. If configured for the UINT8 data type, both the input data element XIN and the weight data element W include 8 bits, and no bit represents a sign. When configured for the UINT4 data type, each of the input data element XIN and the weight data element W includes 4 bits, and no bit represents a sign. When configured for the FP16 data type, each of the input data element XIN and the weight data element W includes 1 sign bit, 5 exponent bits, and 10 mantissa bits. When configured for the BF16 data type, each of the input data element XIN and the weight data element W includes 1 sign bit, 8 exponent bits, and 7 mantissa bits.
[0043] Still referring to Figure 1 , the input circuit 104 is configured to output all of the input data element XIN and the weight data element W to the computing circuit 106. In some embodiments of the present disclosure, the computing circuit 106 may include a plurality of computing blocks corresponding to the number of bits of the input data element XIN. Each computing block may include a Booth encoder, a plurality of sign-aware multiplexers, and a Booth decoder, which are jointly configured to generate at least one partial product, which will be discussed in further detail below with reference to Figure 5 In some other embodiments of the present disclosure, the computing circuit 106 may include a plurality of Booth encoders and a corresponding number of Booth decoders. In such embodiments, the computing circuit 106 may further include a plurality of logic gates configured to process the input data element XIN and the weight data element W regardless of whether they are provided as signed or unsigned in order to determine whether to perform sign extension on the weight data element W and / or the input data element XIN. Details of these embodiments will be discussed below with reference to Figure 10 The adder tree 108 may receive the partial products from the computing circuit 106 and add them together to generate the final product (P) of the input data element XIN and the weight data element W.
[0044] Figure 2 FIG. 200 shows a block diagram of one of the computing blocks of the computing circuit 106 (hereinafter referred to as "computing block 200") according to various embodiments of the present disclosure. As described above, the computing block 200 (or the computing block of the computing circuit 106) may receive the input data element XIN and the weight data element W from the input circuit 104, generate a plurality of partial products based on the Booth algorithm, and provide the partial products to the adder tree 108 to generate the final product. It should be understood that Figure 2The block diagram of the computing circuit 200 depicted has been simplified; thus, the computing circuit 200 can include any of a variety of other components (e.g., sign-aware multiplexers) while still being within the scope of the present disclosure.
[0045] As shown, the computing circuit 200 includes a Booth encoder 210 and a Booth decoder 220. The Booth encoder 210 can receive a multiplicand (e.g., the input data element XIN and / or a subset of the input data element XIN). Each of the Booth encoder 210 and the Booth decoder 220 can be a combination of circuits or logic components (e.g., Figure 17 and Figure 18 ). The Booth encoder 210 can generate and output a plurality of Booth-encoded signals from the multiplicand (e.g., can include an enable bit, Booth-encoded bits, and selection bits). Different combinations of the logical states of the Booth-encoded signals can correspond to respective Booth-encoded values. The Booth decoder 220 can receive a multiplier (e.g., the weight data element W and / or a subset of the weight data element W). The Booth decoder 220 can also receive the Booth-encoded signals from the Booth encoder 210 and multiply the multiplier by the respective Booth-encoded values to generate partial products (PP). In one aspect of the present disclosure (e.g., Figure 5 ), the Booth-encoded values received by the Booth decoder 220 can be forwarded or selected by a plurality of sign-aware multiplexers coupled between the Booth encoder 210 and the Booth decoder 200. In another aspect of the present disclosure (e.g., Figure 10 ), the Booth-encoded values can be directly received by the Booth decoder 220, e.g., without passing through a sign-aware multiplexer.
[0046] Figure 3 Illustrates an example of Booth encoding of input data elements for Booth multiplication in a CIM circuit (e.g., Figure 1 's 100) according to various embodiments of the present disclosure. As shown, the Booth encoder 300 (e.g., Figure 2 's implementation of the Booth encoder 210) can encode or otherwise convert the data element 310 into a plurality of Booth-encoded signals 320 corresponding to one of a plurality of Booth-encoded values (e.g., 0, -1, 1, -2, 2).
[0047] In some embodiments, the data element 310 can include one or more input data elements XIN, which serve as the multiplicand of the CIM circuit, while one or more corresponding weight data elements W can serve as the multiplier. In some other embodiments, the data element 310 can include one or more weight data elements W, which serve as the multiplicand of the CIM circuit, while one or more corresponding input data elements XIN can serve as the multiplier. The following discussion will focus on examples of encoding the input data element XIN (i.e., the input data element XIN serves as the multiplicand and the weight data element W serves as the multiplier).
[0048] The Booth encoder 300 can encode the input data element XIN 310 in various cycles, where the Booth encoder 30 can encode subsets 302, 304 of the input data element XIN 310. Encoding the input data element XIN 310 by converting it into a Booth encoded signal 320 associated with a finite number of operations for performing Booth multiplication in the CIM circuit can simplify the input data element XIN 310. As further described herein, the Booth encoder 300 can convert each of subsets 302 and 304 into a plurality of Booth encoded signals 320 that together correspond to the respective Booth encoded values. The Booth encoded signals 320 can be configured to control other parts of the corresponding CIM circuit (the Booth decoder, such as Figure 2 220) such that the Booth decoder multiplies the weight data element W by the corresponding Booth encoded value to generate a partial product.
[0049] In some embodiments, subsets 302 and 304 of the input data element XIN 310 can overlap. In some embodiments, subsets 302 and 304 can be centered on a bit position and include the bit position immediately before and the bit position immediately after the bit position. For subset 302 centered on the least significant bit (LSB) of the input data element XIN 310, a "0" bit can be added to the input data element ZIN 310 to fill the bit position immediately before the least significant bit.
[0050] Figure 3 A non - limiting example of 3 - bit Booth encoding is shown, encoding 3 - bit subsets 302, 304 of the input data element XIN 310. The multiplication operation performed by a part of the CIM circuit (e.g., Figure 2 the Booth decoder 220) can be the multiplication of the input data element XIN and the weight data element W. The input data element XIN 310 can be of any bit length "p" such that the input data element XIN 310 can include bits X p-1 , ……, X0.
[0051] In Figure 3In the example shown, the input data element XIN 310 has 4 bits, i.e., p = 4. The Booth encoder 300 can encode subsets 302, 304 of the input data element XIN 310 in various loops, where each of the subsets 302 and 304 has 3 bits. Each subset 302, 304 can be used to generate a corresponding number of Booth encoded signals 320. For example, the input data element XIN 310 can include bits X3, X2, X1, X0. A "0" bit can be added to the input data element XIN 310, e.g., appended to the least significant bit X0, such that the input data element XIN 310 can include bits X3, X2, X1, X0, 0. The "0" bit can be added to fill the subset 302 centered on the least significant bit X0. In this example, each of the subsets 302, 304 for 3-bit Booth encoding can include bits centered on a bit position and including the bit position immediately before and the bit position immediately after that bit position. Each consecutive subset 302, 304 can be centered on a bit position consecutive to the previous subset 302, 304. For example, the subsets 302, 304 can be represented as bits X 2i+1 , X 2i and X 2i-1 , where "i" can be the number of loop iterations. For the first loop, e.g., i = 0, there may be no X 2i-1 bit, since there may be no significant bit lower than the least significant bit X0, and a "0" bit can be appended to the least significant bit X0. Since the centers of consecutive subsets 302, 304 are located at bit positions consecutive to the previous subset 302, 204, the least significant bits of consecutive subsets 302 and 304 may overlap with the most significant bits of the previous subset 302 and 304. In other words, the X 2i-1 bits of consecutive subsets 302, 304 and the X 2i+1 bits of the previous subset 302, 304 can overlap in consecutive iterations (e.g., the bit X 2i-1 for i = 1 and the bit X 2i+1 for i = 0 are both X1 bits). Thus, the Booth encoder 300 can encode 2 previously unencoded bits (e.g., bits X 2i+1 , X 2i ) of the input data element XIN 310 and 1 previously encoded bit (e.g., bit X 2i+1 ) of the input data element XIN 310 in consecutive iterations.
[0052] For example, from subsets 302, 304 of bits “111” and / or “000”, Booth encoder 300 can generate Booth encoded signal 320, which represents a “0” Booth encoded value for multiplication with a corresponding weight data element W, e.g., by indicating a logic gating operation to achieve the multiplication result. The logic gating can prevent the bits of weight data element W from propagating in the CIM circuit, thereby generating a “low” or “0” signal in place of weight data element W, effectively multiplying weight data element W by a “0” value.
[0053] From subsets 302, 304 of bits “001” and / or “010”, Booth encoder 300 can generate Booth encoded signal 320, which represents a “1” Booth encoded value for multiplication with a corresponding weight data element W, e.g., by indicating a direct mapping operation of weight data element W in the CIM circuit to achieve the multiplication result. The direct mapping in the CIM circuit can enable the bits of weight data element W to propagate unchanged in the CIM circuit, thereby generating a signal representing the unchanged weight data, effectively multiplying weight data element W by a “1” value.
[0054] From subset 302, 304 of bits “011”, Booth encoder 300 can generate Booth encoded signals 320, which represent “2” Booth encoded values for multiplication with corresponding weight data elements W, e.g., by indicating a direct mapping operation of weight data element W and a left shift operation of weight data element W (e.g., a 1-bit left shift in an adder) in the CIM circuit to achieve the multiplication result. The left shift of the directly mapped weight data element W in the CIM circuit can shift the bits of weight data element W by an amount that changes the bits of weight data element W, thereby generating a signal representing weight data element W multiplied by a “2” value.
[0055] From subsets 302, 304 of bit “100”, Booth encoder 300 can generate a Booth encoded signal 320 that represents a “-2” Booth encoded value for multiplication with a corresponding weight data element W. For example, the multiplication result is achieved by indicating an inversion operation on the weight data element W in the CIM circuit, an operation of adding a “1” value at the least significant bit of the inverted weight data element, and a left shift operation on the sum (e.g., left shift by 1 bit in an adder). In the CIM circuit, inverting the bits of the weight data element W and adding a “1” value at the least significant bit of the inverted bits can generate a signal representing the negative signed version of the weight data element W, effectively multiplying the weight data element W by a “-1” value. Left shifting the negative signed version of the weight data element W in the CIM circuit can shift the bits of the negative signed version of the weight data element W by an amount that changes the bits of the negative signed version of the weight data element W, resulting in a signal representing the negative signed version of the weight data element W multiplied by a “2” value. These operations may result in a signal representing the weight data element W multiplied by a “-2” value.
[0056] From subsets 302, 304 of bit “101” and / or “110”, Booth encoder 300 can generate a Booth encoded signal 320 that represents a “-1” Booth encoded value for multiplication with a corresponding weight data element W. For example, the multiplication result is achieved by indicating an inversion operation on the weight data element W in the CIM circuit and an operation of adding a “1” value at the least significant bit of the inverted weight data element W. In the CIM circuit, inverting the bits of the weight data element W and adding a “1” value at the least significant bit of the inverted bits can generate a signal representing the negative signed version of the weight data element W, effectively multiplying the weight data element W by a “-1” value.
[0057] Figure 4 Table 400 of a non - limiting example of Booth encoder 300 in accordance with various embodiments of the present disclosure is shown. The table 400 encodes one of subsets 302 and 304 of the input data element XIN 310 (e.g., X 2i+1 、X 2i and X 2i-1 ) to generate a Booth encoded signal 320. As a non - limiting example, the Booth encoded signal 320 includes an enable bit (“ENB”), a Booth encoded bit (“BE”), and a select bit (“S”). Different combinations of the logical states of these bits, ENB, BE, and S can correspond to corresponding Booth encoded values. Additionally, the bits, ENB, BE, and S can be provided to a Booth decoder and used as control bits for the Booth decoder. Upon receiving the control bits, the Booth decoder can multiply the received weight data element W by the Booth encoded value.
[0058] As a representative example, a Booth encoder 300 that receives subsets 302, 304 of bits "000" and / or "111" can generate and output a Booth encoded signal 320 (e.g., ENB, BE, S) as bits "100", which can be configured to cause a corresponding Booth decoder to multiply a weight data element W by a "0" value. The Booth decoder can be configured to interpret a Booth encoded signal 320 as bits "100" / be controlled by a Booth encoded signal 320 as bits "100" to perform a logic gating on the weight data element W. As another representative example, a Booth encoder 300 that receives subsets 302, 304 of bits "001" and / or "110" can generate and output a Booth encoded signal 320 (e.g., ENB, BE, S) as bits "000", which can be configured to cause a corresponding Booth decoder to multiply a weight data element W by a "1" value. The Booth decoder can be configured to interpret a Booth encoded signal 320 as bits "000" / be controlled by a Booth encoded signal 320 as bits "000" to perform a direct mapping on the weight data element W. Other combinations of the logic states of ENB, BE, and S, and the corresponding Booth decoded values (or operations performed by the corresponding Booth decoder) are summarized in Table 400.
[0059] Figure 5 illustrates an example implementation of a Figure 2 computation block 200 (hereinafter referred to as "computation block 500") according to various embodiments of the present disclosure. The computation block 500 can be configured to process (e.g., encode) one of multiple subsets of an input data element XIN and multiply a weight data element W by the encoded input data element XIN. Generally, the input data element XIN and the weight data element W can be provided as signed data elements. It should be understood that Figure 5 the schematic diagram has been simplified, and thus, the computation block 500 can include any of various other components while still being within the scope of the present disclosure.
[0060] As shown, the computation block 500 includes a Booth encoder 510 (e.g., Figure 2 210 of Figure 2 ), and a Booth decoder 520 (e.g., Figure 3For the encoder 300 shown, the number of sign-aware multiplexers can be equal to 4. These 4 sign-aware multiplexers can respectively correspond to the Booth encoding values 1, -1, -2, and 2 provided by the Booth encoder 510. In other words, the Booth encoder 510 can operably (e.g., non-physically) have four sign outputs or operation outputs, respectively corresponding to (or otherwise providing) the Booth encoding values 1, -1, -2, and 2. Additionally, the Booth encoder 510 can be implemented as any one of various other Booth encoders (e.g., radix-2 Booth encoder, radix-8 Booth encoder), which can change the number of corresponding sign-aware multiplexers while still being within the scope of the present disclosure.
[0061] The Booth encoder 510 is configured to encode one of the subsets of the received input data element XIN based on the Booth algorithm and provide a Booth encoding signal within each cycle. The Booth decoder 520 is configured to receive the weight data element W (or one of multiple subsets of the weight data element W) and multiply the weight data element W by the Booth encoding value determined based on the Booth encoding signal (provided by the Booth encoder 510) to provide multiple partial products. In various embodiments, the sign-aware multiplexers 530 to 560 are operably coupled between the Booth encoder 510 and the Booth decoder 520.
[0062] The input data element XIN and the weight data element W processed by the computing block 500 can be of integer data type or floating-point data type, and each data type can have a sign bit. That is, each of the input data element XIN and the weight data element W is provided as a signed data element. Therefore, the sign-aware multiplexers 530 to 560 can receive the Booth encoding signal and operably adjust the Booth-encoded signal based on the logical processing signal of the sign bit of the input data element XIN (sometimes referred to as "XINsign") and the sign bit of the weight data element W (sometimes referred to as "Wsign"). However, in some other embodiments, the computing block 500 can multiply an unsigned input data element by an unsigned weight data element while still being within the scope of the present disclosure. For example, when unsigned data elements are provided, the computing block 500 can deactivate the sign-aware multiplexers 530 to 560; and when signed data elements are provided, the computing block 500 can activate the sign-aware multiplexers 530 to 560.
[0063] Each of symbol-aware multiplexers 530 to 560 has a first input terminal, a second input terminal, and an output terminal. The first input terminal of the symbol-aware multiplexer can receive a first combination of respective logical states of the Booth-encoded signal, and the second input terminal of the symbol-aware multiplexer can receive a second combination of respective logical states of the Booth-encoded signal. Equivalently, the first combination of logical states of the Booth-encoded signal can correspond to a first Booth-encoded value, and the second combination of logical states of the Booth-encoded signal can correspond to a second Booth-encoded value. In various embodiments, the first Booth-encoded value and the second Booth-encoded value equivalently received by the first and second input terminals of each of symbol-aware multiplexers 530 to 560 have opposite polarities but the same magnitude. For example, in Figure 5 , the symbol-aware multiplexer 530 can receive Booth-encoded values 1 and -1 at its first and second input terminals respectively; the symbol-aware multiplexer 540 can receive Booth-encoded values -1 and 1 at its first and second input terminals respectively; the symbol-aware multiplexer 550 can receive Booth-encoded values -2 and 2 at its first and second input terminals respectively; the symbol-aware multiplexer 560 can receive Booth-encoded values 2 and -2 at its first and second input terminals respectively.
[0064] In some embodiments, each of symbol-aware multiplexers 530 to 560 can be controlled by the XOR signal of XINsign and Wsign, sometimes referred to as "XOR(Wsign, XINsign)". When XINsign and Wsign are set to "00" or "11", the XOR signal equals logical "0"; when XINsign and Wsign are set to "01" or "10", the XOR signal equals logical "1". That is, when the signs of the input data element XIN and the weight data element W are the same as each other, the XOR signal equals logical "0"; when the signs of the input data element XIN and the weight data element W are different from each other, the XOR signal equals logical "1".
[0065] Based on the signal XOR(Wsign, XINsign) being equal to logic "0", the sign-aware multiplexers 530 to 560 can each select the signal (or equivalent Booth-encoded value) received at their first input; when the signal XOR(Wsign, XINsign) is equal to logic "1", the sign-aware multiplexers 530 to 560 can each select the signal (or equivalent Booth-encoded value) received at their second input. In other words, when the input data element XIN and the weight data element W have the same sign, the sign-aware multiplexers 530 to 560 can each select the first Booth-encoded value; and when the input data element XIN and the weight data element W have different signs, select the second Booth-encoded value. Equivalently, the sign-aware multiplexers 530 to 560 can determine whether to adjust the Booth-encoded signal based on whether the signs of the input data element XIN and the weight data element W are the same (positive product) or different (negative product).
[0066] As a representative example, when the signal XOR(Wsign, XINsign) is "0" and the Booth-encoded signal provided by the Booth encoder 510 corresponds to the Booth-encoded value "1", the sign-aware multiplexer 530 can select the Booth-encoded value "1" and provide it to the Booth decoder 520. That is, when the signal XOR(Wsign, XINsign) is "0", the sign-aware multiplexer 530 can directly forward the Booth-encoded value provided by the Booth encoder 510 to the Booth decoder 520. As another representative example, when the signal XOR(Wsign, XINsign) is "1" and the Booth-encoded signal provided by the Booth encoder 520 corresponds to the Booth-encoded value "1", the signal-aware multiplexer 530 can select the Booth-encoded value "-1" and provide it to the Booth decoder 520. Equivalently, upon identifying that the signal XOR(Wsign, XINsign) is equal to "1", the sign-aware multiplexers 530 to 560 can "adjust" the Booth-encoded value provided by the Booth encoder 510 by selecting a Booth-encoded value with the opposite polarity and provide the adjusted Booth-encoded value to the Booth decoder 520.
[0067] Figure 6 Illustrates a non-limiting example of a table 600 according to various embodiments of the present disclosure, the table 600 summarizing the computational block 500( Figure 5 ) for a subset of the input data elements XIN (e.g., X 2i+1 、X 2i and X 2i-1)Encode, generate a Booth encoding value (or Booth encoding signal), selectively adjust the generated Booth encoding value based on the signs of the input data element XIN and the weight data element W, and multiply the weight data element W by the selectively adjusted Booth encoding value.
[0068] Figure 7 FIG. shows an example circuit diagram of each symbol-aware multiplexer 530 to 560 (hereinafter referred to as "multiplexer 700") according to various embodiments of the present disclosure. In Figure 7 the example, the multiplexer 700 is implemented as a dual-input single-output multiplexer (sometimes referred to as a 2-to-1 MUX or 2:1 MUX) having AND-OR-INVERT (AOI) logic gates. That is, the multiplexer 700 is configured to select one of two input signals based on a control signal. It should be understood that the multiplexer 700 can be implemented in any of various other configurations (e.g., having OR-AND-INVERT (OAI) logic gates) while still being within the scope of the present disclosure.
[0069] As shown, the multiplexer 700 includes a first AND logic gate 710, a second AND logic gate 720, and an OR logic gate 730. The multiplexer 700 can have: (i) a first input terminal connected to one input terminal of the AND logic gate 710, and the other input terminal of the AND logic gate 710 is configured to directly receive the signal XOR(Wsign, XINsign); and (ii) a second input terminal connected to one input terminal of the AND logic gate 720, and the other input terminal of the AND logic gate 720 is configured to receive the signal XOR(Wsign, XINsign) via an inverter. The output terminals of the AND logic gate 710 and the AND logic gate 720 can be connected to the OR logic gate 730. In an example where the symbol-aware multiplexer 530 is implemented as the multiplexer 700, the first input terminal and the second input terminal of the multiplexer 700 are configured to receive a first Booth encoding value "1" and a second Booth encoding value "-1". Therefore, when the signal XOR(Wsign, XINsign) is equal to "0", the multiplexer 700 (or 530) selects a first combination of the logical states of the Booth encoding signal corresponding to the Booth encoding value "1"; when the signal XOR(Wsign, XINsign) is equal to "1", the multiplexer 700 (or 530) selects a second combination of the logical states in the Booth encoding signal corresponding to the Booth encoding value "-1".
[0070] Figure 8 FIG. shows an example block diagram 800 of a computing circuit 106 (hereinafter referred to as "computing circuit 800") according to various embodiments of the present disclosure. In Figure 8In an illustrative example, the computing circuit 800 may be configured to process (e.g., encode) input data elements XIN having 12 bits (X 12 、X 11 、X 10 、X9, X8, X7, X6, X5, X4, X3, X2, X1) and multiply the weight data element W by the encoded input data element XIN to generate a plurality of partial products.
[0071] As shown, the computing circuit 800 may have six computing blocks 810A, 810B, 810C, 810D, 810E, and 810F. Each of the computing blocks 810A through 810F may be configured to Figure 5 compute block 500 of, e.g., encode a 3-bit subset of the input data element XIN to generate a Booth encoded value and multiply the weight data element W by the corresponding selected Booth encoded value to generate a partial product. However, it should be understood that the computing circuit 800 may process data elements having any number of bits. Accordingly, the number of computing blocks included in the computing circuit 800 may change. For example, to process data elements having 8 bits, the computing circuit 800 may have four computing blocks, each configured to generate a partial product. Generally, the number of computing blocks (N1) of the computing circuit 800 is equal to half the number of bits (N2) of the data elements received by the computing circuit 800.
[0072] For example, computing block 810A may encode a subset of (X2, X1, 0) to generate a first Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the weight data element W by the first Booth encoded value to generate a first partial product; computing block 810B may encode a subset of (X4, X3, and X2) to generate a second Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the weight data element W by the second Booth encoded value to generate a second partial product; computing block 810C may encode a subset of (X6, X5, and X4) to generate a third Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the weight data element W by the third Booth encoded value to generate a third partial product; computing block 810D may encode a subset of (X8, X7, and X6) to generate a fourth Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the weight data element W by the fourth Booth encoded value to generate a fourth partial product; computing block 810E may encode a subset of (X 10 、X9, and X8) to generate a fifth Booth encoded value (e.g., 0, 1, -1, -2, or 2) and multiply the weight data element W by the fifth Booth encoded value to generate a fifth partial product; computing block 810F may encode a subset of (X 12 、X 11 、and X 10) subset is encoded to generate a sixth Booth encoding value (e.g., 0, 1, -1, -2, or 2), and the weight data element W is multiplied by the sixth Booth encoding value to generate a sixth partial product. Then, these 6 partial products can be summed (by an adder tree, such as Figure 1 of 108) to obtain the final product of the input data element XIN and the weight data element W.
[0073] Figure 9 FIG. shows a flowchart of an example method 900 for performing a MAC operation on an input data element XIN and a weight data element W according to various embodiments of the present disclosure. In some embodiments, the input data element XIN and the weight data element W can each be provided as signed data elements. The operations of method 900 can be performed by, for example, Figure 5 the components described in, and thus, some of the reference numbers used above can be reused in the following discussion of method 900. Additionally, it can be understood that method 900 has been simplified, and thus, additional operations can be provided before, during, and after Figure 9 the method 900 of, and only some other operations may be briefly described here.
[0074] Method 900 begins with operation 910 of receiving a first data element and a second data element. The first data element can be the input data element XIN, and the second data element can be the weight data element W. In some embodiments, each of the input data element XIN and the weight data element W can be received as a signed data element, which can be of a floating-point data type or an integer data type. Thus, the input data element XIN has a first sign bit and a plurality of first data bits, and the weight data element W has a second sign bit and a plurality of second data bits. Using Figure 5 the computing block 500 of as a non-limiting example, the Booth encoder 510 can receive the input data element XIN, and the Booth decoder 520 can receive the weight data element W.
[0075] Method 900 continues to operation 920 of encoding the first data bits of the first data element to generate a plurality of encoded values. Continuing with the above example, the Booth encoder 510 implemented as a 3-bit Booth encoder can encode a 3-bit subset of the first data bits within each cycle. In an example where the number of first data bits is equal to 4 (e.g., X3, X2, X1, X0), the Booth encoder 510 can generate a first combination of the logical states of the Booth encoding signals corresponding to a first Booth encoding value (e.g., "1") within the first cycle, and a second combination of the logical states of the Booth encoding signals corresponding to a second Booth encoding value (e.g., "-1") within the second cycle.
[0076] Method 900 proceeds to operation 930 where, based on the logical processing signal of the first sign bit of the first data element and the second sign bit of the second data element, one is selected from a pair of Booth encoding values that are opposite to each other. This pair of Booth encoding values are opposite to each other, with opposite polarities but the same magnitude. Continuing with the above example, after the Booth encoder 510 generates the first Booth encoding value "1" and provides it to the corresponding sign-aware multiplexer (e.g., 530), the multiplexer 530 can determine whether to directly forward the first Booth encoding value "1" to the Booth decoder 520 or select another Booth encoding value opposite to "1", i.e., "-1", based on the XOR signal of the first sign bit and the second sign bit. If the XOR signal is equal to "0", which indicates that the input data element XIN and the weight data element W have the same sign, the multiplexer 530 can directly forward (select) the first Booth encoding value "-1" to the Booth decoder 520; if the XOR signal is equal to "1", which indicates that the input data element XIN and the weight data element W have different signs, the multiplexer 530 can invert the first Booth encoding value to "-1" and provide (select) it to the Booth decoder 520.
[0077] Method 900 proceeds to operation 940 where the second data bit of the second data element is multiplied by the selected encoded value. After receiving the selected Booth encoding value, the Booth decoder 520 can multiply the weight data element W by the selected Booth encoding value to generate a partial product. Using the same example as above, if the XOR signal is equal to "0", then within the first cycle (where the first Booth encoding value is provided as "1"), the Booth decoder 520 then multiplies the weight data element W by 1; if the XOR signal is equal to "1", then within the first cycle (where the first Booth encoding value is provided as "1"), the Booth decoder 520 then multiplies the weight data element W by -1. After generating partial products in each cycle, all the partial products can be added together to generate the final product. In the above example where the input data element XIN has 4 bits, two partial products can be added together to generate the final product of the input data element XIN and the weight data element W.
[0078] Figure 10 illustrates a Figure 1 computing circuit 106 according to various embodiments of the present disclosure or Figure 2Schematic diagram of an example implementation of a plurality of computing blocks 200 (hereinafter referred to as "computing circuit 1000"). The computing circuit 1000 can be configured to process (e.g., encode) an input data element XIN and multiply a weight data element W by the encoded input data element XIN. In various embodiments, the input data element XIN and the weight data element W can be provided as signed or unsigned data elements. Thus, the computing circuit 1000 can have control pins to respectively indicate two signals (e.g., two bits), one of which (XSIGNED) indicates whether the input data element XIN is signed or unsigned, and the other (WSIGNED) indicates whether the weight data element W is signed or unsigned. It should be understood that Figure 10 the schematic diagram has been simplified, and thus, the computing circuit 1000 can include any one of various other components while still being within the scope of the present disclosure.
[0079] As shown, the computing circuit 1000 includes a plurality of Booth encoders 1010A to 1010F (e.g., each Booth encoder can correspond to Figure 2 210 of) and a plurality of Booth decoders 1020A to 1020F (e.g., each Booth decoder can correspond to Figure 2 220 of), and a plurality of logic components 1030, 1040, and 1050. In Figure 10 an illustrative example, the data elements (e.g., XIN and W) received by the computing circuit 1000 both have 12 bits (e.g., XIN[11:0] and W[11:0]). In such an example, the computing circuit 1000 can include 6 Booth encoders 1010A to 1010F and 6 corresponding Booth decoders 1020A to 1020F. It should be understood that the data elements processed by the computing circuit 1000 can have any other number of bits while still being within the scope of the present disclosure. The computing circuit 1000 can be operatively coupled to an adder tree 1060 ( Figure 1 an example implementation of adder tree 108 of), and the adder tree 1060 can include a plurality of full adders 1061, 1062, 1063, 1064, 1065, and 1066.
[0080] Each of the Booth encoders 1010A to 1010F can be implemented as a 3-bit Booth encoder (e.g., Figure 3as shown in Encoder 300), and each of Booth encoders 1010A to 1010F can be operatively coupled to a corresponding one of Booth decoders 1020A to 1020F. In an example where the input data element XIN has 12 bits (e.g., the signal 1001, which can be represented as XIN[11:0]), each Booth encoder can encode one of multiple subsets of the signal 1001 (XIN[11:0]) and provide the Booth-encoded value to the corresponding Booth decoder.
[0081] For example, Booth encoder 1010A can encode a first subset of the signal 1001 (XIN[11:0]) to generate a first Booth-encoded value and provide the first Booth-decoded value to Booth decoder 1020A; Booth encoder 1010B can encode a second subset of the signal 1001 (XIN[11:0]) to generate a second Booth-encoded value and provide the second Booth-decoded value to Booth decoder 1020B; Booth encoder 1010C can encode a third subset of the signal 1001 (XIN[11:0]) to generate a third Booth-encoded value and provide the third Booth-decoded value to Booth decoder 1020C; Booth encoder 1010D can encode a fourth subset of the signal 1001 (XIN[11:0]) to generate a fourth Booth-encoded value and provide the fourth Booth-decoded value to Booth decoder 1020D; Booth encoder 1010E can encode a fifth subset of the signal 1001 (XIN[11:0]) to generate a fifth Booth-encoded value and provide the fifth Booth-decoded value to Booth decoder 1020E; Booth encoder 1010F can encode a sixth subset of the signal 1001 (XIN[11:0]) to generate a sixth Booth-encoded value and provide the sixth Booth-decoded value to Booth decoder 1020F.
[0082] In various embodiments of the present disclosure, the computing circuit 1000 can use logic components 1030, 1040, and 1050 to process the input data element XIN and the weight data element W regardless of whether the input data element XIN and the weight data element W are both provided as unsigned or signed. For example, logic component 1030 can be a 2-input NAND (negative AND) gate, logic component 1040 can be a 2-input NOR (negative OR) gate, and logic component 1050 can be a half adder. Logic component 1030 can perform a NAND operation on signals 1003 and 1005 to provide signal 1017; logic component 1040 can perform a NOR operation on signals 1011 and 1017 to provide signal 1019; and logic component 1050 can add a bit to signal 1013 to provide signal 1015. Each of these logic components and signals will be described in detail below.
[0083] The signal 1003 received at one input of the logic component 1030 may represent the most significant bit of the signal 1001, e.g., XIN
[11] . The signal 1005 received at another input of the logic component 1030 may represent the logically inverted version of the signal indicated at one of the control pins (e.g., XSIGNEDB). In some embodiments, the logic component 1030 may provide NAND(XIN
[11] , XSIGNEDB) as the signal 1017.
[0084] The signal 1011 received at one input of the logic component 1040 may represent the logically inverted version WB[11:0] of the weight data element. In some embodiments, after receiving the signal 1017 at another input of the logic component 1030, the logic component 1040 may provide NOR(NAND(XIN
[11] , XSIGNEDB), WB[11:0]) as the signal 1019, where NAND(XIN
[11] , XSIGNEDB) represents the signal 1017. The signal 1019 may represent a partial product of one of the subsets of the signal 1001 (XIN[11:0]), which subset includes its most significant bit and one or more bits appended to the left of the most significant bit.
[0085] The signal 1013 received by the logic component 1050 may represent the weight data element -W with an opposite polarity. In various embodiments, to generate the signal 1015 (e.g., -W), the logic component 1050 may receive the signal 1013 represented as NAND(WSIGNED, W
[11] ), WB[11:0] and add the signal 1013 to a unit binary integer (not shown). Specifically, the signal 1013 (NAND(WSIGNED, W
[11] ), WB[11:0]) may represent performing sign extension on WB[11:0]. For example, when the weight data element W is provided as signed (i.e., WSIGNED = 1), the signal 1013 becomes NAND(1, W
[11] ), WB[11:0], which in turn becomes WB
[11] , WB[11:0]. As disclosed herein, WB
[11] , WB[11:0] refers to appending the most significant bit of WB[11:0] to its left. In another example, when the weight data element W is provided as unsigned (i.e., WSIGNED = 0), the signal 1013 becomes NAND(0, W
[11] ), W[11:0], which in turn becomes 1, WB[11:0]. As disclosed herein, 1, WB[11:0] refers to appending "1" to the left of WB[11:0]. Thus, the signal 1015 (-W) may be represented as WN[12:0].
[0086] Each of the Booth decoders 1020A to 1020F may receive two signals 1007 and 1009, which represent W and -W with sign extension, respectively. In various embodiments, the signal 1007 may be represented as NOR (WSIGNEDB, WB
[11] ), W [11: 0], and the signal 1009 may be represented as WN
[12] , WN [12: 0]. Each of the Booth decoders 1020A to 1020F may generate a partial product by multiplying the weight data element W by a corresponding Booth encoding value (e.g., provided by a corresponding one of the Booth encoders 1010A to 1010F). Specifically, each of the Booth decoders 1020A to 1020F may selectively adjust the received W and -W based on the corresponding Booth encoding value. Taking the Booth decoder 1020F as a representative example, when the Booth encoding value "2" is received from the Booth encoder 1020F, the Booth decoder 1020F may perform a left shift operation on W. Using Booth decoder 1020A as another representative example, upon receiving a Booth encoded value of "-2" from Booth encoder 1010A, Booth decoder 1020A may perform a left shift operation on -W.
[0087] With this configuration, logic components 1030 and 1040 can jointly determine how to process the partial product of the most significant bit of signal 1001 (e.g., XIN
[11] or signal 1003) based on whether signal 1001 (XIN[11:0]) is provided as signed or unsigned. In general, when signal 1001 (XIN[11:0]) is provided as unsigned, logic component 1040 can output signal 1019 based on a logically inverted version of the most significant bit of signal 1001 (e.g., XINB
[11] ) because all of its bits are equal to "0" or equal to the weight data element (W[11:0]). Equivalently, when the input data element XIN is provided as unsigned, the partial product corresponding to the most significant bit of the input data element (signal 1001 or XIN[11:0]) and the weight data element W is "0" or "W". When signal 1001 (XIN[11:0]) is provided as a signed number, logic component 1040 can output signal 1019 as all "0s", regardless of whether the most significant bit (e.g., XINB
[11] ) of signal 1001 is "1" or "0". Equivalently, when input data element XIN is provided as a signed number, the partial product corresponding to the most significant bit of the input data element (signal 1001 or XIN[11:0]) and the weight data element W is always "0". Advantageously, even with the ability to process signed or unsigned data elements, the computational load of computing circuit 1000 (and the corresponding circuit design) does not increase accordingly.
[0088] Figure 11 , Figure 12 , Figure 13 andFigure 14 illustrates an example of four different combinations of a computing circuit 1000 processing signed or unsigned input data elements XIN and signed or unsigned weight data elements W. In Figures 11 to 14 the example, each of the input data element XIN and the weight data element W has 12 bits. However, it should be understood that the number of bits of each of the input data element XIN and the weight data element W processed by the computing circuit 1000 can vary (e.g., Figure 16 ), while remaining within the scope of the present disclosure.
[0089] In Figure 11 an example is shown where the input data element XIN is provided as unsigned and the weight data element W is provided as unsigned (i.e., XSIGNED = 0 and WSIGNED = 0). Thus, the signal 1005 is XSIGNEDB = 1, which causes the logic component 1030 to output the signal 1017 as XINB
[11] by performing a NAND operation on 1 and XIN
[11] . In response, the logic component 1040 outputs the signal 1019 as all bits equal to "0" or W[11:0] by performing a NOR operation on XINB
[11] and WB[11:0]. For example, when XINB
[11] = 1, the signal 1019 is output as 12-bit "0", which refers to a partial product of a subset including the most significant bit XIN
[11] of the input data element and a weight data element W equal to 0. When XINB
[11] = 0, the signal 1019 is output as W[11:0], which refers to a partial product of a subset including the most significant bit XIN
[11] of the input data element and a weight data element W equal to W. In the current example, it is worth noting that each of the Booth decoders 1020A to 1020F receives the signal 1007 (W) and the signal 1009 (-W). The signals 1007 and 1009 can be represented as NOR(1, WB
[11] ), W[11:0] and WN
[12] , WN[12:0], respectively, where NOR(1, WB
[11] ), W[11:0] means appending a "0" bit to the left of the most significant bit of the weight data element W[11:0].
[0090] In Figure 12In the example shown, the input data element XIN is provided as unsigned and the weight data element W is provided as signed (i.e., XSIGNED = 0 and WSIGNED = 1). Thus, the signal 1005 is XSIGNEDB = 1, which causes the logic component 1030 to output the signal 1017 as XINB
[11] by performing a NAND operation on 1 and XIN
[11] . In response, the logic component 1040 outputs the signal 1019 as all bits equal to "0" or W[11:0] by performing a NOR on XINB
[11] and WB[11:0]. For example, when XINB
[11] = 1, the signal 1019 is output as 12 - bit "0", which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to 0. When XINB
[11] = 0, the signal 1019 is output as W[11:0], which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to W. In the current example, it is worth noting that each of the Booth decoders 1020A to 1020F receives the signal 1007 (W) and the signal 1009 (-W). The signals 1007 and 1009 can be represented as NOR(0, WB
[11] ), W[11:0] and WN
[12] , WN[12:0] respectively, where NOR(1, WB
[11] ), W[11:0] means appending an extra most significant bit to the left of the most significant bit of the weight data element W[11:0].
[0091] In Figure 13In the example shown, the input data element XIN is provided as signed and the weight data element W is provided as unsigned (i.e., XSIGNED = 1 and WSIGNED = 0). Thus, the signal 1005 is XSIGNEDB = 0, which causes the logic component 1030 to output the signal 1017 as logic 1 by performing a NAND operation on 0 and XIN
[11] . In response, the logic component 1040 outputs the signal 1019 as all “0” by performing a NOR operation on “1” and WB[11:0], regardless of whether XINB
[11] equals logic 1 or 0. For example, when XINB
[11] = 1, the signal 1019 is output as 12-bit “0”, which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to 0. When XINB
[11] = 0, the signal 1019 is still output as 12-bit “0”, which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to 0. In the current example, it should be noted that each of the Booth decoders 1020A to 1020F receives the signal 1007 (W) and the signal 1009 (-W). The signals 1007 and 1009 can be represented as NOR(1, WB
[11] ), W[11:0] and WN
[12] , WN[12:0], respectively, where NOR(1, WB
[11] ), W[11:0] means appending a “0” bit to the left of the most significant bit of the weight data element W[11:0].
[0092] In Figure 14In the example, it is shown that the input data element XIN is provided as signed and the weight data element W is provided as signed (i.e., XSIGNED = 1 and WSIGNED = 1). Thus, the signal 1005 is XSIGNEDB = 0, which causes the logic component 1030 to output the signal 1017 as logic 1 by performing a NAND operation on 0 and XIN
[11] . In response, the logic component 1040 outputs the signal 1019 as all "0"s by performing a NOR operation on "1" and WB[11:0], regardless of whether XINB
[11] equals logic 1 or 0. For example, when XINB
[11] = 1, the signal 1019 is output as 12-bit "0"s, which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to 0. When XINB
[11] = 0, the signal 1019 is still output as 12-bit "0"s, which refers to the partial product of a subset including the most significant bit XIN
[11] of the input data element and the weight data element W equal to 0. In the current example, it should be noted that each of the Booth decoders 1020A to 1020F receives the signal 1007 (W) and the signal 1009 (-W). The signals 1007 and 1009 can be represented as NOR(0, WB
[11] ), W[11:0] and WN
[12] , WN[12:0], respectively, where NOR(1, WB[11), W[11:0] means appending an extra most significant bit to the left of the most significant bit of the weight data element W[11:0].
[0093] Figure 15 FIG. 1500 is a flowchart showing an example method 1500 for performing a MAC operation on an input data element XIN and a weight data element W in accordance with various embodiments of the present disclosure. In some embodiments, the input data element XIN and the weight data element W can each be provided as signed or unsigned data elements. The operations of method 1500 can be performed by, for example Figures 10 - 14 the components described in FIG. 1, and thus, some of the reference numerals used above can be reused in the following discussion of method 1500. In addition, it can be understood that method 1500 has been simplified, and thus, additional operations can be provided before, during, and after Figure 15 the method 1500 of FIG. 1, and only some of the other operations may be briefly described herein.
[0094] Method 1500 begins with an operation 1510 of receiving a first data element and a second data element. The first data element can be the input data element XIN, and the second data element can be the weight data element W. Using Figure 10As a non - limiting example, in the computing circuit 1000, where the input data element XIN and the weight data element W both have 12 bits, the Booth encoders 1010A to 1010F can respectively receive subsets of the input data element XIN (or signal 1001, such as XIN[11:0]), and the Booth decoders 1020A to 1020F can receive the weight data element W (or signal 1007, such as W[11:0]) and its inverted version -W (or signal 1009).
[0095] The method 1500 proceeds to operation 1520 to identify whether the first data element is signed or unsigned, and whether the second data element is signed or unsigned. In some embodiments, the input data element XIN and the weight data element W can be received as one of the following combinations: an unsigned input data element and an unsigned weight data element; an unsigned input data element and a signed weight data element; a signed input data element and an unsigned weight data element; and a signed input data element and a signed weight data element. The signed / unsigned input data element can be represented by XSIGNED, and the signed / unsigned weight data element can be represented by WSIGNED. For example, whether the input data element is signed or unsigned can be identified by XSIGNED, and whether the weight data element is signed or unsigned can be identified by WSIGNED.
[0096] When it is identified whether each of the input data element XIN and the weight data element W is signed or unsigned (operation 1520), the method 1500 can proceed to one of the following operations 1532, 1534, 1536, and 1538. Each of operations 1532 to 1538 will be discussed in more detail below.
[0097] Operation 1532 includes selectively generating a partial product of a subset of the most significant bits of the first data element and the second data element (equal to "0" or exactly equal to the second data element) in response to identifying that the first data element is unsigned and the second data element is unsigned. Continuing with the same example, when it is identified that the input data element XIN is unsigned and the weight data element W is unsigned (e.g., XSIGNED = 0 and WSIGNED = 0), the inputs of the logic component 1030 (e.g., a 2 - input NAND gate) are XSIGNEDB and XIN
[11] respectively, and can output a signal 1017 representing XINB
[11] , which causes the logic component 1040 (e.g., a 2 - input NOR gate) to output a signal 1019 whose all bits are equal to "0" or the weight data element W[11:0]. In various embodiments, the signal 1019 can represent a partial product of a subset of the input data element including its most significant bit and the weight data element.
[0098] Additionally, operation 1532 includes providing each of Booth decoders 1020A through 1020F with an input (signal 1007) operably equal to W and an output (signal 1009) operably equal to -W. In some embodiments, computing circuit 1000 may use another NOR to generate signal 1007. In operation 1532 (where WSIGNEDB = 1), signal 1007 may be generated as NOR(1, WB
[11] ), W[11:0], which is equal to 0, W[11:0]. Thus, at least one "0" bit is appended to the left of the weighted data element W[11:0]. Signal 1009 may be generated as WN
[12] , WN[12:0], where WN[12:0] is signal 1015. Computing circuit 1000 may first use another NAND and logic component 1050 (e.g., a half adder) to generate signal 1015. In operation 1532 (where WSIGNED = 0), signal 1015 (WN[12,0]) may be generated as a bit added to NAND(0, W
[11] ), WB[11:0], which is equal to 1, WB[11:0].
[0099] Operation 1534 includes selectively generating a partial product of a subset of the most significant bits of a first data element and a second data element (equal to "0" or exactly equal to the second data element) in response to identifying that the first data element is unsigned and the second data element is signed. Continuing with the same example, when it is identified that the input data element XIN is unsigned and the weighted data element W is signed (e.g., XSIGNED = 0 and WSIGNED = 1), the inputs to logic component 1030 (e.g., a 2-input NAND gate) are XSIGNEDB and XIN
[11] , respectively, and may output signal 1017 representing XINB
[11] , such that logic component 1040 (e.g., a 2-input NOR gate) outputs signal 1019, all of whose bits are equal to "0" or equal to the weighted data element W[11:0]. In various embodiments, signal 1019 may represent a partial product of a subset of the input data element including its most significant bits and the weighted data element.
[0100] In addition, operation 1534 includes providing each of Booth decoders 1020A through 1020F with an input (signal 1007) operably equal to W and an output (signal 1009) operably equal to -W. In some embodiments, computing circuit 1000 may use another NOR to generate signal 1007. In operation 1534 where WSIGNEDB = 0, signal 1007 may be generated as NOR(0, WB
[11] ), W[11:0], which is equal to W
[11] , W[11:0]. Thus, at least one most significant bit is appended to the left of the weight data element W[11:0]. Signal 1009 may be generated as WN
[12] , WN[12:0], where WN[12:0] is signal 1015. Computing circuit 1000 may first use another NAND and logic component 1050 (e.g., a half adder) to generate signal 1015. In operation 1534 where WSIGNED = 1, signal 1015 (WN[12,0]) may be generated as a bit added to NAND(1, W
[11] ) WB[11:0], which is equal to WB
[11] , WB[11:0].
[0101] Operation 1536 includes generating a partial product of a subset of the most significant bits of a first data element and a second data element equal to "0" in response to identifying that the first data element is signed and the second data element is unsigned. Continuing with the same example, when it is identified that the input data element XIN is signed and the weight data element W is unsigned (e.g., XSIGNED = 1 and WSIGNED = 0), the inputs to logic component 1030 (e.g., a 2-input NAND gate) are XSIGNEDB and XIN
[11] , and signal 1017 may be output as "1", such that logic component 1040 (e.g., a 2-input NOR gate) outputs signal 1019 with all its bits equal to "0". In various embodiments, signal 1019 may represent a partial product of a subset of the input data element including its most significant bits and the weight data element.
[0102] Additionally, operation 1536 includes providing each of Booth decoders 1020A through 1020F with an input (signal 1007) operably equal to W and an output (signal 1009) operably equal to -W. In some embodiments, computing circuit 1000 may use another NOR to generate signal 1007. In operation 1532 where WSIGNEDB = 1, signal 1007 may be generated as NOR(1, WB
[11] ), W[11:0], which is equal to 0, W[11:0]. Thus, at least one "0" bit is appended to the left side of the weight data element W[11:0]. Signal 1009 may be generated as WN
[12] , WN[12:0], where WN[12:0] is signal 1015. Computing circuit 1000 may first use another NAND and logic component 1050 (e.g., half adder) to generate signal 1015. In operation 1532 where WSIGNED = 0, signal 1015 (WN[12,0]) may be generated as a bit added to NAND(0, W
[11] ), WB[11:0], which is equal to 1, WB[11:0].
[0103] Operation 1538 includes generating a partial product of a subset of the most significant bits of a first data element and a second data element equal to "0" in response to identifying that the first data element is signed and the second data element is signed. Continuing with the same example, in identifying that input data element XIN is signed and weight data element W is signed (e.g., XSIGNED = 1 and WSIGNED = 1), the inputs to logic component 1030 (e.g., 2-input NAND gate) are XSIGNEDB and XIN
[11] respectively, and signal 1017 may be output as "1", such that logic component 1040 (e.g., 2-input NOR gate) outputs signal 1019 with all its bits equal to "0". In various embodiments, signal 1019 may represent the partial product of a subset of the input data element including its most significant bit and the weight data element.
[0104] In addition, operation 1538 includes providing each of Booth decoders 1020A through 1020F with an input (signal 1007) operably equal to W and an output (signal 1009) operably equal to -W. In some embodiments, computing circuit 1000 may use another NOR to generate signal 1007. In operation 1538 (where WSIGNEDB = 0), signal 1007 may be generated as NOR(0, WB
[11] ), W[11:0], which is equal to W
[11] , W[11:0]. Thus, at least one most significant bit is appended to the left side of weight data element W[11:0]. Signal 1009 may be generated as WN
[12] , WN[12:0], where WN[12:0] is signal 1015. Computing circuit 1000 may first use another NAND and logic component 1050 (e.g., a half adder) to generate signal 1015. In operation 1538 (where WSIGNED = 1), signal 1015 (WN[12,0]) may be generated as a bit added to NAND(1, W
[11] ), WB[11:0], which is equal to WB
[11] , WB[11:0].
[0105] Concurrent with or subsequent to any of operations 1532 through 1538, method 1500 may also include one or more operations (not shown for simplicity) Figure 15 to sum all of the partial products generated by the Booth decoder (e.g., Booth decoders 1020A through 1020F). Next, adder tree 1060 of computing circuit 1000 may add these partial products to generate the final product of input data element XIN and weight data element W.
[0106] Figure 16 An example of a computing circuit 1600 that processes signed or unsigned input data element XIN and signed or unsigned weight data element W is shown. Computing circuit 1600 is substantially similar to Figure 10 computing circuit 1000. In Figure 16 the example, each of input data element XIN and weight data element W is provided with k bits. Thus, the number of Booth encoders and Booth decoders of computing circuit 1600 may vary accordingly. For example, computing circuit 1600 may include k / 2 Booth encoders 1610 and k / 2 Booth decoders 1620. In addition, computing circuit 1600 may include Figure 10Other components that are substantially similar to the components shown. For example, the computing circuit 1600 also includes a 2-input NAND gate 1630, a 2-input NOR gate 1640, a half adder 1650, and a plurality of full adders 1661, 1662, 1663, 1664, 1665, and 1666. By providing k bits for the data element, the corresponding bits of the signals received or otherwise processed by the computing circuit 1600 can change accordingly. These signals (1601, 1603, 1605, 1607, 1609, 1611, 1613, 1615, 1619) are all in the Figure 16 form shown. Signals 1601 to 1619 are substantially similar to signals 1001 to 1019 ( Figure 10 ), and therefore, the corresponding discussions will not be repeated.
[0107] Figure 17 FIG. shows an example circuit diagram 1700 of a Booth encoder (e.g., Figure 2 210 of, Figure 3 300 of, Figure 5 510 of, Figures 10 - 14 1010A - 1010F of) according to various embodiments of the present disclosure. Hereinafter, Figure 17 the circuit diagram is referred to as Booth encoder 1700. It should be understood that Figure 17 the circuit diagram is a non-limiting implementation of the Booth encoder and is not intended to limit the scope of the present disclosure.
[0108] In some embodiments, the Booth encoder 1700 can perform 3-bit Booth encoding on a 3-bit subset of the data element (e.g., X 2i+1 , X 2i , and X 2i-1 ). As shown, a first input bit line carrying a first signal representing the first bit of the subset (e.g., X 2i-1 ) and a second input bit line carrying a second signal representing the second bit of the subset (e.g., X 2i ) can be coupled to the input terminals of an exclusive OR (“XOR”) gate 1702. The XOR gate 1702 can receive the first signal and the second signal as inputs and produce an output as a first intermediate signal (“1x”). The second bit line and a third bit line carrying a third signal representing the third bit of the subset (e.g., X 2i+1 ) can be coupled to the input terminals of an exclusive NOR (“XNOR”) gate 1708. The XNOR gate 1708 can receive the second signal and the third signal as inputs and produce an output as a second intermediate signal (“2x”).
[0109] The first NOR gate 1704 can be coupled to the output of the XOR gate 1702 and the output of the XNOR gate 1708 to receive inputs for the first NOR gate 1704. Thus, the first NOR gate 1704 can receive a first intermediate signal 1x from the XOR gate 1702 and a second intermediate signal 2x from the XNOR gate 1708 as inputs. The first NOR gate 1704 can generate an output that is a Booth encoded bit (“BE”).
[0110] The second NOR gate 1706 can be coupled to the output of the XOR gate 1702 to receive the first intermediate signal 1x as an input, and can also be coupled to the output of the first NOR gate 1704 to receive the Booth encoded bit BE as an input to the second NOR gate 1706. Thus, the second NOR gate 1706 can receive the first intermediate signal 1x from the XOR gate 1702 and the Booth encoded bit BE from the first NOR gate 1704 as inputs. The second NOR gate 1706 can generate an output that is an enable bit (“ENB”).
[0111] The third NOR gate 1710 can be coupled at its input to the output of the second NOR gate 1706 to receive ENB as an input. The third NOR gate 1710 can also be coupled at an inverted input to the third bit line to receive the inversion of the third bit line as an input. For example, an inverter can be coupled between the third bit line and the input of the third NOR gate 1710. Thus, the third NOR gate 1710 can receive the enable bit ENB from the second NOR gate 1706 and a third signal representing the inversion of the third bit of the subset as an input from the third bit line. In some embodiments, the third NOR gate 1710 can invert the third signal. In some embodiments, the third NOR gate 1710 can receive the inverted third signal from an inverter. The third NOR gate 1710 can generate an output that is a select bit (“S”).
[0112] Figure 18 An example circuit diagram of a Booth decoder (e.g., Figure 2 220 of Figure 5 520 of Figures 10 - 14 1020A - 1020F of Figure 18 is shown according to various embodiments of the present disclosure. Hereinafter, Figure 18 the circuit diagram is referred to as the Booth decoder 1800. It should be understood that
[0113] In some embodiments, the Booth decoder 1800 may be operatively coupled to a corresponding 3-bit Booth encoder (e.g., Booth encoder 1700) to receive a Booth encoded signal, such as Booth encoded bits (BE), enable bits (ENB), and select bits (S). As shown, the Booth decoder 1800 includes a multiplexer 1810 and an adder 1850.
[0114] The multiplexer 1810 may be coupled at an input to any number of input lines configured to carry weight data elements. For example, the multiplexer 1810 may be coupled to four input lines configured to carry 4-bit weight data elements (e.g., W[3], W[2], W[1], W[0]). The multiplexer 1810 may include a plurality of inverters 1812 and 1814, which may be configured to act as buffers for temporarily storing the weight data elements. For example, one of the inverters 1812 may be configured to temporarily store the weight data element, and a corresponding one of the inverters 1814 may be configured to temporarily store the inversion of the weight data element.
[0115] The multiplexer 1810 may be coupled at a select line to a select signal (e.g., select bit “S”) output by the corresponding Booth encoder. The multiplexer 1810 may include a plurality of transmission gates 1816 coupled between the inverters 1812, 1814 and the output of the multiplexer 1810. The transmission gates 1816 may also be coupled at an input to the select signal. The select signal may determine which of the input signal or the inversion of the input signal for each input weight data element (e.g., W[3], W[2], W[1], W[0]) is output from the multiplexer 1810. In some embodiments, pairs of transmission gates 1816 coupled to the same output of the multiplexer 1810 may be differently configured in response to the select signal. For example, for the same select signal, one transmission gate 1810 may effect the transmission of the weight data and / or the inversion of the weight data element stored at the inverter 1812, while the other transmission gate 1816 may prevent the transmission of the weight data element stored at the inverter 1814 and / or its inversion. Vice versa. The multiplexer 1810 may output the weight data element and / or the inversion of the weight data element at an output controlled by the select signal.
[0116] The adder 1850 may receive, at an input, weight data and / or an inversion of a weight data element output by the multiplexer 1810 (collectively referred to herein as the weight data element of the adder 1850). The adder 1850 may be coupled to an enable signal (e.g., enable bit “ENB”) that may be output from a corresponding Booth encoder. The enable signal may trigger the adder 1850 to add the signal received at the input to a value stored in an adder component 1870 (e.g., a shift register). The adder 1850 may include a plurality of NOR gates 1852A, 1852B, and 1852C configured to receive the weight data element at one input of the NOR gates 1852A - 1852C and the enable signal at a second input. The NOR gates 1852A - 1852C may be configured to perform a NOR of the weight data element and the enable signal such that the enable signal may control the logic gating operation of the adder 1850. For example, for an enable signal configured to enable logic gating (e.g., the enable signal has a value of “1”), the NOR gates 1852A - 1852C may output only a value of “0”, regardless of the value of the weight data. Otherwise, the NOR gates 1852A - 1852C may output the weight data at the input and output an enable signal configured to disable logic gating (e.g., the enable signal has a value of “0”).
[0117] The control of adder 1850 can be coupled to the Booth-coded bits output by the corresponding Booth encoder (e.g., the Booth-coded bit “BE”). The Booth-coded bits can be configured to control whether adder 1850 performs a left shift operation (e.g., shift left by 1 bit). The output of each of NOR gates 1852A - 1852C can be coupled to shifter 1856. Shifter 1856 can include a plurality of transmission gates 1858, which are configured to couple the output of each NOR gate to a plurality of inverters 1860. Additionally, shifter 1856 can be configured to directly couple inverter 1862 to the output of NOR gate 1852A, and can include one of the transmission gates 1858 that is configured to couple the output of NOR gate 1852A to one of the inverters 1860. NOR gate 1852A can be associated with the input of the most significant bit of the weighted data element. The inverter 1860 coupled to NOR gate 1852A can correspond to the most significant bit position of the weighted data element, and the inverter 1862 coupled to NOR gate 1852A can correspond to a more significant bit position than the most significant bit position in the weighted data element. Shifter 1856 can include another transmission gate 1858 that is configured to couple the output of NOR gate 1852C to one of the inverters 1860, and another transmission gate 1858 that is configured to couple the output of NOR gate 1852C to inverter 1864. NOR gate 1852C can be associated with the input of the least significant bit of the weighted data element. The inverter 1864 coupled to NOR gate 1852C can correspond to the least significant bit position of the weighted data element. Adder 1850 can also be coupled to a power supply voltage (VDD). Shifter 1856 can include a transmission gate 1866 that is configured to couple the power supply voltage VDD to inverter 1864.
[0118] Transmission gates 1858 and 1866 can also be coupled to the Booth code (BE) bits. Transmission gate 1858 can be configured to enable and / or block the transmission of the output from NOR gates 1852A - 1852C to inverters 1860 and 1864. Transmission gate 1866 can be configured to enable and / or prevent the power supply voltage from being transmitted to inverter 1864. In some embodiments, the paired transmission gates 1858, 1866 coupled to the same inverters 1860, 1864 can be configured differently in response to the Booth code bits.
[0119] In one aspect of the present disclosure, a memory circuit is disclosed. The memory circuit includes a Booth encoder configured to receive a first data element including a first symbol portion and a first data portion. The memory circuit includes a Booth decoder configured to receive a second data element including a second symbol portion and a second data portion and provide a product based on the first data element and the second data element. The memory circuit includes a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder. The plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and change corresponding logical states of the plurality of encoded signals based on the first symbol portion and the second symbol portion, such that the Booth decoder provides the product.
[0120] In some embodiments, each multiplexer of the plurality of multiplexers is controlled by an XOR signal of the first symbol portion and the second symbol portion.
[0121] In some embodiments, each multiplexer of the plurality of multiplexers has a first input terminal and a second input terminal, the first input terminal and the second input terminal being configured to receive a first combination of logical states of the encoded signals and a second combination of logical states of the encoded signals, respectively.
[0122] In some embodiments, the first combination of the encoded signals corresponds to a first encoded value multiplied by the first data portion, and the second combination corresponds to a second encoded value multiplied by the second data portion.
[0123] In some embodiments, the first encoded value and the second encoded value are opposite to each other.
[0124] In some embodiments, each multiplexer of the plurality of multiplexers is configured to select the first combination in response to receiving an XOR signal of the first symbol portion and the second symbol portion equal to a first logical state.
[0125] In some embodiments, each multiplexer of the plurality of multiplexers is configured to select the second combination in response to receiving an XOR signal of the first symbol portion and the second symbol portion equal to a second logical state.
[0126] In some embodiments, the number of multiplexers corresponds to the number of the first data portions.
[0127] In some embodiments, the first data element represents a plurality of input activations received by a memory array, and the second data element represents a plurality of weights stored in the memory array.
[0128] In some embodiments, the first data portion represents a plurality of first trailing bits of a first signal, and the second data portion represents a plurality of trailing bits of a second signal.
[0129] In another aspect of the present disclosure, a memory circuit is disclosed. The memory circuit includes a memory array. The memory circuit includes a computing circuit coupled to the memory array. The computing circuit includes: a Booth encoder configured to receive a first data element including a first sign bit and a plurality of first data bits, and configured to provide a plurality of encoded values based on the plurality of first data bits; a Booth decoder configured to retrieve a second data element including a second sign bit and a plurality of second data bits from the memory array, and provide a plurality of partial products based on multiplying the first data element by the second data element; and a plurality of multiplexers operatively coupled between the Booth encoder and the Booth decoder. Each of the plurality of multiplexers is configured to select a first encoded value among the encoded values or a second encoded value among the encoded values based on a logical processing signal of the first sign bit and the second sign bit.
[0130] In some embodiments, the Booth decoder is further configured to multiply the second data element by the selected first encoded value or second encoded value for a corresponding one of the plurality of partial products.
[0131] In some embodiments, a first multiplexer among the multiplexers is configured to: (i) select a first encoded value among the encoded values corresponding to a first combination of logical states of a subset of the first data bits when identifying that an XOR signal of the first sign bit and the second sign bit is equal to logic 0; and (ii) select a second encoded value among the encoded values corresponding to a second combination of logical states of a subset of the first data bits when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 1.
[0132] In some embodiments, a second multiplexer among the multiplexers is configured to: (i) select the second encoded value when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 0; and (ii) select the first encoded value when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 1.
[0133] In some embodiments, a third multiplexer among the multiplexers is configured to: (i) select a third encoded value among the encoded values corresponding to a third combination of logical states of a subset of the first data bits when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 0; and (ii) select a fourth encoded value among the encoded values corresponding to a fourth combination of logical states of a subset of the first data bits when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 1.
[0134] In some embodiments, the third multiplexer in the multiplexer is configured to: (i) select a fourth encoded value when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 0; and (ii) select a third encoded value when identifying that the XOR signal of the first sign bit and the second sign bit is equal to logic 1.
[0135] In some embodiments, the number of multiplexers corresponds to the number of first data bits.
[0136] In some embodiments, the first data bits represent multiple first tail bits of a first data element, and the second data bits represent multiple second tail bits of a second data element.
[0137] In another aspect of the present disclosure, a method of operating a memory circuit is disclosed. The method includes receiving a first data element and a second data element, where the first data element includes a first sign bit and multiple first data bits, and the second data element includes a second sign bit and multiple second data bits. The method includes encoding the multiple first data bits to generate multiple encoded values, where each of the encoded values corresponds to a respective combination of the logical states of a subset of the first data bits. The method includes selecting between a first encoded value and a second encoded value that are opposite to each other among the multiple encoded values based on a logical processing signal of the first sign bit and the second sign bit. The method includes multiplying the second data bits by the selected first encoded value or second encoded value.
[0138] In some embodiments, the first data bits represent multiple first tail bits of a first data element, and the second data bits represent multiple second tail bits of a second data element.
[0139] As used herein, the terms “about” and “approximate” generally denote a value of a given quantity that may vary depending on the particular technology node associated with the subject semiconductor device. Based on a particular technology node, the term “about” may denote a value of a given quantity that varies, for example, within a range of 10 - 30% of the value (e.g., ±10%, ±20%, or ±30% of the value).
[0140] The features of several embodiments are outlined above so that those skilled in the art can better understand the various aspects of the present disclosure. Those skilled in the art should understand that they can readily use the present disclosure as a basis for designing or modifying other processes and structures for achieving the same purposes and / or achieving the same advantages as those introduced herein. Those skilled in the art should also recognize that such equivalent structures do not depart from the spirit and scope of the present disclosure, and that they can be made various changes, substitutions, and alterations within the present disclosure without departing from the spirit and scope of the present disclosure.
Claims
1. A memory circuit, comprising: A Booth encoder configured to receive a first data element including a first symbol part and a first data part; A Booth decoder configured to receive a second data element including a second symbol part and a second data part, and provide a product based on the first data element and the second data element; And A plurality of multiplexers operably coupled between the Booth encoder and the Booth decoder; Wherein the plurality of multiplexers are configured to receive a plurality of encoded signals from the Booth encoder and change corresponding logical states of the plurality of encoded signals based on the first symbol part and the second symbol part, so that the Booth decoder provides the product.
2. The memory circuit according to claim 1, wherein, Each multiplexer among the multiplexers is controlled by an exclusive - or signal of the first symbol part and the second symbol part.
3. The memory circuit according to claim 1, wherein, Each multiplexer among the multiplexers has a first input terminal and a second input terminal, and the first input terminal and the second input terminal are configured to receive a first combination of the logical states of the encoded signals and a second combination of the logical states of the encoded signals respectively.
4. The memory circuit according to claim 3, wherein, The first combination of the encoded signals corresponds to a first encoded value multiplied by the first data part, and the second combination corresponds to a second encoded value multiplied by the second data part.
5. The memory circuit according to claim 3, wherein, Each multiplexer among the multiplexers is configured to select the first combination in response to receiving an exclusive - or signal of the first symbol part and the second symbol part equal to a first logical state.
6. The memory circuit according to claim 1, wherein, The first data element represents a plurality of input activations received by a memory array, and the second data element represents a plurality of weights stored in the memory array.
7. The memory circuit according to claim 1, wherein, The first data part represents a plurality of first trailing bits of the first signal, and the second data part represents a plurality of trailing bits of the second signal.
8. A memory circuit, comprising: A memory array; And A computing circuit coupled to the memory array, wherein the computing circuit includes: A Booth encoder configured to receive a first data element including a first symbol bit and a plurality of first data bits, and configured to provide a plurality of encoded values based on the plurality of first data bits; A Booth decoder configured to retrieve a second data element including a second symbol bit and a plurality of second data bits from the memory array, and provide a plurality of partial products based on multiplying the first data element by the second data element; and A plurality of multiplexers operably coupled between the Booth encoder and the Booth decoder, wherein each of the plurality of multiplexers is configured to select a first encoded value among the encoded values or a second encoded value among the encoded values based on a logical processing signal of the first symbol bit and the second symbol bit.
9. The memory circuit according to claim 8, wherein, The first multiplexer among the multiplexers is configured to: (i) When identifying that an exclusive - or signal of the first symbol bit and the second symbol bit is equal to logical 0, select the first encoded value among the encoded values corresponding to a first combination of the logical states of a subset of the first data bits; And (ii) When the exclusive - OR signal identifying the first sign bit and the second sign bit is equal to logic 1, select the second coding value among the coding values corresponding to the second combination of the logical states of the subset of the first data bits.
10. A method of operating a memory circuit, comprising: Receiving a first data element and a second data element, wherein the first data element includes a first sign bit and a plurality of first data bits, and the second data element includes a second sign bit and a plurality of second data bits; Encoding the plurality of first data bits to generate a plurality of coding values, wherein each of the coding values corresponds to a respective combination of the logical states of a subset of the first data bits; Selecting between a first coding value and a second coding value that are opposite to each other among the plurality of coding values based on a logical processing signal of the first sign bit and the second sign bit; and Multiplying the second data bits by the selected first coding value or the second coding value.