Memory internal operation circuit and operation method thereof
By dividing floating point numbers into multiple phases in the memory operation circuit, and phasing and grouping these phases based on the exponent and range, the backmultiple alignment is achieved, and the problems of high resource, high power consumption and low accuracy in the prior art floating point number multiplication and accumulation operation are solved, and the efficiency and accuracy of the operation are improved.
Patent Information
- Application Number
- CN202411671410.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
When performing floating-point number multiplication and accumulation operations in machine learning, the prior art has problems such as high computing resources, high power consumption and low accuracy, especially in large-scale deep neural network computing.
A memory in-memory operation circuit is proposed. By dividing floating-point numbers into multiple phases, and phasing and grouping these phases based on the exponents and ranges, the multiplication alignment is achieved, reducing the consumption of computing resources and power consumption, and improving the accuracy of the operations.
Through phase and grouping technology, the computing resource consumption and power consumption in floating-point number operations are reduced, and the accuracy of operations is improved, solving the problems of high computing resources, large power consumption and low accuracy in the prior art.
Smart Images

Figure CN120067038A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an in-memory computing circuit and an operation method thereof. Background Art
[0002] For example, computer artificial intelligence (AI) has been established on the basis of machine learning using deep learning techniques. Using machine learning, a computing system organized as a neural network calculates the statistical likelihood that input data matches previously calculated data. A neural network refers to multiple interconnected processing nodes that enable the analysis of data to compare input data with "training" data. Training data refers to the computational analysis of the characteristics of known data to develop a model for comparing input data. Application examples of AI and data training are found in object recognition, where the system analyzes the characteristics of many (e.g., thousands or more) images to determine patterns that can be used for statistical analysis to identify input objects. Summary of the Invention
[0003] An embodiment of the present disclosure provides an in-memory computing circuit, which is characterized in that it includes an input circuit, N adder circuits, a first selector circuit, a phase-determining circuit, and N subtractor circuits. The input circuit is used to receive: (i) N first inputs and (ii) N second inputs. These first inputs are composed of at least N first exponents and N first mantissas. These second inputs are composed of at least N second exponents and N second mantissas. Each of these second inputs and the corresponding one of the N first inputs forms one of the N input pairs. Each of the N adder circuits is used to combine the corresponding first exponent and the corresponding second exponent of the corresponding one of the N input pairs to generate the corresponding one of the N exponent sums. The first selector circuit is used to select the largest one among the N exponent sums as the largest exponent sum. The phase-determining circuit is used to divide at least a part of the N exponent sums starting from the largest exponent sum into N phases. Each of the N phases is associated with a respective exponent subset among the N exponent subsets. Each of the N subtractor circuits is used to calculate the corresponding one of the N exponent differences in each of the N phases, and each of the N exponent differences is equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the largest exponent sum. The N exponent differences are used for shift and accumulation operations.
[0004] An embodiment of the present disclosure provides a memory-integrated operation circuit, which is characterized in that it includes an input circuit, N adder circuits, a first selector circuit, a second selector circuit, a phase determination circuit, and N subtractor circuits. The input circuit is configured to receive: (i) N first inputs and (ii) N second inputs. These first inputs are composed of at least N first exponents and N first mantissas. These second inputs are composed of at least N second exponents and N second mantissas. Each of these second inputs and the corresponding one of the N first inputs forms one of the N input pairs. Each of the N adder circuits is configured to combine the corresponding first exponent and the corresponding second exponent of the corresponding one of the N input pairs to generate the corresponding one of the N exponent sums. The first selector circuit is configured to select the largest one among the N exponent sums as the largest exponent sum. The second selector circuit is configured to select the smallest one among the N exponent sums as the smallest exponent sum. The phase determination circuit is configured to: (i) determine an exponent sum range based on the difference between the largest exponent sum and the smallest exponent sum, and (ii) divide the N exponent sums into N phases based on the exponent sum range. Each of the N phases is associated with a respective exponent subset among the N exponent subsets. Each of the N subtractor circuits is configured to calculate the corresponding one of the N exponent differences of each of the N phases. Each of the N exponent differences is equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the largest exponent sum. The N exponent differences are used for shift and accumulation operations.
[0005] An embodiment of the present disclosure provides a method of operating an in-memory computing circuit, which is characterized by including the following operations. Receive, by the in-memory computing circuit, (i) N first inputs and (ii) N second inputs, where the first inputs are composed of at least N first exponents and N first mantissas, and the second inputs are composed of at least N second exponents and N second mantissas, and where each of the second inputs and the corresponding one of the N first inputs form one of N input pairs. Generate, by the in-memory computing circuit, the corresponding one of N exponent sums based on combining the corresponding first exponent and the corresponding second exponent of the corresponding ones of the N input pairs. Generate, by the in-memory computing circuit, the corresponding one of N mantissa products based on selectively multiplying the corresponding first mantissa and the corresponding second mantissa of the corresponding ones of the N input pairs. Select, by the in-memory computing circuit, the largest one among the N exponent sums as the maximum exponent sum. Divide, by the in-memory computing circuit, at least a part of the N exponent sums and at least one corresponding part of the N mantissa products into N phases, with each of the N phases associated with a respective exponent subset among N exponent subsets and a respective mantissa subset among N mantissa subsets. Determine, by the in-memory computing circuit, the corresponding one of N exponent differences for each of the N phases, where each of the N exponent differences is equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the maximum exponent sum. Shift, by the in-memory computing circuit, the respective mantissa subsets based on the corresponding N exponent differences for each of the N phases. Determine, by the in-memory computing circuit, the corresponding one of N sums based on adding the shifted respective mantissa subsets. Combine, by the in-memory computing circuit, the N sums to generate an addition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various aspects of the present disclosure can be best understood when the following detailed description is read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, the various features are not drawn to scale. In fact, for clarity of discussion, the dimensions of the various features may be arbitrarily increased or decreased.
[0007] Figure 1 is a block diagram of a data computing circuit for performing multiplication and accumulation (MAC) operations on floating-point numbers using post-multiplication alignment pairs according to some embodiments;
[0008] Figure 2 illustrates an example process of performing a MAC operation by a circuit according to some embodiments Figure 1 of;
[0009] Figure 3 illustrates an example process of a method for selecting to divide exponent sums into multiple phases for Figure 2 a MAC operation according to some embodiments;
[0010] Figure 4 Illustrates an example process of the step size of MAC operations according to some embodiments Figure 2 ; a flowchart of an example process of MAC operations
[0011] Figure 5 Is a graph of an example distribution and grouping of the sum of exponents of a first phase - determination method (e.g., Method I) performed by a circuit of Figure 1 ; an example distribution and grouping of the sum of exponents of a first phase - determination method (e.g., Method I) performed by a circuit of
[0012] Figure 6 Illustrates an example process of performing MAC operations using a first phase - determination method performed by a circuit of Figure 1 ; a flowchart of an example process of performing MAC operations using a first phase - determination method performed by a circuit of
[0013] Figure 7 Illustrates an example process of performing MAC operations using a first phase - determination method with a local maximum performed by a circuit of Figure 1 ; a flowchart of an example process of performing MAC operations using a first phase - determination method with a local maximum performed by a circuit of
[0014] Figure 8 Is a graph of an example distribution and grouping of the sum of exponents of a second phase - determination method (e.g., Method II) performed by a circuit of Figure 1 ; an example distribution and grouping of the sum of exponents of a second phase - determination method (e.g., Method II) performed by a circuit of
[0015] Figure 9 Illustrates an example process of performing MAC operations using a second phase - determination method performed by a circuit of Figure 1 ; a flowchart of an example process of performing MAC operations using a second phase - determination method performed by a circuit of
[0016] Figure 10 Illustrates an example process of performing MAC operations using a second phase - determination method with a local maximum performed by a circuit of Figure 1 ; a flowchart of an example process of performing MAC operations using a second phase - determination method with a local maximum performed by a circuit of
[0017] Figure 11 Illustrates an example process of performing MAC operations using multiple devices for multiple phases performed by a circuit of Figure 1 ; a flowchart of an example process of performing MAC operations using multiple devices for multiple phases performed by a circuit of
[0018]
Symbol Explanation
[0019] 100: Circuit
[0020] 102: Memory Circuit
[0021] 103: Storage Component
[0022] 104: Input Circuit
[0023] 106: Multiplier Circuit
[0024] 108: Adder circuit
[0025] 110: Difference circuit / Subtractor circuit
[0026] 111A: Selector circuit
[0027] 112: Shift circuit
[0028] 114, 116: Adder circuit / Adder tree
[0029] 614, 1106A, 1106B: Adder tree
[0030] 115, PS, PSSM, PSTC: Sum
[0031] 115S: Sum PSTC
[0032] 118: Converter
[0033] 120: Converter
[0034] 122: Phase - setting circuit
[0035] 200, 400, 600, 700, 900, 1000, 1100: Process
[0036] 202, 204, 206, 208, 210, 212, 214, 302, 304, 306, 308, 310, 402, 404, 406, 408, 410: Operation
[0037] 500, 800: Figure
[0038] 502A~502D, 802A~802C: Phase
[0039] 602: Input activation latch
[0040] 604: Weight buffer
[0041] 606: Multiplier
[0042] 608: Adder
[0043] 610, 702, 902, 1102A, 1102B: Phaser
[0044] 612, 1104A, 1104B: Aligner
[0045] 616: Buffer and accumulator Detailed implementation mode
[0046] The following disclosure provides many different embodiments or examples for implementing different features of the provided subject matter. Specific examples of components and arrangements are described below to simplify the disclosure. Of course, these are merely examples and are not intended to be limiting. For example, in the following description, forming a first feature over or on a second feature may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which additional features may be formed between the first and second features such that the first and second features are not in direct contact. Additionally, the disclosure may repeat reference numerals and / or letters in the various examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate a relationship between the various embodiments and / or configurations discussed.
[0047] In addition, for ease of description, spatially relative terms such as "under", "below", "lower", "above", "upper", "top", "bottom", and the like may be used herein to describe the relationship of one component or feature to another as illustrated in the figures. In addition to the orientation depicted in the figures, the spatially relative terms are also intended to encompass different orientations of the device in use or operation. The device may be oriented otherwise (rotated 90 degrees or at other orientations), and the spatially relative descriptors used herein may be interpreted accordingly.
[0048] A neural network calculates "weights" to perform calculations on new data (input data "words"). A neural network uses multiple layers of computational nodes, where deeper layers perform calculations based on the results of calculations performed by higher layers. Machine learning currently relies on the calculation of dot products and absolute differences of vectors, which are typically performed as multiply–accumulate (MAC) operations on parameters, input data, and weights. Calculations for large deep neural networks typically involve too many data elements, and thus it is not practical to store them in the processor cache memory. Therefore, these data elements are typically stored in memory.
[0049] Therefore, machine learning is extremely computationally intensive, which requires the calculation and comparison of many different data elements. Calculations within the processor are several orders of magnitude faster than the transfer of data elements between the processor and main memory resources. Due to the memory size required to store the data elements, for most practical systems, the cost of placing all data elements closer to the processor in the cache memory is extremely high. Therefore, transferring data elements becomes the main bottleneck in AI computing. As the dataset increases, the time and power / energy that the computing system spends moving data elements can ultimately be several times the time and power available for actual calculations.
[0050] In this regard, computing-in-memory (CIM) circuits have been proposed to perform such MAC operations. CIM circuits perform data processing in situ within a suitable memory circuit. CIM circuits suppress the latency of data / program fetching and output result uploading in the corresponding memory (such as a memory array), thus solving the memory (or von Neumann) bottleneck of conventional computers. Another major advantage of CIM circuits is high computational parallelism, which is due to the specific architecture of the memory array, where computations can be performed simultaneously along several current paths. CIM circuits also benefit from multi-memory arrays with a high density of computing devices, which are typically characterized by excellent scalability and 3D integration capabilities. As a non-limiting example, CIM circuits for various machine learning applications can perform MAC operations locally in memory (i.e., without sending data elements to the host processor) to achieve a higher throughput dot product of neuron activations and weight matrices, while still providing higher performance and lower energy compared to computations performed by the host processor.
[0051] The data elements processed by CIM circuits have various types or forms, such as integers and floating-point numbers. A floating-point number is typically represented by a sign part, an exponent part, and a significand (mantissa) part, and the significand part consists of the significant bits of the number. For example, the floating-point format specified by the Institute of Electrical and Electronics Engineers ( ) has a size of thirty-two bits and includes twenty-three mantissa bits, eight exponent bits, and one sign bit. Another floating-point format has a size of sixteen bits, which includes ten mantissa bits, five exponent bits, and one sign bit.
[0052] In machine learning applications, CIM circuits are typically used to process dot product multiplications based on performing MAC operations on a large number of data elements (such as input word vectors and weight matrices), and then to process the addition (or accumulation) of such dot products, where each of these data elements can be in the form of a floating-point number. The multiplication of each pair of floating-point numbers typically involves the addition of the respective exponent parts (producing an exponent sum) and the multiplication of the respective mantissa parts (producing a mantissa product). Additionally, the exponent sum of each pair of floating-point numbers is compared with the maximum exponent sum among multiple pairs of floating-point numbers to produce an exponent difference. This exponent difference is used to align the exponent parts of different pairs of floating-point numbers, thereby shifting the corresponding mantissa products. The shifted mantissa products are added to the exponent of the maximum exponent sum to arrive at the final sum.
[0053] Using this method, without considering the exponent sum and range, floating-point numbers (in MAC operations) can be aligned with the maximum exponent sum for use in, for example, shift and accumulate operations. For example, in some scenarios, the maximum exponent sum can be significantly larger than other exponent sums, potentially leading to reduced accuracy because some mantissa products shifted corresponding to relatively small exponent sums can be truncated or shortened when shifted based on the maximum exponent sum. In such cases, and to potentially improve accuracy, some systems can retain additional bits for accumulation purposes, which can increase power and resource (e.g., memory space) consumption. Thus, shifting / aligning the mantissa product to the maximum exponent sum can at least reduce the accuracy of the MAC operation due to truncated mantissa products, or increase power consumption and computing resources when retaining additional bits for accumulating the shifted mantissa products.
[0054] The present disclosure provides various embodiments of compute-in-memory (CIM) circuits that can partition / phase / group floating-point numbers into multiple phases for post-multiplication alignment. Post-multiplication can refer to operations after multiplying floating-point numbers (e.g., multiplying one or more pairs of mantissas). The disclosed CIM circuits can include features or elements for detecting the exponent sum range. The disclosed CIM circuits can determine to partition floating-point numbers into multiple phases for MAC operations based on the exponent sum range. Partitioning the floating-point numbers can include grouping the mantissa products and / or exponent sums associated with respective pairs of floating-point numbers into multiple phases.
[0055] In one aspect, the disclosed CIM circuits can partition all floating-point numbers (or pairs of floating-point numbers) (e.g., the entire exponent sum range and corresponding mantissa products) into multiple phases based on the exponent sum range. In another aspect, the disclosed CIM circuits can partition a portion of the floating-point numbers into multiple phases and ignore or discard the remaining portion of the floating-point numbers (e.g., a portion of the exponent sum range and the corresponding mantissa products outside the desired range from the maximum exponent sum). The disclosed CIM circuits can align or shift the mantissa products according to the maximum exponent sum or the respective local maxima of the exponent sums in each of the phases at respective time periods. For example, the disclosed CIM circuits can add the shifted mantissa products in each of the phases or combine these shifted mantissa products, and accumulate the sum of the mantissa products to output the (final) result of the MAC operation. By phasing the floating-point numbers to shift the mantissa products as discussed herein and accumulating the mantissa products, the systems and methods of the technical solution can reduce computing resources, minimize power consumption, and provide accuracy for MAC operations on floating-point numbers.
[0056] Although the various operations, aspects, or features performed by the CIM circuits herein may be described with respect to post-multiplication, it should be noted that similar operations may be performed for pre-multiplication quantization, such as phasing floating-point numbers before multiplying them. The features or operations of the technical solutions discussed herein may be applied to CIM applications and / or near-memory computing (NMC) applications, etc., or be directed to CIM applications and / or near-memory computing applications, etc.
[0057] Figure 1 FIG. shows a block diagram of a data computing circuit 100 according to some embodiments of the present disclosure. In Figure 1 the illustrated embodiment depicted, the data computing circuit 100 (also referred to as, e.g., a CIM circuit 100 or a memory circuit 100) includes various elements that together are used to perform in-memory computations (such as multiply-accumulate (MAC) operations) on an input word vector and a weight matrix. The input word vector may include a plurality (N) of input data elements InDE, and the weight matrix may include a plurality (Nd) of weight data elements WtDE. In various embodiments, each of the input data elements InDE and the weight data elements WtDE may include floating-point numbers. The circuit 100 may perform in-memory computations using post-multiplication alignment.
[0058] As shown, the circuit 100 includes a memory circuit 102, an input circuit 104, a plurality of multiplier circuits 106, a plurality of adder circuits 108, a difference circuit 110 (e.g., sometimes referred to as a subtractor circuit 110), one or more selector circuits 111A - B (e.g., sometimes referred to as a selector circuit 111), a shift circuit 112, an adder circuit (or adder tree) 114, an adder circuit (or adder tree) 116 (e.g., sometimes referred to as an accumulator or aggregation circuit), a converter 118, a converter 120, and a phasing circuit 122 (e.g., sometimes referred to as a grouping, division, or distribution circuit). In some embodiments, the number of multiplier circuits 106 may correspond to the number of adder circuits 108. For example, the circuit 100 may include N (the number of weight / input data elements WtDE / InDE) multiplier circuits 106 and N (the number of weight / input data elements WtDE / InDE) adder circuits 108. In some implementations, the circuit 100 may include a plurality of shift circuits 112, a plurality of adder circuits 114, and / or a plurality of phasing circuits 122, etc. It should be understood that Figure 1 the block diagram of the circuit 100 depicted in is simplified, and thus, the circuit 100 may include any of various other elements while remaining within the scope of the present disclosure. It should be noted that Figure 1The block diagram of the circuit 100 depicted herein may include more or fewer components, or may include other components, not limited to those shown herein, to perform MAC operations for floating-point numbers using post-multiplication alignment, such as described in conjunction with but not limited to Figures 2 to 11 at least one of
[0059] The memory circuit 102 may include one or more memory arrays and one or more corresponding circuits. Each memory array is a storage device including a plurality of storage components 103, and each of the storage components 103 includes an electrical, electromechanical, electromagnetic, or other device for storing one or more data elements, and each data element includes one or more data bits represented by a logical state. In some embodiments, the logical state corresponds to the voltage level of the charge stored in a part or all of the storage component 103. In some embodiments, the logical state corresponds to the physical characteristics (such as resistance or magnetic orientation) of a part or all of the storage component 103.
[0060] In some embodiments, the storage component 103 includes one or more static random-access memory (SRAM) cells. In various embodiments, the SRAM cell includes a plurality of transistors, such as a five-transistor (5T) SRAM cell, a six-transistor (6T) SRAM cell, an eight-transistor (8T) SRAM cell, a nine-transistor (9T) SRAM cell, and the like. In some embodiments, the SRAM cell includes a multi-track SRAM cell. In some embodiments, the length of the SRAM cell is at least twice the width.
[0061] In some embodiments, the memory component 103 includes one or more dynamic random-access memory (DRAM) cells, resistive random-access memory (RRAM) cells, magnetoresistive random-access memory (MRAM) cells, ferroelectric random-access memory (FeRAM) cells, NOR flash cells, NAND flash cells, conductive-bridging random-access memory (CBRAM) cells, data registers, non-volatile memory (NVM) cells, 3D NVM cells, or other types of memory cells capable of storing bit data.
[0062] In addition to the memory array, the memory circuit 102 may include multiple circuits for accessing or otherwise controlling the memory array. For example, the memory circuit 102 may include multiple (e.g., word line) drivers operatively coupled to the memory array. The drivers may apply signals (e.g., voltages) to the corresponding memory components 103 to allow access (e.g., programming, reading, etc.) to these memory components 103. For another example, the memory circuit 102 may include multiple programming circuits and / or reading circuits operatively coupled to the memory array.
[0063] Each memory array of the memory circuit 102 is configured to store multiple weight data elements WtDE. In some embodiments, the programming circuit may write the weight data elements WtDE into the corresponding memory components 103 of the memory array, respectively, and the reading circuit may read the bits written into the memory components 103 to verify or otherwise test whether the written weight data elements WtDE are correct. The drivers of the memory circuit 102 may include or be operatively coupled to multiple input activation latches configured to receive and temporarily store input data elements InDE. In some other embodiments, such input activation latches may be part of the input circuit 104, and the input circuit 104 may further include multiple buffers configured to temporarily store the weight data elements WtDE retrieved from the memory arrays of the memory circuit 102. Thus, the input circuit 104 may receive the input data elements InDE and the weight data elements WtDE.
[0064] In various embodiments of the present disclosure, the input word vectors (including, for example, input data elements InDE) and the weight matrix (including, for example, weight data elements WtDE) for which the circuit 100 performs MAC operations each include a plurality of floating-point numbers. Thus, each of the data elements InDE and the weight data elements WtDE includes a sign bit, a plurality of exponent bits, and a plurality of mantissa bits (sometimes referred to as fractional bits).
[0065] For example, each of the data elements InDE and the weight data elements WtDE has a BF16 format, which is also referred to as the bfloat format or the brain floating-point format in some embodiments, where the first bit represents the sign of the floating-point number, the following eight bits represent the exponent of the floating-point number, and the last seven bits represent the mantissa or fraction of the floating-point number. Since the mantissa is used to start from a non-zero value, the last seven bits of each stored data element represent an eight-bit mantissa with a first most significant bit (MSB) equal to one.
[0066] In some embodiments, each of the data elements InDE and the weight data elements WtDE has an FP16 format, which is also referred to as the half-precision format, where the first bit represents the sign of the floating-point number, the following five bits represent the exponent of the floating-point number, and the last ten bits represent the mantissa or fraction of the floating-point number. In this case, the last ten bits of each stored data element represent an eleven-bit mantissa with a first MSB equal to one. In some other embodiments, each of the data elements InDE and the weight data elements WtDE has a floating-point format other than the BF16 or FP16 format, such as another 16-bit format, 32-bit, 64-bit, 128-bit, or 256-bit format, or a 40-bit or 80-bit extended precision format. The sign and mantissa of the data element representing the floating-point number are collectively referred to as the signed mantissa of the floating-point number. The MSB of the mantissa is referred to as the hidden bit or the hidden MSB.
[0067] Still referring to Figure 1 , the input circuit 104 is configured to output the entirety of each data element of the data elements InDE and WtDE to each of the multiplier circuit 106 and the adder circuit 108. In some embodiments, the input circuit 104 is configured to output the signed mantissa of each data element to the multiplier circuit 106 and the exponent of each data element to the adder circuit 108, which will be described below.
[0068] The multiplier circuits 106 are each an electronic circuit, such as an integrated circuit (IC), for receiving, for example, the sign bit InS and mantissa InM (collectively referred to as the signed mantissa InS / InM) of each of the N data elements InDE from the input circuit 104 and the sign bit WtS and mantissa WtM (collectively referred to as the signed mantissa WtS / WtM) of each of the N data elements WtDE. The adder circuits 108 are each an electronic circuit, such as an IC, for receiving, for example, the exponent InE of each of the N data elements InDE and the exponent WtE of each of the N data elements WtDE from the input circuit 104.
[0069] The multiplier circuits 106 may each include one or more data registers (not shown) for receiving instances of the signed mantissas InS / InM and WtS / WtM. In the Figure 1 embodiment depicted in, the multiplier circuits 106 are for receiving instances of the signed mantissas InS / InM and WtS / WtM corresponding to the data elements InDE and WtDE. In some other embodiments, the multiplier circuits 106 include one or more data registers for receiving instances of the signed mantissas InS / InM and / or WtS / WtM that include a hidden MSB. In some embodiments, the multiplier circuits 106 include one or more data registers for adding the hidden MSB to the received instances of the signed mantissas InS / InM and / or WtS / WtM.
[0070] The multiplier circuits 106 may include logic circuitry (not shown) for reformatting each instance of the signed mantissa InS / InM into a two's complement mantissa InTC (also referred to as the reformatted mantissa InTC) and reformatting each instance of the signed mantissa WtS / WtM into a two's complement mantissa WtTC (also referred to as the reformatted mantissa WtTC) in an operation. The reformatted mantissa InTC has the same number of bits as the signed mantissa InS / InM, while the reformatted mantissa WtTC has the same number of bits as the signed mantissa WtS / WtM.
[0071] The multiplier circuit 106 may include one or more logic gates M1 for multiplying some or all instances of the reformatted mantissa InTC by some or all instances of the reformatted mantissa WtTC in an operation to produce N products (e.g., P[1] to P[N]). In various embodiments, the one or more logic gates M1 include one or more AND gates or NOR gates or other circuits suitable for performing some or all of the multiplication operations. The one or more logic gates M1 are used to produce each of the products P[1] to P[N] as two's complement data elements in an operation, the two's complement data elements having a number of bits equal to twice the number of bits of the reformatted mantissas InTC and WtTC minus one. The one or more logic gates M1 may be referred to as a multiplier for multiplying some or all instances of the reformatted mantissa InTC by some or all instances of the reformatted mantissa WtTC. In some cases, the multiplier (e.g., the one or more logic gates M1) may receive the signed mantissa InS / InM or the signed mantissa WtS / WtM for the multiplication.
[0072] The multiplier circuit 106 is used to produce N products P[1] to P[N] in an operation. For example, the multiplier circuit 106 may produce N products P[1] to P[N] equal to sixteen. In some other embodiments, the multiplier circuit 106 may produce N products P[1] to P[N] less than or greater than sixteen, such as according to a floating-point format.
[0073] In some embodiments (e.g., embodiments where the data elements InDE and WtDE have the BF16 format), the multiplier circuit 106 is used to produce each of the products P[1] to P[N] having a total of 17 bits based on the signed mantissas InS / InM and WtS / WtM and each of the reformatted mantissas InTC and WtTC having a total of nine bits. In some embodiments (e.g., embodiments where the data elements InDE and WtDE have the FP16 format), the multiplier circuit 106 is used to produce each of the products P[1] to P[N] having a total of 23 bits based on the signed mantissas InS / InM and WtS / WtM and each of the reformatted mantissas InTC and WtTC having a total of 12 bits. Embodiments where the multiplier circuit 106 is used to produce each of the products P[1] to P[N] having other total bit numbers based on the signed mantissas InS / InM and WtS / WtM and each of the reformatted mantissas InTC and WtTC having other total bit numbers are within the scope of this disclosure.
[0074] Accordingly, the multiplier circuit 106 is used to perform multiplication and reformatting operations on the signs and fractional bits of the input data element InDE and the weight data element WtDE during the operation, so as to generate the two's complement products P[1] to P[N]. The multiplier circuit 106 is used to output the products P[1] to P[N] to one or more circuits or components (such as the phase alignment circuit 122 and / or the shift circuit 112) of the circuit 100 on a data bus (not shown).
[0075] Each of the adder circuits 108 includes one or more data registers (not shown) for receiving examples of the exponents InE and WtE corresponding to the number of data elements of the data elements InDE and WtDE discussed above with respect to the multiplier circuit 106. Each of the adder circuits 108 includes one or more logic gates A1 for adding each example of the exponent InE to each example of the exponent WtE during the operation. In various embodiments, the one or more logic gates A1 include one or more full adder gates, half adder gates, ripple carry adder circuits, carry save adder circuits, carry select adder circuits, look-ahead carry adder circuits, or other circuits suitable for performing some or all of the addition operations. The respective logic gates A1 of the adder circuits 108 are used to generate the exponent sums S[1] to S[N] as data elements having a total number of bits equal to the number of bits of each of the exponents InE and WtE plus one.
[0076] The adder circuits 108 are used to generate the exponent sums S[1] to S[N] during the operation, and these exponent sums S[1] to S[N] have the same total number N and data element order as the products P[1] to P[N] discussed above with respect to the multiplier circuit 106. The sorting or arrangement of the exponent sums S[1] to S[N] may be similar to that of the products P[1] to P[N]. Accordingly, for a total of N combinations of the data elements InDE and WtDE, every n combinations correspond to both the nth exponent sum S[n] in the exponent sums S[1] to S[N] and the nth product P[n] in the products P[1] to P[N].
[0077] In some embodiments (e.g., embodiments in which data elements InDE and WtDE have the BF16 format), adder circuit 108 is configured to generate each corresponding one of sums of exponents S[1] to S[N] having a total of nine bits based on each of exponents InE and WtE having a total of eight bits. In some embodiments (e.g., embodiments in which data elements InDE and WtDE have the FP16 format), adder circuit 108 is configured to generate each of sums S[1] to S[N] having a total of six bits based on each of exponents InE and WtE having a total of five bits. Also within the scope of the present disclosure is adder circuit 108 configured to generate each of sums of exponents S[1] to S[N] having a different total number of bits based on each of exponents InE and WtE having a different total number of bits. Adder circuit 108 is configured to output sums of exponents S[1] to S[N] on a data bus (not shown) to other elements of circuit 100, including at least one of selector circuit 111, phasing circuit 122, or difference circuit 110.
[0078] Selector circuit 111 (e.g., 111A and / or 111B) is an electronic circuit that includes or corresponds to one or more logic gates (e.g., L1 and / or L2), such as an IC, the one or more logic gates being configured to receive sums of exponents S[1] to S[N] from adder circuit 108. In some cases, circuit 100 may include one selector circuit 111, such as including a first selector circuit 111A and not including a second selector circuit 111B. In some other cases, circuit 100 may include multiple selector circuits 111, such as including a first selector circuit 111A and a second selector circuit 111B. For the purpose of providing examples herein, circuit 100 may include a first selector circuit 111A (e.g., sometimes commonly referred to as selector circuit 111A) and a second selector circuit 111B (e.g., sometimes commonly referred to as selector circuit 111B).
[0079] Each of selector circuits 111A, 111B may include one or more logic gates. For example, first selector circuit 111A may include one or more logic gates L1. The one or more logic gates L1 are configured to generate, in an operation, a maximum sum of exponents MaxExp as a data element having a value equal to the maximum value of the data elements of sums of exponents S[1] to S[N] and having a number of bits equal to the number of bits of the data elements of sums of exponents S[1] to S[N]. The one or more logic gates L1 are configured to select the maximum sum of exponents MaxExp from sums of exponents S[1] to S[N]. As discussed below, the one or more logic gates L1 are configured to output the maximum sum of exponents MaxExp to one or more other circuits or elements of circuit 100, such as but not limited to phasing circuit 122, difference circuit 110, and / or converter 120.
[0080] In other instances, the second selector circuit 111B may include one or more logic gates L2. The one or more logic gates L2 are used to generate the minimum exponent sum MinExp as a data element in an operation, the data element having a value equal to the minimum value of the data elements of the exponent sums S[1] to S[N] and having a number of bits equal to the number of bits of the data elements of the exponent sums S[1] to S[N]. The one or more logic gates L2 can be used to select the minimum exponent sum MinExp from the exponent sums S[1] to S[N]. As discussed below, the one or more logic gates L2 are used to output the minimum exponent sum MinExp to one or more other circuits or components of the circuit 100, such as but not limited to the phase circuit 122 and / or the converter 120.
[0081] The phase circuit 122 is an electronic circuit that includes one or more logic gates or components, such as an IC, the one or more logic gates or components being used to receive at least one of the exponent sums S[1] to S[N] from the addition circuit 108 and / or the products P[1] to P[N] from the multiplier circuit 106. In some cases, the phase circuit 122 may include or correspond to multiple circuits that are electrically coupled to each other for managing the arrangement of the exponent sums S[1] to S[N] and the products P[1] to P[N]. In some embodiments, the phase circuit 122 may sort, rank, or otherwise arrange the exponent sums S[1] to S[N] and the products P[1] to P[N] according to the magnitude of the exponent sums S[1] to S[N]. For example, the phase circuit 122 may sort or arrange the exponent sums S[1] to S[N] from the highest exponent sum to the lowest exponent sum, or vice versa. The phase circuit 122 may sort or arrange the products P[1] to P[N] based on the arrangement of the exponent sums S[1] to S[N], such that the products P[1] to P[N] can be shifted or aligned (e.g., by the shift circuit 112) according to the corresponding exponent sums S[1] to S[N]. In some settings, there may be certain exponent sums that may be the same as each other. In such cases, the phase circuit 122 may group the similar (or identical) exponent sums S[1] to S[N] together and sort these groups of the exponent sums S[1] to S[N] according to the magnitude of the exponent sums S[n] in each group. For the purpose of providing a descriptive example, each of the exponent sums S[1] to S[N] may be different from each other (e.g., non-repeating), but it should be noted that sorting the exponent sums S[1] to S[N] may include sorting groups of the exponent sums S[1] to S[N], where each group includes individual exponent sums. The phase circuit 122 may perform a similar grouping or arrangement for the products P[1] to P[N].
[0082] The phase - determination circuit 122 is configured to receive the maximum exponent sum MaxExp from the selector circuit 111A. In some cases, the phase - determination circuit 122 may be configured to receive the minimum exponent sum MinExp from the selector circuit 111B. The phase - determination circuit 122 may include one or more logic gates (such as subtractors), and the one or more logic gates are used to determine the difference between the maximum exponent sum MaxExp and the minimum exponent sum MinExp. The phase - determination circuit 122 may generate an exponent - sum range indicating the difference between the maximum exponent sum MaxExp and the minimum exponent sum MinExp, such as the variation between the minimum exponent sum and the maximum exponent sum. In some cases, the phase - determination circuit 122 may add one to the difference between the maximum exponent sum MaxExp and the minimum exponent sum MinExp to generate the exponent - sum range, for example, in order to divide the exponent sum into phases. In some cases, the exponent - sum range may represent or indicate the range of the exponent sums S[1]~S[N]. The exponent - sum range may include multiple bits based on the format, such as a total of nine bits for the BE16 format or a total of six bits for the FP16 format.
[0083] The phase - determination circuit 122 is configured to phase - align, divide, or otherwise group the exponent sums S[1]~S[N] and / or the products P[1]~P[N] into multiple phases (such as N phases). Each respective pair of the exponent sum S[n] and the product P[n] may be in the same phase. In some embodiments, the exponent - sum range may be an indicator for the phase - determination circuit 122 to determine whether to divide the exponent sums S[1]~S[N] and / or the products P[1]~P[N] into N phases. For example, the phase - determination circuit 122 may compare the exponent - sum range with a predetermined threshold. If the exponent - sum range is greater than or equal to the predetermined threshold (for example, indicating a relatively large deviation between the maximum exponent sum MaxExp and the minimum exponent sum MinExp), then the phase - determination circuit 122 may decide to divide at least a portion of the exponent sums S[1]~S[N] and the products P[1]~P[N] into N phases. A relatively large deviation may indicate excessive bit expansion, which increases power and resource consumption, or indicates excessive truncation of relatively small mantissa products during shift operations, which reduces the accuracy of MAC operations. In this case, the shift - and - accumulate operations may be performed in multiple phases, each in a respective time period. Therefore, dividing the exponent sums S[1]~S[N] and the products P[1]~P[N] into N phases for shift - and - accumulate operations can reduce power and resource consumption or improve the accuracy of MAC operations by minimizing the truncation of mantissa products during shift operations.
[0084] If the exponent sum range is less than a predetermined threshold (e.g., indicating a relatively small deviation or no deviation between the maximum exponent sum MaxExp and the minimum exponent sum MinExp), the phase - determining circuit 122 may decide to forward the exponent sums S[1]~S[N] to the difference circuit 110 and forward the products P[1]~P[N] to the shift circuit 112. In this case, the phase - determining circuit 122 may not divide the exponent sums S[1]~S[N] and the products P[1]~P[N] into N phases (e.g., performing the shift and accumulation operations in one phase). In some embodiments, the phase - determining circuit 122 may be used to divide at least a portion of the exponent sums S[1]~S[N] and at least a portion of the products P[1]~P[N] into N phases regardless of the exponent sum range (or without comparing the exponent sum range with the predetermined threshold). For the purpose of providing examples herein, the phase - determining circuit 122 is used to divide the exponent sums S[1]~S[N] and / or the products P[1]~P[N] into N phases in response to receiving the exponent sums S[1]~S[N] and / or the products P[1]~P[N] (e.g., performing phase - determination without comparing the exponent sum range with the predetermined threshold).
[0085] For the purpose of providing examples herein, dividing the exponent sums S[1]~S[N] into N phases may be associated with dividing the corresponding products P[1]~P[N] into N phases, and vice versa. Thus, it should be noted that the operation of dividing the exponent sums S[1]~S[N] into N phases may similarly reflect the operation of dividing the corresponding products P[1]~P[N] into N phases, and vice versa. Additionally, it should be noted that each of the N phases includes a subset of the exponent sums S[1]~S[N] (e.g., sometimes referred to as an exponent subset or sum subset of each individual phase) and / or a subset of the products P[1]~P[N] (e.g., sometimes referred to as a mantissa subset or product subset of each individual phase).
[0086] As discussed herein, the phase - determining circuit 122 may use one of at least two methods (e.g., a first phase - determination method / method I and a second phase - determination method / method II) to divide at least a portion of the exponent sums S[1]~S[N] (and at least a portion of the products P[1]~P[N]) into N phases. Although two (phase - determination) methods are discussed herein for the purpose of providing examples, it should be noted that other techniques or embodiments for dividing the exponent sums S[1]~S[N] (and / or the products P[1]~P[N]) into N phases may be similarly performed by the phase - determining circuit 122 and are not limited to these two methods.
[0087] In some configurations, the phase - determining circuit 122 may determine whether to use the first phase - determination method or the second phase - determination method based on the exponent sum range, such as in combination with at least Figure 3As described. For example, the phasing circuit 122 may compare the sum-of-exponents range (or the difference between the maximum sum-of-exponents and the minimum sum-of-exponents) with a difference threshold. The difference threshold may include a predetermined value corresponding to, for example, the range (or maximum number) of sum-of-exponents that are allowed to be grouped into N phases. The difference threshold may be adjusted based on the configuration of the circuit 100. A sum-of-exponents range greater than or equal to the threshold may indicate a relatively large sum-of-exponents range, where a portion of the sum-of-exponents (e.g., starting from the minimum sum-of-exponents) may have a minimal impact on the result of the MAC operation and may be ignored / discarded. The phasing circuit 122 may decide to use the second phasing method based on the sum-of-exponents range (or difference) being greater than or equal to the difference threshold. Otherwise, the phasing circuit 122 may determine to use either the first phasing method or the second phasing method based on the sum-of-exponents range (or difference) being less than the difference threshold.
[0088] In various embodiments, the phasing circuit 122 may use the first phasing method to divide all the sum-of-exponents S[1] to S[N] (e.g., the entire range of the sum-of-exponents S[1] to S[N]) into N individual phases. The characteristics or operations of the first phasing method for dividing the sum-of-exponents S[1] to S[N] (and / or the products P[1] to P[N]) may be described in conjunction with but not limited to Figures 5 to 7 or Figure 11 at least one of. For example, the phasing circuit 122 may be used to divide the sum-of-exponents S[1] to S[N] into a predefined N number of phases, such as two phases, three phases, four phases, etc. For the purpose of providing an example herein, the predefined N number of phases may be four phases. Using the sum-of-exponents range, the phasing circuit 122 may determine the size of each phase by dividing the sum-of-exponents range by the predefined N number of phases (e.g., four in this case).
[0089] For example, if the sum-of-exponents range is eight (e.g., the range from the highest sum-of-exponents to the lowest sum-of-exponents), the phasing circuit 122 may determine the phase size to be two (e.g., eight divided by four). The phasing circuit 122 may divide the sum-of-exponents S[1] to S[N] starting from the highest sum-of-exponents to the lowest sum-of-exponents or from the lowest sum-of-exponents to the highest sum-of-exponents. In this case, depending on the configuration of the phasing circuit 122, each of the N phases may include two different sum-of-exponents starting from the highest sum-of-exponents or from the lowest sum-of-exponents. It should be noted that although the sum-of-exponents S[1] to S[n] are binary numbers, for the purpose of simplifying the example, the value of each sum-of-exponents S[n] discussed herein may be described in decimal format.
[0090] In some embodiments, depending on the distribution of the exponent sums S[1] to S[N], certain phases may contain more or fewer exponent sums. In some cases, depending on the distribution of the exponent sums S[1] to S[N], certain phases may not contain an exponent sum or be associated with an exponent sum. For example, if there are four exponent sums (or four groups of exponent sums) including the values one, two, seven, and eight, the exponent sum range can be from one to eight (e.g., the range is eight). In this example, if the phasing circuit 122 is used to divide the four exponent sums into four phases evenly distributed over the exponent sum range, the first phase may contain the exponent sums of seven and eight, the second and third phases may not be associated with any exponent sum, and the fourth phase may contain the exponent sums of one and two. Thus, each of the N phases may contain a different number of exponent sums as part of the respective exponent subsets in the individual phases. Similarly, each of the N phases may contain a different number of corresponding products as part of the respective mantissa (product) subsets in the individual phases.
[0091] In some arrangements, the N phases (e.g., four phases) may not be evenly distributed over the exponent sum range. For example, the phasing circuit 122 may determine an exponent sum range of ten, e.g., including an exponent sum of one and an exponent sum of ten. The phasing circuit 122 may determine a phase size of approximately three by dividing ten by four and rounding up (e.g., taking the nearest integer). In this case, assuming the exponent sum range is between one and ten, the phasing circuit 122 may, for example, divide the exponent sums S[1] to S[N] into a first phase associated with a subset of the exponent sum range from eight to ten (e.g., the three largest exponent sums, which represent the largest phase), a second phase associated with a subset of the exponent sum range from five to seven (e.g., the next three largest exponent sums, which represent the second largest phase), a third phase associated with a subset of the exponent sum range from two to four (e.g., the next three largest exponent sums, which represent the second smallest phase), and a fourth phase associated with a subset of the exponent sum range including the exponent sum of one (e.g., the remaining exponent sum, which represents the smallest phase). Thus, each of the N phases may contain a different subset size of the exponent sum range. It should be noted that the phasing circuit 122 may divide the exponent sums S[1] to S[N] (and products P[1] to P[N]) into more or fewer phases, not limited to the number of phases discussed herein.
[0092] In some cases, the phasing circuit 122 may use a predefined phase size (e.g., sometimes referred to as a step or range size) to divide the exponent sums S[1] to S[N] (and corresponding products P[1] to P[N]) into N phases. For example, the phasing circuit 122 can be used to divide an exponent sum range of 20 using a phase size of three. In such cases, the phasing circuit 122 may divide or distribute at least a portion of the exponent sums S[1] to S[N] into seven phases. It should be noted that a phase size of three is used as an example, and other phase sizes can be similarly used for any exponent sum range.
[0093] In some embodiments, the phasing circuit 122 may divide the exponent sums S[1] to S[N] into N phases based on the distribution (or clustering) of the exponent sums S[1] to S[N]. In this case, the number of phases can vary based on the clustering of the exponent sums S[1] to S[N]. For example, the phasing circuit 122 can increase or decrease the size of one or more individual phases. Taking an exponent sum range with an exponent sum of eight and including exponent sums from one to two and from seven to eight as an example, the phasing circuit 122 can divide the exponent sum from seven to eight into a first phase and the exponent sum from one to two into a second phase according to the distribution of the exponent sums S[1] to S[N] (e.g., data points). In another example, for an exponent sum range with an exponent sum of ten and including exponent sums from one to two and from six to ten, the phasing circuit 122 can divide the exponent sum from nine to ten into a first phase, the exponent sum from six to eight into a second phase, and the exponent sum from one to two into a third phase according to the distribution of the exponent sums S[1] to S[N] and the configuration of the phasing circuit 122.
[0094] In other examples, the phasing circuit 122 can divide each of the exponent sums S[1] to S[N] into N phases until a predetermined number of data points (e.g., phase size) is satisfied / met. For example, the phase size can be predetermined or set to 10 (and other values). The phasing circuit 122 can group or divide the exponent sums starting from the largest exponent sum into individual phases until each individual phase contains at least 10 exponent sums. The same (or repeated) exponent sum values can be grouped into the same phase. Thus, the number of phases and / or the number of exponent sums in each of the phases can vary or depend on the distribution of the exponent sums S[1] to S[N].
[0095] The grouping / phasing / division of the entire exponent sum range (and the corresponding products P[1] to P[N] of the exponent sums S[1] to S[N] within the exponent sum range) can be performed as part of a first phasing method. One or more of the features discussed herein for dividing or phasing the exponent sums S[1] to S[N] and the corresponding products P[1] to P[N] can be performed in combination with each other, additionally, or alternatively.
[0096] In various embodiments, the phasing circuit 122 may use a second phasing method to divide a portion of the exponent sums S[1] to S[N] (e.g., a portion or subset of the range of the exponent sums S[1] to S[N]) into N respective phases. The second phasing method may include one or more operations or features similar to the first phasing method. The features of the second phasing method may be described in conjunction with but not limited to Figures 8 to 11 at least one of.
[0097] For example, as part of configuring the phasing circuit 122, the phasing circuit 122 may receive or be configured with a step size (e.g., sometimes referred to as a phase size). The step size may be a predetermined or preset value. The phasing circuit 122 may receive or be configured with a predetermined number of constant step sizes (e.g., N constant step sizes) for dividing the exponent sum range into the corresponding N phases, where the size of each of the N phases may correspond to the step size. For example, to divide the exponent sums S[1] to S[N], the phasing circuit 122 may apply N constant step sizes with the step size to the exponent sum range starting from the largest exponent sum, such as at least Figure 8 as shown.
[0098] As an example, in the case where the exponent sum range is 20, the step size is five, and the number of constant step sizes is three, the phasing circuit 122 may divide the exponent sums S[1] to S[N] starting from the largest exponent sum into three phases associated with the three constant step sizes. In this example, the first phase includes the largest five exponent sums, the second phase includes the second largest five exponent sums continuing from the exponent sums of the first phase, and the third phase includes the third largest five exponent sums continuing from the exponent sums of the second phase. Each of the N phases may include a respective subset of exponents, where in this example, the subset of exponents includes the corresponding five exponent sums. Each of the N phases may include a respective subset of mantissas corresponding to the respective subset of exponents. When using the second phasing method, the last (smallest) five exponent sums may be ignored or discarded from the MAC operation because one or more pairs of exponent sums and products associated with one or more exponent sums outside the N constant step sizes or N phases are considered to be negligible in magnitude, relatively small, or have a relatively large difference from the largest exponent sum. By discarding a portion of the exponent sum range outside the N phases, the phasing circuit 122 may improve (reduce) power consumption and resource consumption while maintaining or improving the accuracy of the MAC operation.
[0099] In some embodiments, certain features of the second phasing method may be applied or performed in combination with or in conjunction with features of the first phasing method. For example, the phasing circuit 122 may use the first phasing method to divide the exponent sums S[1] to S[N] into N phases. As part of the second phasing method, in response to using the first phasing method to obtain N phases, the phasing circuit 122 may discard or ignore a portion of the exponent sums S[1] to S[N] (and a portion of the corresponding products P[1] to P[N]) within at least a portion of the last phase associated with the smallest exponent sum. Additionally or alternatively, the phasing circuit 122 may perform other techniques or features for dividing / phasing / grouping the exponent sums S[1] to S[N] and the corresponding products P[1] to P[N], not limited to the phasing methods discussed herein.
[0100] The phasing circuit 122 may output one or more of the exponent sums of each of the N phases on one or more data buses (not shown) to the difference circuit 110. The exponent sum of each phase may be referred to as a local sum LS[1] to LS[N] (e.g., representing one or more exponent sums local to each respective phase). The phasing circuit 122 may output the local sums LS[1] to LS[N] of each phase in a different time period, time frame, or clock cycle than the other phases of the N phases. For example, the phasing circuit 122 may output the local sums LS[1] to LS[N] of the first phase in a first time period, and output the local sums LS[1] to LS[N] of the second phase in a second time period, and so on.
[0101] The phasing circuit 122 may output one or more of the products of each of the N phases on one or more data buses (not shown) to the shift circuit 112. The product of each phase may be referred to as a local product LP[1] to LP[N] (e.g., representing one or more products local to each respective phase). The phasing circuit 122 may output the local products LP[1] to LP[N] of each phase in different time periods. For example, the phasing circuit 122 may output the local products LP[1] to LP[N] of the first phase in a first time period, and output the local products LP[1] to LP[N] of the second phase in a second time period, and so on. In some arrangements, the phasing circuit 122 may output the local products LP[1] to LP[N] in time periods similar to the corresponding local sums LS[1] to LS[N].
[0102] In some other settings, the phasing circuit 122 may output the partial products LP[1] to LP[N] in a different time period (or different clock cycle or different time instance) than the output corresponding partial sums LS[1] to LS[N]. For example, the phasing circuit 122 may output the partial sums LS[1] to LS[N] to the difference circuit 110 at a first time instance, and output the partial products LP[1] to LP[N] to the shift circuit 112 at a second time instance after the first time instance. In this example, as described herein, the difference circuit 110 may output the exponential sum differences (e.g., D[1] to D[N]) generated based on the corresponding partial sums LS[1] to LS[N] to the shift circuit 112 for shifting / aligning the corresponding partial products LP[1] to LP[N] at the second time instance, which is similar to, for example, outputting the corresponding partial products LP[1] to LP[N]. In some other cases, the phasing circuit 122 may output the partial products LP[1] to LP[N] at a first time instance, and output the corresponding partial sums LS[1] to LS[N] at a second time instance after the first time instance.
[0103] In some configurations, the phasing circuit 122 may include one or more logic gates that are used to select the respective maximum exponential sum from each of the N phases as the respective local maximum exponential sum for the corresponding phase (e.g., sometimes commonly referred to as the local maximum or LocalMax). The one or more logic gates for selecting the respective maximum exponential sum may be similar to the one or more logic gates of the selector circuit 111 (e.g., L1 and / or L2). The phasing circuit 122 may output the local maximum of each respective phase and the partial sums LS[1] to LS[N] to the difference circuit 110 at similar time periods. The local maximum of the first phase associated with the maximum exponential sum MaxExp may be the maximum exponential sum MaxExp. In this case, the phasing circuit 122 may forward the maximum exponential sum MaxExp to the difference circuit 110. In some embodiments, the phasing circuit 122 may output or forward the maximum exponential sum MaxExp for each of the N phases to the difference circuit 110. The phasing circuit 122 may output other values or data to other circuits or elements of the circuit 100 (not limited to the difference circuit 110 or the shift circuit 112).
[0104] It should be noted that variables or values (such as the number of phases or constant steps, the magnitude of each phase or constant step, one or more thresholds, exponents, and ranges, etc.) are not limited to the examples provided herein, and circuit 100 or other devices or components of circuit 100 may similarly use other variables or values (such as different formats, thresholds, number of phases, magnitude of one or more phases, etc.) to perform MAC operations on floating-point numbers with reduced computational resources and power consumption and improved accuracy. Additionally, it should be noted that more or fewer components and / or different settings of one or more components may be implemented to perform the features, operations, or processes discussed herein.
[0105] Difference circuit 110 is an electronic circuit including one or more logic gates B1, such as an IC, and each of the one or more logic gates B1 is configured to receive local sums LS[1] to LS[N] from phasing circuit 122. The one or more logic gates B1 may sometimes be referred to as subtractors. In some cases, one or more logic gates L1 or one or more logic gates L2 may be part of difference circuit 110. As discussed below, the one or more logic gates B1 are configured to receive at least one of the maximum exponents MaxExp.
[0106] The one or more logic gates B1 are configured to generate differences D[1] to D[N] in an operation by subtracting each data element of local sums LS[1] to LS[N] from the maximum exponent sum MaxExp or local maximum (for each of the N phases). For the first phasing method, differences D[1] to D[N] may have a total number N and data element order corresponding to the total number and data element order of exponent sums S[1] to S[N] and products P[1] to P[N] discussed above. For the second phasing method, the total number of data elements N of differences D[1] to D[N] may be less than the total number of data elements of exponent sums S[1] to S[N] and products P[1] to P[N] discussed above because one or more pairs of exponent sums and products outside the N phases are removed. In this case, differences D[1] to D[N] may have a total number N and data element order corresponding to the total number and data element order of local sums LS[1] to LS[N] and local products LP[1] to LP[N] from each of the N phases. In Figure 1 the depicted embodiment, the one or more logic gates B1 are configured to output differences D[1] to D[N] to shift circuit 112 on one or more data buses (not shown).
[0107] In various settings, the operation of at least one of adder circuit 108, difference circuit 110, selector circuits 111A, 111B, and / or phasing circuit 122 can be performed before, after, or in parallel with multiplier circuit 106. In some settings, for example, one or more operations of individual adder circuit 108, difference circuit 110, selector circuits 111A, 111B, and / or phasing circuit 122 can be performed sequentially or in parallel.
[0108] Shift circuit 112 is an electronic circuit that includes one or more registers and / or logic gates, such as an IC, and the one or more registers and / or logic gates are configured to perform a shift operation on each instance LP[n] of local products LP[1] to LP[N] from respective phases based on the value of the corresponding instance D[n] of differences D[1] to D[N] generated according to the corresponding partial sums LS[1] to LS[N]. Each instance P[n] of products P[1] to P[N] (or each instance LP[n] of local products LP[1] to LP[N]) is based on the signs and mantissas of corresponding combinations of data elements InDE and WtDE, and each instance D[n] of differences D[1] to D[N] is based on the sum of exponents of the same combination. Shift circuit 112 is configured to right-shift each instance LP[n] of local products LP[1] to LP[N] by an amount equal to the corresponding difference D[n] of respective phases during the operation, thereby generating shifted products SP[1] to SP[N], where the sign and mantissa bits are aligned according to the added exponents used to generate differences D[1] to D[N]. Based on the alignment, shift circuit 112 is configured to generate each instance SP[n] of shifted products SP[1] to SP[N] having the same exponent using the maximum exponent sum MaxExp as the baseline for the N phases. In some cases, based on the alignment, shift circuit 112 is configured to generate each instance SP[n] of shifted products SP[1] to SP[N] having the same exponent using the local maximum as the baseline for each respective phase among the N phases.
[0109] In some embodiments, to compensate for the right-shift operation, shift circuit 112 can add an instance of the sign bit (zero or one) of each product P[n] (or local product LP[n]) as the leftmost bit of the corresponding shifted product SP[n]. The number of added instances of the sign bit is equal to the right-shift amount determined by the corresponding difference D[n].
[0110] In Figure 1In the illustrated embodiments, as discussed above, the multiplier circuit 106 may generate corresponding instances P[n] of the products P[1] to P[N] by performing multiplication operations. The phasing circuit 122 may divide the products P[1] to P[N] into N phases, where each phase includes respective partial products LP[1] to LP[N] corresponding to the partial sums LS[1] to LS[N]. The shift circuit 112 may include one or more shifters for receiving the partial products LP[1] to LP[N] from the phasing circuit 122 in each of the N phases and selectively outputting one or more (e.g., shifting) of the shifted products SP[1] to SP[N] to the adder circuit 114 based on the respective differences D[1] to D[N]. For example, in Figure 1 , the shifted products output to the adder circuit 114 may include SP[w] to SP[z], where "w" to "z" may each be one of the integers from 1 to N. In one aspect of the disclosure, the sum of the number of SP[w] to SP[z] may be equal to N. In another aspect of the disclosure, the sum of the number of SP[w] to SP[z] may be less than N.
[0111] The shift circuit 112 may shift any one of the partial products LP[1] to LP[N] in each of the N phases and output the shifted products SP[1] to SP[N] to the adder circuit 114 in response to the shift operation. As in the first phasing method, for example, the N phases may be associated with respective portions of the products P[1] to P[N]. Thus, the sum of the number of SP[w] to SP[z] may be equal to N. In some configurations, the shift circuit 112 may detect that at least one of the partial products LP[1] to LP[N] from the phasing circuit 122 or at least one of the products P[1] to P[N] from the multiplier circuit 106 is zero. In such cases, the shift circuit 112 may not perform the shift to the corresponding product with a zero value and / or output the product to the adder circuit 114. Accordingly, the sum of the number of SP[w] to SP[z] may be less than N. In another example, a portion of the products P[1] to P[N] may be discarded based on the corresponding exponents and outside of the N phases (e.g., as part of a second phasing method). In such cases, the sum of the number of SP[w] to SP[z] may be less than N.
[0112] In addition, to generate SP[w] to SP[z], the shift circuit 112 may right-shift each instance P[n] of the products P[w] to P[z] (or each instance LP[n] of the partial products LP[1] to LP[N]) by an amount equal to the corresponding difference DA[n], thereby aligning the sign and mantissa bits according to the addition exponent. In some embodiments, the difference DA[n] may be generated (e.g., by the difference circuit 110) based on subtracting each data element of the local sums LS[w] to LS[z] from the maximum exponent sum MaxExp or the local maximum of the respective phases. The maximum exponent sum MaxExp may correspond to the maximum value of the data elements of the sums S[w] to S[z]. The local maximum (e.g., LocalMax) may correspond to the maximum value of the data elements of the local sums LS[w] to LS[z] in the respective phases of the N phases. Based on the alignment, the shift circuit 112 may use the maximum exponent sum MaxExp or the local maximum as a baseline to generate each instance SP[n] of the shifted products SP[w] to SP[z] having the same exponent.
[0113] In some embodiments (e.g., embodiments where the data elements InDE and WtDE have the BF16 format), the shift circuit 112 is configured to generate each of the shifted products (e.g., SP[1] to SP[N]) having a total of 21 bits based on each of the products P[1] to P[N] or the partial products LP[1] to LP[N] having a total of 17 bits. In some embodiments (e.g., embodiments where the data elements InDE and WtDE have the FP16 format), the shift circuit 112 is configured to generate each of the shifted products (e.g., SP[1] to SP[N]) having a total of 27 bits based on each of the products P[1] to P[N] or the partial products LP[1] to LP[N] having a total of 23 bits. Within the scope of the present disclosure, the shift circuit 112 may be configured to generate each of the shifted products SP[1] to SP[N] having other total bit numbers based on each of the products P[1] to P[N] or the partial products LP[1] to LP[N] having other total bit numbers.
[0114] Based on the products P[1] to P[N] having the two's complement format (and thus the partial products LP[1] to LP[N] having the two's complement format), the shift circuit 112 is configured to generate the shifted products having the two's complement format, such as SP[1] to SP[N]. As discussed above, in Figure 1 the illustrated example, the shift circuit 112 is configured to output the shifted products SP[w] to SP[z] to the adder circuit (tree) 114 on a data bus (not shown).
[0115] The adder tree 114 is an electronic circuit of multiple layers that includes one or more logic gates (not shown), such as an IC, for example, as discussed above with respect to one or more logic gates A1 (of the adder circuit 108). For example, the adder tree 114 may include a first layer for receiving the shifted products SP[w] to SP[z] and a last layer for generating the sum 115 as a data element corresponding to the sum of the shifted products SP[w] to SP[z]. The adder tree 114 may generate the sum 115 for each respective phase among N phases. For example, the adder tree 114 may generate the sum 115 of the first phase at a first time period and another sum 115 of the second phase at a second time period, and so on. In some embodiments, each of one or more consecutive layers between the first layer and the last layer is configured to receive a first number of addition data elements generated by the previous layer and generate a second number of addition data elements based on the first number of addition data elements, where the second number is half of the first number. Thus, the total number of layers includes the first layer, the last layer, and each subsequent layer (if any).
[0116] The adder tree 114 may output the sum 115 for each of the N phases on one or more data buses (not shown) to the adder circuit 116. The adder circuit 116 is an electronic circuit of multiple layers that includes one or more logic gates (not shown), such as an IC, for example, as discussed above with respect to one or more logic gates A1 (of the adder circuit 108). The adder circuit 116 may be configured to perform operations similar to those of the adder tree 114, such as for adding multiple sums 115 of the shifted products SP[w] to SP[z] from different phases. For example, the adder circuit 116 may include at least one memory array (not shown) or be electrically coupled thereto, and the memory array is configured to store one or more sums 115 from the adder tree 114. The adder circuit 116 may receive the sum 115 from the adder tree 114 at each respective phase among the N phases, and the sum 115 may be stored in the memory array. For one or more remaining phases among the N phases, the adder circuit 116 may include one or more other sums 115 from the adder tree 114. The adder circuit 116 may combine each of the sums 115 of the N phases to generate a combined result (e.g., which may represent the result of a shift and accumulate operation). In some cases, the adder circuit 116 may receive an indication of the number of phases from the phasing circuit 122.
[0117] In some embodiments, the adder circuit 116 may include one or more elements or features similar to the shift circuit 112. For example, the adder circuit 116 may receive a plurality of sums 115 each aligned with a different exponent. In this example, the adder circuit 116 may perform a shift operation on the sums 115, such as shifting one or more of the sums 115 to the right to align the sums 115 with the maximum exponent sum MaxExp. In response to the alignment, the adder circuit 116 may combine the aligned sums to produce a combined sum. The adder circuit 116 may repeatedly combine the sums 115 with other sums 115 according to N phases and output the combined sum (e.g., sum PSTC 115S) to the converter 118 on one or more data buses (not shown).
[0118] In some embodiments, the sum PSTC 115S is sometimes referred to as a partial sum PSTC or a mantissa sum PSTC and has a total number of bits corresponding to the number of bits and the number of data elements of the shifted products SP[w] to SP[z]. In some embodiments, the number of bits of the sum PSTC 115S is equal to the number of bits of the shifted products SP[w] to SP[z] plus the number of bits capable of representing the number of data elements of the shifted products SP[w] to SP[z]. In some embodiments, the number of bits of the sum PSTC 115S is equal to the number of bits of the shifted products SP[w] to SP[z] plus four bits capable of representing 16 data elements of the shifted products SP[w] to SP[z].
[0119] In some embodiments (e.g., embodiments in which the data elements InDE and WtDE have the BF16 format), the adder tree 114 is used to generate a sum PSTC 115S having a total of 25 bits based on each of the shifted products SP[w] to SP[z] having a total of 21 bits. In some embodiments (e.g., embodiments in which the data elements InDE and WtDE have the FP16 format), the adder tree 114 is used to generate a sum PSTC 115S having a total of 31 bits based on each of the shifted products SP[w] to SP[z] having a total of 27 bits. The adder tree 114 for generating the sum PSTC 115S based on each of the shifted products SP[w] to SP[z] having other total numbers of bits is within the scope of the present disclosure.
[0120] According to various embodiments of the present disclosure, based on the shifted products SP[w] to SP[z] in two's complement format, adder tree 114 is used to generate a sum PSTC 115S in two's complement format. Thus, adder tree 114 is used to output sum PSTC 115S to converter 118 on a data bus (not shown). In some other embodiments, adder tree 114 may output sum PSTC 115S to a circuit (not shown) external to circuit 100.
[0121] Converter 118 is an electronic circuit including logic circuitry, such as an IC, the logic circuitry being used to receive sum PSTC 115S from adder circuit 116 (or in some cases from adder tree 114) in an operation and convert sum PSTC 115S from two's complement to a sum PSSM in sign plus mantissa format. Converter 118 is used to generate a sum PSSM having the same number of bits as the number of bits of sum PSTC. In the Figure 1 embodiment depicted in, converter 118 is further used to output sum PSSM to converter 120 on a data bus (not shown). In some other embodiments, converter 118 may output sum PSSM to a circuit (not shown) external to circuit 100.
[0122] Converter 120 is an electronic circuit including logic circuitry, such as an IC, the logic circuitry being used to receive sum PSSM from converter 118 and the maximum exponent MaxExp from difference circuit 110 or selector circuit 111A in an operation and convert sum PSSM from sign plus mantissa format to a sum PS, the sum PS having an output format based on sum PSSM and MaxExp and different from the sign plus mantissa format, for example, the floating point format as discussed above. In various embodiments of the present disclosure, converter 120 may generate a sum PS compatible with a circuit (not shown) external to circuit 100. For example, converter 120 is used to output sum PS to a circuit (not shown) external to circuit 100 (such as a memory array or other instance of circuit 100 that is part of a convolutional neural network (CNN)). In some settings, converter 118 may be part of converter 120 and vice versa.
[0123] Figure 2 Illustrates an example process 200 of performing a MAC operation executed by Figure 1 circuit 100 according to some embodiments. Example process 200 may be performed by circuit 100 or one or more elements of circuit 100. Thus, the following embodiments of process 200 may be described in conjunction with but not limited to Figure 1 In addition, the following embodiments of process 200 may be described in conjunction with but not limited to Figures 5 to 11At least one of the following embodiments of process 200 is described. The illustrated embodiments of process 200 are provided as examples and do not limit the scope of the present disclosure. Accordingly, it should be understood that any of the various operations of process 200 may be omitted, reordered, and / or added while remaining within the scope of the present disclosure.
[0124] According to various embodiments, process 200 begins with an operation 202 for detecting an exponent sum and a range. For example, the exponent sum and range may correspond to the range between the maximum / highest (maximum value) exponent sum and the minimum / lowest (minimum value) exponent sum among the exponent sums S[1] to S[N]. For example, circuit 100 (e.g., the first selector circuit 111A) may select the largest among the exponent sums S[1] to S[N] as (or to generate) the maximum exponent sum. Circuit 100 (e.g., the second selector circuit 111B) may select the smallest among the exponent sums S[1] to S[N] as (or to generate) the minimum exponent sum. The exponent sums S[1] to S[N] may be generated by the adder circuit 108. The exponent sums S[1] to S[N] may correspond to the respective products P[1] to P[N] (generated by the multiplier circuit 106) of pairs of exponent sums and products (e.g., input pairs). Circuit 100 (e.g., the phase determination circuit 122) may detect or determine the exponent sum range based on the difference between the maximum exponent sum and the minimum exponent.
[0125] Process 200 continues to an operation 204 for determining data (e.g., MAC data including the exponent sums S[1] to S[N] and / or the corresponding products P[1] to P[N]) for partitioning MAC operations based on the exponent sum range. Circuit 100 (e.g., the phase determination circuit 122) may determine to partition the exponent sums S[1] to S[N] (and the corresponding products P[1] to P[N]) into phases based on an exponent sum range that meets a predetermined threshold. The exponent sum range that meets the predetermined threshold may include an exponent sum range that is greater than or equal to the predetermined threshold. For example, the predetermined threshold may include a value that, if the exponent sum range is met, may indicate excessive power consumption or resource consumption and / or potential inaccuracy during the execution of the MAC operation. In such cases, if the exponent sum range is at or above the predetermined threshold, circuit 100 may determine to partition the data of the MAC operation into multiple phases (e.g., using at least one of the phase determination methods discussed herein) to minimize resource consumption, reduce power consumption, and improve the accuracy of the MAC operation. In some embodiments, if the exponent sum range is less than the predetermined threshold, circuit 100 may not partition the data of the MAC operation into phases. For the purpose of providing an example, circuit 100 may determine to partition the data of the MAC operation as described herein.
[0126] Process 200 continues to an operation 206 for selecting a first phase determination method or a second phase determination method in response to determining to partition the data of the MAC operation. Operation 206 may be combined with but not limited toFigure 1 or Figure 3 described by at least one of. For example, Figure 3 illustrates an example process (e.g., associated with operation 206) of a (phasing) method selected according to some embodiments for dividing an exponent sum into multiple phases for Figure 2 MAC operations. The example process of operation 206 can be performed by circuit 100 or one or more elements of circuit 100 and other elements or circuits (not limited to Figure 1 those elements or circuits in).
[0127] The process of operation 206 may include or begin with an operation 302 for receiving a difference threshold. Circuit 100 may receive or be configured with a difference threshold. As described herein, the difference threshold may include a predetermined value representing an upper limit of the difference between the maximum exponent sum and the minimum exponent sum, thereby indicating whether a portion of the MAC data will be ignored or discarded (e.g., using a second phasing method).
[0128] The process of operation 206 continues to an operation 304 for comparing the difference between the maximum value (e.g., the maximum exponent sum) and the minimum value (e.g., the minimum exponent sum) with the difference threshold. The difference may correspond to an exponent sum range. The process of operation 206 continues to an operation 306 for determining whether the difference is greater than or equal to the difference threshold. For example, circuit 100 (e.g., phasing circuit 122) may compare the difference between the maximum value and the minimum value with the difference threshold.
[0129] If the difference is greater than or equal to the difference threshold, circuit 100 may decide to proceed to an operation 308 for selecting a second phasing method to be used for dividing the MAC data (e.g., exponent sums S[1] to S[N] and products P[1] to P[N]) into N phases. By using the second phasing method, circuit 100 (e.g., phasing circuit 122) may divide at least a portion of the MAC data into N phases and discard other portions of the MAC data outside the N phases.
[0130] If the difference is less than the difference threshold, circuit 100 may decide to proceed to an operation 308 for selecting either the first phasing method or the second phasing method to be used for dividing the MAC data into N phases. At operation 310, circuit 100 may be configured such that the first phasing method is selected when the difference is less than the difference threshold. By using the first phasing method, circuit 100 (e.g., phasing circuit 122) may divide all the MAC data into N phases such that all data elements of the exponent sums S[1] to S[N] and the corresponding products P[1] to P[N] are considered during the MAC operation. In some configurations, circuit 100 (e.g., phasing circuit 122) may be used to utilize the second phasing method regardless of the difference. In such cases, for example, operation 206 may not be performed, and process 200 may continue to operation 210.
[0131] Return reference Figure 2 , based on whether the first phase - determination method or the second phase - determination method is selected, process 200 proceeds to operation 208 or operation 210. For example, in response to selecting the first phase - determination method, process 200 proceeds to operation 208. In response to selecting the second phase - determination method, process 200 proceeds to operation 210.
[0132] At operation 208, circuit 100 (e.g., phase - determination circuit 122) may use the first phase - determination method to divide the exponent sums S[1]~S[N] into N phases. It should be noted that dividing the exponent sums S[1]~S[N] may include dividing the corresponding products P[1]~P[N] (e.g., similarly applied to the corresponding products P[1]~P[N]), such that any permutation of the exponent sums S[1]~S[N] is similarly reflected on, for example, the corresponding products P[1]~P[N]. For example, circuit 100 may receive or determine an exponent - sum range. Circuit 100 may divide the exponent sums S[1]~S[N] into N (multiple) phases based on the exponent - sum range, where each of the N phases is associated with a respective subset of the N exponent subsets. Each exponent subset may include one or more of the exponent sums S[1]~S[N] respectively. Each phase is associated with a respective mantissa (product) subset of the N mantissa subsets (including one or more of the products P[1]~P[N] respectively). Each mantissa subset may correspond to a respective exponent subset.
[0133] For example, the N phases may correspond to or be a predetermined / pre - defined number of phases. Using the first phase - determination method, circuit 100 may determine the size of each phase based on the exponent - sum range divided by the predetermined number of phases. Circuit 100 may divide the exponent sums S[1]~S[N] into one or more of the N phases that are evenly segmented between the minimum exponent sum and the maximum exponent sum.
[0134] In some embodiments, each of the N phases may include respective phase sizes based on the distribution of the exponential sums S[1] to S[N]. For example, an exponential sum range (e.g., including one or more exponential sums) having a relatively high number (e.g., concentrated or clustered) of data points (or relatively high repetition of certain exponential sums) may be divided into at least one phase having a relatively small phase size, e.g., to reduce power consumption or avoid reserving more bits for accumulation. For example, an exponential sum range having a relatively low number of data points (or relatively low repetition of certain exponential sums) may be divided into at least one phase having a relatively large phase size. Alternatively, in some embodiments, the number of phases may be based on the distribution of the exponential sums S[1] to S[N], such as described above. In response to dividing the exponential sums S[1] to S[N] (and products P[1] to P[N]) into N phases, the circuit 100 (e.g., the difference circuit 110) may calculate the exponential sum differences D[1] to D[N] based on the exponential sums S[1] to S[N] and one of the maximum exponential sum MaxExp or the local maximum (LocalMax).
[0135] At operation 210, the circuit 100 (e.g., the phasing circuit 122) may use a second phasing method to divide at least a portion of the exponential sums S[1] to S[N] (and products P[1] to P[N]) into N phases. Each of the N phases may include respective exponential subsets of the N exponential subsets and corresponding mantissa subsets of the N mantissa subsets or be associated with them. Each exponential subset includes or corresponds to at least one exponential sum S[n] of the exponential sums S[1] to S[N]. Each mantissa subset includes or corresponds to at least one product P[n] of the products P[1] to P[N].
[0136] For example, using the second phasing method, the circuit 100 may divide at least a portion of the exponential sums S[1] to S[N] into N phases starting from the maximum exponential sum. The N phases may correspond to N constant steps within the exponential sum range, e.g., the range of the N constant steps is less than or equal to the exponential sum range. The N constant steps may have a predefined step size. Each constant step may indicate different exponential sums (e.g., different magnitudes) to be grouped into one phase. For example, the circuit 100 may group a first portion of the exponential sums S[1] to S[N] within a first constant step starting from the maximum exponential sum (e.g., the maximum exponential sum MaxExp) into a first phase. The circuit 100 may group a second portion of the exponential sums S[1] to S[N] within a second constant step continuing from the first constant step into a second phase. The circuit 100 may group other portions of the exponential sums S[1] to S[N] based on the N constant steps and / or the N phases. Thus, the first phase may include one or more maximum exponential sums, the second phase may include one or more second-largest exponential sums, and so on.
[0137] In response to partitioning portions of the exponent sums S[1] to S[N] (and corresponding products P[1] to P[N]) into N phases using a constant step size with a step, at least one other portion of the exponent sums S[1] to S[N] (and corresponding products P[1] to P[N]) may not be grouped into any of the N phases. In this case, the other portion that is not grouped into the N phases includes one or more exponent sums having the largest difference from the largest exponent sum. In other words, circuit 100 may determine that the other portion includes one or more exponent sums and one or more corresponding products that, when used as part of a MAC operation, will be negligible or provide / have a minimal (or no) impact on the result of the MAC operation, thereby potentially increasing power consumption or resource consumption, for example. In this example and continuing to operation 212, as part of the second phasing method, circuit 100 is used to discard or ignore data elements (such as exponent sums S[1] to S[N] and corresponding products P[1] to P[N]) of another portion of the MAC data that have a relatively large difference from the largest exponent sum (e.g., outside of the N phases).
[0138] In some embodiments, circuit 100 (such as phasing circuit 122) may receive an updated step. Circuit 100 may update the step of the constant step size to modify the phases of at least the exponent sums S[1] to S[N]. In some other embodiments, circuit 100 (such as phasing circuit 122) may determine to adjust the step, such as in conjunction with at least Figure 4 as described.
[0139] For example, Figure 4 illustrates an example process 400 of adjusting the step of a MAC operation in accordance with some embodiments Figure 2 The process 400 may be part of one or more operations of process 200, such as part of operation 210. Process 400 begins with an operation 402 for receiving a step or determining a preset (or predefined) step. For example, the step may be a preset step predefined according to the configuration or specification of circuit 100. Process 400 continues to an operation 404 for adding the MAC data in each of the constant step sizes with the step.
[0140] Process 400 continues to an operation 406 for determining whether the addition result has a carry. If the addition result has a carry, then process 400 proceeds to an operation 408 for reducing the step and loops back to operation 404. Operations 404, 406, and 408 may be repeated until the addition result has no carry. If the addition result has no carry, then process 400 continues to an operation 410 for applying the step (such as the received step or the adjusted (or reduced) step) to the constant step size of the second phasing method (such as setting the preset step to the step to be applied to the constant step size).
[0141] In some embodiments, in order to partition at least a portion of the exponent sums S[1] to S[N] into N phases (or prior to partitioning), circuit 100 (e.g., phasing circuit 122 or other element) may arrange or sort the exponent sums S[1] to S[N] based on the magnitudes of the exponent sums S[1] to S[N]. For example, circuit 100 may sort the exponent sums S[1] to S[N] from the largest exponent sum to the smallest exponent sum, or vice versa. The sorted list of exponent sums S[1] to S[N] may be used for phasing or grouping the MAC data. Circuit 100 may align the corresponding products P[1] to P[N] with the exponent sums S[1] to S[N].
[0142] After partitioning the exponent sums S[1] to S[N] into N phases, circuit 100 (e.g., difference circuit 110) may compute or calculate N exponent differences for each of the N phases. Each of the N exponent differences may be equal to the difference between the corresponding one of the exponent sums S[1] to S[N] (e.g., local sums LS[1] to LS[N]) of the respective exponent subsets from a particular phase and the largest exponent sum (e.g., largest exponent sum MaxExp). The exponent differences may be used for shift and add operations. As described herein, each of the N exponent differences may be equal to the difference between the corresponding one of the local sums LS[1] to LS[N] of the respective exponent subsets from a particular phase and the local maximum of the particular phase. The smallest exponent sum may be associated with the largest exponent difference among the N exponent differences. The largest exponent sum may be associated with the smallest exponent difference among the N exponent differences (e.g., a difference of zero compared to the largest exponent sum).
[0143] In some embodiments, circuit 100 (e.g., phasing circuit 122) is used to determine the local maximum of the respective exponent subsets in each of the N phases by selecting the largest of the exponent sums in the respective phases. In this case, each of the N exponent differences may be equal to the difference between the corresponding one of the N exponent sums from the respective exponent subsets and the local maximum of the respective exponent subsets.
[0144] Process 200 continues to operation 214 for performing shift and add operations (e.g., from operation 208 or operation 212). One or more of the corresponding exponent sums of the exponent subsets are used to shift one or more of the products of the mantissa subsets in each of the N phases. The partial sums obtained by adding the shifted products in each phase are accumulated.
[0145] For example, a circuit 100 (such as a shift circuit 112) may perform a shift operation to shift local products LP[1] to LP[N] of respective mantissa subsets based on corresponding exponential differences of each phase. As an example, the respective mantissa subsets may each include at least a first mantissa subset and a second mantissa subset of products P[1] to P[N] of a first phase and a second phase among N phases. Each of the mantissa subsets is associated with one or more local products LP[1] to LP[N] of a specific phase. Data elements of the local products LP[1] to LP[N] and the corresponding differences D[1] to D[N] may be provided from a phase alignment circuit 122 to the shift circuit 112 at different time periods or time instances.
[0146] After shifting the local products LP[1] to LP[N] (such as the shifted mantissa subsets), a circuit 100 (such as an adder tree 114 or a first adder circuit) may add the shifted (mantissa) products of each respective phase. For example, a circuit 100 (such as a shift circuit 112) may generate a shifted first mantissa subset at a first time period, the shifted first mantissa subset including shifted products SP[1] to SP[N] in a first phase. The circuit 100 may generate a shifted second mantissa subset at a second time period, the shifted second mantissa subset including shifted products SP[1] to SP[N] in a second phase. The circuit 100 may add together the shifted products SP[1] to SP[N] of the shifted first mantissa subset in the first phase at the first time period to, for example, generate a first sum. The circuit 100 may add together the shifted products SP[1] to SP[N] of the shifted second mantissa subset in the second phase at the second time period to, for example, generate a second sum. For example, the circuit 100 may add the other shifted products SP[1] to SP[N] of the remaining N phases.
[0147] After adding the shifted products SP[1] to SP[N] of N phases, a circuit 100 (such as an adder circuit 116 or a second adder circuit) may accumulate or combine the sums associated with the N phases. For example, the circuit 100 may combine the first sum and the second sum and the other sums associated with the N phases to generate a result of the shift and accumulation operation (such as a combined sum or sum PSTC 115S). It should be noted that the operations of process 200 are provided as examples and are not intended to limit the features or functionality of process 200. Additionally, it should be noted that more or fewer operations may be implemented for process 200 to, for example, perform MAC operations.
[0148] Figure 5 is according to some embodiments using by Figure 1FIG. 500 is a graph of an example distribution and grouping of the sum of exponents of a first phasing method (e.g., Method I) performed by circuit 100. FIG. 500 illustrates an example range of sum of exponents on the x-axis (e.g., from a minimum sum of exponents MinExp to a maximum sum of exponents MaxExp). FIG. 500 illustrates an example probability of the distribution of the sum of exponents (e.g., sum of exponents S[1] to S[N]) across the sum of exponent range on the y-axis. In this case, for example, there may be relatively more sums of exponents (or more repetitions, groups, or clusters of similar sums of exponents) received from input circuit 104 or memory circuit 102, near the median of the sum of exponent range. Using the first phasing method, circuit 100 may divide the sums of exponents within the sum of exponent range into N phases. For purposes of providing an example, the N phases may include four phases 502A - D, as Figure 5 shown, but other numbers of phases may be implemented. Dividing the sums of exponents into phases and utilizing the data in the phases with the first phasing method may be described in conjunction with but not limited to Figure 6 and Figure 7 at least one of.
[0149] Figure 6 FIG. 600 is a flowchart of an example process for performing a MAC operation using a first phasing method performed by circuit 100 according to some embodiments. The elements or operations of example process 600 may correspond to or be part of circuit 100. For example, one or more elements of example process 600 may include or correspond to at least a portion of circuit 100. For example, one or more operations of one or more elements of example process 600 may be described in conjunction with Figure 1 at least one of etc. Figures 1 to 3 etc.
[0150] For example, process 600 includes at least one input activation latch 602, at least one weight buffer 604, at least one multiplier 606, at least one adder 608, at least one circuit, at least one aligner 612, at least one adder tree 614, and at least one buffer and accumulator 616. Input activation latch 602 and weight buffer 604 may correspond to or perform operations similar to input circuit 104 to, for example, provide input data elements InDE and weight data elements WtDE to one or more circuits or elements of circuit 100. Each input data element InDE and corresponding weight data element WtDE may correspond to a pair of data elements (or a pair of inputs) in a plurality of pairs. Although Figure 6 16 input data elements InDE and 16 weight data elements WtDE are shown, other numbers of data elements may be included as part of the MAC operation.
[0151] Multiplier 606 may correspond to or perform operations similar to at least one of the multiplier circuits 106. Multiplier 606 may receive the sign parts and mantissa parts of the input data element InDE and the weight data element WtDE. For example, multiplier 606 may multiply the first mantissa (of the input data element InDE) and the second mantissa (of the weight data element WtDE) in each respective input pair to produce one of the N (mantissa) products (such as product P[n]). Multiplier 606 may multiply other mantissa pairs (and sign parts) to produce N products (such as products P[1] to P[N]).
[0152] Adder 608 may correspond to or perform operations similar to at least one of the adder circuits 108. Adder 608 may receive the exponent parts of the input data element InDE and the weight data element WtDE. For example, adder 608 may add the first exponent (of the input data element InDE) and the second exponent (of the weight data element WtDE) in each respective input pair to produce one of the N exponent sums (such as exponent sum S[n]). Adder 608 may add other exponent pairs to produce N exponent sums (such as exponent sums S[1] to S[N]).
[0153] Phaser 610 may correspond to or perform operations similar to phasing circuit 122. For example, phaser 610 may identify, find, or receive the maximum exponent sum (such as MaxExp) and the minimum exponent sum (such as MinExp) (e.g., from selector circuit 111). Phaser 610 may determine the exponent sum range based on the difference between the maximum exponent sum and the minimum exponent sum, such as Figure 5 shown. Using the first phasing method, phaser 610 may divide the exponent sum range into N phases (such as four phases 502A - D). For the purpose of providing an example, and with reference to Figure 5 , phaser 610 may divide the exponent sums within the exponent sum range into four phases 502A - D with similar phase sizes. As shown, as an example, there may be seven exponent sums within the exponent sum range. Each of the four phases 502A - D may include respective exponent sum ranges, where phaser 610 is used to determine which phase each exponent sum belongs to.
[0154] For example, phaser 610 may divide the first largest sum of exponents and the second largest sum of exponents into a first phase 502A. Phaser 610 may divide the third largest sum of exponents and the fourth largest sum of exponents into a second phase 502B. Phaser 610 may divide the fifth largest sum of exponents and the sixth largest sum of exponents into the second phase 502B. Phaser 610 may divide the seventh largest sum of exponents (the smallest sum of exponents in this case) into the second phase 502B. It should be noted that FIG. 500 illustrates the sums of exponents sorted from the smallest / minimum value to the largest / maximum value. One or more sums of exponents in each phase may be referred to as an exponent subset (e.g., a subset of the sums of exponents that have been divided into respective phases) or may be part of an exponent subset. In some embodiments, phaser 610 may provide the sums of exponents of the exponent subset to at least one subtractor (e.g., difference circuit 110) to obtain one or more exponent differences between each of the sums of exponents and the largest sum of exponents for aligning corresponding mantissa products.
[0155] Aligner 612 may correspond to or perform operations similar to shift circuit 112. For example, aligner 612 may receive one or more products from phaser 610 at respective ones of the four phases 502A - D. The products may be local products of the respective phases. Aligner 612 may receive the exponent differences corresponding to the products of the respective phases. In this example, aligner 612 may perform a shift operation to align the products with the largest sum of exponents, e.g., the alignment of the products is based on the corresponding exponent sum positions. Aligner 612 may generate or output one or more shifted products of the respective phases.
[0156] Adder tree 614 may correspond to or perform operations similar to adder tree 114. For example, adder tree 614 may receive the shifted products of each of the phases from aligner 612. Adder tree 614 may add the shifted products to each other to produce the sum (e.g., product sum) of each phase.
[0157] Buffer and accumulator 616 may correspond to or perform operations similar to adder circuit 116. For example, buffer and accumulator 616 may receive a first sum from adder tree 614 in the first phase and store the first sum in a buffer. Buffer and accumulator 616 may receive a second sum from adder tree 614 in the second phase and store the second sum in a buffer. Buffer and accumulator 616 may accumulate, add, or combine the sums associated with the respective phases. Aligner 612, adder tree 614, and buffer and accumulator 616 may perform operations for N phases repeatedly. Buffer and accumulator 616 may output the combined sum of the N phases as the result of a shift and accumulate operation (or MAC operation).
[0158] In some embodiments, aligner 612 may shift the products based on the local maximum of each of the N phases, such as in combination with at leastFigure 7 as described. Figure 7 illustrates an example process 700 for performing MAC operations using a first phasing method with local maxima performed by circuit 100 of Figure 1 . Process 700 may include various operations or elements similar to process 600, such as input activation latch 602, weight buffer 604, multiplier 606, adder 608, phaser 610, aligner 612, adder tree 614, and buffer and accumulator 616. Process 700 may include other elements from circuit 100, not limited to those described herein.
[0159] In this case, process 700 may include at least one phaser 702 to find local maximum exponents (e.g., the sum of maximum exponents for each respective phase local). Phaser 702 may be part of phaser 610. For example, as described in connection with Figure 5 , phaser 702 may select the maximum of the sum of exponents in each of the phases as the respective local maximum. The maximum sum of exponents may be the local maximum of the first phase 502A. In such cases, the exponent difference may be the difference between each (local) sum of exponents and the local maximum of the respective phase. Aligner 612 may use the exponent difference determined based on the local maximum to align the products.
[0160] Figure 8 is a graph 800 of an example distribution and grouping of sums of exponents using a second phasing method (e.g., Method II) performed by circuit 100 of Figure 1 . Similar to Figure 5 , graph 800 illustrates an example probability of the example sum of exponents range on the x-axis and the sum of exponents distribution on the y-axis. Graph 800 may include one or more features similar to Figure 5 's graph 500. Using the second phasing method, circuit 100 may divide the sums of exponents within the sum of exponents range into N phases associated with N constant steps. For illustrative purposes, the N phases may include three phases 802A - C associated with three constant steps, as shown in Figure 8 , but other numbers of phases may be implemented. Dividing the sums of exponents into phases and leveraging the data in the phases with the second phasing method may be described in conjunction with but not limited to Figure 9 and Figure 10 at least one of.
[0161] Figure 9 illustrates an example process 900 for performing MAC operations using a second phasing method performed by circuit 100 of Figure 1 . Process 900 may include similar to Figure 6 or Figure 7Various operations or elements of at least one of processes 600 or 700, such as input activation latch 602, weight buffer 604, multiplier 606, adder 608, aligner 612, adder tree 614, and buffer and accumulator 616. In this case, at least one phaser 902 can be used to perform a second phasing method. Process 900 can include other elements from circuit 100, not limited to those described herein.
[0162] For example, phaser 902 can correspond to or perform operations similar to phasing circuit 122 (or phaser 610). Using the second phasing method, phaser 902 can find the maximum exponent sum corresponding to the largest one in the exponent sums. Phaser 902 can divide the exponent sums (from adder 608) into N phases with a constant step size. The constant step size can include a predefined step. In some cases, the step can be adjusted or updated. The constant step size can start from the maximum / maximum exponent sum for dividing the exponent sums into the associated N phases.
[0163] For example, there are Figure 8 shown three constant step sizes, and the three constant step sizes correspond to three phases 802A - C. Phaser 902 can start from the maximum exponent sum and use the first constant step size to divide the first exponent sum and the second exponent sum into the first phase. Phaser 902 can use the second constant step size continuing from the first constant step size to divide the third exponent sum and the fourth exponent sum into the second phase. Phaser 902 can use the third constant step size continuing from the second constant step size to divide the fifth exponent sum and the sixth exponent sum into the third phase. It should be noted that other numbers of constant step sizes can be used, not limited to three constant step sizes. In this example, one or more exponent sums outside the range of the three phases 802A - C (or three constant step sizes) (e.g., not divided into any phase or having a relatively large difference from the maximum exponent sum) can be ignored or discarded from the MAC operation. Thus, for example, in the absence of another part of the exponent sums outside the three phases 802A - C, phaser 902 can provide a part of the exponent sums in the three phases 802A - C for shifting and accumulating the corresponding products.
[0164] In some embodiments, aligner 612 can align the products of the mantissa subsets in each phase according to the maximum exponent sum. In some other embodiments, aligner 612 can align the products of the mantissa subsets in each phase according to the local maximum (e.g., the maximum exponent sum in each individual phase), such as in conjunction with Figure 10 described.
[0165] For example, Figure 10 illustrates, according to some embodiments, the use of Figure 1Flowchart of an example process 1000 for performing MAC operations using a second phasing method with local maxima performed by circuit 100. Process 1000 may include various operations or elements similar to at least one of processes 600, 700, or 900, such as input activation latch 602, weight buffer 604, multiplier 606, adder 608, phaser 702, phaser 902, aligner 612, adder tree 614, and buffer and accumulator 616. Process 1000 may include other elements from circuit 100, not limited to those described herein.
[0166] In this case, process 1000 may include at least one phaser 702, such as described in connection with at least Figure 7 to find local maximum exponents (e.g., the sum of the maximum exponents of each individual phase local). In this case, phaser 702 may be part of phaser 902. For example, as described in connection with Figure 9 , phaser 702 may select the maximum of the exponent sums in each of the phases as the respective local maximum. The maximum exponent sum may be the local maximum of the first phase 802A. In such cases, the exponent difference may be the difference between each (local) exponent sum and the local maximum of the respective phase. Aligner 612 may use the exponent difference determined based on the local maximum to align the products.
[0167] In some embodiments, circuit 100 may include N devices, elements, or circuits (or N groups of devices) corresponding to N phases, such that alignment and addition may be performed in parallel. Circuit 100 including N devices for N phases may be described in connection with Figure 11 . For example, Figure 11 illustrates a flowchart of an example process 1100 for performing MAC operations using multiple devices (e.g., N devices) for multiple phases (e.g., N phases) performed by circuit 100 of Figure 1 . Process 1100 may include various operations or elements that are similar to but not limited to at least one of processes 600, 700, 900, or 1000 of Figure 6 , Figure 7 , Figure 9 or Figure 10 , such as input activation latch 602, weight buffer 604, multiplier 606, adder 608, phaser 610, aligner 612, adder tree 614, buffer and accumulator 616, etc. Process 900 may include other elements from circuit 100, not limited to those described herein.
[0168] For purposes of providing an example, Process 1100 can be performed for two phases, thereby including respective phasers 1102A - B, aligners 1104A - B, and adder trees 1106A - B associated with the two phases. It should be noted that N phases can include more than two phases, thereby including more elements or circuits as described herein. Each of phasers 1102A - B, aligners 1104A - B, and adder trees 1106A - B can perform features or operations similar to phaser 702, aligner 612, and adder tree 614, respectively.
[0169] Taking as an example a first phase - setting method (e.g., similar to Process 700) for aligning respective local maxima, phaser 610 can provide respective exponential subsets of N phases to phasers 1102A - B, respectively. Additionally, phaser 610 can provide respective mantissa (product) subsets corresponding to the exponential subsets of N phases to aligners 1104A - B, respectively. Phaser 1102A can determine or find the largest of the exponential sums in the first phase, which corresponds to the largest exponential sum. The largest exponential sum (e.g., a local maximum in this case) can be used to determine an exponential difference for the exponential subset. Aligner 1104A can align or shift the product of the mantissa subsets according to the exponential difference of the first phase. Adder tree 1106A can add the shifted products to produce a first sum of the first phase (among N phases) for buffer and accumulator 616.
[0170] In other examples, phaser 1102B can find the largest of the exponential sums in the second phase. The local maximum can be used to determine an exponential difference for the exponential subset. Aligner 1104B can align or shift the product of the mantissa subsets according to the exponential difference of the second phase. Adder tree 1106B can add the shifted products to produce a second sum of the second phase for buffer and accumulator 616. In response to receiving the sums (e.g., the first sum and the second sum), buffer and accumulator 616 (e.g., another adder circuit) can combine the sums of the two phases (e.g., N phases) or the sums associated therewith to produce the result of a shift - and - accumulate operation (or MAC operation). Thus, in addition to other elements discussed herein, circuit 100 can be used to reduce resource and power consumption and improve the accuracy of the result of a MAC operation during the execution of the MAC operation.
[0171] In one aspect of the present disclosure, a computing-in-memory (CIM) circuit is disclosed. The CIM circuit includes an input circuit configured to receive: (i) a plurality (N) of first inputs and (ii) N second inputs, where the first inputs are composed of at least N first exponents and N first mantissas, and the second inputs are composed of at least N second exponents and N second mantissas, and where each of the second inputs and the corresponding one of the N first inputs forms one of N input pairs. The CIM circuit includes N adder circuits, each of the N adder circuits configured to combine the corresponding first exponent and the corresponding second exponent of the corresponding one of the N input pairs to produce the corresponding one of N exponent sums. The CIM circuit includes a first selector circuit configured to select the largest one among the N exponent sums as the largest exponent sum. The CIM circuit includes a phasing circuit configured to partition at least a portion of the N exponent sums starting from the largest exponent sum into N phases, each of the N phases being associated with a respective subset of the N exponent subsets. The CIM circuit includes N subtractor circuits, each of the N subtractor circuits configured to compute the corresponding one of N exponent differences for each of the N phases, each of the N exponent differences being equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the largest exponent sum, where the N exponent differences are used for shift-and-add operations.
[0172] In some embodiments, the N phases are associated with N constant steps starting from the largest exponent sum, and where another portion of the N exponent sums outside the N phases is discarded.
[0173] In some embodiments, the phasing circuit is configured to determine a local maximum of the respective exponent subset in each of the N phases, and where each of the N exponent differences is equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the local maximum of the respective exponent subset.
[0174] In some embodiments, the computing-in-memory circuit further includes N multiplier circuits. Each of the N multiplier circuits is configured to selectively multiply the corresponding first mantissa of the corresponding input pair by the corresponding second mantissa, thereby producing the corresponding one of N mantissa products. The phasing circuit is configured to partition a portion of the N mantissa products corresponding to the portion of the N exponent sums into N phases. The respective mantissa subsets of the N mantissa products and the corresponding exponent subsets of each of the N phases are used for shift-and-add operations.
[0175] In some embodiments, the computing-in-memory circuit further includes a shifter circuit. The shifter circuit is configured to shift the respective mantissa subsets based on the corresponding N exponent differences for each of the N phases. The respective mantissa subsets include a first mantissa subset and a second mantissa subset of the N mantissa products for the first phase and the second phase, respectively, among the N phases.
[0176] In some embodiments, a phasing circuit is configured to provide a first phase and a second phase to a shifter circuit at different time periods.
[0177] In some embodiments, the in-memory arithmetic circuit further includes a first adder circuit. The first adder circuit is configured to: (i) add a shifted first subset of mantissa products for a first phase to generate a first sum; and (ii) add a shifted second subset of mantissa products for a second phase to generate a second sum.
[0178] In some embodiments, the in-memory arithmetic circuit further includes a second adder circuit. The second adder circuit is configured to combine the first sum and the second sum of N phases to generate a result of a shift-and-add operation.
[0179] In some embodiments, to divide a portion of N exponent sums into N phases, the phasing circuit is configured to: sort the N exponent sums from the largest exponent sum to the smallest exponent sum, where the smallest exponent sum is associated with the largest exponent difference among the N exponent differences, and where the largest exponent sum is associated with the smallest exponent difference among the N exponent differences; and divide the sorted portion of the N exponent sums into a first phase and a second phase respectively corresponding to a first constant step and a second constant step among N constant steps. The N constant steps have a predefined step size. The first phase includes a plurality of first exponent sums within the first constant step starting from the largest exponent sum. The second phase includes a plurality of second exponent sums within the second constant step continuing from the first constant step.
[0180] In some embodiments, the in-memory arithmetic circuit further includes a second selector circuit and a second subtractor circuit. The second selector circuit is configured to select the smallest one among the N exponent sums as the smallest exponent sum. The second subtractor circuit is configured to calculate a difference between the largest exponent sum and the smallest exponent sum to generate an exponent sum range. The phasing circuit is configured to: compare the exponent sum range with a threshold; and divide all N exponent sums into N phases based on the exponent sum range being less than the threshold; or divide a portion of the N exponent sums into N phases based on the exponent sum range being greater than or equal to the threshold.
[0181] In some embodiments, the in-memory arithmetic circuit further includes N phasing circuits, N shifter circuits, N adder circuits, and another adder circuit. The N phasing circuits are for N respective phases. Each of the N phasing circuits is used to determine the local maximum of the respective index subset of the corresponding one of the N phases. Each of the N index differences is equal to the difference between the corresponding one of the N index sums from the respective index subsets and the local maximum of the respective index subset. The phasing circuit is one of the N phasing circuits. The N shifter circuits are associated with the N respective phasing circuits, and each of the N shifter circuits is used to shift the mantissa subset based on the corresponding one of the N index differences of the corresponding one of the N phases. The N adder circuits are associated with the N respective shifter circuits, and each of the N adder circuits is used to add the shifted mantissa subsets of the corresponding one of the N phases to generate the corresponding one of the N sums. Another adder circuit is used to combine the N sums of the N phases to generate the result of the shift and accumulate operation.
[0182] In another aspect of the present disclosure, a computing-in-memory (CIM) circuit is disclosed. The CIM circuit includes an input circuit for receiving: (i) a plurality (N) of first inputs and (ii) N second inputs, where the first inputs are composed of at least N first exponents and N first mantissas, and the second inputs are composed of at least N second exponents and N second mantissas, and where each of the second inputs and the corresponding one of the N first inputs forms one of the N input pairs. The CIM circuit includes N adder circuits, and each of the N adder circuits is used to combine the corresponding first exponent and the corresponding second exponent of the corresponding one of the N input pairs to generate the corresponding one of the N exponent sums. The CIM circuit includes a first selector circuit for selecting the largest one among the N exponent sums as the maximum exponent sum. The CIM circuit includes a second selector circuit for selecting the smallest one among the N exponent sums as the minimum exponent sum. The CIM circuit includes a phasing circuit for: (i) determining an exponent sum range based on the difference between the maximum exponent sum and the minimum exponent sum, and (ii) dividing the N exponent sums into N phases based on the exponent sum range, and each of the N phases is associated with a respective index subset among the N index subsets. The CIM circuit includes N subtractor circuits, and each of the N subtractor circuits is used to calculate the corresponding one of the N index differences of each of the N phases, and each of the N index differences is equal to the difference between the corresponding one of the N index sums from the respective index subsets and the maximum exponent sum, where the N index differences are used for the shift and accumulate operation.
[0183] In some embodiments, a phasing circuit is used to determine local maxima of respective subsets of exponents in each of N phases. Each of the N exponent differences is equal to the difference between a corresponding one of the N exponent sums from respective subsets of exponents and the local maximum of the respective subset of exponents.
[0184] In some embodiments, the in-memory arithmetic circuit further includes N multiplier circuits. Each of the N multiplier circuits is used to selectively multiply a corresponding first mantissa of a corresponding input pair by a corresponding second mantissa, thereby generating a corresponding one of the N mantissa products. The phasing circuit is used to partition the N mantissa products corresponding to the N exponent sums into N phases. Respective subsets of mantissas of the N mantissa products and corresponding subsets of exponents in each of the N phases are used for shift and accumulate operations.
[0185] In some embodiments, the in-memory arithmetic circuit further includes a shifter circuit, a first adder circuit, and a second adder circuit. The shifter circuit is used to shift respective subsets of mantissas based on corresponding ones of the N exponent differences in each of the N phases, where the respective subsets of mantissas include a first subset of mantissas and a second subset of mantissas of the N mantissa products for a first phase and a second phase, respectively, in the N phases. The first adder circuit is used to: (i) add the shifted first subset of mantissas of the N mantissa products for the first phase, thereby generating a first sum; and (ii) add the shifted second subset of mantissas of the N mantissa products for the second phase, thereby generating a second sum. The second adder circuit is used to combine the first sum and the second sum of the N phases, thereby generating a result of the shift and accumulate operation.
[0186] In some embodiments, the phasing circuit is used to provide the first phase and the second phase to the shifter circuit at a plurality of different time periods. The first adder circuit is used to add the shifted first subset of mantissas of the first phase and add the shifted second subset of mantissas of the second phase at different time periods.
[0187] In some embodiments, the N phases correspond to a predefined number, or wherein the N phases correspond to N constant steps within an exponent sum range, and the N constant steps have a predefined step size.
[0188] In another aspect of the present disclosure, a method is disclosed. The method includes receiving, by a computing-in-memory (CIM) circuit, (i) a plurality (N) of first inputs and (ii) N second inputs, wherein the first inputs are composed of at least N first exponents and N first mantissas, and the second inputs are composed of at least N second exponents and N second mantissas, and wherein each of the second inputs and the corresponding one of the N first inputs form one of N input pairs. The method includes generating, by the CIM circuit, a corresponding one of N exponent sums based on combining the corresponding first exponent and the corresponding second exponent of the corresponding ones of the N input pairs. The method includes generating, by the CIM circuit, a corresponding one of N mantissa products based on selectively multiplying the corresponding first mantissa and the corresponding second mantissa of the corresponding ones of the N input pairs. The method includes selecting, by the CIM circuit, the largest one among the N exponent sums as the largest exponent sum. The method includes dividing, by the CIM circuit, at least a portion of the N exponent sums and at least one corresponding portion of the N mantissa products into N phases, each of the N phases being associated with a respective subset of the N exponent subsets and a respective subset of the N mantissa subsets. The method includes determining, by the CIM circuit, a corresponding one of N exponent differences for each of the N phases, each of the N exponent differences being equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the largest exponent sum. The method includes shifting, by the CIM circuit, the respective mantissa subsets based on the corresponding N exponent differences for each of the N phases. The method includes determining, by the CIM circuit, a corresponding one of N sums based on adding the respective shifted mantissa subsets. The method includes combining, by the CIM circuit, the N sums to produce an addition result.
[0189] In some embodiments, the method further includes selecting, by the computing-in-memory circuit, the largest one among the corresponding N exponent sums of the respective exponent subsets as respective local maxima. Each of the N exponent differences is equal to the difference between the corresponding one of the N exponent sums from the respective exponent subset and the local maximum, and the respective mantissa subsets of each of the N phases are shifted and added at respective time intervals.
[0190] In some embodiments, the method further includes: selecting, by the in-memory computing circuit, the smallest one among N exponential sums as the minimum exponential sum; determining, by the in-memory computing circuit, an exponential sum range based on a difference between the maximum exponential sum and the minimum exponential sum; comparing, by the in-memory computing circuit, the exponential sum range with a threshold; and in response to the comparison: dividing, by the in-memory computing circuit, all N exponential sums into N phases based on the exponential sum range being less than the threshold, where the N phases correspond to a predefined number, or where the N phases correspond to N constant steps within the exponential sum range; or dividing, by the in-memory computing circuit, a portion of the N exponential sums into N phases based on the exponential sum range being greater than or equal to the threshold, where the N phases are associated with N constant steps starting from the maximum exponential sum.
[0191] As used herein, the terms “about” and “approximately” generally indicate a value of a given quantity that may vary based on a particular technology node associated with the subject semiconductor device. Based on a particular technology node, the term “about” may indicate a value of a given quantity that varies, for example, within 10% to 30% of the value (e.g., +10%, ±20%, or ±30% of the value).
[0192] The foregoing has outlined features of several embodiments so that those skilled in the art may better understand various aspects of the disclosure. Those skilled in the art should appreciate that they may readily use the disclosure as a basis for designing or modifying other processes and structures for achieving the same purposes and / or achieving the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of the disclosure, and that various changes, substitutions, and alterations may be made therein without departing from the spirit and scope of the disclosure.
Claims
1. A memory in-memory computing circuit, characterized in that: include: An input circuit for receiving: (i) N first inputs and (ii) N second inputs, wherein the first inputs consist of at least N first exponents and N first mantissas, and the second inputs consist of at least N second exponents and N second mantissas, and wherein each of the second inputs and a corresponding one of the N first inputs form one of N input pairs; N adding circuits, each of the N adding circuits being configured to combine the corresponding first exponent and the corresponding second exponent of a corresponding one of the N input pairs to generate a corresponding one of the N exponential sums; a first selector circuit for selecting a maximum exponential sum among the N exponential sums as a maximum exponential sum; a phasing circuit for dividing at least a portion of the N exponential sums into N phases starting from the maximum exponential sum, each of the N phases being associated with a respective one of the N exponential subsets; as well as N subtractor circuits, each of the N subtractor circuits being configured to calculate a corresponding one of N exponential differences for each of the N phases, each of the N exponential differences being equal to a difference between a corresponding one of the N exponential sums from the respective exponent subset and the maximum exponential sum, wherein the N exponential differences are used for a shift and accumulate operation.
2. The in-memory computing circuit according to claim 1, wherein: Also includes: N multiplier circuits, each of the N multiplier circuits being configured to selectively multiply the corresponding first mantissa of the corresponding input pair by the corresponding second mantissa to generate a corresponding one of the N mantissa products, wherein the phasing circuit is used to divide at least a portion of the N mantissa products corresponding to at least the portion of the N exponential sums into the N phases, wherein a respective subset of mantissas of the N mantissa products and the corresponding subset of exponents of each of the N phases are used for the shift and accumulate operation; as well as a shifter circuit for shifting the respective mantissa subsets based on the corresponding N exponent differences for each of the N phases, the respective mantissa subsets comprising at least a first mantissa subset and a second mantissa subset of the N mantissa products for a first phase and a second phase of the N phases, respectively, The phasing circuit is used to provide the first phase and the second phase to the shifter circuit at different time periods.
3. The in-memory computing circuit according to claim 2, wherein: Also includes: a first adder circuit for: (i) adding the shifted first mantissa subset of the N mantissa products for the first phase to produce a first sum; and (ii) adding the shifted second mantissa subset of the N mantissa products for the second phase to produce a second sum; as well as A second adder circuit is used to combine the first sum and the second sum of the N phases to generate a result of the shift and accumulate operation.
4. The in-memory computing circuit according to claim 1, wherein: Wherein, in order to divide at least the portion of the N exponential sums into the N phases, the phasing circuit is used to: ordering the N exponential sums from the largest exponential sum to a smallest exponential sum, wherein the smallest exponential sum is associated with a largest exponential difference among the N exponential differences, and wherein the largest exponential sum is associated with a smallest exponential difference among the N exponential differences; as well as dividing at least the portion of the sorted N exponential sums into at least one first phase and one second phase of the N phases corresponding to a first constant step size and a second constant step size of the N constant step sizes, respectively, wherein the N constant steps have a predefined step size, wherein the first phase comprises a plurality of first exponential sums within the first constant step starting from the maximum exponential sum, and wherein the second phase comprises a plurality of second exponential sums within the second constant step continuing from the first constant step.
5. The in-memory computing circuit according to claim 1, wherein: Also includes: a second selector circuit for selecting a smallest one among the N exponential sums as a minimum exponential sum; as well as a second subtractor circuit for calculating a difference between the maximum exponential sum and the minimum exponential sum to generate an exponential sum range, The phasing circuit is used to: comparing the index and the range to a threshold; and dividing all of the N index sums into the N phases based on the index sum range being less than the threshold; or At least the portion of the N exponential sums is divided into the N phases based on the exponential sum range being greater than or equal to the threshold.
6. A memory in-memory computing circuit, characterized in that: include: An input circuit for receiving: (i) N first inputs and (ii) N second inputs, wherein the first inputs consist of at least N first exponents and N first mantissas, and the second inputs consist of at least N second exponents and N second mantissas, and wherein each of the second inputs and a corresponding one of the N first inputs form one of N input pairs; N adding circuits, each of the N adding circuits being configured to combine the corresponding first exponent and the corresponding second exponent of a corresponding one of the N input pairs to generate a corresponding one of the N exponential sums; a first selector circuit for selecting a maximum exponential sum among the N exponential sums as a maximum exponential sum; a second selector circuit for selecting a smallest one among the N exponential sums as a minimum exponential sum; a phasing circuit to: (i) determine an exponential sum range based on a difference between the maximum exponential sum and the minimum exponential sum, and (ii) divide the N exponential sums into N phases based on the exponential sum range, each of the N phases being associated with a respective one of the N exponential subsets; as well as N subtractor circuits, each of the N subtractor circuits being configured to calculate a corresponding one of N exponential differences for each of the N phases, each of the N exponential differences being equal to a difference between a corresponding one of the N exponential sums from the respective exponent subset and the maximum exponential sum, wherein the N exponential differences are used for a shift and accumulate operation.
7. The in-memory computing circuit according to claim 6, wherein: wherein the phasing circuit is used to determine a local maximum of the respective index subset in each of the N phases, and wherein each of the N index differences is equal to a difference between the corresponding one of the N index sums from the respective index subset and the local maximum of the respective index subset.
8. A method for operating a computing circuit in a memory, characterized in that: include: receiving, by an in-memory operation circuit, (i) N first inputs and (ii) N second inputs, wherein the first inputs consist of at least N first exponents and N first mantissas, and the second inputs consist of at least N second exponents and N second mantissas, and wherein each of the second inputs and a corresponding one of the N first inputs form one of N input pairs; generating, by the in-memory operation circuit, a corresponding one of the N exponential sums based on combining the corresponding first exponent and the corresponding second exponent of a corresponding one of the N input pairs; generating, by the in-memory operation circuit, a corresponding one of N mantissa products based on selectively multiplying the corresponding first mantissa of the corresponding one of the N input pairs by the corresponding second mantissa; The memory-internal operation circuit selects a maximum exponential sum from the N exponential sums as a maximum exponential sum; dividing, by the in-memory operation circuit, at least a portion of the N exponent sums and at least a corresponding portion of the N mantissa products into N phases, each of the N phases being associated with a respective subset of the N exponents and a respective subset of the N mantissas; determining, by the in-memory operation circuit, a corresponding one of the N exponential differences for each of the N phases, each of the N exponential differences being equal to a difference between a corresponding one of the N exponential sums from the respective exponential subset and the maximum exponential sum; shifting, by the in-memory operation circuit, the respective subsets of mantissas based on the corresponding N exponent differences for each of the N phases; determining, by the in-memory arithmetic circuit, a corresponding one of the N sums based on adding the shifted respective subsets of mantissas; and The N sums are combined by the in-memory operation circuit to generate an addition result.
9. The method according to claim 8, characterized in that Also includes: The in-memory operation circuit selects a maximum one among the N index sums corresponding to the respective index subset as a respective local maximum value, Wherein each of the N exponent differences is equal to a difference between the corresponding one of the N exponent sums from the respective exponent subset and the local maximum, and wherein the respective mantissa subset for each of the N phases is shifted and added at a respective time period.
10. The method according to claim 8, characterized in that Also includes: The in-memory operation circuit selects a smallest one among the N exponential sums as a minimum exponential sum; determining an exponential sum range by the in-memory operation circuit based on a difference between the maximum exponential sum and the minimum exponential sum; comparing the index and the range with a threshold by the in-memory operation circuit; and In response to this comparison: The memory operation circuit divides all the N index sums into the N phases based on the index sum range being less than the threshold, wherein the N phases correspond to a predefined number, or wherein the N phases correspond to N constant steps within the index and range; or At least the portion of the N exponential sums is divided into the N phases by the in-memory operation circuit based on the exponential sum range being greater than or equal to the threshold, wherein the N phases are associated with N constant steps starting from the maximum exponential sum.