Calculation method and calculation device

By aligning the mantissa of floating-point numbers and generating common indexes, the problem of low floating-point arithmetic operations in existing computing devices is solved, and more efficient data processing and computing capabilities are achieved.

CN119937981APending Publication Date: 2025-05-06TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Patent Information

Application Number
CN202510017620.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-24
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Floating point arithmetic operations in existing computing devices are relatively low, especially in in-memory computing systems, which affects the speed and efficiency of data processing.

Method used

By aligning the mantissa of floating point numbers, a common exponent is generated, and the aligned mantissa is stored in the memory device, used to generate the mantissa product, and the output floating point number is generated through the accumulation step.

Benefits of technology

It improves the efficiency of floating-point arithmetic operations, reduces data movement and energy consumption, and enhances the processing power of computing devices, especially in neural network systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937981A_ABST
    Figure CN119937981A_ABST
Patent Text Reader

Abstract

Some embodiments of the present invention provide a computing method comprising generating a set of products for respective pairs of a first floating point operand and a second floating point operand, where each of the first floating point operand and the second floating point operand has a respective mantissa and exponent, and generating a set of products for respective pairs of the first floating point operand and the second floating point operand. Aligning a mantissa of the first operand based on a maximum index of the first operand to generate a shared index; modifying mantissas of the first operands based on the shared index to generate respective adjusted mantissas of the first operands; generating mantissa products, each mantissa product based on a mantissa of a respective one of the second operands retrieved from the memory device and a respective one of the adjusted first mantissa; summing the mantissa products to generate a mantissa product partial sum; and combining the shared index with the product mantissa part sum. The adjusted mantissa of the first operand may be saved in the memory device and may also be retrieved from the memory device to generate a mantissa product. The embodiment of the invention further provides a calculation device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of electronic circuits, and more particularly, to computing methods and computing devices. Background Art

[0002] The present disclosure generally relates to floating-point arithmetic operations in computing devices, such as computing-in-memory or computing-in-memory (CIM) devices and application-specific integrated circuits (ASICs), and also to methods and devices used for data processing (such as multiply-accumulate (MAC) operations). Computing-in-memory or computing-in-memory systems store information in the main random access memory (RAM) of a computer and perform calculations at the memory cell level, rather than moving large amounts of data between the main RAM and data storage to perform each calculation step. Because the stored data can be accessed more quickly when it is stored in RAM, computing-in-memory enables data to be analyzed in real time. ASICs, including digital ASICs, are designed so that data processing can be optimized for specific computing needs. Improved computing performance can achieve faster returns and decisions in business and machine learning applications. Much effort has been invested in improving the performance of such computing memory systems, and more specifically, the performance of floating-point arithmetic operations in such systems. Summary of the invention

[0003] An embodiment of the present invention provides a calculation method, including: for a first plurality of floating-point numbers and a second plurality of floating-point numbers, each floating-point number has a corresponding mantissa and exponent, aligning the mantissas of the first plurality of floating-point numbers based on the maximum exponent of the first plurality of floating-point numbers to generate a first common exponent; storing the aligned first plurality of mantissas in a storage device; generating a first plurality of mantissa products, each mantissa product is based on the mantissa of a corresponding one of the second plurality of floating-point numbers and a corresponding one of the aligned first plurality of mantissas retrieved from the storage device; an accumulation step, including summing the first plurality of mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the first common exponent and the exponent of the second plurality of floating-point numbers; and combining the first product partial sum exponent and the first mantissa product partial sum to form an output floating-point number.

[0004] Another embodiment of the present invention provides a calculation method, including: for a first plurality of weight values, each weight value has its own weight mantissa and weight exponent, aligning the weight mantissas based on the maximum weight exponent of the first plurality of weight values ​​to generate a first common weight exponent; storing the aligned weight mantissas in corresponding first plurality of storage units in an artificial neural network; providing a first plurality of input activations to corresponding input terminals of a first multiplication circuit in the artificial neural network, each of the first plurality of input activations having its own input mantissa and input exponent; using the first multiplication circuit to generate a first plurality of mantissa products, each product being based on its own weight mantissa and its own input mantissa; an accumulation step, including summing the first mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the first common weight exponent and the exponents of the first plurality of input activations; and combining the first product partial sum exponent and the first mantissa product partial sum to form a first output floating-point number.

[0005] Another embodiment of the present invention provides a computing device, comprising: a memory array, comprising a plurality of memory cells, each memory cell being configured to store a corresponding mantissa of a corresponding weight value having a common exponent; a first memory, configured to store the common exponent; a first digital circuit, configured to receive a plurality of input activations, each input activation having its own mantissa and exponent, and; a multiplication circuit, configured to retrieve the mantissa of the corresponding weight value from the memory array, and to generate a product of the retrieved mantissa and the mantissa of the corresponding received input activation; a summation circuit, configured to add the products to generate a product and a mantissa, and to generate a product and an exponent based on the exponent of the received input activation and the common exponent stored in the first memory; and a second memory, having a mantissa portion configured to store the product and mantissa and an exponent portion configured to store the exponent of the product and exponent. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various aspects of the present invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, the various components are not drawn to scale. In fact, the sizes of the various components may be arbitrarily increased or reduced for clarity of discussion.

[0007] Figure 1 A method for multiply-accumulate (MAC) operations according to some embodiments is outlined.

[0008] Figure 2A Methods for processing floating point operands, such as weight values ​​used in MAC operations, prior to multiplication are outlined in accordance with some embodiments.

[0009] Figure 2B schematically illustrates floating point operands, such as weight values ​​used in MAC operations, and their corresponding storage according to some embodiments;

[0010] Figure 3A and Figure 3B Example MAC operations on floating point operands, such as weight values ​​and input activations, are schematically illustrated according to some embodiments.

[0011] Figure 4 illustrates a reduction in storage bits due to use of pre-multiplication mantissa alignment in MAC operations according to some embodiments;

[0012] Figure 5 A method of multiply-accumulate (MAC) operation according to some embodiments is summarized, wherein a MAC operation on a set of input activation-weight value pairs is divided into MAC operations on two or more subgroups, wherein the MAC operation on each subgroup is as follows: Figure 1 The MAC operations are performed as outlined in and the outputs of the MAC operations are also aligned and summed.

[0013] Figure 6 A CIM device for performing a MAC operation according to some embodiments is schematically illustrated.

[0014] Fig. 7A Methods for processing floating point operands prior to multiplication, such as for input activation in MAC operations, are outlined in accordance with some embodiments;

[0015] Figure 7B schematically illustrates floating point operands, such as input activations used in a MAC operation and their corresponding storage, according to some embodiments; and

[0016] Figure 8 A method of multiply-accumulate (MAC) operation according to some embodiments is outlined. DETAILED DESCRIPTION

[0017] The present invention provides many different embodiments or examples for realizing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present invention. Of course, these are merely examples and are not intended to limit the present invention. For example, in the following description, forming a first component above or on a second component may include an embodiment in which the first component and the second component are formed in direct contact, and may also include an embodiment in which an additional component may be formed between the first component and the second component so that the first component and the second component may not be in direct contact. In addition, the present invention may repeat reference numerals and / or characters in various examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or configurations discussed.

[0018] Furthermore, for ease of description, spatially relative terms such as "below," "beneath," "lower," "above," "upper," etc. may be used herein to describe the relationship of one element or component to another (or additional) elements or components as shown in the figures. The spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein should likewise be interpreted accordingly.

[0019] The present disclosure generally relates to floating point arithmetic operations in computing devices, such as computing-in-memory or computing-in-memory (CIM) devices and application-specific integrated circuits (ASICs), and also to methods and devices used for data processing, such as multiply-accumulate (MAC) operations. Computer artificial intelligence (AI) uses deep learning techniques, where computing systems can be organized as neural networks. For example, a neural network represents multiple interconnected processing nodes that enable data analysis. A neural network uses "weights" to perform calculations on new input data. A neural network uses multiple layers of computing nodes, where deeper layers perform calculations based on the results of calculations performed by higher layers.

[0020] The CIM circuit performs operations locally in the memory without sending data to the host processor. This reduces the amount of data transferred between the memory and the host processor, thereby achieving higher throughput and performance. The reduction in data movement also reduces the energy consumption of overall data movement within the computing device.

[0021] Alternatively, the MAC operation may be implemented in other types of systems, such as a computer system that is programmed to perform the MAC operation.

[0022] In certain embodiments disclosed in the present disclosure, a computation method includes: for a set of products of each pair of corresponding number pairs, such as a first floating-point operand and a second floating-point operand, each of which has a corresponding mantissa and exponent, in a multiply-accumulate operation, aligning the exponent of the first floating-point operand based on the maximum exponent of the first floating-point operand to generate a shared exponent; modifying the mantissa of the first floating-point operand based on the shared exponent to generate a corresponding adjusted mantissa of the first floating-point operand; generating mantissa products based on the mantissa of a corresponding one of the second floating-point operands and the corresponding one of the adjusted first mantissas, respectively; summing the mantissa products to generate a mantissa product partial sum; and combining the shared exponent and the product mantissa partial sum. The adjusted mantissa of the first floating-point operand may be stored in a storage device and retrieved from the storage device for mantissa product generation. The product mantissa partial sum may be used in a neural network system. The multiplication may be performed in a computer-in-memory macro (CIM macro), and the alignment of the mantissa of at least one of the first floating point operand and the second floating point operand may be performed offline, and the adjusted (aligned) mantissa is pre-stored in the CIM macro.

[0023] In other embodiments, a computer method includes the steps described above and further includes: prior to the multiplication step, aligning the exponent of the second floating point operand based on the maximum exponent of the second floating point operand to generate a shared exponent; modifying the mantissa of the second floating point operand based on the shared exponent to produce a corresponding adjusted mantissa of the second floating point operand. The summation of the mantissa products is then performed by directly summing the mantissas of the mantissa products without any other alignment (because the input weight products have the same exponent).

[0024] In some embodiments, a computation method includes: for a product of each pair of corresponding pairs, such as a first floating point operand and a second floating point operand, each of the floating point operands having a corresponding mantissa and exponent, in a multiply-accumulate operation, aligning the exponent of the first floating point operand based on the maximum exponent of the first floating point operand to generate a shared exponent; modifying the mantissa of the first floating point operand based on the shared exponent to generate a corresponding adjusted mantissa of the first floating point operand; generating mantissa products based on the mantissa of a corresponding one of the second floating point operands and a corresponding one of the adjusted first mantissas, respectively; modifying the mantissa products based on the shared exponent to generate corresponding adjusted mantissa products; summing the adjusted mantissa products to generate a mantissa product partial sum; and combining the shared exponent and the product mantissa partial sum. The adjusted mantissa of the first floating point operand may be stored in a storage device and retrieved from the storage device for mantissa product generation. The product mantissa partial sum may be used in a neural network system.

[0025] According to some embodiments, a device for performing the above method comprises one or more digital circuits, such as a microprocessor, a shift register, a binary multiplier and adder, and a comparator, for performing the steps of the method; and a memory device for storing the output of the digital circuit. In some embodiments, the mantissa adjustment is performed by shifting the mantissa stored in the register. In some embodiments, the multiplication is performed in a CIM macro, where the mantissa of the first floating point number (e.g., a weight value) is stored in a CIM memory array connected to a logic circuit, and the mantissa of the second floating point number (e.g., an input activation) is applied to the logic circuit, which outputs a digital signal indicating the product of the mantissas.

[0026] In a MAC operation, a set of input numbers are each multiplied by a corresponding one of a set of weight values ​​(or weights), which may be stored in a memory array. The products are then accumulated, e.g., added together, to form an output number. In certain applications, such as neural networks used in machine learning in AI, the output produced by a MAC operation can be used as a new input value in the next iteration of a MAC operation in a subsequent layer of the neural network. An example of a mathematical description of a MAC operation is shown below.

[0027]

[0028] Among them A I is the Ith input, W IJ is the weight corresponding to the I-th input and the J-th weight column. J is the MAC output of the Jth weight column, and h is the accumulated number.

[0029] In a floating point (FP) MAC operation, an FP number may be expressed as a sign, a mantissa or significand, and an exponent, where the exponent is an integer power of a base. The product of two FP numbers or factors may be represented by the product of the mantissas of the factors (product mantissa) and the sum of the exponents. The sign of the product may be determined based on whether the signs of the factors are the same. In a binary floating point (FP) MAC operation that may be implemented in a digital device such as a digital computer and / or a digital CIM circuit, each FP factor may be stored as a mantissa of bit width (number of bits), a sign (e.g., a single sign bit S (1b for negative; 0 for non-negative), the sign of the mantissa, and (-1) S In some representation schemes, binary FP numbers are normalized or adjusted so that the mantissa is greater than or equal to 1b but less than 10. b That is, the integer part of the normalized binary FP number is 1b. In some hardware implementations, the integer part of the normalized binary FP number (that is, 1 b ) is a hidden bit that is not stored, because 1 bIn some representation schemes, the product of two FP numbers or factors may be represented by the product mantissa, the exponents of the factors and the sign, which may be determined, for example, by comparing the signs of the factors or the sum of the sign bits or the least significant bit (LSB) of the sum.

[0030] To perform the accumulation portion of the MAC operation, in some conventional procedures, the product mantissas are first aligned. That is, if necessary, at least a portion of the product mantissas are modified by an appropriate order of magnitude so that the exponents of the product mantissas are all the same. For example, the product mantissas can be aligned so that all exponents are the maximum product exponents of the pre-aligned product mantissas. The aligned mantissas can then be added together (algebraic sum) to form a mantissa of the MAC output having the maximum exponent of the pre-aligned product mantissas.

[0031] In order to improve the MAC operation, according to some embodiments disclosed in the present disclosure, the weight values ​​or mantissas of the weights used in the MAC operation are adjusted by adjusting the mantissas, such as moving the bit pattern away from at least some of the mantissas, so that the weight values ​​have the same exponent, such as the maximum exponent of the weight values. The aligned mantissas of the weight values ​​are then multiplied with the mantissas of the input values ​​to form a mantissa product. The mantissa products are then aligned and, if necessary, summed to form a partial sum mantissa, which is then combined with the exponent to form a partial sum floating point output to be used for further calculation procedures. In some embodiments, the mantissas of the input values ​​are also aligned before being multiplied with the aligned mantissas of the weight values. In this case, the exponents of the mantissa products are therefore the same, and the mantissa products do not need to be aligned for summing.

[0032] In some embodiments, the weight values ​​may be divided into subgroups, and the MAC operation described above is applied to at least one subgroup together with the mantissas of the weight values ​​aligned before multiplying with the input values. In some embodiments, the MAC operation described above is applied to at least two subgroups, thereby generating at least two corresponding partial and floating point outputs of different exponents. The mantissas of the outputs are then aligned with each other before the partial and floating point outputs are summed together.

[0033] In some embodiments, the alignment mantissa of the weight values ​​is stored in a storage device such as a memory array. The stored alignment mantissa of the weight values ​​is then extracted from the storage device to be multiplied with the corresponding input value. In some embodiments, the alignment mantissa of the weight values ​​is generated "offline", that is, before runtime, for example, before the input activation is applied to the trained neural network. In some AI applications, an AI system, such as a system using an artificial neural network, is first "trained" by correlating training data with output data to iteratively determine the weight values ​​of nodes in the network. Once training is complete, the weight values ​​do not need to be changed and can be pre-stored in memory units in the network. Different input data sets can be applied to a neural network with the same set of weight values. In some embodiments, static weight values ​​can be stored in the form of alignment mantissas in at least one subgroup of weight values.

[0034] Thus, in general, according to some embodiments, a computation method includes, for a set of products, in a multiply-accumulate operation, for each of corresponding pairs of first and second floating-point operands (e.g., a weight value (or "weight") and an input value (or "input activation")), each floating-point operand having a corresponding mantissa and exponent, generating a shared exponent based on aligning the first floating-point operand with a maximum exponent of the first floating-point operand; modifying the mantissa of the first floating-point operand based on the shared exponent to generate a corresponding adjusted mantissa of the first floating-point operand; generating mantissa products, each product based on a mantissa of a corresponding one of the second floating-point operands and a mantissa of a corresponding one of the adjusted first mantissas; summing the mantissa products to generate a mantissa product partial sum; and combining the shared exponent with the product mantissa partial sum. The adjusted mantissa of the first floating-point operand may be stored in a storage device and retrieved from the storage device to generate the mantissa product. The product mantissa partial sum may be used in a neural network system. The multiplication may be performed in a compute-in-memory macro ("CIM macro"), the mantissa alignment of at least one of the first and second floating point operands may be performed off-line, and the adjusted (aligned) mantissa pre-stored in the CIM macro.

[0035] In other embodiments, a computation method includes the above steps and further includes: prior to the multiplication step, aligning the exponent of the second floating-point operand (i.e., input activation) based on the maximum exponent of the second floating-point operand to produce a shared exponent; modifying the mantissa of the second floating-point operand based on the shared exponent to produce a corresponding adjusted mantissa of the second floating-point operand. The summation of the mantissa products is then performed by directly summing the mantissas of the mantissa products without any other alignment.

[0036] In some embodiments, a computation method includes, for a set of products, in a multiply-accumulate operation, for each of corresponding pairs of first and second floating point operands (e.g., input activations and weight values, respectively), each floating point operand having a corresponding mantissa and exponent, aligning the exponent of the first floating point operand based on the maximum exponent of the first floating point operand to generate a shared exponent; modifying the mantissa of the first floating point operand based on the shared exponent to generate a corresponding adjusted mantissa of the first floating point operand; generating mantissa products, each product based on the mantissa of a corresponding one of the second floating point operands and the mantissa of a corresponding one of the adjusted first mantissas; modifying the mantissa products based on the shared exponent to generate a corresponding adjusted mantissa product; summing the adjusted mantissa products to generate a mantissa product partial sum; and combining the shared exponent with the product mantissa partial sum. The adjusted mantissa of the first floating point operand can be stored in a storage device and retrieved from the storage device to generate the mantissa product. The product mantissa partial sum can be used in a neural network system.

[0037] According to some embodiments, a device for performing the above method includes one or more digital circuits, such as a microprocessor, a shift register, a binary multiplier and adder, and a comparator, for performing the steps of the method; and a memory device for storing the output of the digital circuit. In some embodiments, the mantissa adjustment is performed by shifting the mantissa stored in the register. In some embodiments, the multiplication is performed in a CIM macro, where the mantissa of a first floating point number (e.g., a weight value) is stored in a CIM memory array connected to a logic circuit, and the mantissa of a second floating point number (e.g., an input activation) is applied to the logic circuit, which outputs a digital signal indicating the product of the mantissas.

[0038] The specific embodiments are described in more detail below with reference to the accompanying drawings. In one example, Figure 1 As outlined in step 101, the weight mantissa W of the weight value W[n] is M The set of [n] and the corresponding input mantissa XIN of the input activation XIN[n] M [n], and have index W respectively E [n] and XIN E [n], WM[n] are aligned to each other based on the difference between the indices WE[n]. In some embodiments, the maximum index (W E ) MAX and W E The difference between [n] ΔW E [n], that is, based on ΔW E [n]=(W E ) MAX -W E [n] to adjust W M [n] tail number W M[n] multiplied by the base number (i.e. 2) (ΔW E [n]) to the power of the last digit, so that after modification, all weights in the set have the same maximum exponent. The last digit is multiplied by the base (ΔW E [n]) can be raised to the power of, for example, a shift register by shifting the mantissa right by ΔW. E [n] bits to achieve this. That is, the mantissa is multiplied by 2^ΔW E [n], thus effectively increasing ΔW exponentially E [n] and becomes the maximum exponent. Then, the weight tail number is the weight tail number after alignment.

[0039] Similarly, in operation 103, the mantissa XIN is input. M [n] Based on index XIN E In some embodiments, XIN M [n] According to the maximum index (XIN E ) MAX with XIN E The difference between [n] ΔXIN E [n], that is, based on ΔXIN E [n]=(XIN E ) MAX -XIN E [n] to adjust. If you input the mantissa, it is the mantissa after alignment.

[0040] Next, in operation 105, each point alignment weight mantissa is multiplied by the corresponding aligned input mantissa to generate a mantissa product PD M [n] = W M [n]×XIN M [n]. In operation 107, the mantissa products are then summed or accumulated to produce the product and mantissa PS M =∑(PD M [n]). The product sum in this example is an algebraic sum, i.e., the sum of the product mantissas, where the sign of the product mantissa is based on the sign of the corresponding weight mantissa and the input mantissa. In operation 109, the product sum mantissa PS M Then multiply the product and the exponent of PS E The combined method is used to generate a floating point output, which may be, for example, a partial sum, as part of an input activation of a deeper layer in the artificial neural network, such as a hidden layer. In this example, "combined" means providing the product sum mantissa PSM and the product sum exponent PS in a computing system, such as an artificial neural network, in a manner that can be used by the system for subsequent operations. E Both. For example, PS M and PS Ecan be combined to form a floating point number in FP16 format, that is, a 16-bit number consisting of a single sign bit PSS followed by a 5-bit exponent PS E , followed by a 10-digit mantissa PS M In this example, since all weight indices are (W E ) MAX And all input indices are (XIN E ) MAX , so the product and exponent PS E The product is the same for all weighted inputs and is (W E ) MAX +(XIN E ) MAX or (W E +XIN E ) MAX Therefore, for operation 107 of the accumulation step, it is not necessary to perform mantissa alignment.

[0041] In some embodiments, the alignment of the mantissas of at least one of the sets of floating point numbers (such as weight values) is performed "offline", i.e., before runtime, e.g., before input activations are applied to a trained neural network, while the alignment of the mantissas of another of the sets of floating point numbers (such as input activations) is performed during runtime. For example, in certain artificial intelligence (AI) and machine learning (ML) applications, more specifically deep learning applications, the model is implemented in an artificial neural network, where MAC operations are performed in successive layers of nodes, while weight values ​​are stored in the nodes, and each layer generates input activations for the next deeper layer. During the training phase of the ML model, a set of training data is propagated through the layers of the neural network, and the weight values ​​are iteratively adjusted to improve the decision-making ability of the model. Once the model is trained, the weight values ​​can be determined and can be stored in a storage device in the neural network. Independent of the data input, the trained weight values, i.e., those used in the trained neural network, remain fixed. Therefore, in some embodiments, the alignment of the mantissas of the trained weights is performed offline and pre-stored in a storage device in the neural network.

[0042] In some embodiments, Figure 2A As shown, in operation 201, the training weight W E The maximum index W of [n1:n2] E-MAX For example, it is determined by a comparator or a microprocessor. In operation 203, the maximum index W E-MAX To align the trained weights W M[n1:n2] In operation 205, the mantissa and W are aligned. E-MAXThen store or program it into the memory device of the neural network. Programming the aligned training weights into a memory device, such as a CIM memory array or a CIM macro, has the advantage of reducing the number of data transfers between the computational unit and the memory outside the macro. Because all weight values ​​for n=n1 to n2 share the same index W E-MAX , so only W E-MAX Need to be stored, and only stored once in the memory. Figure 2B As shown, the trained and aligned weight values ​​can be stored as the corresponding sign bit 211i, the corresponding aligned mantissa 215i and the shared exponent 213, that is, WE-MAX Since only the shared index is stored, storage space can be saved.

[0043] like Figure 2A and Figure 2B As shown, in some embodiments, the mantissa alignment is not performed for all weight values ​​in a MAC operation (such as a MAC operation of an entire neural network layer) to obtain a single shared index. The mantissa alignment can be performed for a subset of weight values ​​in the MAC operation (i.e., i=n1 to n2). In addition, in some embodiments, the mantissa alignment can be performed for multiple subsets of weight values ​​in the MAC operation to obtain a shared index for each subset. The weight values ​​can be different from each other.

[0044] Figure 3A and Figure 3B An example of weight alignment in operation 101 and input alignment in operation 103 of a MAC operation involving two weight values ​​W[i] and two input activations XIN[i], i=0, 1 is shown. The output of the MAC operation in this example is W[0]×XIN[0]+W[1]×XIN[1]. Initially, the weight values ​​and input activations are stored in a memory device, such as a register, in 16-bit floating point or FP16 format: each number is represented by a single sign bit (S) 311 i , 5-digit exponent (E) 313 i and 10-digit mantissa (M) 315 i . Note that each mantissa also contains a hidden bit 1 b is the most significant bit (MSB). In this example, the maximum weight index is 22, as indicated by the labels (1) and (2), respectively. d =10110 b , whichever is greater of the two weight indices 22d and 20d, and the maximum input index is 20 d =10011b, i.e. two weight indexes 20 d and 19 d Therefore, in order to make the weight index of the two weight values ​​the maximum index 10110 b , the tail number W of the weight value with a smaller initial exponentM [1] Shift right by two bits, which is equivalent to multiplying by 2^ΔWE[n] or 2 2 Similarly, the input index that activates both inputs is the maximum index 10011 b , the mantissa XINM[1] of the input activation with the smaller initial exponent is shifted right by one bit, equivalently multiplied by 2^ΔXINE[n] or 2 1 The result, as shown at (3), is a floating point number with shared weights and input exponents, respectively, and an aligned mantissa, where these mantissas, now shifted to the right, contain the previously hidden bits but the previously least significant bits are truncated, so that any 1 in the truncated bits is included. b In some cases, some data is missing.

[0045] Next, as indicated by label (4), the aligned weights and shared exponents are stored. That is, W[0] and W[1] are each stored as a single sign bit (S) 321i and an 11-bit mantissa (M) 325i, but only a single shared exponent 323 is stored. Note that the aligned mantissa 325i does not have any bits hidden in this example: the hidden bits of the weight values ​​initially stored are now stored because the MSB of any shifted mantissa is now 0, so it can no longer be assumed that the MSBs of all mantissas are 1b. Therefore, the MSBs 325-ai of the weight mantissas must be stored, spending an extra bit or extension 325-bi to store the aligned mantissa. The extension 325-bi is one bit in this example and holds the data in the unshifted mantissa, but can be other lengths. For example, the extension can be two or three bits long to reduce data loss in the shifted mantissa due to truncation.

[0046] Using shared indexes can result in storage space savings. For example, in Figure 3A In the example illustrated by mark (4), the total number of bits used to store the pre-aligned weight value is 16×2=32; the total number of bits used to store the aligned weight value is 17×2-5=29, thereby saving three bits.

[0047] After weight alignment, as in Figure 3B As shown in the figure with marks (5) and (6), the weighted tail number after alignment is multiplied by the corresponding input tail number after alignment (W M [0]×XIN M [0] and W M [1]×XIN M [1]) is performed to generate the corresponding mantissa product PD M [n] The result of each multiplication is truncated to an 11-bit product in this example. In addition, the weight value and the aligned exponent of the input activation are added together (taking into account any exponent offset, 15 in this example), as shown by label (7). Note that in this example,M [1]×XIN M [1] In this case, because XIN M The right shift of [1] results in the unshifted XIN M [1] The 1b at the LSB of the 1b is missing, so the subsequent multiplication step reduces one addition step compared to the multiplication without mantissa alignment. Therefore, the computational efficiency is improved at the expense of data loss. By properly selecting computational parameters, such as the number of extension bits and the size of the FP number subset aligned for pre-multiplication, an optimal or acceptable compromise between computational efficiency and accuracy can be achieved.

[0048] The multiplication between the weight value and the corresponding input activation can be performed in a multiplication circuit, which can be any circuit capable of multiplying two digital numbers. For example, U.S. Patent Application No. 17 / 558,105, published as U.S. Patent Application Publication No. 2022 / 0269483A1, and U.S. Patent Application No. 17 / 387,598, published as U.S. Patent Application Publication No. 2022 / 0244916A1, disclose multiplication circuits for use in CIM devices, both of which are jointly assigned to this application and incorporated herein by reference. In some embodiments, the multiplication circuit includes: a memory array for storing a set of FP numbers, such as weight values; the multiplication circuit also includes a logic circuit, connected to the memory array and for receiving another set of FP numbers, such as input values, and outputting a signal, the output signal indicating the product of the stored number and the corresponding input number based on the stored corresponding number and the input number, respectively.

[0049] Next, as shown at (8), the mantissa products are accumulated or added together to produce the product and mantissa PS M (W M [0]×XIN M [0]+W M [1]×XIN M [1]). Because the weights and input activations are post-aligned, the sum of the products and exponents, i.e. the sum of the exponents, is the same for all products. Therefore, accumulation does not involve any mantissa alignment or shifting.

[0050] Finally, as shown in (9), the product and the mantissa PS M and the product and exponent PS E Assembled in memory as floating point numbers (FP16 in this example) for use in other operations in the AI ​​program. In this example, (162.25×49.25+33.0×18.046875) d The final result is 6240 d, where the error from the exact answer 6243.058594 is 3.058594. The error is the same as the error produced by the MAC operation without premultiplication alignment.

[0051] As described above, due to the use of common exponents in weight values ​​and input activations, a savings in storage space may be obtained. The amount of the savings depends on various factors, including the size of the group of weight values ​​and input activations that share the corresponding exponent and the number of bits in the mantissa extension. For example, the storage bit reduction may be expressed as the ratio between the number of bits with pre-multiplication mantissa alignment ("after") and the number of bits without pre-multiplication mantissa alignment ("before"):

[0052]

[0053] in,

[0054] N GP = the weight value or number of input activations grouped to share the index,

[0055] FP bit = number of digits in floating point,

[0056] EXT bit = the number of digits to extend the mantissa, and

[0057] EXP bit = number of exponent digits.

[0058] Equation (1) can be reconfigured to give the following:

[0059]

[0060] Figure 4 The storage bit reduction vs. NGP curves for the examples of FP16 and BF16 floating point numbers in illustrate the dependence of storage bit reduction on group size NGP. As is evident from the curves, the storage used decreases as the group size increases and approaches a lower limit as the group size becomes extremely large. Figure 4 In the example in , the storage bit reduction is close to 0.75 for FP16 and close to 0.5625 for BP16.

[0061] In some embodiments, Figure 5 In the example shown, the weight values ​​are divided into subgroups, and the above MAC operation is used for at least one subgroup, where the mantissas of the weight values ​​are aligned before being multiplied with the input activations. Figure 5In the example shown, the input activations are also divided into subgroups corresponding to the weight value subgroups, and the above-mentioned MAC operation (aligning the mantissa of the input activations before multiplication) is used at least for the subgroup corresponding to the weight value subgroup on which the pre-multiplication alignment is performed. In some embodiments, the above-mentioned MAC operation is applied to at least two subgroups of weight values ​​and two corresponding subgroups of input activations, thereby generating at least two partial and floating-point outputs of different exponents, one partial and floating-point output for each subgroup. Then, before accumulation, the mantissas of the partial and outputs are aligned with each other in a process similar to the alignment of weight values ​​and input activations.

[0062] exist Figure 5 In the example shown, for each subset of weight values ​​and input activations, the weight alignment step 501, the input alignment step 503, the multiplication step 505 and the accumulation step 507a are performed in parallel. Figure 1 The corresponding steps of the weight alignment step 101, input alignment step 103, multiplication step 105 and accumulation step 107 in are the same, except that Figure 5 Instead of generating the product and mantissa PS for the entire set of weight values ​​and the corresponding input activations, the process shown in M Instead, partial products and mantissas pPS are generated for each subset of weight values ​​and the corresponding output activations in the partial accumulation step 507a. M Then, the pPSMs are aligned to each other in step 507b, following a procedure similar to the alignment steps 501, 503 of the weight values ​​and input activations. The alignment step 507b also results in the maximum index of all subsets; thus, the maximum index is the maximum index of the entire set (W E +XIN E ) MAX Then, in a process similar to the partial accumulation step 507a for each subset, the aligned partial product sum mantissa pPS is summed. M Perform accumulation step 507c to form the total product sum mantissa PS M Finally, in step 509, combine (W E +XIN E ) MAX and pPSM, with Figure 1 The floating point output is generated in the same manner as step 109 in FIG.

[0063] The above MAC operation can be performed in any suitable computing device. Figure 6600 used in some embodiments is shown in FIG. The computing device 600 may be an on-chip device and includes a shared memory 601 for storing input data (including input activations and other data), a shared output memory 603 for storing output data (including the output of MAC operations), and various processing elements 610. Each processing element 610 in this example includes an activation memory 611, which can receive and store input activations from the input memory 601, and align the stored input activations as described above. In this example, each processing element 610 also includes a CIM macro 613, which includes a CIM memory array 615 for storing weight values, and in some embodiments, the weight values ​​include aligned weight mantissas with shared exponents generated offline. In this example, the CIM macro 613 also includes an arithmetic circuit 617, which may be, for example, a logic circuit 617, which is connected to the CIM memory array 615 and configured to receive the stored aligned weight mantissas, and is connected to the activation memory 611 and configured to receive input activations and output signals, each of which represents the product of a corresponding weight mantissa and an input mantissa. Arithmetic circuit 617 may also include circuits for performing other processing, such as accumulation and combination, as described above. In this example, each processing element 610 also includes an output memory 619, which is connected to CIM macro 613 and is configured to receive outputs, such as product sums, from CIM macro 613 and output to shared output memory 603. In this example, each processing element 610 also includes a processor, such as a microprocessor, which is programmed to perform various computing tasks, such as controlling the operation of other components of processing element 610 and / or performing certain steps of MAC operations, such as alignment, accumulation, and combination. In this example, each processing element 610 also includes a router 623 configured to manage data traffic.

[0064] As shown in some of the examples above, in some embodiments, both weight values ​​and input activations can be aligned prior to multiplication in a MAC operation. In some embodiments, while weight values ​​are aligned offline in a pre-stored memory, such as a memory array in an AI system based on a trained model, input activations can be aligned at runtime. For example, in a multi-layer deep learning neural network, the output of each layer becomes the input activation of the next deeper layer. The newly generated input activations can be aligned at runtime and then multiplied by the weight values ​​stored in the next layer. In some embodiments, the alignment process for input activations is similar to the alignment process for weight values. Fig. 7A In the example shown, in step 701, the activation XIN E[n1:n2] The maximum index XIN E MAX For example, it is determined by a comparator or a microprocessor. As described above, in step 703, the maximum index XIN E-MAXUsed to align input activation XIN M[n1:n2] Then, in step 705, the aligned mantissa and XIN E MAX is transferred to a storage device, such as an activation memory, for subsequent multiplication. Because multiple input activations of n = n1 to n2 share the same exponent XIN E-MAX , so only XIN needs to be stored in memory E MAX , and is stored only once. Figure 7B As shown, the aligned input activations can be stored as corresponding sign bits 711 i , the corresponding alignment tail 715 i and shared index 713, or XIN E MAX . Because only the shared index is stored, storage savings are achieved.

[0065] In some embodiments, only the weight values ​​are aligned, possibly offline and stored in a memory array of a computing device running on the training model, before being multiplied with the input activations. Figure 8 In the example process shown, Figure 5 In a similar process to the alignment step 501 in FIG. 8 , in step 801, the weight values ​​of the weight value subsets are aligned. No input activation alignment is performed on the corresponding input activation subsets. Next, Figure 5 In a similar manner to the multiplication step 505 in , in step 805 , the aligned weight value is multiplied with the input activation to generate the corresponding mantissa product PD M [n] Next, in step 806, based on the maximum index and (W E +XIN E ) MAX PDM[n] is aligned to generate the aligned mantissa product in a similar manner to the weight value alignment step 801. Next, in the partial accumulation step 807a, similar to Figure 5 The partial accumulation step 507a in the above example generates partial products and mantissas pPS for each subset of weight values ​​and the corresponding input activations. M Then, in step 807b, the pPSMs are aligned to each other in a process similar to the alignment step 507b. The alignment step 807b also results in the maximum index of all subsets; therefore, the maximum index is the maximum index of the entire set (W E +XIN E ) MAX Then, in step 807c, the aligned partial products and mantissas pPS are multiplied in a process similar to the partial accumulation step 507c for each subset. MThe sum is accumulated to form the total product and the mantissa PSM. Finally, in step 809, the combination (W E +XIN E ) MAX and pPS M , with Figure 5 The floating point output is generated in the same manner as step 509.

[0066] Due to the enhanced bit-wise sparsity of weight values ​​and input activations after alignment before multiplication in the MAC operation, certain examples described in this disclosure can save energy because the number of 1s in the mantissa of the shift b The reduction in number reduces the number of operations in the multiplication step. Among other things, the use of shared exponents for aligned floating-point weight values ​​and input activations saves storage space. Thus, the efficiency of the computation process can be improved without losing accuracy.

[0067] In summary, according to some embodiments, a computing method includes: for a first group of floating-point numbers and a second group of floating-point numbers, each of which has a respective mantissa and exponent, aligning the mantissas of the first group of floating-point numbers based on the maximum exponent of the first group of floating-point numbers to generate a first common exponent; storing the first group of aligned mantissas in a storage device; generating a first group of mantissa products, each product based on the mantissa of a corresponding one of the second group of floating-point numbers and the mantissa of a corresponding one of the aligned first mantissas retrieved from the storage device; an accumulation step, including summing the first mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the first common exponent and the exponents of the second plurality of floating-point numbers; and combining the first product partial sum exponent and the first mantissa product partial sum to form an output floating-point number.

[0068] In some embodiments, the computation method further includes generating a second plurality of mantissa products, each mantissa product being based on a mantissa of a corresponding one of a third plurality of floating point numbers and a corresponding one of the aligned first plurality of mantissas retrieved from the memory device.

[0069] In some embodiments, aligning the mantissas of the first plurality of floating point numbers includes modifying the mantissas of the first plurality of floating point numbers based on the first public exponent to generate a first plurality of corresponding adjusted mantissas; and generating a first plurality of mantissa products includes generating a first plurality of mantissa products, each mantissa product being based on the mantissa of a corresponding one of the second plurality of floating point numbers and a corresponding one of the aligned first plurality of mantissas retrieved from the memory device.

[0070] In some embodiments, the calculation method also includes: aligning the mantissas of the second plurality of floating-point numbers based on the maximum exponent of the second plurality of floating-point numbers to generate a second public exponent, wherein the step of generating a first product part and an exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers includes generating the first product part and an exponent based on the first public exponent and the second public exponent.

[0071] In some embodiments, the calculation method also includes: storing the first public exponent in a first memory, wherein generating the first product part and the exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers includes generating the first product part and the exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers stored in the first memory.

[0072] In some embodiments, the calculation method further includes: for a third plurality of floating-point numbers and a fourth plurality of floating-point numbers, each of which has a respective mantissa and exponent, aligning the mantissas of the third plurality of floating-point numbers based on the maximum exponent of the third plurality of floating-point numbers to generate a second common exponent; storing the aligned third plurality of mantissas in a storage device; and generating a second plurality of mantissa products, each mantissa product being based on the mantissa of a corresponding one of the fourth plurality of floating-point numbers retrieved from the storage device and a corresponding mantissa of the aligned third plurality of mantissas; the accumulation step further includes: multiplying the second plurality of mantissa products summing to generate a second mantissa product partial sum, and generating a second mantissa product partial sum exponent based on the second public exponent and the exponent of the fourth plurality of floating-point numbers; aligning the mantissas of the first mantissa product partial sum and the second mantissa product partial sum based on the maximum exponent of the first product partial sum and the second product partial sum to generate a public part sum exponent; and adding the aligned mantissas of the first mantissa product partial sum and the second mantissa product partial sum to generate a mantissa product sum; wherein the combining step includes combining the public part sum exponent and the product mantissa sum to form an output floating-point number.

[0073] In some embodiments, the calculation method also includes: aligning the mantissas of the second plurality of floating-point numbers based on the maximum exponent of the second plurality of floating-point numbers to generate a third public exponent, wherein the step of generating the first product part and the exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers includes generating the first product part and the exponent based on the first public exponent and the third public exponent.

[0074] In some embodiments, the calculation method also includes: aligning the mantissas of the second plurality of floating-point numbers based on the maximum exponent of the second plurality of floating-point numbers to generate a second public exponent, wherein the step of generating the first product part and the exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers includes generating the first product part and the exponent based on the first public exponent and the second public exponent: and storing the second public exponent in a second memory, wherein generating the first product part and the exponent based on the first public exponent and the exponents of the second plurality of floating-point numbers includes generating the first product part and the exponent based on the first public exponent stored in the first memory and the second public exponent stored in the second memory.

[0075] According to a further embodiment, a calculation method includes: for a first group of weight values, each weight value has a respective weight mantissa and weight exponent, aligning the weight mantissas based on the maximum weight exponent of the first group of weights to generate a first common weight exponent; storing the aligned weight mantissas in corresponding first group of storage units in an artificial neural network; providing a first group of input activations to respective input terminals of a first multiplication circuit in the artificial neural network, each of the first group of input triggers having a respective input mantissa and input exponent; using the first multiplication circuit to generate a first group of mantissa products, each product being based on a respective weight mantissa and a respective input mantissa; an accumulation step, including summing the first mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the common weight exponent and the exponent of the first group of input activations; and combining the first product partial sum exponent and the first mantissa product partial sum to form a first output floating point number.

[0076] In some embodiments, the calculation method also includes: storing the first common weight index in a first memory, wherein generating the first product part and index based on the first common weight index and the first plurality of input activations includes generating the first product part and index based on the first common index stored in the first memory and the exponents of the first plurality of input activations.

[0077] In some embodiments, the calculation method also includes: aligning the input mantissas of the first plurality of input activations based on the maximum input exponent of the first plurality of input activations to generate a common input exponent, wherein generating the first product part and exponent based on the first common weight exponent and the input exponent of the first plurality of input activations includes generating the first product part and exponent based on the first common weight exponent and the common input exponent.

[0078] In some embodiments, the computing method further includes: providing a second plurality of input activations to corresponding input terminals of the first multiplication circuit in the artificial neural network, each of the second plurality of output activations having a respective input mantissa and input exponent; and generating a second plurality of mantissa products using the first multiplication circuit, each product being based on a corresponding weight mantissa and a corresponding input mantissa of a corresponding one of the second plurality of input activations.

[0079] In some embodiments, the calculation method also includes: for a second plurality of weight values, each weight value has its own weight mantissa and weight exponent, aligning the weight mantissa based on the maximum weight exponent of the first plurality of weight values ​​to generate a second common weight index; storing the aligned weight mantissas of the second plurality of weight values ​​in corresponding second plurality of storage units in the artificial neural network; providing a second plurality of input activations to corresponding input terminals of a second multiplication circuit in the artificial neural network, each of the second plurality of output activations having its own input mantissa and input exponent, and each of the second plurality of input activations being the first output floating point number; using the second multiplication circuit to generate a second plurality of mantissa products, each product being based on the corresponding aligned weight mantissa of the second weight value and the corresponding input mantissa of the second plurality of input activations.

[0080] In some embodiments, aligning the weight mantissas based on the maximum weight exponent of the first plurality of weight values ​​to generate the first common weight exponent includes: aligning a first subset of the weight mantissas based on the maximum weight exponent of a corresponding first subset of the first plurality of weight values ​​to generate the first common weight exponent; aligning a second subset of the weight mantissas based on the maximum weight exponent of a corresponding second subset of the first plurality of weight values ​​to generate a second common weight exponent; and generating a first plurality of mantissa products and a second plurality of mantissa products using the first multiplication circuit, each product of the first plurality of mantissa products being based on a corresponding weight mantissa of a first subset of the first weight values ​​and a corresponding input mantissa, and each product of the second plurality of mantissa products being based on a corresponding weight mantissa of a second subset of the second weight values ​​and a corresponding output mantissa; wherein the accumulation step includes: adding the first plurality of mantissa products to generate the first mantissa product partial sum, and adding the second plurality of mantissa products to generate a second mantissa product partial sum; and adding the first mantissa product partial sum and the second mantissa product partial sum to generate a mantissa product sum.

[0081] In some embodiments, summing the first mantissa product portion sum and the second mantissa product portion sum includes aligning the first mantissa product portion sum and the second mantissa product portion sum based on a maximum exponent of the first product portion sum and the second product portion sum.

[0082] According to a further embodiment, a computing device includes: a memory array including a set of memory cells, each memory cell being configured to store a corresponding mantissa of a corresponding weight value having a common exponent; a first memory configured to store the common exponent; a first digital circuit configured to receive a set of input activations, each input activation having a respective mantissa and exponent, and a multiplication circuit configured to retrieve the mantissa of the respective weight value from the memory array and generate a product of the retrieved mantissa and the mantissa of the respective received input activation; a summation circuit configured to add the products to generate a product and a mantissa, and to generate the product and the exponent based on the exponents of the received input activations and the common exponent stored in the first memory; and a second memory having a mantissa portion and an exponent portion, the mantissa portion being configured to store the product sum mantissa, and the exponent portion being configured to store the exponent of the product exponent.

[0083] In some embodiments, the first digital circuit is also configured to adjust the mantissa of the received input activation so that the received output activation has a common exponent; the computing device also includes a third memory, which is configured to store the common exponent of the received input activation; and the summation circuit is configured to add the products to generate a product and a mantissa, and generate the product and exponent based on the common exponent of the received input activation stored in the third memory and the common exponent of the weight value stored in the first memory.

[0084] In some embodiments, the computing device also includes: a second digital circuit configured to receive the product from the multiplication circuit and adjust the mantissa of the product so that the product has a common exponent, wherein the summation circuit is configured to add the adjusted mantissas of the product to generate a product and mantissa, and generate the product and exponent based on the common exponent of the product and the common exponent of the weight value stored in the first memory.

[0085] In some embodiments, the computing device also includes: a second digital circuit configured to receive the product from the multiplication circuit and adjust the mantissa of the product so that the product has a public exponent, wherein the summation circuit is configured to add the adjusted mantissas of the product to generate a product and mantissa, and generate the product and exponent based on the public exponent of the product and the public exponent of the weight value stored in the first memory.

[0086] In some embodiments, the memory array is configured to maintain mantissas of respective weight values; the first digital circuit is configured to receive a first plurality of input activations and a second plurality of input activations, each of the first input activations having a respective mantissa and exponent, and each of the second input activations having a respective mantissa and exponent; and the multiplication circuit is configured to retrieve the mantissas corresponding to the weight values ​​maintained from the memory array and generate: a first plurality of products of the retrieved mantissas and the received mantissas of the corresponding first plurality of input activations; and a second plurality of products of the retrieved mantissas and the received mantissas of the corresponding second plurality of input activations.

[0087] The foregoing summarizes the features of several embodiments so that those skilled in the art can better understand aspects of the present disclosure. Those skilled in the art should understand that they can easily use the present disclosure as a basis for designing or modifying other processes and structures to achieve the same purposes and / or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they can be subjected to various changes, substitutions and modifications without departing from the spirit and scope of the present disclosure.

Claims

1. A calculation method, comprising: For a first plurality of floating point numbers and a second plurality of floating point numbers, each floating point number having a corresponding mantissa and exponent, aligning the mantissas of the first plurality of floating point numbers based on a maximum exponent of the first plurality of floating point numbers to generate a first common exponent; storing the aligned first plurality of mantissas in a storage device; generating a first plurality of mantissa products, each mantissa product being based on a mantissa of a corresponding one of the second plurality of floating point numbers and a corresponding one of the first plurality of aligned mantissas retrieved from the memory device; an accumulating step, including summing the first plurality of mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the first public exponent and the exponents of the second plurality of floating point numbers; as well as The first product portion and exponent and the first mantissa product portion and are combined to form an output floating point number.

2. The calculation method according to claim 1, further comprising: A second plurality of mantissa products is generated, each mantissa product being based on a mantissa of a corresponding one of a third plurality of floating point numbers and a corresponding one of the first plurality of aligned mantissas retrieved from the memory device.

3. The calculation method according to claim 1, wherein: aligning the mantissas of the first plurality of floating point numbers comprises modifying the mantissas of the first plurality of floating point numbers based on the first common exponent to generate a first plurality of corresponding adjusted mantissas; as well as Generating a first plurality of mantissa products includes generating a first plurality of mantissa products, each mantissa product being based on a mantissa of a corresponding one of the second plurality of floating point numbers and a corresponding one of the first plurality of aligned mantissas retrieved from the memory device.

4. A calculation method comprising: For a first plurality of weight values, each weight value having a respective weight mantissa and a weight exponent, aligning the weight mantissas based on a maximum weight exponent of the first plurality of weight values ​​to generate a first common weight exponent; storing the aligned weight mantissas in corresponding first plurality of storage units in the artificial neural network; providing a first plurality of input activations to corresponding inputs of a first multiplication circuit in the artificial neural network, each of the first plurality of input activations having a respective input mantissa and input exponent; generating a first plurality of mantissa products using the first multiplication circuit, each product being based on a respective weight mantissa and a respective input mantissa; an accumulating step including summing the first mantissa products to generate a first mantissa product partial sum, and generating a first product partial sum exponent based on the first common weight exponent and the exponents of the first plurality of input activations; as well as The first product portion sum exponent and the first mantissa product portion sum are combined to form a first output floating point number.

5. The calculation method according to claim 4, further comprising: The first common weight index is stored in a first memory, wherein generating the first product portion and index based on the first common weight index and the first plurality of input activations includes generating the first product portion and index based on the first common index stored in the first memory and the exponents of the first plurality of input activations.

6. The calculation method according to claim 4, further comprising: aligning input mantissas of the first plurality of input activations based on a maximum input exponent of the first plurality of input activations to generate a common input exponent, Wherein, generating the first product portion and index based on the first common weight index and the input index of the first plurality of input activations includes generating the first product portion and index based on the first common weight index and the common input index.

7. The calculation method according to claim 4, further comprising: providing a second plurality of input activations to corresponding inputs of the first multiplication circuit in the artificial neural network, each of the second plurality of output activations having a respective input mantissa and input exponent; A second plurality of mantissa products are generated using the first multiplication circuit, each product being based on a respective weight mantissa and a respective input mantissa of a respective one of the second plurality of input activations.

8. A computing device comprising: a memory array including a plurality of memory cells, each memory cell configured to store a respective mantissa of a respective weight value having a common exponent; A first memory configured to store the public index; a first digital circuit configured to receive a plurality of input activations, each input activation having a respective mantissa and exponent, and; a multiplication circuit configured to retrieve a mantissa of a corresponding weight value from the memory array and generate a product of the retrieved mantissa and a mantissa of the corresponding received input activation; a summing circuit configured to add the products to generate a product and a mantissa, and to generate a product and an exponent based on the exponent activated by the received input and the public exponent stored in the first memory; as well as A second memory has a mantissa portion configured to store the product and the mantissa and an exponent portion configured to store the exponent of the product and the exponent.

9. The computing device of claim 8, wherein: The first digital circuit is further configured to adjust the mantissa of the received input activations so that the received output activations have a common exponent; The computing device further comprises a third memory configured to store the public index activated by the received input; as well as The summing circuit is configured to add the products to generate a product sum mantissa, and the product sum exponent is generated based on the common exponents of the received input activations stored in the third memory and the common exponents of the weight values ​​stored in the first memory.

10. The computing device of claim 8, further comprising: a second digital circuit configured to receive the product from the multiplication circuit and adjust the mantissa of the product so that the product has a common exponent, The summing circuit is configured to add the adjusted mantissas of the products to generate a product sum mantissa, and to generate the product sum exponent based on the public exponent of the products and the public exponent of the weight value stored in the first memory.

Citation Information

Patent Citations

  • Compute in memory

    US20220244916A1

  • Compute in memory accumulator

    US20220269483A1

Cited By

  • Floating point data text output dynamic optimization method and system based on multiprocessor system

    CN120523759A

  • Method and system for dynamic optimization of floating point data text output based on multiprocessor system

    CN120523759B