Calculation device and method in memory

By selecting a subset based on the floating-point exponent and generating the product portion sum using multiplication circuits, the challenge of improving floating-point arithmetic operations in memory is solved, and more efficient computing performance is achieved, especially suitable for artificial intelligence and machine learning applications.

CN119937980APending Publication Date: 2025-05-06TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Patent Information

Application Number
CN202510010390.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art challenges in improving the performance of floating point arithmetic operations in memory, especially in artificial intelligence and machine learning applications, where more efficient computing methods are needed to accelerate reporting and decision-making processes.

Method used

The performance of floating-point arithmetic operations is improved by selecting a subset based on the floating-point exponent and generating products using a multiplication circuit, which accumulates these products to generate the sum of the product.

Benefits of technology

This method significantly improves the performance of floating point arithmetic operations in memory, reduces data movement and energy consumption, and enhances the efficiency of computing devices, especially in artificial intelligence and machine learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937980A_ABST
    Figure CN119937980A_ABST
Patent Text Reader

Abstract

Some embodiments disclose in-memory computing devices and methods including providing mantissas of a subset of a plurality of pairs of first and second floating-point numbers to a respective one of the multiplication circuits for the plurality of pairs of first and second floating-point numbers each having corresponding mantissas and indices, the plurality of pairs of subsets of the first and second floating-point numbers each have a sum of indices of the first and second floating-point numbers that meet a predetermined criterion (e.g., the sum is less than a predetermined threshold); generating, using each multiplication circuit, a product of mantissas of the respective first and second floating-point number pairs; accumulating the product mantissas to generate a product mantissa partial sum; combining the product mantissa part sum with the maximum product index to generate an output floating-point number; for each of the remaining pairs of first and second floating-point numbers: not providing mantissas to the respective multiplication circuits; forbidding the corresponding multiplication circuit; or both. The trained AI model may be used to determine a threshold. For digital pairs that do not comply with the standard, various components of the multiplication and accumulation steps may be disabled by control signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to a computing device and method in memory. Background Art

[0002] The present disclosure relates generally to floating-point arithmetic operations in computing devices, such as compute-in-memory or compute-in-memory ("CIM") devices and application-specific integrated circuits ("ASICs"), and to methods and devices using data processing, such as multiply-accumulate ("MAC") operations. Computational memory or memory computing systems store information in a computer's main random access memory (RAM) and perform computations at the main RAM level, rather than moving large amounts of data between the main RAM and data storage for each computational step. Computation-in-memory allows data to be analyzed in real time because data stored in RAM can be accessed more quickly. ASICs, including digital ASICs, are designed to optimize data processing to meet specific computing needs. Improved computing performance enables faster reporting and decision making for business and machine learning applications in applications such as artificial intelligence ("AI") accelerators. Efforts are underway to improve the performance of such computational storage systems, and more specifically, to improve the performance of floating-point arithmetic operations in such systems. Summary of the invention

[0003] According to one aspect of an embodiment of the present application, a calculation method is provided, comprising: for a first plurality of floating-point numbers and a corresponding second plurality of floating-point numbers, each having a corresponding mantissa and exponent, selecting a subset of the first plurality of floating-point numbers and a corresponding subset of the second plurality of floating-point numbers based at least in part on the exponents of the first plurality of floating-point numbers and the corresponding second plurality of floating-point numbers; generating a product between each of the subset of the first plurality of floating-point numbers and a corresponding one of the subset of the second plurality of floating-point numbers using a multiplication circuit; and accumulating the products to generate a partial sum of products.

[0004] According to another aspect of an embodiment of the present application, a method for calculating in a memory is provided, comprising: for a plurality of pairs of first floating-point numbers and second floating-point numbers, each of the first floating-point numbers and the second floating-point numbers having a corresponding mantissa and exponent, providing mantissas of a subset of the plurality of pairs of first floating-point numbers and second floating-point numbers to a corresponding one of a plurality of multiplication circuits, the subset of the plurality of pairs of first floating-point numbers and second floating-point numbers each having a sum of an exponent of a first floating-point number and an exponent of a second floating-point number that satisfies a predetermined standard; using each of the plurality of multiplication circuits, generating a product of the mantissas of the corresponding pair of first floating-point numbers and second floating-point numbers; accumulating the mantissas of the products to generate a partial sum of product mantissas; combining the partial sum of product mantissas with a maximum product exponent to generate an output floating-point number; and for each of the remaining pairs of first floating-point numbers and second floating-point numbers: not providing the mantissa to the corresponding multiplication circuit; disabling the corresponding multiplication circuit; or both.

[0005] According to another aspect of an embodiment of the present application, a computing device in a memory is provided, comprising: a plurality of multiplication circuits, each of which is configured to receive a corresponding first binary number and a second binary number pair as input, and generate a product of the received first binary number and the second binary number; a plurality of multiplexers, each of which has a first data input terminal, a second data input terminal and a selection input terminal, and is configured to receive the product generated by a corresponding one of the multiplication circuits at the first data input terminal, and receive the second input at the second data output terminal, and selectively output the received product or the second input; an accumulator, configured to generate a sum of a plurality of binary numbers, each binary number representing an output of a corresponding one of the plurality of multiplexers; and a plurality of comparators, each of which has a first input terminal, a second input terminal and an output terminal, and is configured to receive a corresponding input signal at the first input terminal, and receive a common input signal of all comparators at the second input terminal, the selection input terminal of the multiplexer being connected to the output terminal of a corresponding comparator among the plurality of comparators. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various aspects of the present invention are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, in accordance with standard practice in the industry, the various components are not drawn to scale and are for illustrative purposes only. In fact, the size of the various components may be increased or decreased arbitrarily for clarity of discussion. In addition, the accompanying drawings are illustrated as examples of embodiments of the present invention and are not intended to be limiting.

[0007] Figure 1 A method of multiply-accumulate ("MAC") operation according to some embodiments is outlined.

[0008] Figure 2 A process for determining a threshold for excluding numbers from a MAC operation in accordance with some embodiments is outlined.

[0009] Figure 3 A device for performing a MAC operation according to some embodiments is schematically illustrated.

[0010] Figure 4 A reduction in storage bits as a result of employing pre-multiplication mantissa alignment in a MAC operation is shown in accordance with some embodiments.

[0011] Figure 5 A device for performing a MAC operation according to some embodiments is schematically illustrated.

[0012] Figure 6 A device for performing a MAC operation according to some embodiments is schematically illustrated.

[0013] Figure 7 A device for performing a MAC operation according to some embodiments is schematically illustrated.

[0014] Figure 8 is a block diagram illustrating a computer system that is programmed to perform computing operations according to some embodiments. DETAILED DESCRIPTION

[0015] The following disclosure provides many different embodiments or examples for realizing different features of the present invention. Specific embodiments or examples of components and arrangements are described below to simplify the present invention. Of course, these are only examples and are not intended to be limiting. For example, in the following description, forming a first component above or on a second component may include an embodiment in which the first component and the second component are directly contacted, and may also include an embodiment in which an additional component may be formed between the first component and the second component so that the first component and the second component may not be in direct contact. In addition, the present invention may repeat reference numbers and / or letters in various examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or configurations discussed.

[0016] Additionally, for ease of description, spacing relation terms such as "below," "beneath," "lower," "above," "upper," etc. may be used herein to describe the relationship of one element or component to another element or component as shown in the figures. The spacing relation terms are intended to encompass different orientations of the device in use or in the process of operation in addition to the orientation shown in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spacing relation descriptors used herein may likewise be interpreted accordingly.

[0017] The present disclosure relates generally to floating point arithmetic operations in computing devices, such as compute-in-memory or compute-in-memory ("CIM") devices and application specific integrated circuits ("ASICs"), and to methods and devices using data processing, such as multiply-accumulate ("MAC") operations. Computer artificial intelligence ("AI") uses deep learning techniques, where computing systems can be organized as neural networks. For example, a neural network refers to multiple interconnected processing nodes that can analyze data. A neural network uses "weights" to perform calculations on new input data. A neural network uses multiple layers of computing nodes, where deeper layers perform calculations based on the results of calculations performed by higher layers.

[0018] The CIM circuit performs operations locally within the memory without sending data to the host processor. This may reduce the amount of data transferred between the memory and the host processor, thereby achieving higher throughput and performance. The reduction in data movement also reduces the energy consumption of overall data movement within the computing device. Alternatively, the MAC operation can be implemented in other types of systems, such as computer systems that are programmed to perform MAC operations.

[0019] In a MAC operation, a set of input numbers are multiplied by a corresponding one of a set of weight values ​​(or multiple weights), which can be stored in a memory array. These products are then accumulated, that is, added together to form an output number. In some applications, such as neural networks used in artificial intelligence machine learning, the output generated by the MAC operation can be used as a new input value for subsequent layers of the neural network. An example of a mathematical description of a MAC operation is as follows.

[0020]

[0021] Among them A I is the Ith input, W IJ is the weight corresponding to the I-th output and the J-th weight column. J is the MAC output of the Jth weight column, and h is the cumulative number.

[0022] In a floating point ("FP") MAC operation, an FP number may be represented as a sign, a mantissa or significand, and an exponent, the exponent being an integer power of a base. The product of two FP numbers or factors may be represented as the product of a mantissa ("product mantissa") and the sum of the exponents of the factors. The sign of the product may be determined by whether the signs of the factors are the same. In a binary floating point ("FP") MAC operation, which may be implemented in a digital device such as a digital computer and / or a digital CIM circuit, each FP factor may be stored as a mantissa of a bit width (number of bits), a sign (e.g., a single sign bit, S(1 b represents negative; 0 represents non-negative), the sign of the mantissa and the floating point number are (-1) S ), and integer powers of the base (i.e., 2). In some representation schemes, binary FP numbers are normalized or adjusted so that the mantissa is greater than or equal to 1. b But less than 10 b That is, the integer part of the normalized binary FP number is 1 b In some hardware implementations, the integer part of the normalized binary FP number (i.e., 1 b ) is a hidden bit, that is, not stored, because it is assumed to be 1 b In some representation schemes, the product of two FP numbers or factors may be represented by the product mantissa, the sum of the factor exponents, and a sign, which may be determined by comparing the signs of the factors or the sum of the sign bits or the least significant bit ("LSB") of the sum.

[0023] To implement the accumulation portion of the MAC operation, in some processes, the product mantissas are first aligned. That is, if necessary, at least some of the product mantissas are modified by an appropriate order of magnitude so that the exponents of the product mantissas are the same. For example, the product mantissas may be aligned by reducing at least some of the product mantissas by an appropriate order of magnitude, such as by right-shifting the mantissas so that all exponents are the maximum exponent of the pre-aligned product mantissas. (i) Product mantissa PD M [i] The order of magnitude of the reduction is the pre-alignment index PD E [i] and maximum index PD E-MAX The difference between Δ [i](“delta index”)(E Δ [i]=PD E-MAX -PDE[i]). The aligned product mantissas can then be added together (algebraic sum) to form the mantissa of the MAC output, where the pre-aligned product mantissa is the largest exponent.

[0024] According to certain aspects of the present disclosure, product mantissas having pre-aligned exponents that are significantly less than the maximum exponent are excluded (or "skipped") from the accumulation portion of the MAC operation. In some embodiments, pre-aligned product mantissas having delta exponents equal to or greater than a predetermined threshold T are excluded. In some embodiments, the threshold T is determined at least in part based on its impact on the inference accuracy of the AI ​​model training weight values ​​applied to the test data (similar to the training data used to build the AI ​​model).

[0025] refer to Figure 1 In one example embodiment, during a MAC process of a set of FP number pairs (e.g., weight values ​​and input activations), the exponents of each pair of FP numbers are summed (SUM) 101 to generate a corresponding product exponent. Next, a maximum exponent of the product exponents is determined 103, for example, by one or more comparators or microprocessors.

[0026] The maximum product exponent is then used to determine 105 the value to be passed in the MAC process. In this example, for each pair of FP numbers, a determination is made 107 whether to exclude the product mantissa from the MAC operation. In some embodiments, the determination 107 is made based on a delta exponent, depending on the maximum product exponent. If the result of the determination 107 is negative, the product mantissa of the FP number pair with the associated sign is generated, i.e., the product of the mantissas; if the result of the determination 107 is positive, a null output is generated 111, such as 0. b, without performing mantissa multiplication. The maximum product exponent in this example is also used as a basis (e.g., by a delta exponent) to select 113 an output (product mantissa or zero) for subsequent steps in the MAC process. The selection 113 can be accomplished, for example, using a multiplexer (MUX) where a signal indicative of a delta exponent relative to a threshold is applied to a select input, and the product mantissa and zero are applied to corresponding data inputs.

[0027] Next, the non-zero product mantissas generated in step 105 and passed forward 113 are aligned to each other 115 using the maximum product exponent as described above. The aligned mantissas are accumulated 117 to generate a partial sum mantissa. The partial sum mantissa is then combined 119 with the maximum product exponent. In this example, "combination" means providing the partial sum mantissa and the maximum product exponent in the computing system so that the system can use it in subsequent operations. For example, the combination can include a l-bit sign, followed by an m-bit exponent, followed by an n-bit mantissa, where l, m, and n are predetermined based on the format of the FP number used. Finally, the combination is output 121 as a floating point number. In some embodiments, the output step 121 includes normalization as described above.

[0028] In some embodiments, a decision 107 is made as to whether to exclude the product mantissa from the MAC process based on the delta exponent relative to a threshold. The threshold T is determined at least in part based on the impact on the inference accuracy of the trained AI model applied to the test data. Figure 2 An example process for determining a threshold is outlined in . In this example, a threshold is to be determined for an AI model for object classification. Inference run 201 is performed using a trained AI model, i.e., an AI model with trained weight values. For example, the input data can be images of various categories of objects, such as dogs, cats, cars, etc., and the output run is a label generated by the AI ​​model. For example, multiple (e.g., 1000) input images can be used, one input image for each object category.

[0029] In this example, the process of determining the threshold is based on algorithm-hardware co-optimization, where the threshold is predetermined at the algorithm level, by examining the distribution of product delta exponents and verifying that the inference accuracy is not degraded by MAC skipping (i.e., MAC operations that set the product mantissa to zero for FP number pairs where the product exponent is equal to or greater than the threshold) compared to the baseline accuracy, which can be established by software inference running on the GPU or CPU using FP32 or FP16 data formats without utilizing any MAC skipping. Figure 2In the example shown, a software-based AI model is used to calculate the weight input product delta exponent of all layers of the AI ​​model, and an initial threshold for MAC skipping is determined 205 using a delta exponential distribution (an example of a specific input image is shown at label 203). For example, the initial threshold may be selected at a point where a large percentage (e.g., 75% or 80%) of the products and / or beyond the trailing edge of the dominant peak are included in the MAC operation. In other examples, a sufficiently large threshold may be selected empirically.

[0030] In this example, the initial threshold is then used to verify 207 the accuracy of the AI ​​model with MAC skipping. The accuracy of MAC skipping based on the initial threshold is compared to the software baseline accuracy 209. If the inference accuracy of MAC skipping is lower than the baseline accuracy by more than an acceptable amount, the threshold is slightly increased 211 (e.g., by 1 or 2), and the AI ​​model with MAC skipping is run again to verify 207 the accuracy. The verification process is repeated until the inference accuracy is acceptable. The final threshold can then be selected for the hardware implementation of the AI ​​model.

[0031] Conversely, in some embodiments, if an initial threshold results in acceptable inference accuracy, smaller thresholds may be tested until the accuracy decreases to an unacceptable level, and then a maximum threshold that still results in an acceptable level of accuracy may be selected for hardware implementation of the AI ​​model. In either case, a threshold greater than a barely acceptable threshold may be selected to implement the AI ​​model.

[0032] The analysis shows that using a sufficiently large delta exponent threshold that also reduces a large number of MAC operations can achieve inference accuracy that is essentially the same as the software baseline accuracy. In one example, as shown in the following table, 10 d The threshold level results in a 20% reduction in MAC operations; 8 d The threshold level of 1 leads to a 25% reduction in MAC operations. In both cases, the inference accuracy measured by top-1 and top-5 accuracy is essentially the same as the software baseline accuracy.

[0033]

[0034] The selected threshold resulting from the above process may be sent 213 to hardware implementing the AI ​​model with MAC skipping, or otherwise used in hardware.

[0035] Figure 3 An example of a computing device capable of performing MAC skipping MAC operations is shown in FIG. The device in this example includes a set of n (64 in this example) adders 301 i , where i = 0 to n-1. Each adder 301 i Receive the corresponding pair of input indices EX [i] and the weight exponent E W [i], and generate the sum of each pair of exponents. The computing device further includes a set of circuits connected to the output of the adder 301 i for receiving the product exponents and determining the maximum product sum. In some embodiments, the circuits include n - 1 circuits 303 in the log2n layers i . Each circuit 303 in this example i receives a pair of product exponents and outputs the maximum (larger value) of the two exponents. The first layer of n / 2 circuits 303 i receives the input from the adder 301 i ; each successive layer of circuits 303 i has half the number of circuits 303 in the previous layer i and each outputs the maximum of the two received product exponents. The last layer has a single circuit 303 i and outputs the maximum product exponent of all the product exponents to the adder 301 i . Any suitable circuits for selection, sorting, or other data processing based on the relative values of numbers can be used. For example, a digital comparator can be used to compare a pair of product exponents, and the output of the comparator can be applied to the selection line of a multiplexer to select the larger of the two product exponents received at the input of the multiplexer.

[0036] The computing device in this example further includes a set of subtracters 305 i , each subtracter receiving the corresponding product exponent E SUM [i] and the maximum product exponent as inputs and outputting the difference E Δ [i] between the product exponent and the maximum product exponent or the Δ exponent. The computing device in this example further includes a set of comparators 307 i , each comparator receiving the corresponding Δ exponent E Δ [i] and a threshold T of the Δ exponent as inputs and outputting a control signal indicating the relationship between E Δ [i] and T. For example, the control signal can be a unit binary number, where 0 indicates E Δ [i] < T and 1 indicates E Δ [i] ≥ T. The computing device in this example further includes a set of registers 309 i , each register receiving the corresponding Δ exponent E Δ [i] and the control signal from the corresponding comparator 307 i for the exponent as inputs and storing the Δ exponent or zero according to the output of the comparator. Each register 309 i also stores the control signal from the respective comparators 307 i .

[0037] Other devices are capable of producing different outputs depending on the relative values ​​of the delta index and the threshold. For example, a subtractor can be used to subtract the threshold from the delta index, and the sign bit of the result can be used as a control signal. Alternatively, the threshold can be added to the product exponent and a subtractor 305 can be used. i The sum is subtracted from the maximum product exponent. The sign bit of the difference can be used as a control signal. As another option, a threshold value can be subtracted from the maximum product exponent and a subtractor 305 can be used. i The difference is used to subtract the product exponent. The sign bit of the difference can be used as a control signal. This alternative has the advantage of using a single subtractor instead of multiple subtractors or comparators, thus reducing the number of components and associated operations. For both alternatives, register 309 i The product exponent input can be directly taken from adder 301 i rather than from the output of subtractor 305 i The output is obtained.

[0038] The computing device in this example also includes a register 311 i , each register receives a corresponding pair of input mantissas M X [i] and weight digit M W [i] and corresponding comparator 307 i The output signal of the comparator 307 is used as input. i The comparator control signal outputs each register 311 i Stores the input and weight mantissa or zero. In some embodiments, if the delta index is equal to or greater than the threshold T, register 311 i Store zero; if the Δ index is less than the threshold T, register 311 i Stores input and weight mantissas.

[0039] The computing device in this example also includes a multiplication circuit 313 i Each multiplication circuit receives the value stored in the corresponding register 311 i The corresponding pairs of input tail numbers M in X [i] and weight tail number M W [i] as input. Multiplication circuit 313 i Each output in the respective product mantissa M PROD [i], which is stored in the corresponding register 311 i The input mantissa M in X [i] and weight digit M WThe product of [i]. The multiplication between the weight value and the corresponding input activation can be performed in a multiplication circuit, which can be any circuit capable of multiplying two numbers. For example, the multiplication circuits for CIM devices disclosed in U.S. Patent Application No. 17 / 558,105 published as U.S. Patent Application Publication No. 2022 / 0269483A1 and U.S. Patent Application No. 17 / 387,598 published as U.S. Patent Application Publication No. 2022 / 0244916A1, which are assigned with and incorporated herein by reference with this application. In some embodiments, the multiplication circuit includes a memory array configured to store a set of FP numbers, such as weight values; the multiplication circuit further includes a logic circuit coupled to the memory array and configured to receive another set of FP numbers, such as input values, and output signals, each signal being based on the corresponding stored number and input number and indicating the product of the stored number and the corresponding input number.

[0040] The computing device in this example further includes a selection circuit, such as multiplexer 315 i , each multiplexer receiving the product mantissa M i [i] and zero as data inputs and receiving the control signal stored in the corresponding register 309 PROD as the selection input. Each multiplexer 315 i outputs the input selected by the control signal. For example, if E i [i] ≥ T, then zero is selected as the output; if E Δ [i] < T, then M Δ [i] is selected as the output. Then, the outputs from each multiplexer 315 PROD are stored in register 317 i . i

[0041] The computing device in this example further includes a product mantissa alignment circuit, such as shifter 319i, each shifter receiving the product mantissa M i stored in register 317 Δ and the Δ exponent E PROD [i] or zero as an input and shifting M PROD [i] to the right by E △ [i] bits to generate the corresponding aligned product mantissa, which is stored in the corresponding register 321 i . The aligned product mantissas are accumulated or summed by an accumulator (such as adder tree 323 i ). The sum of the product mantissas (now excluding E Δ[i] ≥ T) is stored in register 325. Finally, the product mantissa stored in register 325 is then combined with the maximum product exponent in normalization circuit 327 to form the floating point MAC output.

[0042] Therefore, according to some embodiments, the MAC operation is performed without generating a product mantissa based on the comparison result between the product exponent and the maximum product exponent, such as Figure 4 The example timing diagram shown is shown. In the first part of the timing diagram, "Case-A", the calculated product delta index of the input and weight pair is less than or equal to the threshold. In this case, the output of the comparator is 0, indicating that the product mantissa is not excluded from the MAC operation and operates according to the conventional MAC process: First, the input and weight mantissas are loaded into the multiplication circuit or multiplier, and the product of the two (i.e., the product mantissa) is generated by the multiplier. Next, the product mantissa is selected by the multiplexer and shifted in the mantissa alignment operation. Next, the adder tree accumulates the aligned product mantissas. Finally, the accumulated product mantissas are normalized.

[0043] In the second part of the timing diagram, "Case-B", the calculated product delta exponent of the input and weight pair is greater than the threshold. In this case, the comparator output is 1, indicating that the product mantissa is excluded from the MAC operation. Therefore, the input and weight mantissas are prohibited from being loaded from the registers into the multiplier; the multiplier itself is disabled; the multiplexer selects zero; the product mantissa with a value of zero is not aligned (shifted); the input to the adder tree is zero; and the product mantissas that are not skipped are normalized.

[0044] In some embodiments, Figure 5 As shown in the example shown, a computer device for implementing a MAC operation with MAC skipping is similar to Figure 3 The device shown, but the comparator 307 i The output of the multiplication circuit 313 is connected to i To disable multiplication, instead connect to register 311 i To disable loading of input and weight mantissas into multiplication circuit 313 i , for delta exponents greater than or equal to the threshold. In this example, the input mantissa M X [i] and weight digit M W [i] is directly input to the multiplication circuit 313 i , and the output of the multiplication circuit is stored in the corresponding register 311 i middle.

[0045] In some embodiments, Figure 6 As shown in the example shown, a computer device for implementing a MAC operation with MAC skipping is similar to Figure 3The device shown, but the comparator 307 i The output of the shifter 319 is also connected to i , to disable mantissa alignment for delta exponents greater than or equal to the threshold.

[0046] In some embodiments, Figure 7 As shown in the example shown, a computer device for implementing a MAC operation with MAC skipping is similar to Figure 3 The device shown, but the comparator 307 i The output of the multiplexer 315 is also connected to i , to provide a selected input for a delta index greater than or equal to a threshold. In this particular example, the comparator output provided to the multiplexer input is 0 when the delta index is greater than or equal to the threshold. The 0 value can be provided directly by the comparator, or can be provided through an inverter in the case where the output of the comparator is 1 for a delta index greater than or equal to the threshold.

[0047] The above calculation method can be implemented by the above specific computing system, but can also be implemented by any suitable system. For example, as an alternative to performing the mantissa multiplication in the CIM memory, for example, processor-based operations can be used in a computer program to perform the above algorithm. For example, Figure 8800 shown in . In this example, the computer 800 includes a processor 810, which may include a register 812 and is connected to other components of the computer through a data communication path such as a bus 820. These components include a system memory 830, which is loaded with instructions for the processor 810 to perform the above method. A large-capacity storage device is also included, which includes a computer-readable storage medium 840. The large-capacity storage device is an electronic, magnetic, optical, electromagnetic, infrared and / or semiconductor system (or device or device). For example, the computer-readable storage medium 840 includes a semiconductor or solid-state memory, a magnetic tape, a removable computer disk, a random access memory (RAM), a read-only memory (ROM), a rigid disk and / or an optical disk. In one or more embodiments using an optical disk, the computer-readable storage medium 840 includes a compact disk read-only memory (CD-ROM), a compact disk read / write (CD-R / W) and / or a digital video disk (DVD). The mass storage device 840 stores: an operating system 842, etc.; programs 844, including programs that, when read into the system memory 820 and executed by the processor 810, cause the computer 800 to perform the above-described processes; and data 846. The computer 800 also includes an I / O controller 850, which inputs and outputs to a user interface 852. The user interface 852 may include, for example, various parts of a vehicle dashboard, audio devices, a video display, input devices such as buttons, dials, touch screen input, keyboards, mice, trackballs, and any other suitable user interface devices. The I / O controller 850 may have other input / output ports for input from and / or output to devices such as external devices 854, which may include sensors, actuators, external storage devices, etc. The computer 800 may also include a network interface 860, which enables the computer to receive and send data to a remote network 862 (such as a cellular or satellite data network), which may be used for tasks such as remote monitoring and control of the vehicle and software / firmware updates.

[0048] Certain examples described in this disclosure omit resource-intensive computational steps, such as multiplication, that produce results that have a negligible impact on the accuracy of the overall result of the entire computational process (such as a MAC). This omission can significantly reduce the overall computational steps without sacrificing accuracy. This reduction can significantly improve the efficiency of computing devices, such as general-purpose digital ASIC AI accelerators and digital CIM or near memory computing ("NMC") macros.

[0049] For example, in the application of U.S. patent application No. 17 / 558,105, a method configured to perform bit-serial multiplication calculation in a memory computing device includes: determining at least one input according to the type of application; determining at least one weight according to the training result or the configuration of the user; performing bit-serial multiplication based on the input and the weight from the most significant bit (MSB) of the input to the least significant bit of the input by the memory computing device to obtain a result according to a plurality of partial products, wherein the first partial sum of the first bit of the input is shifted left by one bit and then added to the second partial product of the second bit of the input to obtain the second partial sum of the second bit, the second bit being one bit after the first bit; and outputting the result by the memory computing device. Wherein, performing bit-serial multiplication includes: determining the first partial product of the first bit by multiplying the MSB I[N] (N>0) of the input by each bit of the weight by a multiplication circuit. Wherein, the input includes a plurality of inputs, and wherein performing bit-serial multiplication includes: determining a plurality of first partial products for the first bit by multiplying the MSB of each input in the plurality of inputs by each bit of the weight by a multiplication circuit; and summing the plurality of first partial products. Wherein, performing the bit-serial multiplication includes: shifting the first partial sum by one position to the left through the accumulator circuit; and determining the second partial product of the second position by multiplying the next position I[N-1] of the input by each position of the weight through the multiplication circuit. Wherein, performing the bit-serial multiplication includes: adding the first partial sum shifted to the left and the second partial product through the accumulator circuit to obtain the first partial sum of the next position I[N-1]. Wherein, performing the bit-serial multiplication includes: shifting the first partial sum of the next position I[N-1] obtained by one position to the left through the accumulator circuit; and determining the second partial product of the next position I[N-2] by multiplying the next position I[N-2] of the input by each position of the weight through the multiplication circuit; and adding the first partial sum of the next position I[N-1] shifted to the left and the second partial product of the next position I[N-2] obtained by the accumulator circuit to obtain the first partial sum of the next position I[N-2]. The bit serial multiplication includes: shifting the first partial sum of the next bit I[N-1] obtained by one position to the left through an accumulator circuit; determining the second partial product of LSBI[0] by multiplying the input LSBI[0] by each bit of the weight through a multiplication circuit; and adding the first partial sum of the next bit I[N-1] obtained by the left shift and the second partial product of LSBI[0] through an accumulator circuit to obtain a total.A device includes: an adder; a shifter having an output terminal operably connected to a first input terminal of the adder, the shifter being configured to shift left by one bit; a first register having an output terminal operably connected to an input terminal of the shifter; a second register having an output terminal operably connected to a second input terminal of the adder; a multiplier configured to perform bit-serial multiplication based on an input signal and a weight signal to obtain a plurality of partial products; wherein an input terminal of the second register is operable to receive a first partial product of the plurality of partial products based on a most significant bit (MSB) of the input signal; and wherein an input terminal of the first register is operable to receive an output of the adder.

[0050] For example, in the application of U.S. patent application No. 17 / 387,598, a computing device in memory includes: a storage array, including a plurality of storage cells arranged in rows and columns, the plurality of storage cells include a first storage cell in the first row and the first column of the storage array, and a second storage cell in the first row and the second column of the storage array, the first storage cell and the second storage cell are configured to store respective first weight signals and second weight signals; an input driver, configured to provide a plurality of input signals; a first logic circuit, coupled to the first storage cell and configured to provide a first output signal based on the first weight signal and the first input signal from the input driver; and a second logic circuit, coupled to the second storage cell and configured to provide a second output signal based on the second weight signal and the second input signal from the input driver. Wherein, the first logic circuit and the second logic circuit each include a multiplication circuit. Wherein, the multiplication circuit includes a NOR gate. Wherein, the multiplication circuit includes an AND gate. Wherein, at least one of the first weight signal and the second weight signal is a signed weight. The device also includes: a third storage unit in the second row and the first column of the storage array, and a fourth storage unit in the second row and the second column of the storage array, the third storage unit and the fourth storage unit are configured to store respective third weight signals and fourth weight signals; a third logic circuit coupled to the third storage unit and configured to provide a third output signal based on the third weight signal and a third input signal from the input driver; and a fourth logic circuit coupled to the fourth storage unit and configured to provide a fourth output signal based on the fourth weight signal and a fourth input signal from the input driver. The device also includes: an adder circuit configured to add the first output signal, the second output signal, the third output signal, and the fourth output signal. A method for computing in a memory, comprising: storing a plurality of weight signals in a plurality of storage cells, wherein each of the weight signals has w bits, w is a positive integer, and wherein each of the storage cells stores one bit of the w-bit weight signal; providing a plurality of logic circuits, the plurality of logic circuits being connected to corresponding storage cells in the plurality of storage cells; providing an input signal to the plurality of logic circuits to multiply the weight signal by the input signal to provide a plurality of product signals; outputting the plurality of product signals from the plurality of logic circuits to an adder tree; providing a weight sign signal, the weight sign signal being configured to indicate whether the weight signal is signed; and outputting a partial sum signal based on the product signal and the weight sign signal through the adder tree.

[0051] In summary, in some embodiments, a computation method includes: for a first plurality of floating point numbers and a corresponding second plurality of floating point numbers, each having a corresponding mantissa and exponent, selecting a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers based at least in part on the exponents of the first plurality of floating point numbers and the corresponding second plurality of floating point numbers; generating, using a multiplication circuit, a product between each of the subset of the first plurality of floating point numbers and a corresponding one of the subset of the second plurality of floating point numbers; and accumulating the products to generate a partial sum of products.

[0052] In some embodiments, selecting a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers based at least in part on the exponents of the first plurality of floating point numbers and the corresponding second plurality of floating point numbers includes selecting the subset of the first plurality of floating point numbers and the corresponding subset of the second plurality of floating point numbers based at least in part on a difference between a sum of the exponents of each pair of the first floating point number and the corresponding second floating point number and a maximum of the sums of the exponents, the difference being a delta exponent.

[0053] In some embodiments, the selecting step comprises excluding each pair of a first floating point number and a corresponding second floating point number having a delta exponent greater than a predetermined threshold.

[0054] In some embodiments, the calculation method also includes using a trained artificial neural network to determine the threshold and determine the accuracy of the results of each test threshold. The trained artificial neural network uses training data and one or more test thresholds, and if the corresponding accuracy meets the predetermined criteria, the test threshold is set to a predetermined threshold.

[0055] In some embodiments, the excluding step includes using a control signal to inhibit providing the pair of the first floating point number and the second floating point number to a corresponding one of the multiplication circuits.

[0056] In some embodiments, the excluding step includes disabling a corresponding one of the multiplication circuits using a control signal.

[0057] In some embodiments, the excluding step includes setting the product between the mantissa of the first floating point number and the corresponding second floating point number to zero.

[0058] In some embodiments, setting the product to zero includes: connecting each output terminal of the multiplication circuit for all first plurality of floating point numbers and corresponding second plurality of floating point numbers to a data input terminal of a corresponding multiplexer; providing 0 to another data input terminal of each multiplexer; and operating each multiplexer connected to a corresponding one of the multiplication circuits to select to provide a data input of 0 for each pair of the first floating point number and the corresponding second floating point number having a delta exponent greater than a threshold.

[0059] Furthermore, according to some embodiments, a computing method includes: for a plurality of pairs of first floating-point numbers and second floating-point numbers, each of the first floating-point numbers and the second floating-point numbers having a corresponding mantissa and exponent, providing mantissas of a subset of the plurality of pairs of first floating-point numbers and the second floating-point numbers to a corresponding one of a plurality of multiplication circuits, the subset of the plurality of pairs of first floating-point numbers and the second floating-point numbers each having a sum of an exponent of the first floating-point number and an exponent of the second floating-point number that satisfies a predetermined criterion; using each of the plurality of multiplication circuits, generating a product of the mantissas of the corresponding pair of first floating-point numbers and second floating-point numbers; accumulating the product mantissas to generate a product mantissa partial sum; combining the product mantissa partial sum with a maximum product exponent to generate an output floating-point number; and for each of the remaining pairs of first floating-point numbers and second floating-point numbers: not providing a mantissa to the corresponding multiplication circuit; disabling the corresponding multiplication circuit; or both.

[0060] In some embodiments, the accumulating step includes aligning the mantissa products using a plurality of shifters so that the exponents of all products between the first floating point number and the second floating point number in corresponding pairs are equal to the maximum product exponent.

[0061] In some embodiments, providing a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers comprises providing a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers based at least in part on a difference between a sum of exponents of each pair of the first floating point number and the corresponding second floating point number and a maximum of the sums of the exponents, the difference being a delta exponent.

[0062] In some embodiments, the providing step includes excluding each pair of a first floating point number and a corresponding second floating point number having a delta exponent greater than a predetermined threshold.

[0063] In some embodiments, the excluding step includes: using a control signal to disable a register storing a pair of the first floating point number and the second floating point number connected to a corresponding one of the multiplication circuits.

[0064] In some embodiments, the excluding step includes disabling a corresponding one of the multiplication circuits using a control signal.

[0065] In some embodiments, the excluding step includes setting the product between the mantissa of the first floating point number and the corresponding second floating point number to zero.

[0066] In addition, according to some embodiments, the computing device includes: a plurality of multiplication circuits, each multiplication circuit being configured to receive a corresponding first binary number and a second binary number pair as input, and generating a product of the received first binary number and the second binary number; a plurality of multiplexers, each multiplexer having a first data input terminal and a second data input and a selection input terminal, and being configured to receive a product generated by a corresponding one of the multiplication circuits at the first data input terminal, and to receive the second input at a second data output terminal, and selectively output the received product or the second input; an accumulator, being configured to generate a sum of a plurality of binary numbers, each binary number representing an output of a corresponding one of the plurality of multiplexers; and a plurality of comparators, each comparator having a first input terminal and a second input terminal and an output terminal, and being configured to receive a corresponding input signal at a first input terminal, and to receive a common input signal of all the comparators at a second input terminal, the selection input terminal of the multiplexer being connected to the output terminal of a corresponding comparator among the plurality of comparators.

[0067] In some embodiments, the accumulator includes: a plurality of shifters, each shifter configured to receive an output from a corresponding one of the multiplexers as an input and configured to produce an output; and an adder configured to generate a sum of the outputs from the shifters.

[0068] In some embodiments, the output of each comparator is connected to a corresponding one of the multiplication circuits to enable or disable the corresponding multiplication circuit depending on the state of the output of the comparator.

[0069] In some embodiments, the computing device further includes a plurality of registers, each register being configured to receive, store and output a corresponding first binary number and a second binary number pair as input to a corresponding one of a plurality of multiplication circuits, wherein an output terminal of each comparator is connected to a corresponding register to enable or disable an output of the corresponding register according to a state of the output terminal of the comparator.

[0070] In some embodiments, the output of each comparator is connected to a corresponding shifter to enable or disable the shifter based on the state of the comparator output.

[0071] The features of several embodiments are summarized above so that those skilled in the art can better understand the various aspects of the present disclosure. Those skilled in the art will appreciate that they can easily use the present disclosure as a basis for designing or modifying other processes and structures for realizing the same purpose of the embodiments introduced herein and / or realizing the same advantages thereof. Those skilled in the art will also appreciate that such equivalent structures do not deviate from the spirit and scope of the present invention, and they can make various changes, substitutions and changes in the present invention without deviating from the spirit and scope of the present invention.

Claims

1. A method for computing in memory, comprising: For a first plurality of floating point numbers and a corresponding second plurality of floating point numbers, each having a corresponding mantissa and exponent, selecting a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers based at least in part on the exponents of the first plurality of floating point numbers and the corresponding second plurality of floating point numbers; generating, using multiplication circuitry, a product between each of the subset of the first plurality of floating point numbers and a corresponding one of the subset of the second plurality of floating point numbers; as well as The products are accumulated to generate a partial sum of products.

2. The in-memory computing method according to claim 1, wherein: Selecting the subset of the first plurality of floating point numbers and the corresponding subset of the second plurality of floating point numbers based at least in part on the exponents of the first plurality of floating point numbers and the corresponding second plurality of floating point numbers includes selecting the subset of the first plurality of floating point numbers and the corresponding subset of the second plurality of floating point numbers based at least in part on a difference between a sum of the exponents of each pair of the first floating point number and the corresponding second floating point number and a maximum of the sums of the exponents, the difference being a delta exponent.

3. The in-memory computing method according to claim 2, wherein: The step of selecting comprises excluding each pair of a first floating point number and a corresponding second floating point number having the delta index greater than a predetermined threshold.

4. The in-memory computing method of claim 3 further comprises determining the threshold using a trained artificial neural network and determining the accuracy of the result corresponding to the test threshold, wherein the trained artificial neural network uses training data and one or more test thresholds, and if the corresponding accuracy meets a predetermined criterion, the test threshold is set to the predetermined threshold.

5. The in-memory computing method according to claim 3, wherein: The step of excluding comprises setting the product between the mantissa of the first floating point number and the corresponding second floating point number to zero.

6. The in-memory computing method according to claim 5, wherein: Setting the product to zero comprises: connecting each output of the multiplication circuit for all of the first plurality of floating point numbers and the corresponding second plurality of floating point numbers to a data input of a corresponding multiplexer; providing a 0 to another data input of each of said multiplexers; and Each multiplexer connected to a corresponding one of the multiplication circuits is operative to select a data input providing 0 for each pair of a first floating point number and a corresponding second floating point number having a delta exponent greater than a threshold.

7. A method for in-memory computing, comprising: For a plurality of pairs of first and second floating point numbers, each of the first and second floating point numbers having a corresponding mantissa and exponent, providing mantissas of a subset of the plurality of pairs of first and second floating point numbers to a corresponding one of a plurality of multiplication circuits, the subset of the plurality of pairs of first and second floating point numbers each having a sum of an exponent of the first floating point number and an exponent of the second floating point number that satisfies a predetermined criterion; generating, using each of the plurality of multiplication circuits, a product of a corresponding pair of mantissas of a first floating point number and a second floating point number; Accumulating the product mantissas to generate a product mantissa partial sum; combining the product mantissa partial sum and a maximum product exponent to generate an output floating point number; and For each of the remaining pairs of first and second floats: not providing the mantissa to a corresponding multiplication circuit; disabling the corresponding said multiplication circuit; or Both.

8. The in-memory computing method according to claim 7, wherein: The step of accumulating includes: aligning the mantissa products using a plurality of shifters so that the exponents of all products between corresponding pairs of the first floating point number and the second floating point number are equal to the maximum product exponent.

9. The in-memory computing method according to claim 7, wherein: Providing a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers comprises providing a subset of the first plurality of floating point numbers and a corresponding subset of the second plurality of floating point numbers based at least in part on a difference between a sum of exponents of each pair of the first floating point number and the corresponding second floating point number and a maximum of the sums of the exponents, the difference being a delta exponent.

10. A computing-in-memory device, comprising: a plurality of multiplication circuits, each multiplication circuit being configured to receive as input a corresponding pair of a first binary number and a second binary number and to generate a product of the received first binary number and the second binary number; a plurality of multiplexers, each multiplexer having a first data input and a second data input and a select input, and configured to receive the product generated by a corresponding one of the multiplication circuits at the first data input and receive a second input at the second data output, and selectively output the received product or the second input; an accumulator configured to generate a sum of a plurality of binary numbers, each binary number representing an output of a corresponding one of a plurality of said multiplexers; as well as a plurality of comparators, each comparator having a first input terminal and a second input terminal and an output terminal, and configured to receive a corresponding input signal at the first input terminal and to receive a common input signal of all comparators at the second input terminal, The selection input of the multiplexer is connected to the output of a corresponding comparator of the plurality of comparators.

Citation Information

Patent Citations

  • Compute in memory

    US20220244916A1

  • Compute in memory accumulator

    US20220269483A1

Cited By

  • Floating-point number multiplication method and floating-point number multiplication circuit

    CN120803394A