Floating point multiply-accumulate operation circuit based on SRAM (Static Random Access Memory) and in-memory calculation chip

Through the SRAM-based floating point multiplication and accumulation operation circuit, the index sum is filtered using subtractors and counters, and shifted addition is performed at the tail of the adder tree, the area and power consumption loss problems of in-memory computing circuits or chips during floating point multiplication and accumulation operation are solved, and efficient floating point calculation is realized.

CN120353428APending Publication Date: 2025-07-22ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510425698.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing in-memory computing circuits or chips have large area and power losses during floating point multiplication and accumulation operations, which cannot meet the needs of complex edge artificial intelligence tasks.

Method used

The floating point multiplication and accumulation calculation circuit based on SRAM is used to calculate the exponents and mantissa product of the floating point input value and weight value bit by bit, and the index sum of the same first two bits is filtered, and the index sum is filtered using subtractors and counters, and shifted addition is performed at the tail of the adder tree, saving the loss of the excess registers and adder trees.

Benefits of technology

With small amplitude accuracy loss, floating-point computing is implemented, reducing the power consumption loss of register area and adder tree, improving computing efficiency and throughput, and suitable for high-performance computing architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353428A_ABST
    Figure CN120353428A_ABST
Patent Text Reader

Abstract

The invention discloses a floating point multiply-accumulate operation circuit based on an SRAM (Static Random Access Memory) and an in-memory calculation chip. The operation circuit is used for calculating the exponent sum and mantissa product of a multi-bit input value and a multi-bit weight value which are both floating point type according to bits; the method comprises the following steps: firstly, determining a maximum index, then screening out the sum of the first two indexes which are the same as the maximum index, screening out the mantissa products corresponding to the sum of the first two different indexes, finally calculating the bit difference between the residual indexes and the residual bits of the maximum index, starting counting, and inputting the corresponding mantissa products into an adder tree when the period is the same as the bit difference; adding mantissa products in the adder tree, and performing shift addition at the tail part of the adder tree in each period to obtain a mantissa sum; and combining the maximum exponent sum with the mantissa sum to obtain a multiply-accumulate calculation result. According to the invention, the area loss of redundant registers and the power loss of the adder tree are saved, the calculation loss of the adder tree can be reduced, and the application prospect is very wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an arithmetic circuit in the field of integrated circuit technology, and particularly to a floating-point multiply-accumulate arithmetic circuit based on SRAM, and also to an in-memory computing chip. Background Art

[0002] With the rapid development of artificial intelligence, some application fields such as machine learning and edge computing have developed rapidly, and there are higher requirements for computing speed. However, in the traditional computing architecture, the computing unit is separated from the memory, and data needs to be frequently transferred between the two, resulting in high energy consumption and low efficiency. Due to the rapid development of Moore's Law, the operating speed of the memory is out of sync with the processor speed, and the access speed of the memory lags far behind the computing speed of the processor. Memory performance has become an important bottleneck in the overall computer performance, and this bottleneck is particularly obvious in fields with large amounts of computation such as machine learning and image recognition. To overcome the drawbacks brought by these traditional von Neumann architectures, in-memory computing has become a hot topic to solve this problem. In-memory computing does not require data to be transferred to the processor and directly performs operations in the memory, thus greatly reducing the energy consumption of data access during the computing process and improving the computing speed and energy efficiency at the same time.

[0003] However, with the continuous development of edge artificial intelligence applications, they are facing increasingly complex tasks such as object detection and device model fine-tuning. Existing in-memory computing structures based on integer (INT) quantization integer (INT) NN models (such as INT8) are not sufficient to meet the requirements of these complex tasks, and more powerful neural network models based on floating-point (FP) operations are needed. Floating-point has an additional exponent part compared to integer type, so there is an additional step of exponent bit alignment in floating-point operations. Existing in-memory computing circuits or chips require a large number of barrel shifters during alignment, which results in very large area and power consumption losses. Or, serial shifting can be performed on each result, which also consumes a lot of registers and computing cycles, and a very large adder tree is also required in floating-point operations, which also occupies a very large overhead. Summary of the Invention

[0004] To solve the technical problem of large losses in the floating-point multiply-accumulate operation of existing in-memory computing circuits or chips, the present invention provides a floating-point multiply-accumulate arithmetic circuit based on SRAM and an in-memory computing chip.

[0005] The present invention is implemented by the following technical solutions: A floating-point multiply-accumulate arithmetic circuit based on SRAM, which is used for:

[0006] (1) Calculate the exponent sum and the mantissa product of multi-bit input values and multi-bit weight values that are both floating-point types bit by bit respectively;

[0007] (2) First, sort the exponent sum calculation results and determine the maximum exponent. Then, filter out the exponent sums that have the same first two digits as the maximum exponent, and eliminate the mantissa products corresponding to the exponent sums with different first two digits. Finally, calculate the digit difference between the remaining exponent sum and the remaining digits of the maximum exponent, and start counting. When reaching the same period as the digit difference, input the corresponding mantissa product into the adder tree;

[0008] (3) Add the mantissa products in the adder tree, and perform a shifted addition on the final result at the tail of the adder tree every period. After the counting is completed, obtain the total mantissa;

[0009] (4) Combine the maximum exponent sum and the total mantissa to obtain the multiplication and accumulation calculation result of the multi-bit input value and the multi-bit weight value.

[0010] The present invention proposes a new screening and alignment operation, which can achieve floating-point calculations with very little precision loss. At the same time, it can also eliminate the area loss of redundant registers and the power consumption loss of the adder tree. Therefore, it has a very broad application prospect. The circuit structure is simple and has great advantages in terms of area and power consumption, solving the technical problem of large losses in the existing in-memory computing circuits or chips during floating-point multiplication and accumulation operations.

[0011] As a further improvement of the above solution, the arithmetic circuit includes:

[0012] An array circuit, which includes a weight exponent array, an exponent sum array, and a weight mantissa array; each row of the weight exponent array is used to pre-store the exponent part of the multi-bit weight value bit by bit; the exponent sum array is used to store the exponent sum calculation result; each row of the weight mantissa array is used to pre-store the mantissa part of the multi-bit weight value bit by bit, and perform a multiplication operation on the mantissa parts of the multi-bit input value and the multi-bit weight value to obtain the mantissa product.

[0013] Further, the arithmetic circuit further includes:

[0014] An exponent input module;

[0015] An exponent adder module, which is used to first receive the exponent part of the multi-bit input value input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the exponent sum of the two exponent parts, and finally store the exponent sum calculation result in the exponent sum array.

[0016] Still further, the arithmetic circuit further includes:

[0017] A maximum value finding module, which is used to first read the exponent and calculation result of the exponent adder module, then sequentially screen out the maximum exponent from the high bit to the low bit, and write back the exponent sum to the exponent sum array.

[0018] Furthermore, the arithmetic circuit further includes:

[0019] A subtractor and counter module, which includes a subtractor and a counter; the subtractor and counter module is used to match the first two bits of the exponent sum in each row of the exponent sum array with the first two bits of the maximum exponent, screen out the exponent sum having the same first two bits as the maximum exponent, and screen out the mantissa products corresponding to the exponent sums with different first two bits, cause the counter to start counting, and input the corresponding mantissa product into the adder tree every time the period is the same as the bit difference.

[0020] Furthermore, the arithmetic circuit further includes:

[0021] An adder tree module; the subtractor and counter module inputs the corresponding mantissa product into the adder tree module every time the count of the counter is the same as the bit difference. The adder tree module is used to add the received mantissa products, and perform shift addition on the final result at the tail of the adder in each counting period, and obtain the mantissa sum after the counting is completed.

[0022] Furthermore, the subtractor includes a judge, a selector, a half-subtractor and three full-subtractors; the judge includes three AND logic gates: a first AND logic gate, a second AND logic gate and a third AND logic gate; the first AND logic gate is used to perform an AND operation on the first bit of the maximum exponent and the first bit of the exponent sum, and the second AND logic gate is used to perform an AND operation on the second bit of the maximum exponent and the second bit of the exponent sum; the third AND logic gate performs an AND operation on the operation results of the first AND logic gate and the second AND logic gate, and inputs the operation result to the selector; the output end of the selector is connected to the power enable ends of the half-subtractor and the three full-subtractors, and the half-subtractor and the three full-subtractors are used to perform subtraction calculations on the remaining four bits of the maximum exponent and the exponent sum respectively.

[0023] Further, the adder tree module includes an AND logic gate, a multi-stage adder tree, and multiple shifters; each stage of the adder tree includes a selector and multiple adders connected in series; at the start of counting for each mantissa product, each mantissa product corresponds to a zero signal Zn = 0. If the counter does not count to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 1. If the counter counts to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 0; the AND logic gate is used to receive the zero signal Zn, and its output terminal is connected to the control terminal of the selector of each stage of the adder tree; in each stage of the adder tree, the enable output terminal of the selector is connected to the enable power supply terminals of multiple adders; multiple shifters are respectively used to perform shift addition on the final calculation result of the multi-stage adder tree in each counting cycle, and the mantissa sum is obtained after counting is completed.

[0024] Further, the arithmetic circuit is used to implement FP 16 calculation, and the column ratio of the weight exponent array, the exponent sum array, and the weight mantissa array is 6:6:10.

[0025] The present invention also provides an in-memory computing chip, which includes any one of the above SRAM-based floating-point multiply-accumulate arithmetic circuits.

[0026] Compared with the existing in-memory computing circuits or chips, the SRAM-based floating-point multiply-accumulate arithmetic circuit and the in-memory computing chip of the present invention have the following beneficial effects:

[0027] 1. The SRAM-based floating-point multiply-accumulate arithmetic circuit can achieve floating-point calculation with very little precision loss by proposing a new screening and alignment operation. At the same time, it can also save the area loss of redundant registers and the power consumption loss of the adder tree. Therefore, it has a very broad application prospect. The circuit structure is simple and has great advantages in terms of area and power consumption, solving the technical problem of large losses in the existing in-memory computing circuits or chips during floating-point multiply-accumulate operations.

[0028] 2. The SRAM-based floating-point multiply-accumulate arithmetic circuit uses a subtractor to screen the exponent sum, and uses a counter and an adder tree to align and sum the mantissa, and further improves the adder tree. It can not only save the consumption of a large number of shift registers, but also reduce the calculation loss of the adder tree. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is the framework diagram of the SRAM-based floating-point multiply-accumulate arithmetic circuit according to Embodiment 2 of the present invention;

[0030] Figure 2 is Figure 1 the structural schematic diagram of the subtractor and counter module of the SRAM-based floating-point multiply-accumulate arithmetic circuit in

[0031] Figure 3 is Figure 1 a schematic structural diagram of the adder tree module of the SRAM-based floating-point multiply-accumulate operation circuit in

[0032] Figure 4 is Figure 1 the alignment flowchart of the SRAM-based floating-point multiply-accumulate operation circuit in FP16 mode in Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] Embodiment 1

[0035] This embodiment provides an SRAM-based floating-point multiply-accumulate operation circuit, which is designed based on a static random access memory array and its peripheral circuits, and the operation circuit is mainly used to implement the following functions. It should be noted here that in other embodiments, these functions do not necessarily have to be divided into four, and can be adjusted according to actual needs.

[0036] (1) Calculate the exponent sum and the mantissa product of multi-bit input values and multi-bit weight values that are both floating-point types bit by bit. In this embodiment, multi-bit input values and multi-bit weight values in the standard floating-point format (such as IEEE 754 FP16 / FP32) are used. The entire operation process can adopt a pipeline architecture for parallel processing of multiple floating-point data pairs, and finally output the normalized exponent sum and mantissa product to provide intermediate calculation results for subsequent floating-point result recombination (including sign bit determination, result normalization, etc.).

[0037] (2) First, sort the exponent sum calculation results and determine the maximum exponent, then filter out the exponent sums that have the same first two bits as the maximum exponent, and filter out the mantissa products corresponding to the exponent sums with different first two bits. Finally, calculate the bit difference between the remaining exponent sum and the remaining number of bits of the maximum exponent, and start counting. Every time the period is the same as the bit difference, the corresponding mantissa product is input to the adder tree.

[0038] In the design and implementation of this embodiment, first, all the exponent sum results obtained from parallel computing are sorted, and the maximum exponent value in the current operation batch is quickly determined through a comparator array. Subsequently, a two-stage screening mechanism is initiated: the first stage of screening retains the exponent sum results that have the same first two digits (the most significant bit MSB and the second most significant bit) as the maximum exponent, and the second stage screens out the mantissa product data corresponding to the exponent sums whose first two digits do not match. For the exponent sums that pass the screening, the system precisely calculates the difference in the remaining digits between them and the maximum exponent (i.e., the bit difference from the third digit to the lowest digit). This process uses a synchronous counting mechanism. Whenever the count value of the counter reaches the bit difference corresponding to a certain exponent sum, the system sends the mantissa product data paired with that exponent sum into the multi-stage pipelined adder tree for cumulative operation. This design effectively avoids the waste of computing resources caused by exponent alignment in the traditional method by dynamically adjusting the input timing of the mantissa products, while ensuring the accuracy requirements of floating-point addition operations. The entire processing flow specifically optimizes the tensor operation scenarios commonly found in AI inference, and can significantly improve the throughput while maintaining the computing accuracy, and is applicable to high-performance computing architectures such as GPUs and TPUs.

[0039] (3) Add the mantissa products in the adder tree, and perform a shifted addition on the final result at the tail of the adder tree every cycle. After the counting is completed, the total mantissa is obtained. In this embodiment, during the floating-point operation mantissa accumulation stage, the screened and timing-adjusted mantissa product data are sequentially input into the multi-stage pipelined adder tree for cumulative operation. This adder tree adopts a hierarchical structure design and can complete the parallel calculation of partial sums within a single cycle. During the operation process, the system performs a dynamic shifted addition operation on the intermediate result output by the adder tree at the end of each cycle: according to the currently accumulated exponent difference information, the partial sum is precisely shifted and aligned through a barrel shifter and then sent to the next-level accumulation unit.

[0040] This design realizes the progressive alignment and accumulation of mantissa products, effectively avoiding the performance bottleneck in the traditional method that requires waiting for all mantissa products to be aligned before starting the accumulation. When all mantissa products have completed input and the counter indicates the end of the alignment cycle, the system will obtain the fully aligned total mantissa at the final output stage of the adder tree. This innovative shifted accumulation architecture is particularly suitable for processing large-scale matrix operations in deep learning, significantly improving the operation throughput rate while ensuring the computing accuracy, and providing an efficient mantissa processing solution.

[0041] (4) Combine the maximum exponent sum and the mantissa sum to obtain the multiplication-accumulation calculation result of the multi-bit input value and the multi-bit weight value. In the final result synthesis stage of the floating-point multiplication-accumulation operation, the maximum exponent value and the mantissa sum obtained through the above steps are intelligently combined to generate the final calculation result that conforms to the IEEE 754 standard. The multiplication-accumulation calculation result not only contains accurate numerical information but also retains the complete floating-point format characteristics, and can be directly used for subsequent neural network calculations or scientific operations. This optimized combination method makes full use of the intermediate results of the previous-stage processing, avoids the performance loss caused by repeated calculations in traditional methods, and is particularly suitable for implementing high-efficiency floating-point matrix operations in scenarios such as AI accelerators and high-performance computing chips. The entire synthesis process adopts a pipeline design, which can complete the result combination within a single cycle, significantly improving the overall throughput of the system.

[0042] Embodiment 2

[0043] Please refer to Figure 1 , this embodiment provides a floating-point multiplication-accumulation operation circuit based on SRAM to implement the multiplication-accumulation operation between multiple groups of multi-bit floating-point input feature numbers and multi-bit floating-point weights. This embodiment further clarifies the structure on the basis of Embodiment 1 and is used to implement the multiplication-accumulation operation between multiple groups of FP16 floating-point input feature values and FP16 floating-point weights, and can perform 64 MAC operations in FP16 format in one calculation. Of course, in other embodiments, the operation circuit can also implement the multiplication-accumulation operation between other types of floating-point input feature values and weights. This embodiment is only illustrated as one implementation manner.

[0044] In this embodiment, the operation circuit has two working modes: storage and calculation, and the function switching of the circuit is controlled by a mode switching module. Among them, the peripheral circuit of the static random access memory array refers to the related circuits and modules related to realizing data reading, writing, and saving using the static random access memory array. The operation circuit may include an array circuit, an exponent input module, an exponent adder module, a maximum value finding module, a subtractor and counter module, and an adder tree module.

[0045] The array circuit includes a Weight Exponent Array, a SumofExponent Array, and a Weight Mantissa Array. Each row of the Weight Exponent Array is used to pre-store the exponent part of the multi-bit weight value bit by bit. The SumofExponent Array is used to store the exponent sum calculation result. Each row of the Weight Mantissa Array is used to pre-store the mantissa part of the multi-bit weight value bit by bit, and uses its own logical function to perform a multiplication operation on the mantissa parts of the multi-bit input value and the multi-bit weight value to obtain the mantissa product.

[0046] The weight exponent array is mainly used to handle tasks related to the exponent part of floating-point numbers during the operation process. The weight mantissa array is used to handle tasks related to the mantissa part of floating-point numbers during the operation process. The exponent sum array is used to store the sum of exponents between the input and the weights; for floating-point numbers in FP16 format, it contains 5 exponent bits and 10 mantissa bits. Considering the overflow situation in the actual operation process, the number of columns of the exponent sum array, the weight exponent array, and the weight mantissa array can be set to 6:6:10 respectively.

[0047] The exponents or mantissas of the same input and weights are input bit by bit into the same row in the array, and one bit is stored in the same cell in the array. Different rows are used to store different multi-bit exponents or multi-bit mantissas. On this basis, the maximum number of groups of inputs or weights for floating-point multiply-accumulation that the arithmetic circuit of this embodiment can support and complete is determined by the number of rows of the static random access memory array. For example, in a circuit where the scales of the exponent sum array, the weight exponent array, and the weight mantissa array are 64x6, 64x6, and 64x10 respectively, it can support and complete the multiply-accumulation operation between 64 floating-point inputs and 64 floating-point weights at most.

[0048] The exponent adder module is used to first receive the exponent part of the multi-bit input value input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the sum of the two exponent parts, and finally store the calculation result of the exponent sum into the exponent sum array.

[0049] The maximum value finding module is used to first read the calculation result of the exponent sum of the exponent adder module, then screen out the maximum exponent from the high bit to the low bit in turn, and write back the exponent sum to the exponent sum array.

[0050] The subtractor and counter module includes a subtractor and a counter. The subtractor and counter module is used to match the first two bits of the exponent sum of each row in the exponent sum array with the first two bits of the maximum exponent, screen out the exponent sum that has the same first two bits as the maximum exponent. If the first two bits are different, it means the difference is relatively large, and the mantissa product corresponding to the exponent sum with different first two bits is screened out. This value has a very small impact on the final calculation result. At the same time, the counter starts to count, and when it reaches the same period as the bit difference, the corresponding mantissa product is input into the adder tree.

[0051] Please refer to Figure 2, when the count of the counter is the same as the difference in the number of digits in each cycle, the subtractor and counter module input the corresponding product of the mantissas into the adder tree module. The adder tree module is used to add the received products of the mantissas, and in each counting cycle, the final result is shifted and added at the end of the adder. After the counting is completed, the total mantissa is obtained. The calculation result input into the adder tree in each cycle is shifted and added at the end. After the counter finishes counting, the total mantissa is obtained. At this time, there are many zeros in the result input into the adder tree in each cycle, so a signal can be added to control the enable of the adder, thereby achieving the effect of reducing power consumption; by adding all the mantissas corresponding to the sum of all exponents, the total mantissa can be calculated.

[0052] The subtractor includes a discriminator, a selector, a half-subtractor, and three full-subtractors. The discriminator includes three AND logic gates: the first AND logic gate, the second AND logic gate, and the third AND logic gate. The first AND logic gate is used to perform an AND operation on the first digit of the maximum exponent and the first digit of the sum of exponents. The second AND logic gate is used to perform an AND operation on the second digit of the maximum exponent and the second digit of the sum of exponents. The third AND logic gate performs an AND operation on the operation results of the first AND logic gate and the second AND logic gate, and inputs the operation result into the selector. The output terminal of the selector is connected to the power enable terminals of the half-subtractor and the three full-subtractors. The half-subtractor and the three full-subtractors are used to perform subtraction calculations on the remaining four digits of the maximum exponent and the sum of exponents respectively.

[0053] For FP16 operands, the exponent part is five bits, with a weight exponent WE and an input feature exponent IE. After adding the two using an exponent adder module, there is a six-bit exponent sum SUM<5:0>, and then a maximum value search module is used to find the maximum value MAX<5:0> among all the six-bit exponent sums. Then, the first two bits of the maximum exponent sum are compared with the first two bits of each exponent bit, that is, to judge MAX<5:4> and SUM<5:4>. MAX<5> and SUM<5> are connected to the input of an AND logic gate, and MAX<4> and SUM<4> are also connected to the input of an AND logic gate in the same way. The outputs of the two AND logic gates are then connected to the input of an AND logic gate, and the output Cn signal is connected to a multiplexer, and the power supply VDD of the subsequent four-bit subtractor is controlled by this signal. If the first two bits of the exponents are the same, that is, MAX<5>=SUM<5> and MAX<4>=SUM<4>, at this time Cn = 1, and the multiplexer selects VDD to connect to the power enable of the four-bit subtractor, so as to calculate the result of MAX<3:0>-SUM<3:0> of this row, which is also the result of MAX<5:0>-SUM<5:0> of this row; if they are different, at this time the judgment signal Cn = 0, indicating that the difference between the exponent sum of this row and the maximum exponent sum is large, and the number of bits that need to be aligned for the mantissa is also large. At this time, the impact on the result after multi-cycle alignment is very small, so this directly controls the multiplexer to turn off the power enable of the subsequent four-bit subtractor and does not allow the subtractor to work. Using the large difference in exponents to control the activation of the subtractor saves some power consumption loss of the subtractor and reduces the operating cycle of the adder tree with little loss of precision.

[0054] The adder tree module includes an AND logic gate, a multi-stage adder tree, and multiple shifters; each stage of the adder tree includes a multiplexer and multiple adders connected in series. At the start of each mantissa product waiting count, each mantissa product corresponds to a zero signal Zn = 0. If the counter does not count to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 1. If the counter counts to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 0. The AND logic gate is used to receive the zero signal Zn, and the output end is connected to the control end of the multiplexer of each stage of the adder tree. In each stage of the adder tree, the enabled output end of the multiplexer is connected to the enabled power supply ends of multiple adders. Multiple shifters are respectively used to perform shift addition on the final calculation result of the multi-stage adder tree in each counting cycle, and the total mantissa is obtained after the counting is completed.

[0055] Based on the circuit structure, as Figure 3 shown, which includes two parts, the upper and the lower, to illustrate the low power consumption and alignment method respectively.

[0056] For the mantissa part of the FP16 operand, there are ten bits, including the weight mantissa WM and the input feature mantissa IM. Each row in the weight mantissa array is used to pre-store the mantissa part WMn of the weight bit by bit, and then use its own logical operation function to perform a multiplication operation with the mantissa part of the input feature value to obtain the mantissa product. The first ten bits of the mantissa product Pn are taken into the adder; at this time, each mantissa product waits for the counting to start. Each mantissa product Pn corresponds to a zero signal Zn. If the counter has not counted to this exponent bit, the corresponding mantissa Pn will send out a zero signal Zn = 1. If the counter has counted to this exponent bit, the corresponding mantissa Pn will send out a zero signal Zn = 0; as Figure 3 in the example given above, if P0 and P1 enter the first-level adder tree and the exponent bits corresponding to P0 and P1 have not been counted, Z0 and Z1 are both equal to 1. At this time, through ZI0 = 1, the data selector selects VSS, and at this time, the power supply of the serial adder is turned off, and a zero signal ZI0 = 1 is given to the next level; if P0 and P1 enter the first-level adder tree and one of the exponent bits corresponding to P0 and P1 has been counted, Z0 and Z1 are both equal to 0. At this time, through ZI0 = 0, the data selector selects VDD, and at this time, the power supply of the serial adder is turned on. At this time, the sum of the two ten-bit multiplier products P0<9:0> and P1<9:0> for eleven bits ADD<10:0> is sent to the next level, and a zero signal ZI0 = 0 is given to the next level; and so on. For the addition of 64 10-bit results, that is, 32 pairs of multiplier products in the first level, 32 11-bit results are output, and then added to get 16 12-bit results, followed by 8 13-bit results, until the final 16-bit result; according to verification, only need to set the above zero signal to the third level, and it can be utilized. Due to the large scale of the adder, such a design can reduce the large static power consumption of the adder.

[0057] Moreover, it is different from the traditional method of using a barrel shifter or a serial shifter to shift and sum, as Figure 3 below, the obtained final 16-bit result A0<15:0> is shifted and added. It is shifted and added once per cycle, and each shift is equivalent to calculating the sum of the mantissas corresponding to different exponent bits; when the counting is completed, that is, when the count reaches the maximum value, it is equivalent to ending the operation. The result of the final shift adder is the result of the operation of the embodiment; then it is normalized to the mantissa of FP16, and the maximum exponent value is the exponent of the calculated FP16 result. The combination forms the final result, and the calculation is completed.

[0058] Please refer to Figure 4, first, the four - digit counter starts counting to match the last four digits of the subtraction result. When the same exponent and the corresponding mantissa are matched, they enter the adder tree, and other data are set to zero. After obtaining the result, if the count has not reached the maximum value, the counting continues, and the result is shifted, waiting for the mantissa corresponding to the next exponent bit to come in for addition. This loop continues until the count reaches the maximum value.

[0059] In this embodiment, it is divided into two stages:

[0060] Initial stage: First, the exponent input module, the weight exponent array, and the exponent adder module are used to complete the input and addition of the exponent parts of each input eigenvalue and weight to obtain the exponent sum. Then, the maximum - value search module reads the exponent - sum calculation results of each row in the exponent adder module bit - by - bit in ascending order. At the same time, the mantissa input module and the weight mantissa array complete the input and multiplication of the mantissa parts of each input eigenvalue and weight to obtain each mantissa product.

[0061] Operation stage: The maximum - value search module writes the read exponent - sum calculation results back to the exponent - sum array, and determines the maximum exponent while writing back. Then, in the subtractor and counter module, the maximum exponent sum is matched with each exponent sum in the exponent - sum array, and the exponent sums with the first two digits the same as the maximum exponent sum are screened out. After matching, the subtractor is used to perform subtraction calculation on these qualified exponent sums to obtain the digit difference. At this time, the counter starts counting. When the period is the same as the digit difference, the corresponding mantissa product enters the adder tree. The calculation results entering the adder tree in each period are shifted and added at the tail. After the counter finishes counting, the total mantissa is obtained. Finally, the combination of the maximum exponent sum and the total mantissa calculated by the adder tree is the corresponding multiply - accumulate result.

[0062] Compared with the existing in - memory computing circuits or chips, the SRAM - based floating - point multiply - accumulate operation circuit of this embodiment has the following beneficial effects:

[0063] 1. For the SRAM - based floating - point multiply - accumulate operation circuit, by proposing a new screening and alignment operation, it can achieve floating - point calculations with very little precision loss. At the same time, it can eliminate the area loss of redundant registers and the power consumption loss of the adder tree. Therefore, its application prospect is very broad. The circuit structure is simple and has great advantages in terms of area and power consumption, solving the technical problem of large losses in the existing in - memory computing circuits or chips during floating - point multiply - accumulate operations.

[0064] 2. For the SRAM - based floating - point multiply - accumulate operation circuit, it uses a subtractor to screen the exponent sums, and uses a counter and an adder tree to align and sum the mantissas. By improving the adder tree, it can not only eliminate the consumption of a large number of shift registers but also reduce the calculation loss of the adder tree.

[0065] Example 3

[0066] This embodiment provides an in-memory computing chip (CIM chip), which includes the SRAM-based floating-point multiply-accumulate operation circuit in Embodiment 1 or Embodiment 2, and this circuit can be integrated on the chip. The CIM chip has a storage mode and a computing mode. In the storage mode, the CIM chip is used as a memory. In the computing mode, the CIM chip is used to implement the multiply-accumulate operation between multiple groups of multi-bit floating-point input features and multi-bit floating-point weights.

[0067] Example 4

[0068] This embodiment provides a static random access memory (SRAM), and this memory uses the SRAM-based floating-point multiply-accumulate operation circuit in Embodiment 1 or Embodiment 2 to implement the multiply-accumulate calculation of multi-bit inputs and multi-bit weights.

[0069] Based on the SRAM-based floating-point multiply-accumulate operation circuit in Embodiment 1 or Embodiment 2, the in-memory computing of the SRAM in this embodiment directly completes the multiply-accumulate operation in the storage unit, reduces data movement, and significantly reduces power consumption. The SRAM can simultaneously process the multiply-accumulate operations of multiple inputs and weights, greatly improving the computing efficiency, and is particularly suitable for application scenarios that require high throughput. The read and write speeds of the SRAM are much higher than those of DRAM and flash memory, and can achieve low-latency multiply-accumulate calculations, making it suitable for applications with high real-time requirements, such as edge computing and Internet of Things devices. The SRAM can be integrated with other computing units (such as CPU, GPU) on the same chip to form an efficient in-memory computing architecture.

[0070] The SRAM in this embodiment is applicable to artificial intelligence and machine learning. The inference and training processes of neural networks involve a large number of multiply-accumulate operations, and the in-memory computing of the SRAM can significantly accelerate these operations and improve the overall performance. The multi-bit inputs and weights enable the SRAM to support from simple linear models to complex deep neural networks. The in-memory computing of the SRAM in this embodiment reduces the complex interface between the memory and the processor in the traditional computing architecture and simplifies the system design. By reducing data movement and simplifying the architecture, the SRAM can reduce the overall cost and power consumption of the system.

[0071] Example 5

[0072] This embodiment provides an electronic device, which includes a memory and a processor. Among them, the memory includes the SRAM-based floating-point multiply-accumulate operation circuit in Embodiment 1 or Embodiment 2. Compared with existing electronic devices, this electronic device can significantly improve the computing efficiency, reduce power consumption, and support high-precision computing. It has broad application prospects in fields such as artificial intelligence and edge computing. Although it faces some technical challenges, its advantages make it an important technical direction for in-memory computing.

[0073] Embodiment 6

[0074] This embodiment provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. Among them, the memory is the static random access memory in Embodiment 4.

[0075] This computer device can take various forms. It can be an embedded chip or module, or a general-purpose data processing device, such as an intelligent terminal capable of executing programs, a tablet computer, a laptop computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc.

[0076] The computer device of this embodiment at least includes, but is not limited to, a memory and a processor that can communicate with each other through a system bus. The memory (i.e., a readable storage medium) includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as the hard disk or memory of the computer device.

[0077] In some embodiments, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data.

[0078] Embodiment 7

[0079] This embodiment provides a method for SRAM-based floating-point multiply-accumulate operation. This method can use the circuit in Embodiment 1 or Embodiment 2, and specifically includes the following steps:

[0080] (1) Calculate the sum of exponents and the product of mantissas of multi-bit input values and multi-bit weight values that are all floating-point types bit by bit respectively;

[0081] (2) First, sort the calculation results of the sum of exponents and determine the maximum exponent. Then, screen out the sum of exponents that has the same first two digits as the maximum exponent, and eliminate the product of mantissas corresponding to the sum of exponents with different first two digits. Finally, calculate the bit difference between the remaining sum of exponents and the remaining number of digits of the maximum exponent, and start counting. When reaching the same period as the bit difference, input the corresponding product of mantissas into the adder tree;

[0082] (3) Add the products of mantissas in the adder tree, and perform shift addition on the final result at the tail of the adder tree every period. After the counting is completed, obtain the total mantissa;

[0083] (4) Combine the maximum exponent sum and the total mantissa to obtain the multiply-accumulate calculation result of the multi-bit input value and the multi-bit weight value.

[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An SRAM-based floating-point multiply-accumulate operation circuit, characterized in that, It is used for: (1) Calculating the sum of exponents and the product of mantissas of multi-bit input values and multi-bit weight values that are all floating-point types bit by bit respectively; (2) First, sorting the calculation results of the sum of exponents to determine the maximum exponent, then screening out the sum of exponents that has the same first two bits as the maximum exponent, and screening out the product of mantissas corresponding to the sum of exponents with different first two bits. Finally, calculating the bit difference between the remaining sum of exponents and the remaining bits of the maximum exponent, and starting to count. When reaching the same period as the bit difference, input the corresponding product of mantissas into the adder tree; (3) Adding the products of mantissas in the adder tree, and performing a shifted addition on the final result at the tail of the adder tree every period. After the counting is completed, the total mantissa is obtained; (4) Combining the maximum exponent sum and the total mantissa to obtain the multiply-accumulate calculation result of the multi-bit input value and the multi-bit weight value.

2. The floating-point multiply-accumulate operation circuit based on SRAM according to claim 1, wherein The arithmetic circuit includes: An array circuit, which includes a weight exponent array, a sum-of-exponents array, and a weight mantissa array; each row of the weight exponent array is used to pre-store the exponent part of the multi-bit weight value bit by bit; the sum-of-exponents array is used to store the calculation result of the sum of exponents; each row of the weight mantissa array is used to pre-store the mantissa part of the multi-bit weight value bit by bit, and perform a multiplication operation on the mantissa parts of the multi-bit input value and the multi-bit weight value to obtain the product of mantissas.

3. The floating-point multiply-accumulate operation circuit based on SRAM according to claim 2, wherein The arithmetic circuit further includes: An exponent input module; An exponent adder module, which is used to first receive the exponent part of the multi-bit input value input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the sum of the two exponent parts, and finally store the calculation result of the sum of exponents into the sum-of-exponents array.

4. The SRAM-based floating-point multiply-accumulate operation circuit according to claim 3, wherein The arithmetic circuit further includes: A maximum value finding module, which is used to first read the calculation result of the sum of exponents of the exponent adder module, and then screen out the maximum exponent from the highest bit to the lowest bit in turn, and write back the sum of exponents to the sum-of-exponents array.

5. The floating-point multiply-accumulate operation circuit based on SRAM according to claim 4, wherein The arithmetic circuit further includes: A subtractor and counter module, which includes a subtractor and a counter; the subtractor and counter module is used to match the first two bits of the sum of exponents in each row of the sum-of-exponents array with the first two bits of the maximum exponent, screen out the sum of exponents that has the same first two bits as the maximum exponent, and screen out the product of mantissas corresponding to the sum of exponents with different first two bits, make the counter start to count, and input the corresponding product of mantissas into the adder tree every time it reaches the same period as the bit difference.

6. The SRAM-based floating-point multiply-accumulate operation circuit according to claim 5, wherein The arithmetic circuit further includes: An adder tree module; the subtractor and counter module inputs the corresponding product of mantissas into the adder tree module every time the count of the counter reaches the same period as the bit difference. The adder tree module is used to add the received products of mantissas, and perform a shifted addition on the final result at the tail of the adder every counting period. After the counting is completed, the total mantissa is obtained.

7. The floating-point multiply-accumulate operation circuit based on SRAM according to claim 5, wherein The subtractor includes a judge, a selector, a half-subtractor, and three full-subtractors; the judge includes three AND logic gates: a first AND logic gate, a second AND logic gate, and a third AND logic gate; the first AND logic gate is used to perform an AND operation on the first bit of the maximum exponent and the first bit of the exponent sum, and the second AND logic gate is used to perform an AND operation on the second bit of the maximum exponent and the second bit of the exponent sum; the third AND logic gate performs an AND operation on the operation results of the first AND logic gate and the second AND logic gate, and inputs the operation result to the selector; the output end of the selector is connected to the power enable ends of the half-subtractor and the three full-subtractors, and the half-subtractor and the three full-subtractors are used to perform subtraction calculations on the remaining four bits of the maximum exponent and the exponent sum respectively.

8. The SRAM-based floating-point multiply-accumulate operation circuit according to claim 6, wherein The adder tree module includes an AND logic gate, multiple levels of adder trees, and multiple shifters; each level of adder tree includes a selector and multiple adders connected in series; at the start of counting for each mantissa product waiting count, each mantissa product corresponds to a zero signal Zn = 0. If the counter does not count to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 1. If the counter counts to the corresponding exponent bit, the corresponding mantissa emits a zero signal Zn = 0; the AND logic gate is used to receive the zero signal Zn, and its output end is connected to the control ends of the selectors of each level of adder tree; in each level of adder tree, the enable output end of the selector is connected to the enable power supply ends of multiple adders; multiple shifters are respectively used to perform shift addition on the final calculation result of the multiple levels of adder trees in each counting period, and obtain the mantissa sum after the counting is completed.

9. The floating-point multiply-accumulate operation circuit based on SRAM according to claim 2, wherein The arithmetic circuit is used to implement FP16 calculation, and the column ratio of the weight exponent array, the exponent sum array, and the weight mantissa array is 6:6:

10.

10. An in-memory computing chip, characterized in that, It includes the SRAM-based floating-point multiply-accumulate arithmetic circuit according to any one of claims 1-9.