In-memory calculation circuit and floating-point number multiply-accumulate operation method thereof

By using the floating point multiplication and accumulation operation method in the in-memory calculation circuit, the bit stream form is used to compare and make differences, and in the process of finding the maximum value, it solves the problems of large area overhead, high delay and power consumption of floating point multiplication and accumulation CIM circuits in the prior art, and achieves more efficient operations.

CN119937982APending Publication Date: 2025-05-06ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510034818.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing floating-point number multiplication and accumulation CIM circuits have problems with large area overhead, high latency and high power consumption, especially in the von Neumann bottleneck caused by a large number of floating-point number multiplication and accumulation operations in convolutional neural networks.

Method used

A floating point number multiplication and accumulation operation method of in-memory computing circuit is proposed. By processing the exponent and mantissa separately, comparing and making differences in bitstream form, reducing the power consumption of the adder circuit and operation, and determining whether the data is less than the accuracy in the process of finding the maximum value, it is directly regarded as 0 to reduce unnecessary operations.

Benefits of technology

It effectively reduces the area overhead and operation delay of the circuit, reduces the overall power consumption, and improves the efficiency of floating-point number multiplication and accumulation operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937982A_ABST
    Figure CN119937982A_ABST
Patent Text Reader

Abstract

The invention relates to an in-memory computing circuit and a floating-point number multiply-accumulate operation method thereof in the field of integrated circuit design. The method comprises the following steps: adding an index weight EWE of an index in stored original data to an index input EIN to obtain an index and an ESUM; the method comprises the following steps: taking a high k bit in an ESUM to obtain an index and a high k bit ESUMH, selecting a maximum value from all the ESUMHs as the output of the index and the maximum high k bit EMAXH, judging that an output control signal is 1 when a difference value between the EMAXH and the ESUMH is smaller than 2n, taking a low x bit in the ESUM to obtain an index and a low x bit ESUML, finding out the maximum value from all the ESUMLs in a bit stream form to serve as the maximum low x bit EMAXL to output, the EMAXL and the ESUML are used as alignment shift amount ESUB, MIN * MWE is used as mantissa product MPD, MIN is mantissa input, and MWE is mantissa weight for storing mantissa in original data; and shifting the stored data according to the ESUB to obtain the stored data MOUT. According to the invention, a large amount of power consumption during adder circuits and operation can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an in-memory computing circuit and a floating point exponent processing method thereof in the field of integrated circuit design, and in particular to a floating point multiplication and accumulation operation method of an in-memory computing circuit. Background Art

[0002] With the rapid development and popularization of artificial intelligence, convolutional neural networks (CNN) and deep neural networks (DNN) have become one of the most influential innovations in the field of computer vision. Neural networks such as CNN and DNN need to perform a large number of multiplication and multiply-accumulate (MAC) operations for data processing. When such operations are processed in computers based on the von Neumann architecture, the need to frequently move data between the processor and the memory results in high energy consumption and latency, a problem known as the von Neumann bottleneck or memory wall. Demonstrations of DNN processors and accelerators based on the von Neumann architecture show that energy consumption and latency depend mainly on the input data between the processor and the memory. Therefore, traditional von Neumann computers are not suitable for processing artificial intelligence-related computing tasks such as neural networks.

[0003] In order to overcome the von Neumann bottleneck and process artificial intelligence-related computing tasks such as neural networks, technicians in this field have proposed a memory-based computing-in-memory (CIM) architecture. This new computer architecture directly uses memory to implement logical operations, without the need to move data between the memory and the processor, thus significantly improving data processing efficiency and reducing device power consumption.

[0004] However, convolutional neural networks contain a large number of floating-point multiplication and accumulation operations. The existing floating-point multiplication and accumulation CIM circuits have three main characteristics when implementing floating-point multiplication and accumulation tasks: First, it is necessary to find the maximum exponent sum first, and then subtract the maximum value from all exponent sums to obtain the shift required to align the mantissa part, and then start the shift operation, which will significantly increase the delay of the circuit operation process. Second, the difference of the exponent bits in the floating-point operation requires a large number of adders, and registers are also required to store the difference, which will significantly increase the area overhead of the circuit operation process. Third, the multiplication operation of the mantissa part consumes a lot of power, but the mantissa shift of a large amount of data exceeds the length of the mantissa, which will significantly increase the overall power consumption of the circuit operation. Summary of the invention

[0005] In order to solve the problems of large area overhead, high delay and high power consumption commonly existing in various types of CIM circuits with floating-point multiplication and accumulation functions, the present invention provides an in-memory computing circuit and a floating-point multiplication and accumulation operation method thereof, wherein the input weight bits are configurable.

[0006] The object of the present invention is achieved through the following technical scheme: a floating point multiplication and accumulation operation method of an in-memory computing circuit is used to input exponents of length j and n into E IN and the mantissa input M IN Perform floating-point multiplication and accumulation operations to obtain the exponent and the maximum value high k bits EMAX H , exponent and maximum value low x bits EMAX L , and store data M OUT , k = j + 1 - x, 2 x ≥2n, the method comprises the following steps:

[0007] For E IN Add the exponential weight E of the exponent in the original data WE Get the exponent and E SUM ;

[0008] Take E SUM The high k bits in get the exponent and the high k bits E SUMH , in all E SUMH Select the maximum value as EMAX H Output and determine EMAX H With E SUMH When the difference is less than 2n, the output control signal is 1;

[0009] When the control signal is 1, take E SUM The low x bits in get the exponent and the low x bits E SUML , in bitstream form, in all E SUML Find the maximum value as EMAX L Output and calculate EMAX L With E SUML The difference is the alignment shift E SUB ;

[0010] Multiply the mantissa by M PD When the control signal is 1, take M IN ×M WE , M WE To store the mantissa weight of the mantissa in the original data;

[0011] According to E SUB Shift the stored data to get M OUT .

[0012] As a further improvement of the above solution, the method comprises the following steps: EMAX H With E SUMH If the difference is greater than or equal to 2n, the output control signal is 0, and there is no need to calculate the corresponding alignment shift amount E in the future SUB and the mantissa, M out That is, take 0.

[0013] The present invention also discloses an in-memory computing circuit for inputting exponents of lengths j and n into E IN and the mantissa input M IN Perform floating-point multiplication and accumulation operations to obtain the exponent and the maximum value high k bits EMAX H , exponent and maximum value low x bits EMAX L , and store data M OUT , k = j + 1 - x, 2 x ≥2n, the in-memory computing circuit includes:

[0014] Exponential weight SRAM array, used to store the exponential weight E of the exponent in the original data WE ;

[0015] Addition module, whose exponent and E SUM For E IN +E WE ;

[0016] index and SRAM array to store E SUM ;

[0017] The high k-bit value search judgment module is used to obtain E SUM The high k bits in get the exponent and the high k bits E SUMH , in all E SUMH Select the maximum value as EMAX H Output and determine EMAX H With E SUMH If the difference between and 2n is greater than or equal to , the output control signal is 0, otherwise it is 1;

[0018] The low x bit bit stream conversion and value search module is used to take E when the control signal is 1. SUM The low x bits in get the exponent and the low x bits E SUML , in bitstream form, in all E SUML Find the maximum value as EMAX L Output and calculate EMAX L With E SUML The difference is the alignment shift E SUB ;

[0019] Mantissa weight SRAM array, used to store the mantissa weight M of the mantissa in the original data WE ;

[0020] The multiplication module, whose mantissa product M PD When the control signal is 1, take M IN ×M WE , when the control signal is 0, it takes 0;

[0021] Mantissa product shift register circuit, used to store M PD , and according to E SUB Shift the stored data to get M OUT .

[0022] As a further improvement of the above solution, the in-memory computing circuit further includes a peripheral circuit, and the peripheral circuit includes:

[0023] A word line driver, used to control the opening of the word lines of each row in the SRAM array;

[0024] An address decoder is connected to the word line driver and is used to decode the address signal and transmit it to the word line driver;

[0025] A precharge circuit, used for precharging the bit lines of each column in the SRAM array;

[0026] A timing control module is used to generate various clock signals required to perform data storage tasks or logic operations;

[0027] The read / write selection module is used to select each SRAM cell in the SRAM array that needs to perform a read / write operation.

[0028] Furthermore, the in-memory computing circuit further includes a peripheral circuit, and the peripheral circuit further includes:

[0029] The mode switching circuit is used to switch the working mode of the in-memory computing circuit.

[0030] As a further improvement of the above scheme, k is 4, x is 2, j is 5, and n is 6; the addition module adopts a 1-bit full adder, and the 1-bit full adder performs the addition operation of the 5-bit exponent weight and the 5-bit exponent input within 5 cycles, and the output 6-bit exponent and result are written into the index and SRAM array, the W end of the full adder is connected to the bit line EBL of the exponent weight SRAM array, the I end is input from the outside, the S end is connected to the word line SWL of the exponent and SRAM array, the CI end is connected to the read circuit controlled by CIWL in the high-k-bit value search and judgment module, and the CO end is connected to the write circuit controlled by COWL in the high-k-bit value search and judgment module.

[0031] Furthermore, the high k-bit value search judgment module includes 4 CBLs from high to low, connecting the units representing the same bit in the 128 numbers, and the Q point values ​​of these units and the signal of CWL jointly control the connection between CBL and the low level. The upper part of CBL is connected to the precharge circuit controlled by PCH, and the lower part is connected to an inverter, and then connected to a register for storing the maximum value MAX of the current bit; the QB points of these units are connected to a NAND gate with MAX, and the output is connected to an AND gate with CWL, and the output result serves as the CWL of the next bit; starting from the second highest bit, the QB points of each unit are connected to a NOR gate with MAX, and the output is connected to an OR gate with CWL, as a judgment on whether the low bit of the exponent and the mantissa of the stored data are to be processed subsequently.

[0032] Preferably, the Q and QB point signals of the two units in each column represented by the low x bit bit stream conversion and value search and difference module are connected to a plurality of logic gates with a common signal source, and then each bit stream signal is generated in the same timing sequence, and then combined with the judgment signal generated by the high k bit value search and judgment module, the low 2 bits E are outputted through a circuit composed of a plurality of OR gates and XOR gates. SUML The corresponding bit stream and the mantissa product shift signal corresponding to each number.

[0033] Preferably, the multiplication module includes 4 AND gates and 2 half adders, the weight and the low bit of the input are connected to the AND gate, the low bit of the weight and the low bit of the input, the high bit of the weight and the low bit of the input are respectively connected to the AND gate and then connected to the first half adder, the weight and the high bit of the input are connected to the AND gate, and then connected to the second half adder with the carry of the first half adder.

[0034] Preferably, the mantissa product shift register circuit includes 4 groups of D flip-flops and two-to-one MUX, the input of each D flip-flop is connected to the single output of the two-to-one MUX, one end of the input of the two-to-one selector is connected to the output of the multiplication module, and the other end is connected to the previous level D flip-flop, and the CLK signals of all D flip-flops are generated by the system timing signal and the corresponding shift signal through the AND gate.

[0035] Compared with the traditional in-memory computing circuit, the present invention has the following advantages:

[0036] 1. The traditional calculation of alignment shift amount RSA requires a large number of adders. Using the bit stream form of the present invention to compare sizes and make differences can reduce a large number of adder circuits and power consumption during operations.

[0037] 2. The traditional calculation process needs to calculate RSA first, and then use RSA to generate a shift signal to act on the mantissa shift circuit. Using the bit stream form of the present invention, because the bit stream itself is also a signal represented by high and low levels, it can directly act on the shift circuit at the beginning stage of generating the bit stream.

[0038] 3. In the traditional calculation process, a large amount of initial data will become 0 in the later stage because the actual value is less than the precision, but it still participates in the calculation process, which increases power consumption. In the process of finding the maximum value, the high and low bits are processed separately, and in the process of processing the high bit, it is judged whether the data is less than the precision, and a control signal is obtained, which is directly regarded as 0 without subsequent calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The present invention is a flowchart of a floating-point multiplication and accumulation method of an in-memory computing circuit provided in Embodiment 1 of the present invention.

[0040] Figure 2 To achieve Figure 1 Circuit architecture diagram of the in-memory computing circuit for the floating-point multiplication and accumulation method.

[0041] Figure 3 A circuit architecture diagram of the in-memory computing circuit provided in Example 2 of the present invention.

[0042] Figure 4 for Figure 3 Circuit diagram of the SRAM array based on 6T-SRAM cells in the in-memory computation circuit, with a 5×128 exponent bit array on the top and a 2×128 mantissa bit array on the bottom.

[0043] Figure 5 for Figure 4 Circuit diagram of the 6T-SRAM unit used in the SRAM array.

[0044] Figure 6 for Figure 3 The circuit diagram of the module that completes 5bit+5bit addition within 5 cycles is shown in the figure.

[0045] Figure 7 for Figure 3 Timing diagram of the addition module at work.

[0046] Figure 8 for Figure 3 Circuit diagram of the middle and high 4-bit value search and judgment module.

[0047] Fig. 9 for Figure 3 Bitstream conversion circuit diagram of the low and middle 2-bit bitstream conversion and value search module.

[0048] Fig.10 for Figure 8 Circuit diagram of the Bitstream maximum value search and difference circuit.

[0049] Fig.11 for Figure 3 Circuit diagram of the 2bit×2bit multiplication module implemented in .

[0050] Fig.12 for Figure 3 The circuit diagram of the mantissa product shift register circuit that realizes 4-bit mantissa product shift. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0052] Example 1

[0053] See also Figure 1 , which is a flow chart of a floating-point multiplication and accumulation method of an in-memory computing circuit provided in Embodiment 1 of the present invention. The floating-point multiplication and accumulation method is used to input exponents of lengths j and n into E IN and the mantissa input M IN Perform floating-point multiplication and accumulation operations to obtain the exponent and the maximum value high k bits EMAX H , exponent and maximum value low x bits EMAX L , and store data M OUT , k = j + 1 - x, 2 x ≥2n. In this embodiment, please combine Figure 2 The in-memory calculation circuit includes an exponent weight SRAM array, an addition module, an exponent and SRAM array, a high k-bit value search judgment module, a low x-bit bit stream conversion and value search difference module, a mantissa weight SRAM array, a multiplication module, and a mantissa product shift register circuit, and may also include peripheral circuits, input modules, output modules, etc. This in-memory calculation circuit is one of the upper embodiments of the floating-point multiplication and accumulation operation method of the present invention. For detailed numerical lower embodiments, please refer to Embodiment 2.

[0054] The method comprises the following steps S1 to S2:

[0055] Step S1, input E for the index IN Add the exponential weight E of the exponent in the original data WE Get the exponent and E SUM Step S1 is used to achieve: E IN +E WE =E SUM Step S1 can be performed by an addition module. The exponential weight E used in step S1 WE Can be stored in the exponential weight SRAM array. The exponential weight SRAM array implements the basic SRAM storage function and stores the weight part E of the exponent in the original data. WE .

[0056] Step S2, take the index and E SUM The high k bits in get the exponent and the high k bits E SUMH , at all exponents and high k bits E SUMH Select the maximum value as the index and the maximum k bits EMAX H Output.

[0057] Step S3, determine EMAX H With E SUMH Is the difference less than 2n? If EMAX H With E SUMH When the difference is less than 2n, step S4 is executed: the output control signal is 1. Otherwise, if it is determined that EMAX H With E SUMH When the difference is greater than or equal to 2n, step S5 is executed: the output control signal is 0. Steps S2 and S3 can be executed by the high k bit value search judgment module. The exponent used in step S2 and the high k bit E SUMH It can be stored by index and SRAM array. Index and SRAM array realize basic SRAM storage function, storing calculation intermediate quantity - E SUM The high k-bit value search module takes E SUM The high k bits E SUMH , in all E SUMH Select the maximum value EMAX H And output to the output module, and then judge EMAX H -E SUMH If the value is greater than or equal to 2n, a control signal 0 is generated, and the corresponding low-order exponent and mantissa do not need to be calculated subsequently, and are directly regarded as 0 (i.e., the corresponding alignment shift amount E does not need to be calculated subsequently). SUB and the mantissa, M out If it is less than 0, a control signal 1 is generated, and the corresponding data participates in the calculation normally. Therefore, after step S5, the method executes step S9: EMAX H With E SUMH If the difference is greater than or equal to 2n, the output control signal is 0, and there is no need to calculate the corresponding alignment shift amount E in the future SUB and the mantissa, M out That is, take 0.

[0058] Step S6: When the control signal is 1, take E SUM The low x bits in get the exponent and the low x bits E SUML , in bitstream form, in all E SUML Find the maximum value as EMAX L Output and calculate EMAX L With E SUML The difference is the alignment shift E SUBStep S6 can be executed by the low x bit bit stream conversion and value search and difference module, which processes the data when the control signal is 1 and takes E SUM The low x bit E in SUML , converted into Bitstream format, the data with control signal 0 is directly regarded as a bitstream of all 0s without conversion. Then compare the size in bitstream format to find the maximum value EMAX L And calculate EMAX L -E SUML Get the alignment shift amount E SUB .

[0059] Step S7, multiply the mantissa by M PD When the control signal is 1, take M IN ×M WE , M WE The mantissa weight of the mantissa in the original data is stored. Step S7 can be performed by a multiplication module. The M used by the multiplication module WE Can be stored in the mantissa weight SRAM array, the mantissa weight SRAM array realizes the basic SRAM storage function, and stores the mantissa weight M of the mantissa in the original data WE When the control signal of the multiplication module is 1, M IN ×M WE =M PD ; When the control signal is 0, no calculation is required, corresponding to M PD is 0.

[0060] Step S8, according to E SUB Shift the stored data to get M OUT Step S8 can be implemented by a mantissa product shift register circuit, which stores the calculated intermediate quantity - mantissa product M. PD , and accepts the alignment shift amount E SUB The signal shifts the stored data to obtain the new stored data M OUT .

[0061] The following data should be the same except EMAX is unique, set to m:

[0062] Index Input E IN , length j;

[0063] Index Weight E WE , length j;

[0064] Index and E SUM , length j+1;

[0065] Exponential and high k bits E SUMH , length j+1-x;

[0066] Exponent and maximum value high k bits EMAX H , length j+1-x;

[0067] Control signal, length 1 (detect EMAX H -E SUMH Is it ≥2n? If it is greater than or equal to 2n, there is no need to calculate the low-order exponent and the mantissa part is directly regarded as 0);

[0068] Exponent and low x bits E SUML , length x (x should satisfy 2 x ≥2n);

[0069] Exponent and maximum value low x bits EMAX L , length x;

[0070] Alignment shift amount E SUB (The traditional analog is RSA), E SUB =EMAX L -E SUML ;

[0071] Enter the last digit as M IN , length n;

[0072] Mantissa weight M WE , length n;

[0073] Mantissa product M PD , length 2n;

[0074] Align the mantissa product M OUT , length 2n.

[0075] The RAM array cooperates with the peripheral circuit to realize the data storage function of the SRAM circuit, and the SRAM array cooperates with the other parts to realize the floating-point multiplication and accumulation operation in the FP8 format. The various calculation circuits and SRAM circuits in the present invention can be executed to realize the floating-point multiplication and accumulation calculation processing in the FP8 format. The working principle of the circuit is different from that of the existing circuit, and can overcome the problems of large area overhead, high delay and high power consumption that are common in the existing circuits.

[0076] Advantages compared to traditional structures:

[0077] (1) The traditional calculation of alignment shift amount RSA requires a large number of adders. Using the bit stream form of the present invention to compare sizes and make differences can reduce the power consumption of a large number of adder circuits and operations.

[0078] (2) The traditional calculation process needs to calculate RSA first, and then use RSA to generate a shift signal to act on the mantissa shift circuit. Using the bit stream form of this patent, because the bit stream itself is also a signal represented by high and low levels, it can directly act on the shift circuit at the beginning stage of generating the bit stream.

[0079] (3) In the traditional calculation process, a large amount of initial data will become 0 in the later stage because the actual value is less than the precision, but it still participates in the calculation process, which increases power consumption. In the process of finding the maximum value, the high and low bits are processed separately. In the process of processing the high bit, it is determined whether the data is less than the precision, and a control signal is obtained. It is directly regarded as 0 without subsequent calculation.

[0080] Example 2

[0081] The in-memory computing circuit provided in this embodiment 2 is a specific example of embodiment 1, where k is 4, n is 2, m is 128, and j is 5, and a detailed example is given.

[0082] See also Figure 3 , the in-memory calculation circuit is designed based on the traditional SRAM circuit, which includes the SRAM array in the SRAM circuit and its corresponding various peripheral circuits for realizing the data storage function. And other circuit modules are added on the basis of the SRAM circuit to realize the floating-point multiplication and accumulation (MAC) operation with FP8 format. According to the functional division, in addition to the SRAM array and the peripheral circuit, the in-memory calculation circuit also includes: exponent weight SRAM array, addition module, exponent and SRAM array, high 4-bit value search judgment module, low 2-bit bit stream conversion and value search difference module, mantissa weight SRAM array, multiplication module, mantissa product shift register circuit, peripheral circuit, input module, output module. Among them, the SRAM array and the peripheral circuit can realize the data storage function of the SRAM circuit, and the SRAM array with the other parts can realize the floating-point multiplication and accumulation operation in FP8 format.

[0083] In the present invention, the peripheral circuit mainly includes a word line driver, an address decoder, a precharge circuit, a timing control module, a read-write selection module, etc. Among them, the word line driver is used to control the opening of the word line WL of each row in the SRAM array. The address decoder is connected to the word line driver, and the address decoder is used to decode the address signal and transmit it to the word line driver. The precharge circuit is used to perform precharge operations on signal lines such as bit lines BL and BLB. The timing control module is used to generate various clock signals required to perform data storage tasks or logical operations. The read-write selection module is used to select each SRAM unit in the SRAM array that needs to perform read and write operations. In addition, considering that the in-memory calculation circuit in this embodiment has both data storage and logical operations, the peripheral circuit should also include a mode switching circuit, and the mode switching circuit is used to switch the working mode of the in-memory calculation circuit.

[0084] In the circuit scheme of the present invention, as Figure 4 As shown, the SRAM array can be mainly constructed by 6T-SRAM cells, which are divided into a 5×128 exponent bit array and a 2×128 mantissa bit array, which are used to store the exponential weight E WE and tail weight M WE .

[0085] In the circuit scheme of the present invention, as Figure 5 As shown, the 6T-SRAM cell includes two NMOS tubes N1 and N2, and two inverters INV0 and INV1. The slave device is mainly composed of two PMOS tubes P1-P2 and four NMOS tubes N1-N4. The circuit connection relationship is as follows: P1, P2, N3, and N4 form an anti-phase cross-coupled data latch structure, which includes two anti-phase storage nodes Q and QB; the source of N1 and N2 is connected to Q and QB respectively; the drain of N1 and N2 is connected to the bit line BL and BLB respectively, the gate of N1 and N2 in the index weight array is connected to the word line WL, and the gate of N1 and N2 in the mantissa weight array is connected to the word line WLL, WLR. The word line controls the connection state between the storage nodes Q and QB in the 6T-SRAM cell and the bit lines BL and BLB on the corresponding side. The input end of INV0 and the output end of INV1 are connected to the source of N1 and serve as the storage node Q; the output end of INV0 and the input end of INV1 are connected to the source of N2 and serve as the storage node QB; the drains of N1 and N2 are connected to the bit lines BL and BLB respectively, and the gates of N1 and N2 are connected to the word lines WL in the index weight array, and the gates of N1 and N2 are connected to the word lines WLL and WLR in the mantissa weight array.

[0086] In the in-memory calculation circuit provided in this embodiment, the SRAM array cooperates with the peripheral circuit to realize the reading, writing and holding operations of data on the one hand; on the other hand, each column of SRAM cells contained in it is used to store an operand in the FP8 format in the multiplication operation, and cooperates with the calculation module to realize the floating-point multiplication and accumulation operation in the FP8 format. The SRAM array is used to cooperate with the peripheral circuit to realize the data storage function; the array is divided into a 5-bit exponential weight SRAM array and a 2-bit mantissa weight SRAM array. The transmission tubes on both sides of each SRAM cell located in the same row in the exponential weight SRAM array are connected to the same group of word lines EWL, and the transmission tubes on both sides of each SRAM cell located in the same row in the mantissa weight SRAM array are connected to the same group of word lines MWLL and MWLR. On the one hand, the SRAM array cooperates with the peripheral circuit to realize the reading, writing and holding operations of data; on the other hand, each SRAM cell contained in it is used to store one of the bits of the floating-point weight in the FP8 format.

[0087] like Figure 6 As shown, the addition module can be mainly composed of a 1-bit full adder with the W end connected to the EBL of the corresponding index, the I end connected to the external input, the S end connected to the SWL of the corresponding index and array, the CI connected to the read circuit controlled by CIWL in the index and the highest bit unit, and the CO connected to the write circuit controlled by COWL in the index and the highest bit unit. The unit performs the addition operation of the 5-bit index weight and the 5-bit index input in 5 cycles, and the output 6-bit index and result are written into the index and array. The timing is as follows Figure 7As shown, in the initial state, SUM <5> The internal value is 0. In the first cycle, the weight and external input enter the full adder, and then SWL <0> and COWL are at high level, the gate of the transmission tube connected to the two word lines is turned on, and the carry is written into SUM <5> , and bits are written to SUM <0> ; In the second cycle, the weights and external inputs are updated, CIWL is at a high level first, SUM <5> As the carry of the previous bit, the operand enters the full adder, and then SWL <1> and COWL become high, the gate and the transmission tube connected to the two word lines are turned on, and the carry is written into SUM <5> , and bits are written to SUM <1> ; Repeat this process for five cycles to get the 6-bit addition sum, which is stored in SUM from high to low. <5> to SUM <0> In. Each column of the exponential weight SRAM array is connected to the addition module. Within 5 cycles, the 5-bit exponential weight and the 5-bit exponential input are added from low to high in a 1-bit full adder. The carry generated in each cycle is stored in the highest bit of the exponent and SRAM array and input into the carry of the next cycle. The sum of the 5 cycles is written to the second highest bit to the lowest bit in the exponent and SRAM array in sequence. The exponent and SRAM array store the calculated 6-bit exponent. The 1-bit full adder performs the addition operation of the 5-bit exponential weight and the 5-bit exponent input within 5 cycles, and the output 6-bit exponent and result are written into the exponent and SRAM array. The W end of the addition unit is connected to the EBL of the corresponding exponent column, the I end is input from the outside, the S end is connected to the SWL of the corresponding exponent and SRAM array, the CI is connected to the readout circuit controlled by CIWL in the index and the highest bit unit, and the CO is connected to the write circuit controlled by COWL in the number and the highest bit unit. The 6-bit exponent and result obtained by adding the 5-bit exponential weight and the 5-bit exponent input are written into the index and SRAM array. Each EBL of the exponent weight array is connected to an addition unit. Within 5 cycles, the 5-bit exponent weight and the 5-bit exponent input are added from low to high in a 1-bit full adder. The carry generated in each cycle is stored in the highest bit of the exponent sum array and input into the carry of the next cycle. The sum of 5 cycles is written into the exponent sum SRAM array from the second highest bit to the lowest bit in sequence.

[0088] The high 4 bits are connected to the value-seeking judgment module to find out the high 4 bits of the maximum value EMAX of all exponents and judge whether the low bits of these exponents and the corresponding mantissa parts need to be processed further. The low 2 bits are connected to the Bitstream conversion and comparison difference module, and the low 2 bits of the maximum value EMAX of all exponents and the difference between each number and EMAX are found to obtain the shift signal of the mantissa product. Each column of the mantissa weight SRAM array is connected to a multiplication module, and the 2-bit mantissa weight is multiplied by the 2-bit mantissa input, and the 4-bit mantissa product result is stored in the shift register module, i.e., the mantissa product shift register circuit. The mantissa product shift register circuit stores the 4-bit mantissa multiplication result, receives the corresponding shift signal, performs shift processing, and then outputs it to the next level output module. The input module includes an input unit, a shutdown management unit, a transmission management unit, and a precharge unit. The input unit is connected to each addition or multiplication module. The input unit is used to input a group of floating-point numbers and the weight calculation stored in the array to the addition or multiplication module. The shutdown management signal is used to generate the enable signal of each transmission gate input into the shutdown control module; the transmission management signal is used to generate the enable signal of each transmission gate input into the transmission control module. The precharge unit is used to precharge CBL to a specified potential when performing a calculation task. The output module includes a counter and an adder tree. The counter counts the bitstream of the lower 2 bits of the exponent and EMAX to obtain the binary result of the complete exponent and EMAX. The adder tree calculates the sum of the products of the mantissas after the shift processing.

[0089] like Figure 3 As shown, based on the 6×128 index and array, it is divided into a high 4-bit connection value search judgment module and a low 2-bit connection Bitstream conversion and value search difference module.

[0090] like Figure 8As shown, in the high 4-bit exponent and SRAM array, there are 4 CBLs from high to low, connecting the units representing the same bit in the 128 numbers. The Q point values ​​of these units and the signal of CWL jointly control the connection between CBL and the low level. The upper part of CBL is connected to the precharge circuit controlled by PCH, the lower part is connected to an inverter, and then connected to the register used to store the maximum value MAX of the current bit; the QB points of these units are connected to a NAND gate with MAX, and the output is connected to an AND gate with CWL, and the output result serves as the CWL of the next bit; starting from the second highest bit, the QB points of each unit are connected to a NOR gate with MAX, and the output is connected to an OR gate with CWL, as a judgment on whether the number needs to be processed with the low bit of the exponent and the mantissa in the future. When performing the value search operation, first set PCH to a high level so that CBL is precharged to a high potential, and then set the first-level CWL to a high level, and turn on the transmission tube controlled by it. If there is an SRAM unit with a storage node Q at a high level on this CBL, the transmission tube controlled by it will be turned on, causing CBL to connect to a low level, the voltage to drop, and 1 to be stored in the register MAX through the inverter; if there is no SRAM unit with a storage node Q at a high level on this CBL, CBL will not connect to a low level, maintain a high voltage, and store 0 in the register MAX through the inverter. The significance of the above operation is to determine whether the maximum value at this bit is 1. MAX and the QB point of each SRAM unit in this column perform a negative and operation respectively, and then perform an AND operation with CWL, and its output controls the next level CWL. When the MAX value is 1 and QB is 1 or the CWL of this level is not at a high level, the AND gate will output 0, turning off the transmission tubes of the next level SRAM and CBL, so that its subsequent low bits do not participate in the maximum value search. The above circuit has 4 levels from high to low, and 4 cycles can find the first four bits of the maximum value EMAX. Starting from the second level, MAX and the QB point of each SRAM unit in this column are respectively operated in or-not, and then operated in or-with CWL. If its output is 1, it means that there is no difference between EMAX and the value at this bit, and it still participates in the subsequent difference operation. If its output is 0, it means that there is a difference between EMAX and the value in the previous bit, because this circuit can judge whether there is a difference in the first three bits, and the difference of each of the first three bits is 32, 16, and 8, which exceeds the length of the mantissa product. Therefore, it is judged that there is no need to perform the difference operation of the low bit and the multiplication operation of the mantissa in the subsequent process, and the corresponding data can be discarded and no longer processed. The high 4-bit array contains 4 CBLs from high to low, connecting the units representing the same bit in the 128 numbers. The Q point values ​​of these units and the signal of CWL jointly control the connection between CBL and the low level.CBL is connected to a precharge circuit controlled by PCH above and an inverter below, and then to a register used to store the maximum value MAX of the current bit; the QB points of these units are connected to a NAND gate with MAX, and the output is connected to an AND gate with CWL, and the output result serves as the CWL of the next bit; starting from the second highest bit, the QB points of each unit are connected to a NOR gate with MAX, and the output is connected to an OR gate with CWL, as a judgment on whether the number needs to be processed with the low bit of the exponent and the mantissa later.

[0091] like Fig. 9 As shown, the lower 2-bit array represents two units in each column, the storage nodes of the high-order units are Q1 and QB1, and the low-order units are Q2 and QB2. The high-order and low-order bits in the signal source are IN1 and IN0 respectively, and they perform logical operations continuously in three cycles. QB1 and IN1 perform NAND operations, the exclusive OR results of Q1IN1, QB0 and IN0 perform NAND operations, and then the two NAND results perform AND logic. IN1 IN0 is 01, 10, and 11 in three cycles. If Q1 Q0 is greater than or equal to IN1 IN0, a high-level Bitstream signal is output; if Q1 Q is less than IN1 IN0, a low-level Bitstream signal is output. If 00 is stored in S1S0, the output Bitstream is 000; if 01 is stored in S1S0, the output Bitstream is 100; if 10 is stored in S1S0, the output Bitstream is 110; if 11 is stored in S1S0, the output Bitstream is 111. If the last bit of the search result of the upper 4 bits is 1, a three-cycle Bitstream signal is directly output; if the last bit of the search result of the upper 4 bits is 0, the data with the first four bits as the maximum value first outputs a 4-cycle high-level Bitstream signal and then outputs a 3-cycle Bitstream signal obtained by logic operation. The data with the first four bits not as the maximum value first outputs a 3-cycle Bitstream signal obtained by logic operation and then outputs a 4-cycle low level. The Q and QB point signals of the two units in each column represented by the lower 2-bit array are connected to multiple logic gates with a common signal source, and then each Bitstream signal is generated in the same timing, and then combined with the judgment signal generated by the upper 4-bit array, the Bitstream corresponding to the lower 2 bits EMAX and the mantissa product shift signal corresponding to each number are output through a circuit composed of multiple OR gates and XOR gates.

[0092] like Fig.10As shown, the circuit can be mainly composed of 127 two-input OR gates and 128 XOR gates. Except for the first OR gate, the input of each OR gate is connected to the output of the previous OR gate, and the output of the last OR gate is connected to 128 XOR gates. The OR logic represents the comparison between two numbers, and the XOR logic represents the absolute value of the difference between the two numbers. 128 Bitstream inputs can output the Bitstream representing the maximum value EMAX and the difference between each Bitstream and the Bitstream representing the maximum value. The OR logic represents the comparison between two numbers, and the XOR logic represents the absolute value of the difference between the two numbers.

[0093] like Fig.11 As shown in the figure, the 2bit×2bit multiplication module can be mainly composed of 4 AND gates and 2 half adders. The high bit MW1 and low bit MW0 of the weight, the high bit MI1 and low bit MI0 of the input, MW0 and MI0 are connected to the lowest bit PD0 of the product output by the first AND gate; MW1 and MI0 are connected to the second AND gate, MW0 and MI1 are connected to the third AND gate, the outputs of both are connected to the input of the first half adder, and the sum output is the second lowest bit PD1 of the product; MW1 and MI1 are connected to the fourth AND gate, its output and the carry output of the first half adder are connected to the input of the second half adder, and the sum output is the second highest bit PD2 of the product, and the carry output is the highest bit PD3. The weight and the low bit of the input are connected to the AND gate, the low bit of the weight and the low bit of the input, the high bit of the weight and the low bit of the input are connected to the AND gate respectively, and then connected to the first half adder, the weight and the high bit of the input are connected to the AND gate, and then connected to the second half adder with the carry of the first half adder.

[0094] like Fig.12 As shown, the mantissa product shift register can be mainly composed of 4 groups of D flip-flops and a two-choice MUX. The input of each D flip-flop is connected to the single output of the two-choice MUX. One end of the input of the two-choice selector is connected to the output of the multiplication module, and the other end is connected to the previous level D flip-flop (the highest bit is connected to a low level). The CLK signal of all D flip-flops is generated by the system timing signal and the corresponding shift signal through an AND operation. When the MUX selection signal is a low level, the mantissa product is written into the D flip-flop; when the MUX selection signal is a high level, the output of the D flip-flop is written into the next register, and the highest bit is written to 0, so as to realize the right shift of the mantissa product. The input of each D flip-flop is connected to the single output of the two-choice MUX. One end of the input of the two-choice selector is connected to the output of the multiplication module, and the other end is connected to the previous level D flip-flop (the highest bit is connected to a low level). The CLK signal of all D flip-flops is generated by the system timing signal and the corresponding shift signal through an AND gate.

[0095] In the output module, the bitstream representing the lower 2 bits of EMAX is restored to a binary form by the counter, and combined with the upper 4 bits of EMAX stored in the exponent and MAX register below the array to form a complete 6-bit EMAX. After the mantissa shift register completes the shift operation, the data is input into the adder tree in the output module to obtain the multiplication and accumulation result of the mantissa part. In this way, the complete multiplication and accumulation result is obtained.

[0096] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A floating point multiplication and accumulation method for an in-memory computing circuit, for inputting exponents of lengths j and n into E IN and the mantissa input M IN Perform floating-point multiplication and accumulation operations to obtain the exponent and the maximum value high k bits EMAX H , exponent and maximum value low x bits EMAX L , and store data M OUT , k = j + 1 - x, 2 x ≥2n, characterized in that, The method comprises the following steps: For E IN Add the exponential weight E of the exponent in the original data WE Get the exponent and E SUM ; Take E SUM The high k bits in get the exponent and the high k bits E SUMH , in all E SUMH Select the maximum value as EMAX H Output and determine EMAX H With E SUMH When the difference is less than 2n, the output control signal is 1; When the control signal is 1, take E SUM The low x bits in get the exponent and the low x bits E SUML , in bitstream form, in all E SUML Find the maximum value as EMAX L Output and calculate EMAX L With E SUML The difference is the alignment shift E SUB ; Multiply the mantissa by M PD When the control signal is 1, take M IN ×M WE , M WE To store the mantissa weight of the mantissa in the original data; According to E SUB Shift the stored data to get M OUT .

2. The floating point multiplication and accumulation operation method of the in-memory computing circuit according to claim 1, characterized in that: The method comprises the following steps: EMAX H With E SUMH If the difference is greater than or equal to 2n, the output control signal is 0, and there is no need to calculate the corresponding alignment shift amount E in the subsequent SUB and the mantissa, M out That is, take 0.

3. An in-memory computing circuit for inputting exponentials of length j and n into E IN and the mantissa input M IN Perform floating-point multiplication and accumulation operations to obtain the exponent and the maximum value high k bits EMAX H , exponent and maximum value low x bits EMAX L , and store data M OUT , k = j + 1 - x, 2 x ≥2n, characterized in that, The in-memory computing circuit includes: Exponential weight SRAM array, used to store the exponential weight E of the exponent in the original data WE ; Addition module, whose exponent and E SUM For E IN +E WE ; index and SRAM array to store E SUM ; The high k-bit value search judgment module is used to obtain E SUM The high k bits in get the exponent and the high k bits E SUMH , in all E SUMH Select the maximum value as EMAX H Output and determine EMAX H With E SUMH If the difference between and 2n is greater than or equal to , the output control signal is 0, otherwise it is 1; The low x bit bit stream conversion and value search module is used to take E when the control signal is 1. SUM The low x bits in get the exponent and the low x bits E SUML , in bitstream form, in all E SUML Find the maximum value as EMAX L Output and calculate EMAX L With E SUML The difference is the alignment shift E SUB ; Mantissa weight SRAM array, used to store the mantissa weight M of the mantissa in the original data WE ; The multiplication module, whose mantissa product M PD When the control signal is 1, take M IN ×M WE , when the control signal is 0, it takes 0; Mantissa product shift register circuit, used to store M PD , and according to E SUB Shift the stored data to get M OUT .

4. The in-memory computing circuit according to claim 3, characterized in that: The in-memory computing circuit also includes a peripheral circuit, and the peripheral circuit includes: A word line driver, used to control the opening of the word lines of each row in the SRAM array; An address decoder is connected to the word line driver and is used to decode the address signal and transmit it to the word line driver; A precharge circuit, used for precharging the bit lines of each column in the SRAM array; A timing control module is used to generate various clock signals required to perform data storage tasks or logic operations; The read / write selection module is used to select each SRAM cell in the SRAM array that needs to perform a read / write operation.

5. The in-memory computing circuit according to claim 4, characterized in that: The in-memory computing circuit also includes a peripheral circuit, and the peripheral circuit also includes: The mode switching circuit is used to switch the working mode of the in-memory computing circuit.

6. The in-memory computing circuit according to claim 4, characterized in that: k is 4, x is 2, j is 5, and n is 6; the addition module adopts a 1-bit full adder, and the 1-bit full adder performs the addition operation of the 5-bit exponential weight and the 5-bit exponential input within 5 cycles, and the output 6-bit exponent and result are written into the index and SRAM array, the W end of the full adder is connected to the bit line EBL of the exponential weight SRAM array, the I end is input from the outside, the S end is connected to the word line SWL of the exponent and SRAM array, the CI end is connected to the read circuit controlled by CIWL in the high-k bit value search and judgment module, and the CO end is connected to the write circuit controlled by COWL in the high-k bit value search and judgment module.

7. The in-memory computing circuit according to claim 6, characterized in that: The high k-bit value-seeking judgment module includes 4 CBLs from high to low, connecting the units representing the same bit in the 128 numbers. The Q point values ​​of these units and the CWL signal jointly control the connection between CBL and the low level. The upper part of CBL is connected to the precharge circuit controlled by PCH, the lower part is connected to an inverter, and then connected to the register used to store the maximum value MAX of the current bit; The QB points of these units are connected to a NAND gate with MAX, and the output is connected to an AND gate with CWL, and the output result serves as the CWL of the next bit; starting from the second highest bit, the QB points of each unit are connected to a NOR gate with MAX, and the output is connected to an OR gate with CWL to determine whether the low-order exponent and mantissa of the stored data need to be processed later.

8. The in-memory computing circuit according to claim 7, characterized in that: The Q and QB point signals of the two units in each column represented by the low x-bit bit stream conversion and value search module are connected to multiple logic gates with a common signal source, and then various bit stream signals are generated in the same timing sequence, and then combined with the judgment signal generated by the high k-bit value search judgment module, the circuit composed of multiple OR gates and XOR gates is used to output the low 2-bit E SUML The corresponding bit stream and the mantissa product shift signal corresponding to each number.

9. The in-memory computing circuit according to claim 7, characterized in that: The multiplication module includes 4 AND gates and 2 half adders. The weight and the low bit of the input are connected to the AND gate, the low bit of the weight and the low bit of the input, the high bit of the weight and the low bit of the input are connected to the AND gate respectively and then connected to the first half adder, the weight and the high bit of the input are connected to the AND gate, and then connected to the second half adder with the carry of the first half adder.

10. The in-memory computing circuit according to claim 7, characterized in that: The mantissa product shift register circuit includes 4 groups of D flip-flops and two-to-one MUX. The input of each D flip-flop is connected to the single output of the two-to-one MUX. One end of the input of the two-to-one selector is connected to the output of the multiplication module, and the other end is connected to the previous level D flip-flop. The CLK signals of all D flip-flops are generated by the system timing signal and the corresponding shift signal through the AND gate.