Floating-point number operation circuit based on SRAM (Static Random Access Memory) and in-memory operation chip
Through the SRAM-based floating-point number operation circuit, the global difference decreasing traversal and multiple round data screening mechanisms are used to solve the problem of low calculation efficiency of floating-point number data in the existing technology, and efficient and high-speed floating-point operation is realized, which is suitable for the high computing power requirements of edge devices.
Patent Information
- Application Number
- CN202510425747.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
Existing computing circuits are inefficient when processing floating point digital data, slow speed and low accuracy, making it difficult to meet the needs of high-efficiency computing.
The floating-point number calculation circuit based on SRAM is adopted, and the mantissa shift operation is converted into conditionally triggered mantissa accumulation through global difference decreasing traversal and multi-round row data filtering mechanisms, and the parallelism of the adder tree is used to improve the computing efficiency.
At the hardware level, complex shift circuits are avoided, the calculation efficiency and accuracy of floating point digital data are improved, and the problems of inefficiency and slow speed in the prior art are solved, which is suitable for the high computing power requirements of edge devices.
Smart Images

Figure CN120353429A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an arithmetic circuit in the field of integrated circuit technology, and more particularly to a floating-point arithmetic circuit based on SRAM, and also relates to an in-memory computing chip. Background Art
[0002] In recent years, the rapid development of artificial intelligence (AI) and edge computing has given rise to an urgent need for energy-efficient computing architectures. Intelligent terminal devices need to process massive amounts of data in real time (such as tasks like computer vision and natural language processing), and their core operations rely on floating-point multiply-accumulate (MAC) operations of large-scale neural networks. However, the physical separation of the storage unit and the computing unit in the traditional von Neumann architecture has led to an increasingly severe "memory wall" problem caused by frequent data transfer: research shows that the energy consumption ratio of data access exceeds 60%, and with the iteration of semiconductor processes, the speed difference between the arithmetic unit and the storage unit has further widened, severely restricting the system energy efficiency.
[0003] To break through this bottleneck, the Computing-in-Memory (CIM) architecture embeds computing logic into the memory array and directly performs operations at the data storage location, completely eliminating the data migration overhead. Existing SRAM-based CIM circuits have shown significant advantages in integer MAC tasks, but their design paradigms are difficult to adapt to the requirements of floating-point operations. Floating-point operations require multiple-step operations such as "exponent alignment" and "mantissa multiplication / shift". Traditional CIM architectures need to break them down into multiple-cycle tasks, resulting in a sharp deterioration of computing latency and energy efficiency. At the same time, deep learning models are evolving towards high precision and large parameter scales. The low numerical representation ability of integer networks can no longer meet the requirements of complex scenarios. Although floating-point neural networks can support model performance through high dynamic range and computing precision, due to the low efficiency of floating-point operations in the CIM architecture, it is difficult to achieve efficient deployment at the edge. Therefore, existing arithmetic circuits have problems of low efficiency, slow speed, and low precision when processing floating-point data. Summary of the Invention
[0004] To solve the technical problems that existing arithmetic circuits have low efficiency, slow speed, and low precision when processing floating-point data, the present invention provides a floating-point arithmetic circuit based on SRAM and an in-memory arithmetic chip.
[0005] The present invention is implemented by the following technical solutions: A floating-point arithmetic circuit based on SRAM, which is used to implement the following steps:
[0006] S1: Calculate the exponent sum and the mantissa product of the multi-bit floating-point input value and the multi-bit floating-point weight value bit by bit respectively;
[0007] S2: Search for and determine the maximum exponent bit by bit among all the exponent sums, and then calculate the difference between the maximum exponent and the exponent sums of each row;
[0008] S3: First generate a self - decreasing value with an adjustable initial value, then compare the self - decreasing value with each difference to generate an addition control signal, and finally determine whether to input the mantissa products of each row into the adder tree for addition according to the addition control signal;
[0009] S4: After comparing with the exponent sums of all rows, perform a self - decreasing operation on the self - decreasing value and execute step S3; when the self - decreasing value reaches zero, take the accumulated sum of the mantissas in the adder tree as the total number of bits;
[0010] S5: Generate the multiplication - accumulation operation result of the input value and the weight value according to the maximum exponent and the total number of bits.
[0011] Through the "global difference decreasing traversal + multi - round row data screening" mechanism, the present invention converts the mantissa shift operation into condition - triggered mantissa product accumulation, avoids complex shift circuits at the hardware level, and improves the calculation efficiency by utilizing the parallelism of the adder tree, solving the technical problems of low efficiency, slow speed, and low precision existing in the existing arithmetic circuits when processing floating - point data.
[0012] As a further improvement of the above - mentioned solution, the arithmetic circuit includes:
[0013] An array circuit, which includes a weight exponent array, an exponent sum array, and a weight mantissa array; each row of the weight exponent array is used to pre - store the exponent part of the weight value bit by bit; the exponent sum array is used to store the exponent sums of each row; each row of the weight mantissa array is used to pre - store the mantissa part of the weight value bit by bit and perform a multiplication operation on the mantissa parts of the input value and the weight value to obtain the mantissa product.
[0014] Further, the weight mantissa array includes a plurality of storage units, and each storage unit includes NMOS transistors N1 - N6 and PMOS transistors P1, P2; N1, N2, P1, P2 are cross - coupled in an inverted manner to form a pair of storage nodes Q and QB; the gates of N3 and N4 are connected to the word line WL, N3 is the transfer transistor between the bit line BL and the node Q, and N4 is the transfer transistor between the bit line BLB and the node QB; the gate of N5 is connected to the node Q, the source is grounded, and the drain is connected to the drain of N6; the gate of N6 is connected to the calculation word line LRWL, and the source is connected to the calculation word line LHBL; N5 and N6 are used to calculate the mantissa product, and NMOS transistors N1 - N4 and PMOS transistors P1, P2 are used to pre - store the mantissa part of the weight.
[0015] Still further, the operation logic of the weight mantissa array includes:
[0016] (1) When node Q is at a low level and node QB is at a high level, the corresponding bit pre-stored in the mantissa part of the weight is "0"; when node Q is at a high level and node QB is at a low level, the corresponding bit pre-stored in the mantissa part of the weight is "1".
[0017] (2) First, pre-charge the calculation bit line LHBL to a high level, and then input the mantissa of the operand into the corresponding storage unit through the calculation bit line LRWL; among them, when the calculation word line LRWL is at a low level, it means that the corresponding bit of the mantissa part of the input operand is "0", and when the calculation word line LRWL is at a high level, it means that the corresponding bit of the mantissa part of the input operand is "1".
[0018] (3) Output the calculation result of the final mantissa product according to the level change of the calculation bit line LHBL.
[0019] As a further improvement of the above solution, the operation circuit includes:
[0020] An adjustable subtraction counter, which includes D flip - flops DFF2, DFF3, DFF4, DFF5 with asynchronous set and clear functions, a D flip - flop DFF6 with an asynchronous set function, a D flip - flop DFF1 with an asynchronous clear function, delay elements DEL1, DEL2, DEL3, DEL4, buffers BUFF1, BUFF2, an inverter NOT, AND gate circuits AND1, AND2, NAND gate circuits NAND1, NAND2, NAND3, and OR gate circuits OR1, OR2; the input end of BUFF1 is connected to the controllable signal A<1>, and the output end is connected to the input end of NOT; the output end of NOT is connected to one of the input ends of NAND1; the input end of BUFF2 is connected to the controllable signal A<0>, and the output end is connected to the other input end of NAND1; the input end of DEL3 is connected to the maximum value search end signal, and the output end is connected to one of the input ends of NAND2; the other input end of NAND2 is connected to the inversion of the maximum value search end signal, and the output end is connected to one of the input ends of OR2; the output end of OR2 outputs the signal SET_I<3:0>; the input end of DEL1 is connected to the maximum value search end signal, and the output end is connected to one of the input ends of DFF1; the other input end of DFF1 is connected to the power supply VDD, and the output end is connected to one of the input ends of AND1; the other input end of AND1 is the output end of OR1, and the output end is connected to one of the input ends of AND2; the other input end of AND2 is connected to the CLK signal, and the output end is connected to the input CPN end of DFF2; the input D and the output QN of DFF2 are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF3 is connected to the output QN of DFF2, the input D and the output QN are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF4 is connected to the output QN of DFF3, the input D and the output QN are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF5 is connected to the output QN of DFF4, the input D and the output QN are connected, and the input clear end is connected to the clear signal CLR; one of the input ends of DEL2 is connected to the output end of AND2, the output end is connected to one of the input ends of DFF6, and the input clear end is connected to the clear signal CLR; one of the input ends of DEL4 is connected to the output end of AND2, and the output end is connected to one of the input ends of NAND3; the other input end of NAND3 is connected to the output end of AND2, and the output end is the output end of the adjustable subtraction counter.
[0021] Further, the operation logic of the adjustable subtraction counter includes:
[0022] (1) Controlling the initial value of the signal SET_I<3:0> according to the input values of the controllable signals A<1> and A<0>;
[0023] (2) Connect the signals SET_I<3:0> to the set-1 terminals of DFF2, DFF3, DFF4, and DFF5 respectively, and generate the values of the signal SELF_SUB_I<3:0> whose output initial value is incremented by 1 according to the signals SET_I<3:0>. Among them, the other input terminal of DFF6 is connected to the signal SELF_SUB_I<3:0>, and the QN output terminals of DFF2, DFF3, DFF4, and DFF5 are respectively connected to the signals SELF_SUB_I<0>, SELF_SUB_I<1>, SELF_SUB_I<2>, and SELF_SUB_I<3>.
[0024] (3) Asynchronously set the initial value of the signal SELF_SUB_I<3:0> through DFF6 and obtain the initial value of the signal SELF_SUB<3:0> after a delay.
[0025] Furthermore, the arithmetic circuit further includes:
[0026] A comparator, which includes exclusive-NOR gate circuits XNOR1, XNOR2, XNOR3, XNOR4, and AND gate circuits AND3 and AND4. One input terminal of XNOR1 is connected to the signal SELF_SUB<3>, and the other input terminal is connected to the difference SUB<3> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A1 of AND3. One input terminal of XNOR2 is connected to the signal SELF_SUB<2>, and the other input terminal is connected to the difference SUB<2> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A2 of AND3. One input terminal of XNOR3 is connected to the signal SELF_SUB<1>, and the other input terminal is connected to the difference SUB<1> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A3 of AND3. One input terminal of XNOR4 is connected to the signal SELF_SUB<0>, and the other input terminal is connected to the difference SUB<0> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A4 of AND3. The output terminal of AND3 is connected to the input terminal A1 of AND4. The input terminal A2 of AND4 is connected to the output terminal of the adjustable subtraction counter, the input terminal A3 is connected to the exponent sum subtraction completion signal EN_I, and the output terminal outputs the addition control signal.
[0027] As a further improvement of the above solution, the operation logic of the comparator includes:
[0028] (1) Judge whether the corresponding bits of the signals SELF_SUB<3:0> and the difference SUB<3:0> are equal according to the outputs of XNOR1, XNOR2, XNOR3, and XNOR4.
[0029] (2) Determine whether the signal SELF_SUB<3:0> is exactly equal to the difference SUB<3:0> according to the output of AND3;
[0030] (3) Determine whether to make the addition control signal high through AND3 and AND4. If so, input the corresponding mantissa product into the adder tree for addition; otherwise, control the adder tree not to perform mantissa product accumulation addition.
[0031] Further, the arithmetic circuit further includes:
[0032] A maximum value search module, which is used to search bit by bit in the exponent sum array and determine the maximum exponent;
[0033] An exponent input module;
[0034] An adder, which is used to first receive the exponent part of the input value input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the exponent sum of the two exponent parts, and finally store the exponent sum calculation result in the exponent sum array;
[0035] An adjustable subtraction counter, which is used to generate a self-decrement value with an adjustable initial value, then compare the self-decrement value with each difference through a comparator to generate an addition control signal, and finally determine whether to input the mantissa product of each row into the adder tree for addition according to the addition control signal. After comparing with the exponent sums of all rows, the self-decrement value performs a self-decrement operation until the self-decrement value reaches zero, and the accumulated addition result of the mantissa products in the adder tree is used as the digit sum;
[0036] A mantissa input module, which is used to input the operand to the corresponding row of the weight mantissa array;
[0037] An adder tree and a normalization module, which provide the adder tree and are used to generate the multiplication and accumulation operation result of the input value and the weight value according to the maximum exponent and the digit sum.
[0038] The present invention also provides an in-memory computing chip, which includes any one of the above SRAM-based floating-point arithmetic circuits.
[0039] Compared with the existing floating-point arithmetic circuits, the SRAM-based floating-point arithmetic circuit and the in-memory computing chip of the present invention have the following beneficial effects:
[0040] 1. The SRAM-based floating-point arithmetic circuit converts the mantissa shift operation into condition-triggered mantissa accumulation addition through the "global difference decreasing traversal + multi-round row data screening" mechanism, avoiding complex shift circuits at the hardware level and improving the calculation efficiency by utilizing the parallelism of the adder tree, thus solving the technical problems of low efficiency, slow speed, and low precision existing in the existing arithmetic circuits when processing floating-point data.
[0041] 2. The SRAM-based floating-point arithmetic circuit has relatively simple floating-point operation steps, can be compatible with the parallel computing characteristics of the CIM architecture, can solve the imbalance between the limited resources of edge devices and the high computing power requirements of large-scale floating-point models, can optimize the operation process and hardware resource allocation, can release the potential of CIM, and support the implementation of the next-generation intelligent applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a framework diagram of the SRAM-based floating-point arithmetic circuit according to Embodiment 2 of the present invention.
[0043] Figure 2 It is Figure 1 a circuit diagram of a storage unit in the weight mantissa array of the SRAM-based floating-point arithmetic circuit in
[0044] Figure 3 It is Figure 1 a first partial circuit diagram of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0045] Figure 4 It is Figure 1 a second partial circuit diagram of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0046] Figure 5 It is Figure 1 a third partial circuit diagram of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0047] Figure 6 It is Figure 1 a fourth partial circuit diagram of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0048] Figure 7 It is Figure 1 a fifth partial circuit diagram of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0049] Figure 8 It is Figure 1 an operation flowchart of the adjustable subtraction counter of the SRAM-based floating-point arithmetic circuit in
[0050] Figure 9 For Figure 1 the circuit diagram of the comparator of the SRAM-based floating-point arithmetic circuit in
[0051] Figure 10 For Figure 1 the operation flowchart of the mantissa shift part in the SRAM-based floating-point arithmetic circuit in Specific Embodiments
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Embodiment 1
[0054] This embodiment provides a SRAM-based floating-point arithmetic circuit, which is designed based on an SRAM array and its peripheral circuits to implement the multiply-accumulate operation between multiple groups of multi-bit floating-point inputs and multiple groups of multi-bit floating-point weights. The architecture of the circuit adopts a dual-mode working mechanism: a storage mode and a calculation mode. The two modes are dynamically switched through a timing control unit and a power gating module. Among them, the mode selection signal synchronously regulates the enabling states of the storage interface circuit and the calculation logic unit, so as to achieve an integrated storage and calculation design on the premise of ensuring data integrity. Among them, the floating-point arithmetic circuit is used to implement the following steps (Steps S1-S5).
[0055] S1: Calculate the exponent sum and the mantissa product of the multi-bit floating-point input value (operand) and the multi-bit floating-point weight value bit by bit respectively. The input value and the weight value adopt the floating-point representation format of the IEEE 754 standard, including 1 sign bit, multiple exponent bits and multiple mantissa bits. In this embodiment, for the sign bit, a simple exclusive-OR logic circuit can be used for parallel processing to determine the sign of the final result. The exponent part can be calculated by a dedicated adder array. Considering the offset problem in the floating-point representation, an offset compensation adjustment needs to be performed additionally when calculating the exponent sum. The mantissa part completes the product operation through a group of parallel fixed-point multiplier arrays. In particular, the implicit highest bit "1" needs to be explicitly restored during mantissa processing to ensure the calculation accuracy. All calculation units adopt a pipeline design and can complete the processing of a batch of data within one clock cycle.
[0056] S2: Search for and determine the maximum exponent bit by bit among all the exponent sums, and then calculate the difference between the maximum exponent and the exponent sums of each row. The core of this step is to find the maximum value among all the exponent sums, establishing a benchmark for subsequent alignment operations. After determining the maximum exponent, a set of parallel subtractors is used to calculate the difference between each exponent sum and this maximum value. These differences determine the number of bits by which the corresponding mantissa product needs to be right-shifted. To optimize hardware resources, the subtractors can also adopt a carry-select structure, which can complete the calculation of all differences within one cycle. At the same time, the system also detects special cases, such as the situation where the mantissa product is shifted out of the valid range due to an overly large difference, and sets corresponding flag bits to skip these invalid operations.
[0057] S3: First generate a self-decrement value with an adjustable initial value, then compare the self-decrement value with each difference to generate an addition control signal, and finally, based on the addition control signal, determine whether to input the mantissa products of each row into the adder tree for addition. This step implements an innovative dynamic alignment mechanism. The circuit first generates a self-decrement counter with a configurable initial value, and the initial value of this counter is dynamically set according to the operation precision requirements. In each processing cycle, the current value of the counter is compared with all the pre-calculated differences, and the comparison results generate a set of control signals through combinational logic. These control signals can be connected to a multiplexer array to determine which mantissa products should be gated into the adder tree. The selected mantissa products are precisely shifted and aligned according to their corresponding differences. The entire control logic adopts a fully pipelined design, which can handle multiple parallel alignment operations while maintaining a high clock frequency.
[0058] S4: After comparing with the exponent sums of all rows is completed, the self-decrement value is decremented, and step S3 is executed. When the self-decrement value reaches zero, the accumulated sum of the mantissa products in the adder tree is used as the total number of bits. This step implements an efficient iterative accumulation algorithm. After the system completes a round of comparison and selection of partial mantissa products in each clock cycle, the self-decrement counter automatically decrements to start the next round of processing. In this embodiment, to handle multi-cycle accumulation, the system designs a dedicated accumulation register to save intermediate results. When the counter reaches zero, it means that all valid mantissa products have participated in the accumulation. At this time, the system performs overflow detection and preliminary normalization processing on the accumulation result. The entire accumulation process adopts a deep pipelined design, which can handle the accumulation operations of multiple groups of data simultaneously, significantly improving the throughput.
[0059] S5: Generate the multiplication and accumulation operation result of the input value and the weight value according to the maximum exponent and the total number of digits. The last step is responsible for generating the floating-point result that meets the standard. In this embodiment, first, perform a leading zero detection on the accumulated mantissa sum, and adjust the exponent value corresponding to the mantissa sum according to the detection result. The exponent adjustment unit will handle various boundary conditions, such as overflow, underflow, etc., and generate corresponding exception flags. The sign bit processing unit synthesizes the sign information of all input data to determine the sign of the final result. The normalization unit ensures that the most significant bit of the mantissa is 1, and performs rounding operations if necessary. Finally, all processed parts are recombined in the standard floating-point format to output the final multiplication and accumulation result. The entire output stage also includes result verification and exception handling logic to ensure the reliability of the operation.
[0060] Embodiment 2
[0061] Please refer to Figure 1-10 , this embodiment provides a floating-point arithmetic circuit based on SRAM, which is similar to the floating-point arithmetic circuit in Embodiment 1. This embodiment provides a specific circuit structure. Among them, the floating-point arithmetic circuit includes an array circuit, an adjustable subtraction counter, a comparator, a maximum value finding module, an exponent input module, an adder, a subtractor, a mantissa input module, an adder tree, and a normalization module.
[0062] The array circuit includes a weight exponent array, an exponent sum array, and a weight mantissa array. Each row of the weight exponent array is used to pre-store the exponent part in the weight value bit by bit. The exponent sum array is used to store the exponent sum of each row. Each row of the weight mantissa array is used to pre-store the mantissa part in the weight value bit by bit, and perform a multiplication operation on the mantissa parts of the input value and the weight value to obtain the mantissa product. The weight exponent array mainly processes the relevant tasks of the weight exponent of the floating-point number during the operation process. The exponent sum array mainly stores the exponent sum obtained by adding the weight exponent of the floating-point number and the exponent of the input operand. The weight mantissa array mainly processes the relevant tasks of the weight mantissa of the floating-point number during the operation process.
[0063] These three core memory arrays need to be refined and configured according to the numerical structure of the target floating-point format. Taking a typical format as an example: for the 8-bit exponent and 7-bit mantissa of the BF16 (Bfloat16) format, the system adopts a column allocation ratio of 8:8:8 for the exponent and array, weight exponent array, and weight mantissa array. Among them, the mantissa column is extended by 1 bit to accommodate the overflow protection of the intermediate result in the multiply-accumulate process; for the 8-bit exponent and 23-bit mantissa of the FP32 single-precision format, an 8:8:24 column division strategy is adopted, and an additional column is added to the mantissa array to support the generation of guard bits and sticky bits. This dynamic column width configuration mechanism is implemented through hardware description language parameterization, allowing the SRAM bank structure to be automatically adjusted according to the floating-point format bit width at runtime, achieving an optimal balance between silicon area efficiency and computing accuracy. The BF16 format has a 5-bit exponent and a 10-bit mantissa.
[0064] The function of the exponent input module is to transfer the exponent part of the operand to the adder. The adder is used to first receive the exponent part of the input value (operand) input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the exponent sum of the two exponent parts, and finally store the exponent sum calculation result in the exponent sum array. The mantissa input module is used to input the operand to the corresponding row of the weight mantissa array.
[0065] The floating-point arithmetic architecture of this embodiment adopts a SRAM array design paradigm of bit-parallel storage-row addressing calculation. Its core mechanism is: the exponent bits and mantissa bits of a single floating-point operand / weight are respectively disassembled into independent bit streams and stored bit-by-bit along the row direction of the SRAM array: the storage units in the same row carry the exponent or mantissa bits corresponding to the bit width of the same operand / weight, and the column direction realizes the bit width expansion of different floating-point numbers through a multi-bank structure. Under this architecture, the maximum parallel computing scale of the system is directly determined by the row depth of the SRAM array: taking three independently configured 64-row × 8-column memory arrays (processing exponent sum, weight exponent, and weight mantissa respectively) as an example, each row corresponds to a complete set of floating-point operands / weights (5-bit exponent + 10-bit mantissa), and the bit line calculations of 64 sets of operands and 64 sets of weights are synchronously started through the row activation signal, and the full-parallel multiply-accumulate operation is completed in the analog memory-computation or digital logic unit.
[0066] Please continue to refer to Figure 2, the weight mantissa array includes a plurality of memory cells. In this embodiment, the memory cell is an 8T-SRAM cell. Each memory cell includes NMOS transistors N1 to N6 and PMOS transistors P1 and P2. N1, N2, P1, and P2 are inversely cross-coupled to form a pair of memory nodes Q and QB. The gates of N3 and N4 are connected to the word line WL. N3 is the transfer transistor between the bit line BL and the node Q, and N4 is the transfer transistor between the bit line BLB and the node QB. The gate of N5 is connected to the node Q, the source is grounded, and the drain is connected to the drain of N6. The gate of N6 is connected to the calculation word line LRWL, and the source is connected to the calculation word line LHBL. Among them, P1, P2, N1 to N4 constitute a 6T memory cell, N5 and N6 are used to calculate the mantissa product, and NMOS transistors N1 to N4 and PMOS transistors P1 and P2 are used to pre-store the mantissa part of the weight.
[0067] In this embodiment, the operation logic of the weight mantissa array includes:
[0068] (1) The mantissa part in the weight is first stored in the memory cell: when the node Q is at a low level and the node QB is at a high level, the corresponding bit pre-stored in the mantissa part of the weight is "0"; when the node Q is at a high level and the node QB is at a low level, the corresponding bit pre-stored in the mantissa part of the weight is "1".
[0069] (2) First, pre-charge the calculation bit line LHBL to a high level, and then input the mantissa of the operand into the corresponding memory cell through the calculation bit line LRWL. Among them, when the calculation word line LRWL is at a low level, it means that the corresponding bit of the mantissa part of the input operand is "0", and when the calculation word line LRWL is at a high level, it means that the corresponding bit of the mantissa part of the input operand is "1".
[0070] (3) Output the calculation result of the final mantissa product according to the level change of the calculation bit line LHBL. When CL remains at a high level state, it means that the product result is "0"; when CL drops to a low level, it means that the product result is "1".
[0071] In this embodiment, when LHBL remains at a high level, it means that the mantissa product result is "0", and when LHBL drops to a low level, it means that the mantissa product result is "1".
[0072] Specifically, when LRWL is at a low level, Q is at a low level, and QB is at a high level, both N5 and N6 are turned off. At this time, the calculation bit line LHBL cannot form a discharge path, and LHBL remains at a high level, that is, the operation "0×0 = 0" is completed.
[0073] When LRWL is at a high level, Q is at a low level, and QB is at a high level, N5 is turned off and N6 is turned on. At this time, the calculation bit line LHBL cannot form a discharge path, and LHBL remains at a high level, that is, the operation "1×0 = 0" is completed.
[0074] When LRWL is at a low level, Q is at a high level, and QB is at a low level, N5 conducts and N6 turns off. At this time, the calculation bit line LHBL cannot form a discharge path, and LHBL remains at a high level, that is, the operation "0 × 1 = 0" is completed.
[0075] When LRWL is at a high level, Q is at a high level, and QB is at a low level, both N5 and N6 conduct. At this time, the calculation bit line LHBL forms a discharge path, and LHBL drops to a low level, that is, the operation "1 × 1 = 1" is completed.
[0076] The truth table corresponding to the above process is shown in Table 1.
[0077] Table 1: Truth Table of Multiplication Operation in 8T-SRAM
[0078]
[0079] The maximum value search module is used to search and determine the maximum exponent bit by bit in the exponent sum array. In this embodiment, the subtractor calculates the difference between the maximum exponent sum and the exponent sum of each row.
[0080] The adjustable subtraction counter is used to generate a self-decrement value with an adjustable initial value, then compare the self-decrement value with each difference through a comparator to generate an addition control signal, and finally determine whether to input the mantissa products of each row into the adder tree for addition according to the addition control signal. After comparing with the exponent sums of all rows, the self-decrement value performs a self-decrement operation until the self-decrement value reaches zero, and the accumulated sum of the mantissas in the adder tree is used as the total number of bits.
[0081] Please continue to refer to Figure 3-8 , in this embodiment, the adjustable subtraction counter includes D flip-flops DFF2, DFF3, DFF4, DFF5 with asynchronous set and clear functions, a D flip-flop DFF6 with asynchronous set function, a D flip-flop DFF1 with asynchronous clear function, delay elements DEL1, DEL2, DEL3, DEL4, buffers BUFF1, BUFF2, an inverter NOT, AND gate circuits AND1, AND2, NAND gate circuits NAND1, NAND2, NAND3, and OR gate circuits OR1, OR2.
[0082] The input end of BUFF1 is connected to the controllable signal A<1>, and the output end is connected to the input end of NOT. The output end of NOT is connected to one of the input ends A1 of NAND1, and the output is SET<1>. The input end of BUFF2 is connected to the controllable signal A<0>, and the output end is connected to the other input end A2 of NAND1. The output of NAND1 is SET<2>.
[0083] The input terminal I of DEL3 is connected to the maximum value search end signal ST, and the output terminal Z is connected to one of the input terminals A2 of NAND2. The other input terminal A1 of NAND2 is connected to the inverse of the maximum value search end signal ST, and the output terminal is connected to one of the input terminals A2 of OR2. The input terminal A1 of OR2 is connected to the high level VDD and SET<2:0>, and the output terminal of OR2 outputs the signal SET_I<3:0>.
[0084] The input terminal I of DEL1 is connected to the maximum value search end signal, and the output terminal is connected to one of the input terminals CPN of DFF1. The other input terminal D of DFF1 is connected to the power supply VDD, and the output terminal is connected to one of the input terminals A1 of AND1. The other input terminal A2 of AND1 is the output terminal Z of OR1, and the output terminal is connected to one of the input terminals A1 of AND2. The four input terminals A1 to A4 of OR1 are respectively connected to SELF_SUB_I<0>, SELF_SUB_I<1>, SELF_SUB_I<2>, and SELF_SUB_I<3>. The other input terminal A2 of AND2 is connected to the CLK signal, and the output terminal is connected to the input terminal CPN of DFF2.
[0085] The input terminal D and the output terminal QN of DFF2 are connected, and the input clear terminal is connected to the clear signal CLR. The input terminal of DFF3 is connected to the output terminal QN of DFF2, the input terminal D and the output terminal QN are connected, and the input clear terminal is connected to the clear signal CLR. The input terminal of DFF4 is connected to the output terminal QN of DFF3, the input terminal D and the output terminal QN are connected, and the input clear terminal is connected to the clear signal CLR. The input terminal of DFF5 is connected to the output terminal QN of DFF4, the input terminal D and the output terminal QN are connected, and the input clear terminal is connected to the clear signal CLR.
[0086] One of the input terminals of DEL2 is connected to the output terminal of AND2, the output terminal is connected to one of the input terminals E of DFF6, and the input clear terminal is connected to the clear signal CLR. One of the input terminals of DEL4 is connected to the output terminal of AND2, and the output terminal is connected to one of the input terminals of NAND3. The other input terminal of NAND3 is connected to the output terminal of AND2, and the output terminal is the output terminal of the adjustable subtraction counter, and the output signal is EN_II. Among them, the other input terminal D of DFF6 is connected to the signal SELF_SUB_I<3:0>, the output terminals QN of DFF2, DFF3, DFF4, and DFF5 are respectively connected to the signals SELF_SUB_I<0>, SELF_SUB_I<1>, SELF_SUB_I<2>, and SELF_SUB_I<3>, and the input set terminals SDN of DFF2, DFF3, DFF4, and DFF5 are respectively connected to SET_I<0>, SET_I<1>, SET_I<2>, and SET_I<3>.
[0087] Initialization process: BUFF1, BUFF2, NOT, and NAND1 are used to generate the initial values of SET<0>, SET<1>, and SET<2> based on the input values of A<0> and A<1>; DEL3, NAND2, and OR2 initialize the input initial value SET_I<3:0> based on SET<0>, SET<1>, SET<2>, and the high level; the input CDN clear terminals of DFF2, DFF3, DFF4, and DFF5 are connected to the clear CLR signal to set SELF_SUB_I<3:0> to "1111", and then the input SDN set terminals of DFF2, DFF3, DFF4, and DFF5 are connected to the input initial value SET_I<3:0> to initialize SELF_SUB_I<3:0> to the value of the initial value of the adjustable subtraction counter plus 1. DEL2 and DFF6 initialize the output SELF_SUB<3:0> based on the value of SELF_SUB_I<3:0>.
[0088] The operation logic of the adjustable subtraction counter includes:
[0089] (1) Control the initial value of the signal SET_I<3:0> according to the input values of the controllable signals A<1> and A<0>. Specifically, according to the value range of A<1> and A<0> being "00 - 11", control the range of the input initial value SET_I<3:0> of the four D flip - flops DFF2, DFF3, DFF4, and DFF5 in the adjustable subtraction counter to be "1011 - 1110".
[0090] When A<1> is at a low level and A<0> is at a low level, the initial value of SET<2:0> is "110", and the initial value of SET_I<3:0> is "1110".
[0091] When A<1> is at a low level and A<0> is at a high level, the initial value of SET<2:0> is "011", and the initial value of SET_I<3:0> is "1011".
[0092] When A<1> is at a high level and A<0> is at a high level, the initial value of SET<2:0> is "101", and the initial value of SET_I<3:0> is "1101".
[0093] When A<1> is at a high level and A<0> is at a low level, the initial value of SET<2:0> is "100", and the initial value of SET_I<3:0> is "1100".
[0094] (2) Connect the signals SET_I<3:0> to the set terminals of four D flip-flops DFF2, DFF3, DFF4, and DFF5 with asynchronous set (SET) and clear (CLR) functions respectively, and generate the values of the signals SELF_SUB_I<3:0> with the output initial values incremented by 1, ranging from "1011 - 1110".
[0095] When the value of SET_I<3:0> is "1110", the initial value of SELF_SUB_I<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1110".
[0096] When the value of SET_I<3:0> is "1011", the initial value of SELF_SUB_I<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1011".
[0097] When the value of SET_I<3:0> is "1101", the initial value of SELF_SUB_I<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1101".
[0098] When the value of SET_I<3:0> is "1100", the initial value of SELF_SUB_I<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1100".
[0099] (3) Asynchronously set the initial values of the signals SELF_SUB_I<3:0> through DFF6 and obtain the initial values of the signals SELF_SUB<3:0> after delay. In this embodiment, the initial values of SELF_SUB<3:0> are controlled by a circuit consisting of a D flip-flop with an asynchronous set (SET) function and a delay circuit, and the range of the initial values of SELF_SUB<3:0> is "1010 - 1101".
[0100] When the value of SELF_SUB_I<3:0> is "1110", the initial value of SELF_SUB<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1101".
[0101] When the value of SELF_SUB_I<3:0> is "1011", the initial value of SELF_SUB<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1010".
[0102] When the value of SELF_SUB_I<3:0> is "1101", the initial value of SELF_SUB<3:0> with the output initial value of the adjustable subtraction counter incremented by 1 is "1100".
[0103] When the value of SELF_SUB_I<3:0> is "1100", the initial value of the output of the adjustable subtraction counter is incremented by 1, and the initial value of SELF_SUB<3:0> is "1011".
[0104] The truth table corresponding to the above process is shown in Table 2.
[0105] Table 2: Truth table for the initialization of controllable subtraction counting
[0106] A<1> A<0> SET<2:0> SET<2:0> SELF_SUB_I<3:0> SELF_SUB<3:0> 0 0 110 1110 1110 1101 0 1 011 1011 1011 1010 1 1 101 1101 1101 1100 1 0 100 1100 1100 1011
[0107] After the initialization is completed, the subtraction logic process is executed: A synchronous cyclic shift register composed of four D flip-flops DFF2, DFF3, DFF4, and DFF5 with asynchronous set and clear functions is cascaded. The output of the last-stage flip-flop (DFF5) is fed back to the input of the previous stage (DIFF2) through the SELF_SUB_I feedback network. Combining with the CPN and QN clock phase controls, a cyclic data stream loop is achieved, forming the logic function of the subtraction counter. The value of SELF_SUB_I<3:0> is output from the range "1010 - 1101" through DEL2 and DFF6. DEL4 and NAND3 control the output signal EN_II of the adjustable subtraction counter. When the output signal EN_II of the adjustable subtraction counter is high, it means that the adjustable subtraction counter has an output. After several cycles, SELF_SUB_<3:0> will output "0000". When the value of SELF_SUB_<3:0> is "0000", it controls AND1 to end and makes the output signal EN_II of the adjustable subtraction counter 0. The adjustable subtraction counter completes all functions.
[0108] Please continue to refer to Figure 9, in this embodiment, a comparator is used to determine whether the exponent sum difference and the value of the adjustable subtraction counter are equal. The comparator includes exclusive-NOR gate circuits XNOR1, XNOR2, XNOR3, XNOR4, and AND gate circuits AND3, AND4. One input terminal A1 of XNOR1 is connected to the signal SELF_SUB<3>, and the other input terminal A2 is connected to the difference SUB<3> between the maximum exponent and the exponent sum. The output terminal is connected to the input terminal A1 of AND3. One input terminal A1 of XNOR2 is connected to the signal SELF_SUB<2>, and the other input terminal A2 is connected to the difference SUB<2> between the maximum exponent and the exponent sum. The output terminal is connected to the input terminal A2 of AND3. One input terminal A1 of XNOR3 is connected to the signal SELF_SUB<1>, and the other input terminal A2 is connected to the difference SUB<1> between the maximum exponent and the exponent sum. The output terminal is connected to the input terminal A3 of AND3. One input terminal A1 of XNOR4 is connected to the signal SELF_SUB<0>, and the other input terminal A2 is connected to the difference SUB<0> between the maximum exponent and the exponent sum. The output terminal is connected to the input terminal A4 of AND3. The output terminal of AND3 is connected to the input terminal A1 of AND4. The input terminal A2 of AND4 is connected to the output terminal of the adjustable subtraction counter. The input terminal A3 is connected to the exponent sum subtraction completion signal EN_I. The output terminal Z outputs the addition control signal EN. The four exclusive-NOR gate circuits XNOR1, XNOR2, XNOR3, and XNOR4 and the AND gate circuit AND3 are used to determine whether the output value of the adjustable subtraction counter is equal to the exponent sum difference. The AND gate circuit AND4 is used to determine whether the exponent sum difference operation, the controllable subtraction calculator, and the comparator judgment are all completed.
[0109] The operation logic of the comparator includes:
[0110] (1) Determine whether the corresponding bits of the signals SELF_SUB<3:0> and the difference SUB<3:0> are equal according to the outputs of XNOR1, XNOR2, XNOR3, and XNOR4. Among them, the high level of the outputs of the four exclusive-NOR gates indicates that the four bits of the exponent sum and the output value of the adjustable subtraction counter are respectively the same; the low level of the outputs of the four exclusive-NOR gates indicates that the four bits of the exponent sum and the output value of the adjustable subtraction counter are respectively different.
[0111] When SELF_SUB<3> is at a high level and SUB<3> is at a high level, the output of the exclusive-NOR gate XNOR1 is at a high level.
[0112] When SELF_SUB<3> is at a low level and SUB<3> is at a low level, the output of the exclusive-NOR gate XNOR1 is at a high level.
[0113] When SELF_SUB<3> is at a high level and SUB<3> is at a low level, the output of the exclusive-NOR gate XNOR1 is at a low level.
[0114] When SELF_SUB<3> is at low level and SUB<3> is at high level, the output of the XNOR gate XNOR1 is at low level.
[0115] The logic of SELF_SUB<2> AND SUB<2>, SELF_SUB<1> AND SUB<1>, SELF_SUB<0> AND SUB<0> is exactly the same as that of SELF_SUB<3> AND SUB<3>, so it will not be elaborated here.
[0116] (2) Judge whether the signals SELF_SUB<3:0> and the difference SUB<3:0> are exactly equal according to the output of AND3. Among them, the output of the AND gate AND3 being at high level indicates that the exponent sum is exactly the same as the output value of the adjustable subtraction counter; the output of the AND gate AND3 being at low level indicates that the exponent sum is different from the output value of the adjustable subtraction counter.
[0117] Take SELF_SUB<3:0> being "1001" as an example:
[0118] When SELF_SUB<3:0> is "1001" and SUB<3:0> is "1001", the outputs of the four XNOR gates XNOR1, XNOR2, XNOR3, and XNOR4 are all at high level, that is, all four input terminals of the AND gate AND3 are at high level, so the output of the AND gate AND3 is at high level.
[0119] When SELF_SUB<3:0> is "1001" and SUB<3:0> is not "1001", the outputs of the four XNOR gates XNOR1, XNOR2, XNOR3, and XNOR4 are not all at high level, that is, not all four input terminals of the AND gate AND3 are at high level, so the output of the AND gate AND3 is at low level.
[0120] (3) Judge whether to make the addition control signal at high level through AND3 and AND4. If so, input the corresponding mantissa product into the adder tree for addition; otherwise, control the adder tree not to perform mantissa product accumulation addition. Among them, the output of the AND gate AND4 being at high level, that is, the signal EN_II being at high level and the parallel exponent sum subtraction completion signal EN_I being at high level, represents that the output SELF_SUB<3:0> of the adjustable subtraction counter is exactly equal to the exponent sum difference SUB<3:0>, indicating that the control addition signal is "1" to control the adder tree to execute the addition function. The output of the AND gate AND4 being at low level represents that when the output SELF_SUB<3:0> of the controllable subtraction counter is not equal to the exponent sum difference SUB<3:0>, it indicates that the control addition signal is "0" to control the adder tree not to execute the addition function.
[0121] The output value of the adjustable subtraction counter is transmitted to the comparator circuit. Meanwhile, the subtractor calculates the difference between the global maximum value obtained by the maximum value finding module and the exponent sum of the current row, and synchronously inputs the exponent sum difference into the comparator. By comparing the value of the adjustable subtraction counter with the exponent sum difference, the comparator generates a corresponding addition control signal. This signal is used to control the mantissa product corresponding to this row to enter the adder tree for accumulation operation. After all row exponent sums have been compared, the adjustable subtraction counter performs a self-subtraction operation by 1, and based on the updated value, initiates the next round of comparison and mantissa product accumulation process. The above process is iterated until the adjustable subtraction counter reaches zero. At this time, the adder tree has completed the accumulation calculation of all mantissa products, marking the end of a complete mantissa operation.
[0122] Please continue to refer to Figure 10 , in the solution of this embodiment, an adjustable subtraction counter, a comparator, and an adder tree are used together to complete the process of mantissa shifting and mantissa summation, and the process of mantissa shifting is omitted. The specific process is as follows: (1) Initial matching stage: Obtain the exponent sum difference of the first row and match it with the current output value of the controllable subtraction counter (initially the maximum exponent difference): If the match is successful (the difference is equal to the counter value), then input the mantissa product of this row into the adder tree for accumulation; if the match fails, skip the current row and read the exponent sum difference of the next row until all rows have been traversed. (2) Loop iteration stage: When all rows have been traversed, if the counter value is non-zero, perform the operations of "subtract 1 from the counter → reset the row index → traverse all rows again", and then execute the above matching and accumulation process again. (3) Termination condition: When the controllable subtraction counter reaches zero, the adder tree has completed the accumulation of all mantissa products that meet the matching conditions, and directly outputs the mantissa sum.
[0123] In this embodiment, the floating-point arithmetic circuit can complete the floating-point multiply-accumulate operation task in a pipelined manner in practical applications. The process is as follows:
[0124] Initial stage: Within one cycle, use the exponent input module, weight exponent array, and adder to complete the addition of the exponent parts of each operand and the weight in the first round of operation to obtain the exponent sum. Store the exponent sum in the exponent sum array. Use the mantissa input module, weight mantissa array, and adder to complete the multiplication of the mantissa parts of each operand and the weight in the first round of operation. The maximum value finding module obtains the largest exponent sum in the exponent sum array, and obtains the difference between the largest exponent sum and the exponent sum difference of each row through the subtractor.
[0125] Operation stage:
[0126] In the second cycle, a comparator is used to compare the value of the adjustable subtraction counter with the exponent sum and difference to generate a control signal. This signal is used to control whether the mantissa product enters the adder tree for addition, obtaining the mantissa sum. Finally, the normalization module generates the corresponding multiply-accumulate operation result based on the determined maximum exponent and mantissa sum.
[0127] Loop stage:
[0128] According to the above logic, in each subsequent cycle, the operation result of the previous round is calculated and output, and the subsequent exponent sum addition and storage are completed.
[0129] Compared with the existing floating-point arithmetic circuits, the SRAM-based floating-point arithmetic circuit of this embodiment has the following beneficial effects:
[0130] 1. The SRAM-based floating-point arithmetic circuit, through the "global difference decreasing traversal + multi-round row data screening" mechanism, converts the mantissa shift operation into a condition-triggered mantissa accumulation addition, avoiding complex shift circuits at the hardware level, and at the same time using the parallelism of the adder tree to improve the calculation efficiency, solving the technical problems of low efficiency, slow speed, and low precision existing in the existing arithmetic circuits when processing floating-point data.
[0131] 2. The SRAM-based floating-point arithmetic circuit has relatively simple floating-point arithmetic steps, can be compatible with the parallel computing characteristics of the CIM architecture, can solve the imbalance between the limited resources of edge devices and the high computing power requirements of large-scale floating-point models, can optimize the arithmetic process and hardware resource allocation, can release the potential of CIM, and support the implementation of the next-generation intelligent applications.
[0132] Embodiment 3
[0133] This embodiment provides an in-memory computing chip (CIM chip), which includes the SRAM-based floating-point arithmetic circuit in Embodiment 1 or Embodiment 2, and this circuit can be integrated on the chip. The CIM chip has a storage mode and a computing mode. In the storage mode, the CIM chip is used as a memory. In the computing mode, the CIM chip is used to implement the multiply-accumulate operation between multiple groups of multi-bit floating-point input features and multi-bit floating-point weights.
[0134] Embodiment 4
[0135] This embodiment provides a static random access memory (SRAM), and this memory uses the SRAM-based floating-point arithmetic circuit in Embodiment 1 or Embodiment 2 to implement the multiply-accumulate calculation of multi-bit inputs and multi-bit weights.
[0136] Based on the SRAM-based floating-point arithmetic circuit in Embodiment 1 or Embodiment 2, the in-memory computing of SRAM in this embodiment directly completes the multiply-accumulate operation within the storage unit, reducing data movement and significantly lowering power consumption. SRAM can simultaneously process multiply-accumulate operations of multiple inputs and weights, greatly improving computing efficiency, and is particularly suitable for application scenarios that require high throughput. The read / write speed of SRAM is much higher than that of DRAM and flash memory, enabling low-latency multiply-accumulate calculations, and is suitable for applications with high real-time requirements, such as edge computing and Internet of Things devices. SRAM can be integrated with other computing units (such as CPU, GPU) on the same chip to form an efficient in-memory computing architecture.
[0137] The SRAM in this embodiment is applicable to artificial intelligence and machine learning. The inference and training processes of neural networks involve a large number of multiply-accumulate operations, and the in-memory computing of SRAM can significantly accelerate these operations and improve overall performance. Multi-bit inputs and weights enable SRAM to support from simple linear models to complex deep neural networks. The in-memory computing of the SRAM in this embodiment reduces the complex interface between the memory and the processor in the traditional computing architecture and simplifies the system design. By reducing data movement and simplifying the architecture, SRAM can reduce the overall cost and power consumption of the system.
[0138] Embodiment 5
[0139] This embodiment provides an electronic device, which includes a memory and a processor. Among them, the memory includes the SRAM-based floating-point arithmetic circuit in Embodiment 1 or Embodiment 2. Compared with existing electronic devices, this electronic device can significantly improve computing efficiency, reduce power consumption, and support high-precision computing. It has broad application prospects in the fields of artificial intelligence, edge computing, etc. Although it faces some technical challenges, its advantages make it an important technical direction for in-memory computing.
[0140] Embodiment 6
[0141] This embodiment provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. Among them, the memory is the static random access memory in Embodiment 4.
[0142] This computer device can take various forms. It can either adopt an embedded chip or module, or a general-purpose data processing device, such as an intelligent terminal capable of executing programs, a tablet computer, a laptop computer, a desktop computer, a rack-mounted server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc.
[0143] The computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can communicate with each other through a system bus. The memory (i.e., a readable storage medium) includes flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device.
[0144] In some embodiments, the processor may be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data.
[0145] Embodiment 7
[0146] This embodiment provides a floating-point operation method based on SRAM. This method can use the circuits in Embodiment 1 or Embodiment 2 and specifically includes the following steps:
[0147] S1: Calculate the exponent sum and mantissa product of the multi-bit floating-point input value and the multi-bit floating-point weight value bit by bit respectively;
[0148] S2: Search for and determine the maximum exponent bit by bit among all the exponent sums, and then calculate the difference between the maximum exponent and the exponent sums of each row;
[0149] S3: First generate a self-decrement value with an adjustable initial value, then compare the self-decrement value with each difference to generate an addition control signal, and finally determine whether to input the mantissa product of each row into the adder tree for addition according to the addition control signal;
[0150] S4: After comparing with the exponent sums of all rows, perform a self-decrement operation on the self-decrement value and execute step S3; when the self-decrement value reaches zero, use the accumulated sum of the mantissas in the adder tree as the total number of bits;
[0151] S5: Generate the multiply-accumulate operation result of the input value and the weight value according to the maximum exponent and the total number of bits.
[0152] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A floating-point arithmetic circuit based on SRAM, characterized in that, It is used to implement the following steps: S1: Calculate the sum of exponents and the product of mantissas of the multi-bit floating-point input value and the multi-bit floating-point weight value bit by bit respectively; S2: Search for and determine the maximum exponent bit by bit among all the sums of exponents, and then calculate the difference between the maximum exponent and the sum of exponents of each row; S3: First generate a self-decrement value with an adjustable initial value, then compare the self-decrement value with each difference to generate an addition control signal, and finally determine whether to input the mantissa product of each row into the adder tree for addition according to the addition control signal; S4: After comparing with the sums of exponents of all rows, the self-decrement value performs a self-decrement operation, and step S3 is executed; when the self-decrement value reaches zero, the accumulated sum of mantissas in the adder tree is used as the total number of bits; S5: Generate the multiply-accumulate operation result of the input value and the weight value according to the maximum exponent and the total number of bits.
2. The SRAM-based floating-point arithmetic circuit according to claim 1, wherein The operation circuit includes: An array circuit, which includes a weight exponent array, a sum-of-exponents array, and a weight mantissa array; each row of the weight exponent array is used to pre-store the exponent part of the weight value bit by bit; the sum-of-exponents array is used to store the sum of exponents of each row; each row of the weight mantissa array is used to pre-store the mantissa part of the weight value bit by bit, and perform a multiplication operation on the mantissa parts of the input value and the weight value to obtain the mantissa product.
3. The SRAM-based floating-point arithmetic circuit according to claim 2, wherein The weight mantissa array includes a plurality of storage units, and each storage unit includes NMOS transistors N1 to N6 and PMOS transistors P1, P2; N1, N2, P1, and P2 are inversely cross-coupled to form a pair of storage nodes Q and QB; the gates of N3 and N4 are connected to the word line WL, N3 is the transfer transistor between the bit line BL and the node Q, and N4 is the transfer transistor between the bit line BLB and the node QB; the gate of N5 is connected to the node Q, the source is grounded, and the drain is connected to the drain of N6; the gate of N6 is connected to the calculation word line LRWL, and the source is connected to the calculation word line LHBL; N5 and N6 are used to calculate the mantissa product, and NMOS transistors N1 to N4 and PMOS transistors P1 and P2 are used to pre-store the mantissa part of the weight.
4. The SRAM-based floating-point arithmetic circuit according to claim 3, wherein The operation logic of the weight mantissa array includes: (1) When the node Q is at a low level and the node QB is at a high level, the corresponding bit pre-stored in the mantissa part of the weight is "0"; when the node Q is at a high level and the node QB is at a low level, the corresponding bit pre-stored in the mantissa part of the weight is "1"; (2) First pre-charge the calculation bit line LHBL to a high level, and then input the mantissa of the operand into the corresponding storage unit through the calculation bit line LRWL; among them, when the calculation word line LRWL is at a low level, it means that the corresponding bit of the mantissa part of the input operand is "0", and when the calculation word line LRWL is at a high level, it means that the corresponding bit of the mantissa part of the input operand is "1"; (3) Output the calculation result of the final mantissa product according to the level change of the calculation bit line LHBL.
5. The SRAM-based floating-point arithmetic circuit according to claim 1, wherein The operation circuit includes: An adjustable subtraction counter, which includes D flip-flops DFF2, DFF3, DFF4, DFF5 with asynchronous set and clear functions, a D flip-flop DFF6 with an asynchronous set function, a D flip-flop DFF1 with an asynchronous clear function, delay elements DEL1, DEL2, DEL3, DEL4, buffers BUFF1, BUFF2, an inverter NOT, AND gates AND1, AND2, NAND gates NAND1, NAND2, NAND3, OR gates OR1, OR2; the input end of BUFF1 is connected to the controllable signal A<1>, and the output end is connected to the input end of NOT; the output end of NOT is connected to one of the input ends of NAND1; the input end of BUFF2 is connected to the controllable signal A<0>, and the output end is connected to the other input end of NAND1; the input end of DEL3 is connected to the maximum value search end signal, and the output end is connected to one of the input ends of NAND2; the other input end of NAND2 is connected to the reverse of the maximum value search end signal, and the output end is connected to one of the input ends of OR2; the output end of OR2 outputs the signal SET_I<3:0>; the input end of DEL1 is connected to the maximum value search end signal, and the output end is connected to one of the input ends of DFF1; the other input end of DFF1 is connected to the power supply VDD, and the output end is connected to one of the input ends of AND1; the output end of OR1 at the other input end of AND1 is connected to one of the input ends of AND2; the other input end of AND2 is connected to the CLK signal, and the output end is connected to the input CPN end of DFF2; the input end D and the output end QN of DFF2 are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF3 is connected to the output end QN of DFF2, the input end D and the output end QN are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF4 is connected to the output end QN of DFF3, the input end D and the output end QN are connected, and the input clear end is connected to the clear signal CLR; the input end of DFF5 is connected to the output end QN of DFF4, the input end D and the output end QN are connected, and the input clear end is connected to the clear signal CLR; one of the input ends of DEL2 is connected to the output end of AND2, the output end is connected to one of the input ends of DFF6, and the input clear end is connected to the clear signal CLR; one of the input ends of DEL4 is connected to the output end of AND2, and the output end is connected to one of the input ends of NAND3; the other input end of NAND3 is connected to the output end of AND2, and the output end is the output end of the adjustable subtraction counter.
6. The SRAM-based floating-point arithmetic circuit according to claim 5, wherein The operation logic of the adjustable subtraction counter includes: (1) Controlling the initial value of the signal SET_I<3:0> according to the input values of the controllable signals A<1> and A<0>; (2) Connect the signals SET_I<3:0> to the set-to-1 terminals of DFF2, DFF3, DFF4, and DFF5 respectively, and generate the values of the signal SELF_SUB_I<3:0> with the output initial value incremented by 1 according to the signals SET_I<3:0>; wherein, another input terminal of DFF6 is connected to the signal SELF_SUB_I<3:0>, and the QN output terminals of DFF2, DFF3, DFF4, and DFF5 are respectively connected to the signals SELF_SUB_I<0>, SELF_SUB_I<1>, SELF_SUB_I<2>, and SELF_SUB_I<3>; (3) Asynchronously set the initial value of the signal SELF_SUB_I<3:0> through DFF6 and delay it to obtain the initial value of the signal SELF_SUB<3:0>.
7. The SRAM-based floating-point arithmetic circuit according to claim 6, wherein The arithmetic circuit further includes: A comparator, which includes exclusive-NOR gate circuits XNOR1, XNOR2, XNOR3, XNOR4, and AND gate circuits AND3, AND4; one input terminal of XNOR1 is connected to the signal SELF_SUB<3>, and another input terminal is connected to the difference SUB<3> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A1 of AND3; one input terminal of XNOR2 is connected to the signal SELF_SUB<2>, and another input terminal is connected to the difference SUB<2> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A2 of AND3; one input terminal of XNOR3 is connected to the signal SELF_SUB<1>, and another input terminal is connected to the difference SUB<1> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A3 of AND3; one input terminal of XNOR4 is connected to the signal SELF_SUB<0>, and another input terminal is connected to the difference SUB<0> between the maximum exponent and the sum of exponents, and the output terminal is connected to the input terminal A4 of AND3; the output terminal of AND3 is connected to the input terminal A1 of AND4; the input terminal A2 of AND4 is connected to the output terminal of the adjustable subtraction counter, the input terminal A3 is connected to the exponent sum subtraction completion signal EN_I, and the output terminal outputs the addition control signal.
8. The SRAM-based floating-point arithmetic circuit according to claim 1, wherein The operation logic of the comparator includes: (1) Judge whether the corresponding bits of the signals SELF_SUB<3:0> and the difference SUB<3:0> are equal according to the outputs of XNOR1, XNOR2, XNOR3, and XNOR4; (2) Judge whether the signals SELF_SUB<3:0> and the difference SUB<3:0> are completely equal according to the output of AND3; (3) Judge whether to make the addition control signal high level through AND3 and AND4. If so, input the corresponding mantissa product into the adder tree for addition, otherwise control the adder tree not to perform mantissa product accumulation addition.
9. The SRAM-based floating-point arithmetic circuit according to claim 2, wherein The arithmetic circuit further includes: A maximum value search module, which is used to search for and determine the maximum exponent bit by bit in the exponent sum array; An exponent input module; An adder, which is used to first receive the exponent part of the input value input by the exponent input module, simultaneously read the exponent part stored in the weight exponent array, then calculate the exponent sum of the two exponent parts, and finally store the calculation result of the exponent sum into the exponent sum array; An adjustable subtraction counter, which is used to generate a self-decreasing value with an adjustable initial value, then compare the self-decreasing value with each difference through a comparator to generate an addition control signal, and finally determine whether to input the mantissa products of each row into the adder tree for addition according to the addition control signal. After comparing with the exponent sums of all rows, the self-decreasing value performs a self-decreasing operation until the self-decreasing value reaches zero, and the addition result of the mantissa products in the adder tree is used as the total number of bits; A mantissa input module, which is used to input an operand to the corresponding row of the weight mantissa array; An adder tree and a normalization module, which provide the adder tree and are used to generate the multiplication and accumulation operation result of the input value and the weight value according to the maximum exponent and the total number of bits; 10. An in-memory computing chip, characterized in that, It includes the SRAM-based floating-point arithmetic circuit according to any one of claims 1-9.
Citation Information
Cited By
Sram floating point in-memory computing architecture and computing method
CN121349406B