Floating-point number index maximum value searching circuit and in-memory computing chip

By using SRAM cell matrix and logic control in the floating-point exponential maximum value search circuit, excluding the indices that cannot be the maximum value in real time, solving the power consumption and delay problems of existing circuits, achieving fast and low-power exponential maximum value search, and improving floating-point calculation performance.

CN120388591APending Publication Date: 2025-07-29ANHUI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510485496.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing floating-point exponential maximum value search circuit has high power consumption, complex hardware implementations and long computing delays, which cannot meet the needs of low power consumption and high speed in CIM architectures.

Method used

A floating-point exponential maximum value search circuit is adopted, and the SRAM cell matrix and logic control is used to control the charging and discharging of nodes through word lines and bit lines, detect the signal status line by line, find the index that cannot be the maximum value in real time and exclude it in subsequent calculations. Combined with the efficient storage characteristics of 7T-SRAM, a fast and low-power exponential maximum value search is achieved.

Benefits of technology

It saves the number of search cycles, shortens the calculation time, reduces power consumption, improves the calculation efficiency, solves the power consumption and delay problems of existing circuits, and supports efficient floating-point number operations, especially in CIM architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388591A_ABST
    Figure CN120388591A_ABST
Patent Text Reader

Abstract

The invention discloses a floating-point number index maximum value searching circuit and an in-memory computing chip. The searching circuit comprises an SRAM (Static Random Access Memory) unit matrix which comprises a plurality of rows and columns of SRAM units, wherein each unit comprises NMOS (N-channel Metal Oxide Semiconductor) tubes N1-N5 and PMOS (P-channel Metal Oxide Semiconductor) tubes P1 and P2; n1, N2, P1 and P2 are in reverse cross coupling to form a pair of storage nodes Q and QB; the grid electrode of the N5 is connected with a node QB; the grid electrode of the N3 in the ith row is connected with a word line WLLlt; igt; the grid electrode of the N4 is connected with a word line WLRlt; igt; ; the drain electrode of each column of N5 is connected in series with the source electrode of the next column of N5, and the source electrode of the first column of N5 is connected with a signal SELlt; igt; the N5 drain electrode of the last column is connected with a signal PRE; in the jth column, N3 is a corresponding node Q and a bit line BLlt; jgt; the transmission tube N4 is a corresponding node QB and a bit line BLBlt; jgt; and a transmission tube. According to the invention, the number of searching cycles is saved, the calculation time is shortened, the subsequent calculation amount is reduced, the efficiency is improved, the power consumption is reduced, and the defects of relatively high power consumption, complex hardware implementation, relatively long calculation delay and the like of an existing circuit are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a maximum value finding circuit in the field of integrated circuits, in particular to a floating-point exponent maximum value finding circuit, and also relates to an in-memory computing chip. Background Art

[0002] In-memory computing (CIM, Compute-In-Memory) is an emerging computing architecture aiming to closely integrate data storage and computing processes, thereby reducing data transmission latency and improving computing efficiency. CIM technology embeds computing operations directly into storage units, eliminating the frequent transfer of data between the memory and the central processing unit in the traditional von Neumann architecture, which has important application values in large-scale data processing, artificial intelligence, deep learning and other fields.

[0003] In CIM, many scientific computing, machine learning, and image processing tasks rely on the high precision and wide range of the floating-point data type. The floating-point format can represent very large or very small numerical values and provide sufficient precision in complex mathematical operations. The FP16 (16-bit floating-point) format, as a common floating-point representation method, is widely used in image processing and deep learning and other fields due to its advantages in storage space and operation speed. The FP16 format divides the floating-point number into a sign bit, an exponent bit, and a mantissa bit, where the size of the exponent bit determines the range of the floating-point number. In FP16 calculations, especially when performing addition or multiply-accumulation operations, correctly aligning different mantissas according to the sizes of their respective exponents is a necessary step to ensure the accuracy and efficiency of the operation result. For this reason, quickly finding the maximum exponent value in FP16 numbers and then guiding the mantissa alignment operation has become a key link in FP16 operations. However, existing circuits for finding the maximum exponent value usually rely on methods such as multi-stage comparators, tree structures, or parallel computing architectures, and often face defects such as high power consumption, complex hardware implementation, and long calculation latency. In the CIM architecture, the limitations of these traditional methods are particularly prominent because they cannot effectively meet the requirements of low power consumption and high speed. Summary of the Invention

[0004] To solve the technical problems that existing circuits for finding the maximum exponent number have defects such as high power consumption, complex hardware implementation, and long calculation latency, the present invention provides a floating-point exponent maximum value finding circuit and an in-memory computing chip.

[0005] The present invention is implemented by the following technical solutions: A floating-point exponent maximum value finding circuit, which includes:

[0006] SRAM cell matrix: It includes SRAM cells arranged in multiple rows and columns. Each cell includes NMOS transistors N1 to N5 and PMOS transistors P1 and P2; N1, N2, P1, and P2 are cross-coupled in an inverter configuration to form a pair of storage nodes Q and QB; the gate of N5 is connected to node QB;

[0007] Row control: The gate of N3 in the i-th row is connected to the word line WLL , the N4 gate is connected to the word line WLR ;

[0008] Column cascade: The drain of N5 in each column is connected in series with the source of N5 in the next column, and the source of N5 in the first column is connected to signal SEL , the drain of N5 in the last column is connected to the signal PRE;

[0009] Bit line control: In the j-th column, N3 is for the corresponding node Q and the bit line BL <j>The transmission pipe, and N4 is the corresponding node QB and bit line BLB <j>transmission tube;

[0010] Maximum value judgment logic: (1) Through the word line (WLL , WLR ) and bit line (BL <j>, BLB <j>)Control the charging and discharging of the control nodes Q and QB so that the unit in the i-th row and j-th column stores the i-th bit of the j-th floating-point exponent; (2) Detect the signal SEL row by row Transmission status with signal PRE: If signal PRE can be transmitted to signal SEL At the end, if the transmission is not blocked, determine that this bit is 0; if the transmission is blocked, determine that this bit is 1; (3) Combine the determination results of each row to obtain the maximum exponent value.

[0011] In the process of finding the maximum exponent value in the present invention, the circuit will in real time find out the exponents that can no longer be the maximum value, and exclude and mask them in subsequent calculations. This not only saves the number of cycles for searching, shortens the calculation time, reduces the subsequent calculation amount, improves the efficiency, but also reduces the power consumption, and solves the technical problems of the existing circuit for finding the maximum exponent number, such as high power consumption, complex hardware implementation, and long calculation delay.

[0012] Further, the searching circuit further includes:

[0013] Multiple configuration circuits, respectively corresponding to multiple rows of SRAM cells; each configuration circuit includes a multiplexer MUX1, MUX2, MUX3, and a D flip-flop DFF3; the control end of MUX1 is connected to the corresponding signal SEL Connection, where one output terminal is connected to one input terminal of MUX2, and the other output terminal is connected to the input terminal D of DFF3; another input terminal of MUX2 is connected to the output terminal of DFF3 and one input terminal of MUX3, and the output terminal outputs the signal D<i+1>; another input terminal of MUX3 is connected to the external signal WL , the output terminal of MUX3 is connected to the corresponding word line WLL , the control terminal is connected to an external signal R_C_N; the clock signal input terminal of DFF3 is connected to an external signal CP, and the output terminal outputs a signal WL_S ; external signal WL With word line WLR Connection.

[0014] Furthermore, the searching circuit further includes:

[0015] An initialization circuit, which includes D flip-flops DFF1 and DFF2; the input terminal D of DFF1 is connected to the signal VSS, the clock signal input terminal is connected to the external signal CP, the reset terminal is connected to the external signal SET, and the output terminal is connected to the input terminal D of DFF2 and outputs the signal D_SET; the clock signal input terminal of DFF2 is connected to the external signal CP, and the output terminal is connected to the input terminal of MUX1 and outputs the signal D<0>.

[0016] Furthermore, the searching circuit further includes:

[0017] A selector circuit, which includes multiplexers MUX4 corresponding to multiple rows of SRAM cells respectively; two input terminals of MUX4 are respectively connected to the signals VSS and VDD, the output terminal is connected to the drain of the last N5 in the corresponding row, and the control terminal is connected to the signal PRE.

[0018] Furthermore, the searching circuit further includes:

[0019] A read / write control circuit, which is used to the bit line BL <j>, BLB <j>Writing data and precharging;

[0020] An amplifier circuit, which includes a plurality of sense amplifiers respectively corresponding to a plurality of bit lines; two input ends of each sense amplifier are respectively connected to the corresponding bit line BL <j>, BLB <j>, the control terminal is connected to an external signal SA_EN, and the output terminal outputs a signal OUT_B <j>。

[0021] Furthermore, the searching circuit further includes:

[0022] A column shielding circuit, which includes a plurality of column shielding flip-flops respectively corresponding to a plurality of sense amplifiers, and a plurality of column shielding selectors respectively corresponding to the plurality of column shielding flip-flops; the input terminal D of each column shielding flip-flop is connected to the output terminal of the corresponding sense amplifier, the clock signal input terminal is connected to an external signal SA_EN, the clear terminal is connected to an external signal MAX_CLR, the set terminal is connected to an external signal CLR, and the output terminal Q is connected to the control terminal of the corresponding column shielding selector and outputs a signal IVDD_C <j>; Two input terminals of each column masking selector are respectively connected to signal VSS and VDD, and the output terminal outputs signal IVDD <j>。

[0023] Further, the strategy for the searching circuit to write the floating-point exponent includes:

[0024] In the initial state, the word line WLL , WLR Set to low level, and according to the floating-point exponent, the bit line BL <j>, BLB <j>Perform a set operation; among them, when the data to be written is "1", the bit line BL <j>Set to high level, bit line BLB <j>Set to low level; when the data to be written is "0", the bit line BL <j>Set to low level, bit line BLB <j>Set to high level;

[0025] When writing data, set the word line WLL , WLR Set to high level; among them, when the data to be written is "1", it is passed through the bit line BL <j>Charge node Q to pull node Q high to the high level and via bit line BLB <j>Discharge node QB to pull node QB high to low level; when the data to be written is "0", via bit line BL <j>Discharge node Q to pull node Q low to a low level and through bit line BLB <j>Charge node QB to pull it high to the high level;

[0026] When in the hold state, word line WLL , WLR Set to low level.

[0027] Furthermore, the strategy of the searching circuit for finding the maximum value of multiple floating-point exponents includes:

[0028] (1) Make all nodes QB at high level, and make the MUX4 output at high level through the signal PRE and transmit it to the signal SEL , so that the entire line where each N5 is located is pre-charged to a high level;

[0029] (2) Store multiple multi-bit exponents, all of which are floating-point numbers, in multiple SRAM cells. In each SRAM cell, it is defined that when node Q is at a high level and node QB is at a low level, it indicates that the stored data is "1", and when node Q is at a low level and node QB is at a high level, it indicates that the stored data is "0";

[0030] (3) Make the output of MUX4 a low level through signal PRE;

[0031] (4) Find the maximum exponent in the array: In the initial state, make signals PRE_CHAR, R_C_N, and SA_EN all at a high level, signal SET generates a low-level pulse with a width of half a cycle, and set the signal D_SET of the initialization circuit to a high level; when the rising edge of the second cycle of the external signal CP arrives, through the read / write control circuit, for bit line BL <j>, BLB <j>Precharge to high level; after the external signal CP is delayed by a quarter of a cycle, the signal SA_EN generates a high level with a width of half a cycle, enabling multiple sense amplifiers to start working; in the second half of the second cycle of the external signal CP, the signal R_C_N generates a low level with a width of half a cycle, controlling MUX3 to select WL_S The high level is transmitted to the WLL , turn on the word line WLL , multiple floating-point exponents will be characterized on the bit line BL; the signal R_C_N controls the read / write control circuit to turn on the bit line BL <j>, BLB <j>The path between the multiple sense amplifiers, and the signal SA_EN also remains at a high level. The data in the i-th row of SRAM cells is read out by the corresponding sense amplifier and inverted and output to all signals OUT_B <j>; After a delay of one-quarter of a cycle of the external signal CP, the signal SA_EN drops to a low level, and the falling edge of the signal SA_EN controls the corresponding flip-flop to transfer all the signals OUT_B to the corresponding signal IVDD_C;

[0032] (5) When all the nodes Q in the i-th row are at a low level, the signal SEL is at a low level;

[0033] (6) After searching all the arrays composed of multiple SRAM cells, all the multi-bit exponents that are not the exponential maximum value are cleared, and the exponential maximum value is represented by the signal SEL.

[0034] As a further improvement of the above solution, the searching circuit is used to find the exponential maximum value among 32 six-bit exponents. Each column of 6 SRAM cells is responsible for storing one 6-bit exponent. The first row stores the most significant bit, and the sixth row stores the least significant bit.

[0035] The present invention also provides an in-memory computing chip, which includes any one of the above floating-point exponent maximum value searching circuits.

[0036] Compared with the existing exponent maximum value searching circuits, the floating-point exponent maximum value searching circuit and the in-memory computing chip of the present invention have the following beneficial effects:

[0037] 1. In the process of finding the exponential maximum value, the floating-point exponent maximum value searching circuit can find out the exponents that are no longer likely to be the maximum value in real time, exclude and mask them in subsequent calculations, reduce the subsequent calculation amount, improve the efficiency, reduce the power consumption, and solve the technical problems of the existing exponent maximum value searching circuits, such as high power consumption, complex hardware implementation, and long calculation delay.

[0038] 2. In the process of finding the exponential maximum value, if all Qs of a certain row of cells are 0, then in this cycle, the circuit will skip this row and directly look for "1" in the next row, which not only saves the number of cycles and shortens the calculation time, but also further reduces the power consumption.

[0039] 3. By combining the efficient storage characteristics of 7T-SRAM, the floating-point exponent maximum value searching circuit realizes the function of quickly and low-power finding the exponential maximum value. This design can not only greatly reduce the power consumption and delay, but also improve the calculation efficiency, and can better support efficient floating-point operations, especially in the application of the CIM architecture. This circuit design provides an innovative solution for finding the exponential maximum value in FP16 operations, which helps to improve the performance of floating-point calculations. Description of the Drawings

[0040] Figure 1 is the circuit diagram of the SRAM cell of the floating-point exponent maximum value searching circuit according to Embodiment 1 of the present invention;

[0041] Figure 2 is the circuit diagram of the structure of the first row in the SRAM cell matrix of the floating-point exponent maximum value searching circuit according to Embodiment 1 of the present invention;

[0042] Figure 3 Circuit diagram of the multi - row structure in the SRAM cell matrix of the floating - point exponent maximum finding circuit according to Embodiment 1 of the present invention;

[0043] Figure 4 Circuit diagram of the last row in the SRAM cell matrix of the floating - point exponent maximum finding circuit according to Embodiment 1 of the present invention, as well as the read - write control circuit, the amplification circuit, and the column shielding circuit;

[0044] Figure 5 Waveform diagram of some signals in the floating - point exponent maximum finding circuit according to Embodiment 1 of the present invention. Detailed implementation manners

[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] Embodiment 1

[0047] Please refer to Figures 1-4 This embodiment provides a floating - point exponent maximum finding circuit, which is used to find the maximum exponent, that is, the exponent maximum, among a plurality of multi - bit exponents that are all floating - point numbers. In this embodiment, the finding circuit is used to find the exponent maximum among 32 6 - bit exponents. Each column of 6 SRAM cells is responsible for storing 1 6 - bit exponent. The first row stores the most significant bit, and the sixth row stores the least significant bit. In other embodiments, the number of multi - bit exponents and the number of exponent bits can be different and can be set according to actual needs. Among them, the finding circuit includes an SRAM cell matrix, and may also include a plurality of configuration circuits, initialization circuits, selector circuits, read - write control circuits, amplification circuits, column shielding circuits, etc., and may also include various logics and strategies.

[0048] Please continue to refer to Figure 1 , the SRAM cell matrix includes multiple rows and columns of SRAM cells, and each cell includes NMOS transistors N1 to N5 and PMOS transistors P1 and P2. N1, N2, P1, and P2 are cross-coupled in an inverter configuration to form a pair of storage nodes Q and QB, and the gate of N5 is connected to node QB. In fact, each SRAM cell contains a pair of cross-coupled inverters, forming complementary storage nodes Q and QB. These transistors constitute a 7T-SRAM cell, which not only has the same read, write, and hold functions as a traditional 6T-SRAM cell, but also has the function of clearing the stored data. Among them, P1, P2, and N1 to N4 constitute a 6T storage cell and support the implementation of data storage functions including data read, write, and hold. The circuit formed by the remaining N5 is used to capture the state of QB in the 6T storage cell in real time and provide a path connected to other storage cells in the same row: WL_WN_I and WL_WN_O.

[0049] Specific connection relationships: The source of P1 is connected to locally controlled IVDD, and the source of P2 is connected to VDD. The gates of P2 and N2 and the drains of P1 and N1 are connected and used as storage node Q. The gates of P1 and N1 and the drains of P2 and N2 are connected and used as storage node QB. The sources of N1 and N2 are grounded. N3 serves as a transfer transistor between storage node Q and bit line BL. N4 serves as a transfer transistor between storage node QB and bit line BLB. The gate of N3 is connected to the left word line WLL, and the gate of N4 is connected to the right word line WLR. The gate of N5 is connected to storage node QB of the 6T storage cell, the drain is connected to signal line WL_EN_I, and the source is connected to signal line WL_EN_O. After 32 7T-SRAM cells as described above are arranged horizontally, they together with the configuration circuits at both left and right ends constitute a row of "1" search circuits.

[0050] The search circuit has row control: the gate of N3 in the i-th row is connected to word line WLL , the N4 gate is connected to the word line WLR 。In the i-th row, the gates of all N3s are connected to their respective word lines WLL, and the gates of all N4s are connected to their respective word lines WLR, that is, the first access transistor N3 of the cells in the same row is connected to the row selection line WLL, and the second access transistor N4 is connected to the complementary row selection line WLR. For example, in the 0-th row, the gate of N3 is connected to the word line WLL<0>, and the gate of N4 is connected to the word line WLR<0>.

[0051] The search circuit has column cascading: the drains of N5s in each column are connected in series with the sources of N5s in the next column, and the source of the N5 in the first column is connected to the signal SEL , the drain of N5 in the last column is connected to the signal PRE. That is: the drain of N5 in each column is connected to the source of N5 in the next column, and the source of the first N5 is connected to the signal SEL , the drain of the last N5 receives the signal PRE. The series control transistors of the same row of cells are cascaded in sequence to form a series path. The head of this path is connected to the selection signal SEL, and the end is connected to the precharge signal PRE.

[0052] The search circuit has bit line control: in the j-th column, N3 is for the corresponding node Q and the bit line BL <j>transmission tube, N4 is the corresponding node QB and bit line BLB <j>transmission tube

[0053] The maximum value judgment logic for finding the circuit is: (1) Through the word line (WLL , WLR ) and bit line (BL <j>, BLB <j>)Control the charging and discharging of control nodes Q and QB so that the unit in the i-th row and j-th column stores the i-th bit of the j-th floating-point exponent; (2) Detect the signal SEL row by row Transmission status of signal PRE: If signal PRE can be transmitted to signal SEL At the end, if the transmission is not blocked, determine that this bit is 0; if the transmission is blocked, determine that this bit is 1; (3) Combine the determination results of each row to obtain the maximum exponent value.

[0054] Among them, the search circuit controls the word line WLL , WLR Turn on or off the corresponding N3 and N4, and control the bit line BL <j>, BLB <j>Charge and discharge the corresponding nodes Q and QB so that the SRAM cell in the i-th row and j-th column stores the i-th bit of the j-th floating-point exponent. The search circuit also determines the signal SEL Whether it is the same as the signal PRE. If so, determine that the i-th bit of the maximum exponent is 0; otherwise, determine that the i-th bit of the maximum exponent is 1. Judge all rows in turn and generate the maximum exponent.

[0055] Please continue to refer to Figure 2 and Figure 3 As shown in FIGS. 0000120 and 0000121, multiple configuration circuits respectively correspond to multiple rows of SRAM cells. Each configuration circuit includes a multiplexer MUX1, MUX2, MUX3, and a D flip-flop DFF3. The control terminal of MUX1 is connected to the corresponding signal SEL Connection, where one output terminal is connected to one input terminal of MUX2, and the other output terminal is connected to the input terminal D of DFF3. MUX1 is controlled by the WL_EN_O<0> signal (also named SEL) of the 0th 7T-SRAM cell, and selects whether to transfer the input D<0> to output terminal 0 or output terminal 1. Another input terminal of MUX2 is connected to the output terminal of DFF3 and one input terminal of MUX3, and the output terminal outputs the signal D<i + 1>. Another input terminal of MUX3 is connected to the external signal WL , the output terminal of MUX3 is connected to the corresponding word line WLL , the control terminal is connected to an external signal R_C_N. The clock signal input terminal of DFF3 is connected to an external signal CP, and the output terminal outputs a signal WL_S 。External signal WL With word line WLR Connection

[0056] The initialization circuit includes D flip - flops DFF1 and DFF2. The input terminal D of DFF1 is connected to the signal VSS, the clock signal input terminal is connected to the external signal CP, the reset terminal is connected to the external signal SET, and the output terminal is connected to the input terminal D of DFF2 and outputs the signal D_SET. The clock signal input terminal of DFF2 is connected to the external signal CP, and the output terminal is connected to the input terminal of MUX1 and outputs the signal D<0>.

[0057] In the initialization circuit and the configuration circuit, CP, SET, R_C_N, WL, and W_C are externally input signals, where CP is the clock signal. Each 7T - SRAM cell shares the same left word line WLL and a right word line WLR. Among them, WLR is directly connected to the external signal WL, and WLL is controlled by a two - to - one multiplexer MUX3. The WL_EN_I signal lines of each 7T - SRAM cell in a row are connected to the WL_EN_O signal lines of the adjacent 7T - SRAM cell on the right. Among them, the WL_EN_I<31> of the 31st 7T - SRAM cell on the leftmost side is controlled by a two - to - one multiplexer MUX4. MUX4 selects whether to transfer VSS (low level) or VDD (high level) to WL_EN_I<31> according to the value of the PRE signal.

[0058] The selector circuit includes multiplexers MUX4 corresponding to multiple rows of SRAM cells. The two input terminals of MUX4 are respectively connected to the signals VSS and VDD, the output terminal is connected to the drain of the last N5 in the corresponding row, and the control terminal is connected to the signal PRE. That is, the signal PRE controls MUX4 to select one of the signals VSS and VDD and output it to WL_EN_I<31>.

[0059] In this embodiment, 6 "find 1" circuits are arranged in columns to form a 7T - SRAM cell array of 6 rows and 32 columns. However, only the 0th row is equipped with an initialization circuit, and the following rows 1 to 5 do not have an initialization circuit. The input terminals of their MUX1 are all connected to D of the previous row Connected (i = 1, 2, 3, 4, 5).

[0060] Please continue to refer to Figure 4 , and the read / write control circuit is used to supply to bit line BL <j>, BLB <j>Writing data and precharging, which is no different from the read / write control circuit of traditional 6T-SRAM, is mainly responsible for writing data to BL and BLB, precharging BL and BLB, and controlling the connection between BL and BLB and the sense amplifier (SA). The inverted output terminal of the SA is connected to the input terminal D of the D flip-flop, and the output terminal Q of the flip-flop is named IVDD_C and serves as the selection signal of a 2-to-1 multiplexer. According to the value of IVDD_C, this selector chooses to connect VDD or VSS to the IVDD of the six 7T-SRAM cells in this column.

[0061] The amplification circuit includes multiple sense amplifiers (SA), and the multiple sense amplifiers (SA) respectively correspond to multiple bit lines. Two input terminals of each sense amplifier are respectively connected to the corresponding bit line BL <j>、BLB <j>, the control terminal is connected to an external signal SA_EN, and the output terminal outputs a signal OUT_B <j>The signal SA_EN is the enable signal of SA and also the CLK signal of the D flip-flop. It should be noted that this D flip-flop is triggered on the falling edge and has the functions of setting (Set) and clearing (Clr), which are controlled by the two signals CLR and MAX_CLR respectively.

[0062] The column shielding circuit includes multiple column shielding flip-flops and multiple column shielding selectors. The multiple column shielding flip-flops correspond to multiple sense amplifiers respectively, and the multiple column shielding selectors correspond to the multiple column shielding flip-flops respectively. The input terminal D of each column shielding flip-flop is connected to the output terminal of the corresponding sense amplifier, the clock signal input terminal is connected to the external signal SA_EN, the clear terminal is connected to the external signal MAX_CLR, the set terminal is connected to the external signal CLR, and the output terminal Q is connected to the control terminal of the corresponding column shielding selector and outputs the signal IVDD_C <j>。The two input terminals of each column mask selector are respectively connected to the signals VSS and VDD, and the output terminal outputs the signal IVDD <j>。

[0063] The strategy for writing the floating-point exponent mainly includes the following processes.

[0064] (1) In the initial state, the word line WLL , WLR Set to low level, according to the floating-point exponent to the bit line BL <j>, BLB <j>Perform a set operation. Among them, when the data to be written is "1", the bit line BL <j>Set to high level, bit line BLB <j>Set to low level. When the data to be written is "0", the bit line BL <j>Set to low level, bit line BLB <j>Set to high level.

[0065] (2) When in the data writing state, set the word line WLL , WLR Set it to high level to turn on the transmission transistors N3 and N4. Among them, when the data to be written is "1", it is through the bit line BL <j>Charge node Q to pull node Q high to a high level and via bit line BLB <j>Discharge the node QB to pull the node QB high to a low level. When the data to be written is "0", via the bit line BL <j>Discharge node Q to pull node Q low to a low level and through bit line BLB <j>Charge node QB to pull node QB high to the high level.

[0066] (3) When in the hold state, word line WLL , WLR Set it to low level. At this time, the transmission transistors N3 and N4 are turned off, and the circuit is in the hold state. The latch structure composed of P1, P2, N1, and N2 can stably hold the written data.

[0067] In this embodiment, in the provided exponential maximum value finding circuit, the maximum value of the exponents stored in the 7T-SRAM cell array can be found through the peripheral configuration circuit. The main logic is to store 32 6-bit exponents column by column, start from the 32 highest bits of the 0th row, find "1" row by row, and finally find the maximum value. The strategy for finding the maximum value of multiple floating-point exponents includes the following processes.

[0068] (1) Set all nodes QB to high level, and make the MUX4 output high level through the signal PRE and transmit it to the signal SEL , so that the entire line where each N5 is located is precharged to a high level. In this embodiment, the signal CLR is set to a high level, the signal IVDD_C<31:0> is set to 1, and VSS is passed to all signals IVDD<31:0> in the array through a selector, so that all Qs in the array are cleared. Then the signal CLR returns to a low level, the signal MAX_CLR is set to a high level again, the signal IVDD_C<31:0> is cleared, and the selector passes the signal VDD to the signal IVDD in the array. Finally, the signal MAX_CLR also returns to a low level; at the same time, as Figure 2 shown, the signal PRE remains at a high level, and the signal VDD is passed to the signal WL_EN_I<31> through MUX4. Since the array has been cleared at this time and all node QBs are at a high level, all N5 transistors are in an open state, and the high level output by MUX4 will be directly passed to the leftmost SEL, making the entire line precharged to a high level.

[0069] (2) Store multiple multi-bit exponents that are all floating-point numbers in multiple SRAM cells. In each SRAM cell, it is defined that when the node Q is at a high level and the node QB is at a low level, it means the stored data is "1", and when the node Q is at a low level and the node QB is at a high level, it means the stored data is "0".

[0070] Based on the data storage function of the 6T memory cell, this embodiment pre-stores 32 6-bit exponents in the 6T memory cell: in each 6T memory cell, when the node Q is at a high level and the node QB is at a low level, it means the stored data is "1", and when the node Q is at a low level and the node QB is at a high level, it means the stored data is "0"; each column of 6 memory cells is responsible for storing 1 6-bit exponent, where the first row stores the most significant bit (MSB), and the sixth row stores the least significant bit (LSB). 32 columns of memory cells can store 32 6-bit exponents. Specifically, the signal R_C_N remains at a high level, and the word lines WLL and WLR of the array are both controlled by the WL signal of the same row. First, all word lines WL<5:0> are set to a low level, and the RWC circuit sets the bit lines BL<31:0> and BLB<31:0> of each column according to the specific values of the most significant bits of 32 6-bit exponents; then WL<0> is set to a high level, controlling the word lines WLL<0> and WLR<0> of the first row of memory cells to also be set to a high level, and writing the bit lines BL<31:0> and BLB<31:0> into their respective memory cells. Finally, the signal WL<0> is set to a low level to turn off the word lines of the first row. In this way, the writing of one row of data is completed. The data writing scheme for the latter 5 rows of memory cells is the same and will not be elaborated here.

[0071] (3) The output of MUX4 is set to low level by signal PRE. After all data are written into the array, signal WL<5:0> returns to low level and signal PRE is set to low level. If there is no 1 stored in a certain row of memory cells (i.e., all nodes Q are at low level and node QB is at high level), then the N5 transistors of this row will all turn on, and VSS (i.e., low level) will be directly transmitted to signal SEL. If there is at least one 1 stored in a certain row of memory cells, then the N5 transistor of this cell will turn off, and the low level cannot be transmitted to signal SEL, and signal SEL remains at high level.

[0072] (4) Search for the exponential maximum value in the array: Please refer to Figure 5 , and the process of searching for the maximum value is mainly controlled by several signals, namely SET, PRE_CHAR, R_C_N, and SA_EN (SET, PRE_CHAR, and R_C_N are active low, and SA_EN is active high). In the initial state, make signals PRE_CHAR, R_C_N, and SA_EN all at high level, and signal SET generates a low-level pulse with a width of half a cycle, setting the signal D_SET of the initialization circuit to high level. When the first rising edge of CLK arrives, DFF2 transfers the high level of D_SET to D<0>, and DFF1 transfers the low level of VSS to D_SET. Assume that not all Qs of the memory cells in the 0th row of the array are 0. Then, according to the description in (3), SEL<0> is at high level at this time, controlling MUX1 to transfer the high level of D<0> to its output terminal 1, which is connected to the input terminal D of DFF3.

[0073] When the rising edge of the second cycle of the external signal CP arrives, DFF3 transfers the high level at the input terminal to WL_S<0>. At the same time, signal PRE_CHAR generates a low level with a width of half a cycle, and through the read / write control circuit, the bit line BL <j>, BLB <j>Precharge to high level. After the external signal CP is delayed by a quarter of a cycle, the signal SA_EN generates a high level with a width of half a cycle, enabling multiple sense amplifiers to start working. In the second half of the second cycle of the external signal CP, PRE_CHAR returns to the high level, and the signal R_C_N generates a low level with a width of half a cycle, controlling MUX3 to select WL_S The high level is transferred to WLL , turn on the word line WLL , multiple floating-point exponents will be represented on the bit line BL. For example, the control MUX3 transfers the high level of WL_S<0> to WLL<0>, turning on the left word line of the 0th row of memory cells, and the 32 stored data (i.e., the most significant bits of 32 6-bit exponents) will be represented on the bit lines BL<31:0> through the transfer transistors.

[0074] The signal R_C_N also controls the read / write control circuit to turn on the bit line BL <j>, BLB <j>The path between the plurality of sense amplifiers, and the signal SA_EN also remains at a high level. The data in the i-th row of SRAM cells is read out by the corresponding sense amplifier and inverted and output to all signals OUT_B <j>。After a delay of one - quarter of a cycle of the external signal CP, the signal SA_EN drops to a low level, and the falling edge of the signal SA_EN controls the corresponding flip - flop to transfer all the signals OUT_B to the corresponding signal IVDD_C. For example, when the signal SA_EN remains high, the data in the storage unit of the 0th row is read out by SA and inverted and output to the signal OUT_B<31:0>. After a delay of one - quarter of a cycle, the signal SA_EN drops to a low level, and the falling edge of the signal SA_EN controls the flip - flop to transfer the signal OUT_B<31:0> to the signal IVDD_C<31:0>.

[0075] Assume that the Q of the storage unit at the 0th row and 0th column is 0. Then the exponent stored in this column must not be the maximum value (because the highest bit of the exponent of some other column has 1, while the highest bit of the 0th column is 0), and the data in this column is no longer needed in the subsequent maximum - value search process. At this time, the signal OUT_B<0> = IVDD_C<0> = 1, and the signal IVDD_C<0> controls the selector to transfer the signal VSS to the signal IVDD<0>, pulling down the node Q of all the units in the 0th column to 0, so that all the N5 transistors in the 0th column are turned on to shield the 0th column in the subsequent calculation. On the contrary, if the node Q of the storage unit at the 0th row and 0th column is 1, then the signal IVDD<0> remains high and does not affect the data stored in this column.

[0076] (5) When all the nodes Q in the ith row are at a low level, the signal SEL is at a low level. If all the data in the 0th row are 0, then according to the discussion in (3), SEL<0> is at a low level at this time. When the first rising edge of CLK arrives, the signal SEL<0> controls MUX1 and MUX to directly transfer the high level of D<0> to D<1>. Assuming that not all the data stored in the 1st row are 0, then the subsequent process of finding the maximum value in (4) will directly be carried out in the storage units of the 1st row, skipping the 0th row that is all 0.

[0077] (6) After searching all the arrays composed of multiple SRAM cells, all the multi-bit exponents that are not the exponent maximum value are cleared, and the exponent maximum value is represented by the signal SEL. In this embodiment, after all the 6-row storage units are processed, all the columns that are not the maximum value will be cleared, and the corresponding N5 transistors of these columns will also be turned on. At this time, the signal SEL<5:0> is only controlled by the N5 transistors of the column where the maximum value is located. Therefore, the exponent maximum value will be represented by the signal SEL<5:0>.

[0078] Sometimes, the shielding of a certain column will cause the Q of the subsequent row cells to be all 0, thus triggering the skipping of an entire row, further saving time and reducing power consumption. For example: There are currently 3 6-bit exponents A = 110001, B = 011111, C = 101111. In the first cycle, the highest bit is judged, and it is found that the highest bit of B is 0, so B is shielded and all its bits are pulled down to the low level. Therefore, at this time, B = 000000; in the second cycle, the second highest bit is judged, and it is found that the second highest bit of C is 0, so C is shielded again. In the third cycle, the subsequent 3 bits of the three numbers are all 0. Therefore, the circuit directly skips these three bits and starts judging the last row. Finally, the maximum value 110001 is found.

[0079] Compared with the existing exponent maximum value finding circuit, the floating-point exponent maximum value finding circuit of this embodiment has the following beneficial effects:

[0080] 1. In the process of finding the exponent maximum value, this floating-point exponent maximum value finding circuit can find out the exponents that are no longer likely to be the maximum value in real time, exclude and shield them in the subsequent calculations, reduce the subsequent calculation amount, improve the efficiency, reduce the power consumption, and solve the technical problems of the existing exponent maximum value finding circuit, such as high power consumption, complex hardware implementation, and long calculation delay.

[0081] 2. In the process of finding the exponent maximum value, if the Q of a certain row of cells is all 0, then in this cycle, the circuit will skip this row and directly find "1" in the next row, which not only saves the number of cycles and shortens the calculation time, but also further reduces the power consumption.

[0082] 3. The floating-point exponent maximum finding circuit realizes the function of quickly and low-power finding the exponent maximum by combining the efficient storage characteristics of 7T-SRAM. This design not only significantly reduces power consumption and latency but also improves computational efficiency, enabling better support for efficient floating-point operations, especially in the application of the CIM architecture. This circuit design provides an innovative solution for finding the exponent maximum in FP16 operations, contributing to the improvement of the performance of floating-point calculations.

[0083] Embodiment 2

[0084] This embodiment provides an in-memory computing chip (CIM chip), which includes the floating-point exponent maximum finding circuit in Embodiment 1. The CIM chip in this embodiment uses the circuit in Embodiment 1 to find the maximum value, no longer relying on methods such as multi-stage comparators, tree structures, or parallel computing architectures. The power consumption is significantly reduced, the hardware implementation is simpler, and the computing latency is shorter, which can significantly reflect the computing power of the chip and effectively meet the requirements of low power consumption and high speed. The CIM chip has a storage mode and a computing mode. In the storage mode, the CIM chip is used as a memory. In the computing mode, the CIM chip is used to implement the multiply-accumulate operation between multiple groups of multi-bit floating-point input feature numbers and multi-bit floating-point weights.

[0085] Embodiment 3

[0086] This embodiment provides a floating-point multiply-accumulate operation circuit, which uses the floating-point exponent maximum finding circuit in Embodiment 1 to find the exponent maximum among multiple multi-bit exponents that are all floating-point numbers. This floating-point multiply-accumulate operation circuit is used to implement the multiply-accumulate operation between multiple groups of multi-bit floating-point input feature numbers and multi-bit floating-point weights. During its operation, it is necessary to find the maximum value in the sum of the calculated exponents. This embodiment uses the finding circuit in Embodiment 1, which can significantly improve the finding efficiency and reduce power consumption, and enhance the computing speed.

[0087] Embodiment 4

[0088] This embodiment provides a static random access memory (SRAM), which uses the floating-point multiply-accumulate operation circuit in Embodiment 3 to implement the multiply-accumulate calculation of multi-bit inputs and multi-bit weights.

[0089] Based on the floating-point multiply-accumulate operation circuit in Embodiment 3, the multiply-accumulate operation is directly completed within the storage unit, reducing data movement and significantly lowering power consumption. SRAM can simultaneously process the multiply-accumulate operations of multiple inputs and weights, greatly improving the computing efficiency, and is particularly suitable for application scenarios that require high throughput. The read and write speeds of SRAM are much higher than those of DRAM and flash memory, enabling low-latency multiply-accumulate calculations, and are suitable for applications with high real-time requirements, such as edge computing and Internet of Things devices. SRAM can be integrated with other computing units (such as CPU, GPU) on the same chip to form an efficient memory-computation integrated architecture.

[0090] The SRAM of this embodiment is applicable to artificial intelligence and machine learning. The inference and training processes of neural networks involve a large number of multiply-accumulate operations, and in-memory computing of SRAM can significantly accelerate these operations and improve the overall performance. Multi-bit inputs and weights enable SRAM to support from simple linear models to complex deep neural networks. The in-memory computing of the SRAM in this embodiment reduces the complex interface between the memory and the processor in the traditional computing architecture, simplifying the system design. By reducing data movement and simplifying the architecture, SRAM can reduce the overall cost and power consumption of the system.

[0091] Embodiment 5

[0092] This embodiment provides an electronic device, which includes a memory and a processor. Among them, the memory includes the floating-point multiply-accumulate operation circuit in Embodiment 3. Compared with existing electronic devices, this electronic device can significantly improve the computing efficiency, reduce power consumption, and support high-precision computing. It has broad application prospects in the fields of artificial intelligence, edge computing, etc. Although it faces some technical challenges, its advantages make it an important technical direction for in-memory computing.

[0093] Embodiment 6

[0094] This embodiment provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, the memory is the static random access memory in Embodiment 4.

[0095] This computer device can take various forms. It can either adopt an embedded chip or module, or a general-purpose data processing device, such as an intelligent terminal capable of executing programs, a tablet computer, a laptop computer, a desktop computer, a rack-mounted server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc.

[0096] The computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can be communicatively connected to each other through a system bus. The memory (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device.

[0097] In some embodiments, the processor may be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data.

[0098] The foregoing are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention. < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j> < / j>

Claims

1. A floating-point exponent maximum value search circuit, characterized in that, It includes: SRAM cell matrix: It includes SRAM cells arranged in multiple rows and columns. Each cell includes NMOS transistors N1 to N5 and PMOS transistors P1 and P2; N1, N2, P1, and P2 are cross-coupled in an inverter configuration to form a pair of storage nodes Q and QB; the gate of N5 is connected to node QB; Row control: The gate of N3 in the i-th row is connected to the word line WLL , the N4 gate is connected to the word line WLR ; Column cascade: The drain of N5 in each column is connected in series with the source of N5 in the next column, and the source of N5 in the first column is connected to signal SEL , the drain of N5 in the last column is connected to signal PRE; Bit line control: In the j-th column, N3 is the corresponding node Q and bit line BL <j>transmission tube, N4 is the corresponding node QB and bit line BLB <j>transmission tube; < / j> < / j> Maximum value judgment logic: (1) Through the word line (WLL , WLR ) and bit line (BL <j>, BLB <j>)Control the charging and discharging of the control nodes Q and QB so that the unit in the i-th row and j-th column stores the i-th bit of the j-th floating-point exponent; (2) Detect the signal SEL row by row Transmission status with signal PRE: If signal PRE can be transmitted to signal SEL terminal, determine that the bit is 0; if the transmission is blocked, determine that the bit is 1; (3) Combine the determination results of each row to obtain the maximum exponent value. < / j> < / j> 2. The floating-point exponent maximum value searching circuit according to claim 1, wherein The searching circuit further includes: Multiple configuration circuits, corresponding to multiple rows of SRAM cells respectively; each configuration circuit includes a multiplexer MUX1, MUX2, MUX3, and a D flip-flop DFF3; the control terminal of MUX1 is connected to the corresponding signal SEL Connection, one output terminal is connected to one input terminal of MUX2, and the other output terminal is connected to the input terminal D of DFF3; another input terminal of MUX2 is connected to the output terminal of DFF3 and one input terminal of MUX3, and the output terminal outputs the signal D<i+1>; another input terminal of MUX3 is connected to the external signal WL , the output terminal of MUX3 is connected to the corresponding word line WLL , the control terminal is connected to an external signal R_C_N; the clock signal input terminal of DFF3 is connected to an external signal CP, and the output terminal outputs a signal WL_S ; External signal WL With word line WLR connection.

3. The floating-point exponent maximum value searching circuit according to claim 2, wherein The searching circuit further includes: Initialization circuit, which includes D flip-flops DFF1 and DFF2; the input terminal D of DFF1 is connected to signal VSS, the clock signal input terminal is connected to an external signal CP, the reset terminal is connected to an external signal SET, and the output terminal is connected to the input terminal D of DFF2 and outputs signal D_SET; the clock signal input terminal of DFF2 is connected to an external signal CP, and the output terminal is connected to the input terminal of MUX1 and outputs signal D<0>.

4. The floating-point exponent maximum value finding circuit according to claim 3, wherein, The searching circuit further includes: Selector circuit, which includes multiplexers MUX4 corresponding to multiple rows of SRAM cells respectively; two input terminals of MUX4 are respectively connected to signal VSS and VDD, the output terminal is connected to the drain of the last N5 in the corresponding row, and the control terminal is connected to signal PRE.

5. The floating-point exponent maximum value finding circuit according to claim 4, wherein The searching circuit further includes: A read / write control circuit, which is used to supply to bit line BL <j>, BLB <j>Write data and precharge; < / j> < / j> An amplifying circuit, comprising a plurality of sense amplifiers respectively corresponding to a plurality of bit lines; two input ends of each sense amplifier are respectively connected to a corresponding bit line BL <j>, BLB <j>, the control terminal is connected to the external signal SA_EN, and the output terminal outputs the signal OUT_B <j> 。< / j> < / j> < / j> 6. The floating-point exponent maximum value finding circuit according to claim 5, wherein The searching circuit further includes: Column shielding circuit, which includes a plurality of column shielding flip-flops corresponding to a plurality of sense amplifiers respectively, and a plurality of column shielding selectors corresponding to the plurality of column shielding flip-flops respectively; the input terminal D of each column shielding flip-flop is connected to the output terminal of the corresponding sense amplifier, the clock signal input terminal is connected to an external signal SA_EN, the clear terminal is connected to an external signal MAX_CLR, the set terminal is connected to an external signal CLR, and the output terminal Q is connected to the control terminal of the corresponding column shielding selector and outputs a signal IVDD_C <j>; Two input terminals of each column masking selector are respectively connected to signal VSS and VDD, and the output terminal outputs signal IVDD <j> 。< / j> < / j> 7. The floating-point exponent maximum value finding circuit according to claim 6, characterized in that, The strategy for the searching circuit to write the floating-point exponent includes: When in the initial state, the word line WLL , WLR Set to low level, and according to the floating-point exponent, for the bit line BL <j>, BLB <j>Perform a set operation; wherein, when the data to be written is "1", the bit line BL <j>Set to high level, bit line BLB <j>Set to low level; when the data to be written is "0", the bit line BL <j>Set to low level, bit line BLB <j>Set to high level; < / j> < / j> < / j> < / j> < / j> < / j> When in the data writing state, the word line WLL , WLR Set to high level; among them, when the data to be written is "1", through the bit line BL <j>Charge node Q to pull node Q high to a high level and via bit line BLB <j>Discharge node QB to pull node QB high to low level; when the data to be written is "0", via bit line BL <j>Discharge node Q to pull node Q low to a low level and through bit line BLB <j>Charge node QB to pull node QB high to high level; < / j> < / j> < / j> < / j> When in the hold state, the word line WLL , WLR Set to low level.

8. The floating-point exponent maximum value searching circuit according to claim 7, wherein The strategy for the searching circuit to find the maximum value of multiple floating-point exponents includes: (1) Set all nodes QB to high level. Make the output of MUX4 high level through signal PRE and transmit it to signal SEL , so that the entire line where N5 is located in each row is precharged to high level; (2) Store multiple multi-bit exponents that are all floating-point numbers in multiple SRAM cells. In each SRAM cell, it is defined that when node Q is high level and node QB is low level, it means the stored data is "1", and when node Q is low level and node QB is high level, it means the stored data is "0"; (3) Make MUX4 output low level through signal PRE; (4) Search for the exponential maximum value in the array: In the initial state, make the signals PRE_CHAR, R_C_N, and SA_EN all at high level, the signal SET generates a low-level pulse with a width of half a cycle, and set the signal D_SET of the initialization circuit to high level; when the rising edge of the second cycle of the external signal CP arrives, through the read-write control circuit, for the bit line BL <j>, BLB <j>Precharge to high level; after the external signal CP is delayed by a quarter of a cycle, the signal SA_EN generates a high level with a width of half a cycle, enabling multiple sense amplifiers to start working; in the second half of the second cycle of the external signal CP, the signal R_C_N generates a low level with a width of half a cycle, controlling MUX3 to select WL_S The high level is transmitted to WLL , turn on the word line WLL , multiple floating-point exponents will be characterized on the bit line BL; the signal R_C_N controls the read / write control circuit to turn on the bit line BL <j>, BLB <j>There is a path between the multiple sense amplifiers, and the signal SA_EN remains at a high level. The data in the i-th row of SRAM cells is read out by the corresponding sense amplifier and inverted and output to all signals OUT_B <j>; After a quarter of the cycle of the external signal CP, signal SA_EN drops to low level, and the falling edge of signal SA_EN controls the corresponding flip-flop to transfer all signals OUT_B to the corresponding signal IVDD_C; < / j> < / j> < / j> < / j> < / j> When all nodes Q in the i-th row are at a low level, the signal SEL is low level; (6) After all the arrays composed of multiple SRAM cells are searched, all non-exponent maximum multi-bit exponents are cleared, and the exponent maximum value is characterized by signal SEL.

9. The floating-point exponent maximum value finding circuit according to claim 1, wherein The searching circuit is used to find the maximum exponent value among 32 6-bit exponents. Each column of 6 SRAM cells is responsible for storing 1 6-bit exponent. The highest significant bit is stored in the first row, and the lowest significant bit is stored in the sixth row.

10. An in-memory computing chip, characterized in that, It includes the floating-point exponent maximum value searching circuit as described in any one of claims 1-9.

Citation Information

Cited By

  • Floating point type index comparison circuit based on SRAM (Static Random Access Memory) and chip thereof

    CN120447865A

  • Floating-point exponential comparison circuit and chip based on SRAM

    CN120447865B