Single-bit weight generation unit, multi-bit weight generation unit, array group, and compute macro

CN117153218BActive Publication Date: 2026-09-18ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310968651.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2026-09-18
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

[0005]基于此,有必要针对现有的推理-训练芯片在推理操作时出现速度减慢、后向传播精确度降低的问题,提供单bit权重产生单元、多bit权重产生单元、阵列组及计算宏

Benefits of technology

[0021] 1. Based on the characteristics of inference and training operations, this invention has formulated different quantization schemes, which have been integrated to achieve effective utilization of chip resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117153218B_ABST
    Figure CN117153218B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of dynamic random access memory, and in particular to a single bit weight generating unit, a multi-bit weight generating unit, an array group and a computing macro. The single bit weight generating unit comprises n standard 6T-SRAM units and a transposed XNOR accumulation unit, the transposed XNOR accumulation unit is used as a computing unit, and is externally connected to the standard 6T-SRAM, thereby realizing inference and training operation of multi-bit XNOR accumulation. The multi-bit weight generating unit is composed of four single bit weight generating units, the array group is composed of array-distributed multi-bit weight generating units, and the in-memory computing macro is constructed based on the array group. According to the characteristics of inference and training operation, different quantization schemes are formulated, integration is realized, chip resources are effectively utilized, and the problems of speed reduction and backward propagation accuracy reduction of existing inference-training chips during inference operation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic random access storage technology, and more specifically, to a single-bit weight generation unit, a multi-bit weight generation unit composed of four such single-bit weight generation units, an array group composed of an array of multi-bit weight generation units, and an in-memory computation macro. Background Technology

[0002] With the advent of the "computing power era," large-scale data needs to travel back and forth between memory and processors. However, in the traditional von Neumann architecture, computing units and storage units are separated, and frequent data access consumes a lot of power and time. The emergence of in-memory computing (CIM) technology has broken through the von Neumann bottleneck and overcome the "memory wall" problem in traditional computing architectures, thus having revolutionary significance for the "computing power era." Because SRAM has a fast data read speed and good compatibility with advanced logic processes, SRAM-based in-memory computing has great development potential.

[0003] From image classification to speech recognition, machine learning has driven the widespread application of Artificial Intelligence (AI). While cloud computing provides powerful computing support for AI training and computation, this relies on users' personal data, and many users are unwilling to send their personal data to the cloud to retrain models. Meanwhile, edge applications require real-time network connectivity; without a network connection, edge devices cannot perform real-time retraining to handle new situations encountered in the field. Based on these factors, edge device learning (or on-chip training) is a preferred approach. That is, combining training and inference chips into a single unit allows the chip to incrementally learn from samples after performing a phase of inference operations. This facilitates the training of existing models by updating the knowledge base, learning user-specific features to personalize the model, and mitigating privacy concerns.

[0004] However, most existing chip circuit designs still remain at the inference stage and have not yet integrated with training. Furthermore, existing inference-training chips use the same bit width for both inference and training processes, which leads to slower inference operations and reduced backpropagation accuracy. Summary of the Invention

[0005] Therefore, it is necessary to address the issues of slowed inference operation and reduced backpropagation accuracy in existing inference-training chips by providing single-bit weight generation units, multi-bit weight generation units, array groups, and computational macros.

[0006] This invention is achieved using the following technical solution:

[0007] In a first aspect, the present invention discloses a single-bit weight generation unit, comprising n standard 6T-SRAM units and one transposed XNOR accumulation unit.

[0008] n standard 6T-SRAM cells are used as storage units; the reading of any standard 6T-SRAM cell is controlled by its word line WL, and the weight value stored in it is reflected in its BL and BLB. The bit lines BL of the n standard 6T-SRAM cells are connected to the local bit line LBL, and the bit lines BLB of the n standard 6T-SRAM cells are connected to the local bit line LBLB; n≥1.

[0009] One transposed XNOR accumulator unit serves as the computation unit. The transposed XNOR accumulator unit includes: NMOS transistors N1–N6 and PMOS transistors P1–P6. Specifically, the drain of N1 is connected to signal line C-XACN, and its gate is connected to control signal RWLAN. The drain of N2 is connected to C-XACN, and its gate is connected to control signal RWLBN. The drain of N3 is connected to the source of N1, its gate is connected to control signal CWLAN, and its source is connected to C-XACN. The drain of N4 is connected to the source of N2, its gate is connected to control signal CWLBN, and its source is connected to C-XACN. The drain of N5 is connected to the source of N1 and the drain of N3, its gate is connected to LBL, and its source is connected to signal line R-XACN. The drain of N6 is connected to the source of N2 and the drain of N4, its gate is connected to LBLB, and its source is connected to R-XACN. The gate of P1 is connected to control signal CWLAP, and its source is connected to signal line C-XACP. The gate of P2 is connected to the control signal CWLBP, and its source is connected to C-XACP. The drain of P3 is connected to C-XACN, its gate is connected to the control signal RWLAP, and its source is connected to the drain of P1. The drain of P4 is connected to C-XACN, its gate is connected to the control signal RWLBP, and its source is connected to the drain of P2. The drain of P5 is connected to the drain of P1 and the source of P3, its gate is connected to LBL, and its source is connected to the signal line R-XACP. The drain of P6 is connected to the drain of P2 and the source of P4, its gate is connected to LBLB, and its source is connected to R-XACP.

[0010] The implementation of this single-bit weight generation unit is based on the method or process of an embodiment of this disclosure.

[0011] In a second aspect, the present invention discloses a multi-bit weight generation unit, comprising four single-bit weight generation units as disclosed in the first aspect.

[0012] The four single-bit weight generation units are located in the same row and share the same RWLAN, RWLBN, RWLAP, RWLBP, CWLAP, CWLBP, CWLAN, and CWLBN. Among the four single-bit weight generation units in the same row, the m-th standard 6T-SRAM cell of each single-bit weight generation unit shares the same WL, where m∈[1,n].

[0013] The implementation of this multi-bit weight generation unit is based on the method or process of embodiments of this disclosure.

[0014] Thirdly, this invention discloses an array group comprising N×N multi-bit weight generation units disclosed in the second aspect, arranged in an array; N=2 i , i>0.

[0015] In this context, multi-bit weight generation units located in the same column share the same CWLAP, CWLBP, CWLAN, and CWLBN. Among the N multi-bit weight generation units in the same column, the q-th single-bit weight generation unit of each multi-bit weight generation unit shares the same C-XACN and C-XACP; q∈[1,4]. Multi-bit weight generation units located in the same row share the same RWLAN, RWLBN, RWLAP, and RWLBP. Among the N multi-bit weight generation units in the same row, the q-th single-bit weight generation unit of each multi-bit weight generation unit shares the same R-XACN and R-XACP.

[0016] The implementation of such an array group is based on the method or process of an embodiment of this disclosure.

[0017] Fourthly, the present invention discloses an in-memory computing macro, including the array group, word line driver, backward channel input driver, forward bit line input device, forward channel input driver, backward bit line input device, in-memory computing controller, flash memory analog-to-digital converter, successive approximation analog-to-digital converter, and timing controller disclosed in the third aspect.

[0018] The array group is used for forward or backward propagation. The word line driver controls the WL switch. The backward channel input driver controls the CWLAN, CWLAP, CWLBN, and CWLBP switches. The forward bit line input precharges C-XACN to VDD / 2 during forward propagation and connects C-XACN to VSS and C-XACP to VDD during backward propagation. The forward channel input driver controller controls the RWLAN, RWLAP, RWLBN, and RWLBP switches. The backward bit line input controller precharges R-XACN to VDD and R-XACP to VSS during backward propagation. The backward bit line input precharges R-XACN to VDD and R-XACP to VSS during backward propagation. The in-memory compute controller switches the array group's functions. The flash analog-to-digital converter provides a 4-bit output during forward propagation. A successive approximation analog-to-digital converter is used to obtain an 8-bit output during backward propagation. A timing controller is used to control the clock pulses of each signal.

[0019] The implementation of such in-memory computation macros is based on the methods or processes of embodiments of this disclosure.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. Based on the characteristics of inference and training operations, this invention has formulated different quantization schemes, which have been integrated to achieve effective utilization of chip resources.

[0022] 2. Compared with the existing 4-bit input and 4-bit output circuit structure, the accuracy of the present invention is significantly improved, and it reaches a level similar to that of the existing 8-bit input and 8-bit output circuit structure; and the energy efficiency of the present invention is also improved. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a circuit diagram of the single-bit weight generation unit provided in Embodiment 1 of the present invention;

[0025] Figure 2 for Figure 1 Six cases of forward propagation of the single-bit weight generation unit;

[0026] Figure 3 for Figure 1Six cases of backward propagation of the single-bit weight generation unit;

[0027] Figure 4 This is a structural diagram of the multi-bit weight generation unit provided in Embodiment 2 of the present invention;

[0028] Figure 5 This is a structural diagram of the array group provided in Embodiment 2 of the present invention;

[0029] Figure 6 This is a structural diagram of the in-memory computation macro provided in Embodiment 3 of the present invention;

[0030] Figure 7 for Figure 6 A schematic diagram of the forward propagation of the in-memory computation macro;

[0031] Figure 8 for Figure 7 The signal waveform diagram;

[0032] Figure 9 for Figure 6 A schematic diagram of the backward propagation of the in-memory computation macro;

[0033] Figure 10 for Figure 9 The signal waveform diagram;

[0034] Figure 11 This is a diagram showing the result of comparing the accuracy of the in-memory calculation macro with that of existing arithmetic circuits in Embodiment 3 of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.

[0038] Example 1

[0039] See Figure 1 This is a circuit diagram of the single-bit weight generation unit provided in Embodiment 1. The single-bit weight generation unit (which can be abbreviated as BWPU) includes: n standard 6T-SRAM cells (which can be abbreviated as W) and 1 transposed XNOR accumulator cell (which can be abbreviated as TXAC).

[0040] A standard 6T-SRAM cell is used as the storage unit. n≥1. In this embodiment 1, after simulation, considering various indicators such as area, power consumption, latency, and throughput, n=8 can achieve the best results.

[0041] The structure of a single standard 6T-SRAM cell is well-known. Here's a simplified explanation: A standard 6T-SRAM includes two PMOS transistors (PM1 to PM2) and four NMOS transistors (NM1 to NM4). PM1 and PM2 act as pull-up transistors, NM1 and NM2 as pull-down transistors, and NM3 and NM4 as transmission transistors. The word line WL controls NM3 and NM4, and the bit lines BL and BLB are connected to NM3 and NM4 respectively. PM1, NM1, and NM3 are connected to memory node Q, and PM2, NM2, and NM4 are connected to memory node QB.

[0042] In general, the reading of any standard 6T-SRAM cell is controlled by its word line WL, and the stored weight value is reflected in its BL and BLB. For a given standard 6T-SRAM cell, if the stored weight value is 1 (Q is "1" and QB is "0"), when WL is high, Q is connected to BL, and QB is connected to BLB. Q discharges to BL, and BLB discharges to QB, making BL 1 and BLB 0. If the stored weight value is 0 (Q is "0" and QB is "1"), when WL is high, Q is connected to BL, and QB is connected to BLB. BL discharges to Q, and QB discharges to BLB, making BL 0 and BLB 1.

[0043] The bit lines BL of n standard 6T-SRAM cells are connected together to the local bit line LBL. The bit lines BLB of n standard 6T-SRAM cells are connected together to the local bit line LBLB.

[0044] In this way, each time a standard 6T-SRAM cell is activated, LBL and LBLB will synchronously reflect the BL and BLB of that standard 6T-SRAM cell, so that the weight value stored in the standard 6T-SRAM cell can be input into the transpose XNOR accumulator cell.

[0045] The transposed XNOR accumulator unit serves as the computation unit. It comprises six NMOS transistors (N1 to N6) and six PMOS transistors (P1 to P6). Specifically, the drain of N1 is connected to C-XACN, and its gate is connected to the control signal RWLAN. The drain of N2 is connected to C-XACN, and its gate is connected to the control signal RWLBN. The drain of N3 is connected to the source of N1, its gate is connected to the control signal CWLAN, and its source is connected to the signal line C-XACN. The drain of N4 is connected to the source of N2, its gate is connected to the control signal CWLBN, and its source is connected to the signal line C-XACN. The drain of N5 is connected to the source of N1 and the drain of N3, its gate is connected to the local bit line LBL, and its source is connected to the signal line R-XACN. The drain of N6 is connected to the source of N2 and the drain of N4, its gate is connected to the local bit line LBLB, and its source is connected to the signal line R-XACN. The gate of P1 is connected to the control signal CWLAP, and its source is connected to the signal line C-XACP. The gate of P2 is connected to the control signal CWLBP, and its source is connected to the signal line C-XACP. The drain of P3 is connected to the signal line C-XACN, its gate is connected to the control signal RWLAP, and its source is connected to the drain of P1. The drain of P4 is connected to the signal line C-XACN, its gate is connected to the control signal RWLBP, and its source is connected to the drain of P2. The drain of P5 is connected to the drain of P1 and the source of P3, its gate is connected to the local bit line LBL, and its source is connected to the signal line R-XACP. The drain of P6 is connected to the drain of P2 and the source of P4, its gate is connected to the local bit line LBLB, and its source is connected to the signal line R-XACP.

[0046] In other words, the drain of N1, the drain of N2, the source of N3, and the source of N4 are all connected to C-XACN; the source of N1 and the drain of N3 are all connected to the drain of N5; and the gate of N1 is connected to RWLAN.

[0047] The drain of N2, the drain of N1, the source of N3, and the source of N4 are all connected to C-XACN; the source of N2 and the drain of N4 are all connected to the drain of N6; the gate of N2 is connected to RWLBN.

[0048] The source of N3, the drain of N1, the drain of N2, and the source of N4 are all connected to C-XACN; the drain of N3 and the source of N1 are all connected to the drain of N5; the gate of N3 is connected to CWLAN.

[0049] The source of N4, the drain of N1, the drain of N2, and the source of N3 are all connected to C-XACN; the drain of N4 and the source of N2 are all connected to the drain of N6; the gate of N4 is connected to CWLBN.

[0050] The gate of N5 is connected to LBL, the source is connected to R-XACN, and the drain is connected to the source of N1 and the drain of N3.

[0051] The gate of N6 is connected to LBLB, the source is connected to R-XACN, and the drain is connected to the source of N2 and the drain of N4.

[0052] The source terminals of P1 and P2 are connected to C-XACP; the drain terminals of P2 and P3 are connected to the drain terminals of P5; the gate terminal of P2 is connected to CWLAP.

[0053] The source terminals of P2 and P1 are connected to C-XACP; the drain terminals of P2 and P4 are connected to the drain terminals of P6; the gate terminal of P2 is connected to CWLBP.

[0054] The drains of P3 and P4 are connected to C-XACN; the sources of P3 and P1 are connected to the drain of P5; the gate of P2 is connected to RWLAP.

[0055] The drains of P4 and P3 are connected to C-XACN; the sources of P4 and P2 are connected to the drain of P6; the gate of P4 is connected to RWLBP.

[0056] The gate of P5 is connected to LBL; the source of P5 is connected to R-XACP; the drain of P5 is connected to the drain of P1 and the source of P3.

[0057] The gate of P6 is connected to LBLB; the source of P6 is connected to R-XACP; the drain of P6 is connected to the drain of P2 and the source of P4.

[0058] For the sake of completeness, there is some overlap in the above connections.

[0059] Based on the structure of the single-bit weight generation unit described above, its working methods in forward propagation and backward propagation are explained respectively.

[0060] (1) During forward propagation, C-XACN is precharged to VDD / 2.

[0061] See Figure 2 The six scenarios are shown respectively:

[0062] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VDD, RWLBN is connected to VSS, RWLAP is connected to VSS, and RWLBP is connected to VDD (Input = +1), N1 is on, N2 is off, P3 is on, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN discharges to R-XACN through N1 and N5, and the voltage of C-XACN decreases, which is XNOR = +1.

[0063] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VSS, RWLBN is connected to VDD, RWLAP is connected to VDD, and RWLBP is connected to VSS (Input = -1), N2 is on, N1 is off, P4 is on, and P3 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; R-XACP discharges to C-XACN through P4 and P6, and the voltage of C-XACN rises, which is XNOR = -1.

[0064] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, RWLAP is connected to VDD, and RWLBP is connected to VDD (Input = 0); N1 is off, N2 is off, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN is kept at VDD / 2, and there is no voltage change in C-XACN, which means XNOR = 0.

[0065] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VDD, RWLBN is connected to VSS, RWLAP is connected to VSS, and RWLBP is connected to VDD (Input = +1), N1 is on, N2 is off, P3 is on, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; R-XACP discharges to C-XACN through P5 and P3, and the voltage of C-XACN rises, which is XNOR = -1.

[0066] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, RWLAP is connected to VDD, and RWLBP is connected to VDD (Input = -1), N1 is off, N2 is off, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN discharges to R-XACN through N2 and N6, and the voltage of C-XACN decreases, which is XNOR = +1.

[0067] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, RWLAP is connected to VDD, and RWLBP is connected to VDD (Input = 0); N1 is off, N2 is off, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN is kept at VDD / 2, and there is no voltage change in C-XACN, which means XNOR = 0.

[0068] In summary, during forward propagation, RWLAN, RWLBN, RWLAP, and RWLBP are used as inputs and calculated with the 1-bit weight of the storage unit. The XOR summation result is reflected in C-XACN.

[0069] (2) During backward propagation, R-XACN is precharged to VDD, and R-XACP is precharged to VSS; C-XACN is connected to VSS, and C-XACP is connected to VDD.

[0070] See Figure 3 The six scenarios are shown respectively:

[0071] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; CWLAP is connected to VSS, CWLBP is connected to VDD, CWLAN is connected to VDD, and CWLBN is connected to VSS (Input = +1), P1 is on, P2 is off, N3 is on, and N4 is off. R-XACN discharges to C-XACN through N5 and N3, the voltage of R-XACN decreases, and the voltage of R-XACP remains unchanged, which is XNOR = +1.

[0072] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; CWLAP is connected to VDD, CWLBP is connected to VSS, CWLAN is connected to VSS, and CWLBN is connected to VDD (Input = -1), P2 is on, P1 is off, N4 is on, and N3 is off. C-XACP discharges to R-XACP through P2 and P6. The voltage of R-XACN remains unchanged, while the voltage of R-XACP increases, which is XNOR = -1.

[0073] If LBL is high and LBLB is low (Weight = +1), N5 is on, P5 is off, N6 is off, and P6 is on; CWLAP is connected to VDD, CWLBP is connected to VDD, CWLAN is connected to VSS, and CWLBN is connected to VSS (Input = 0), P1 is off, P2 is off, N3 is off, and N4 is off. R-XACN remains at VDD, and R-XACP remains at VSS. There is no voltage change in R-XACN and R-XACP, which means XNOR = 0.

[0074] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; CWLAP is connected to VSS, CWLBP is connected to VDD, CWLAN is connected to VDD, and CWLBN is connected to VSS (Input = +1), P1 is on, P2 is off, N3 is on, and N4 is off. C-XACP discharges to R-XACP through P1 and P5. The voltage of R-XACN remains unchanged, while the voltage of R-XACP increases, which is XNOR = -1.

[0075] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; CWLAP is connected to VDD, CWLBP is connected to VSS, CWLAN is connected to VSS, and CWLBN is connected to VDD (Input = -1), P2 is on, P1 is off, N4 is on, and N3 is off. R-XACN discharges to C-XACN through N6 and N4. The voltage of R-XACN decreases, while the voltage of R-XACP remains unchanged, which is XNOR = +1.

[0076] If LBL is low and LBLB is high (Weight = -1), P5 is on, N5 is off, P6 is off, and N6 is on; CWLAP is connected to VDD, CWLBP is connected to VDD, CWLAN is connected to VSS, and CWLBN is connected to VSS (Input = 0), P1 is off, P2 is off, N3 is off, and N4 is off. R-XACN remains at VDD, and R-XACP remains at VSS. There is no voltage change in R-XACN and R-XACP, which means XNOR = 0.

[0077] In summary, during backpropagation, CWLAN, CWLBN, CWLAP, and CWLBP are used as inputs and calculated with the 1-bit weight of the storage unit. The XOR summation result is reflected in R-XACN and R-XACP.

[0078] Example 2

[0079] See Figure 4 This is a structural diagram of the multi-bit weight generation unit provided in Embodiment 2. The multi-bit weight generation unit (which can be abbreviated as MWPU) includes four single-bit weight generation units as in Embodiment 1.

[0080] The four single-bit weight generation units are located in the same row and share the same RWLAN, the same RWLBN, the same RWLAP, the same RWLBP, the same CWLAP, the same CWLBP, the same CWLAN, and the same CWLBN.

[0081] Among the four single-bit weight generation units in the same row, the m-th standard 6T-SRAM cell of each single-bit weight generation unit shares the same WL, m∈[1,n].

[0082] Based on the working principle of the single-bit weight generation unit in Embodiment 1, the four single-bit weight generation units in a single multi-bit weight generation unit operate in parallel. Although four standard 6T-SRAM cells can be opened at once through WL, and four 1-bit weights can be input to the four computing units respectively, the final output result can be controlled by controlling the number of units actually involved in the calculation. For example, if only one computing unit participates in the calculation, the output result has only one set of 4 bits, which can also be regarded as only 1 bit weight as input; if only two computing units participate in the calculation, the output result has two sets of 4 bits, which can also be regarded as only 2 bits weight as input; if all four computing units participate in the calculation, the output result has four sets of 4 bits, which can also be regarded as 4 bits weight as input.

[0083] See again Figure 5 This is the array group provided in Embodiment 2. The array group includes N×N multi-bit weight generation units arranged in an array; N=2 i In this embodiment 2, i is 4, meaning the array group consists of 16×16 distributed multi-bit weight generation units.

[0084] Among them, multiple bit weight generation units located in the same column share the same CWLAP, the same CWLBP, the same CWLAN, and the same CWLBN.

[0085] In the N multi-bit weight generation units in the same column, the q-th single-bit weight generation unit of each multi-bit weight generation unit shares the same C-XACN and the same C-XACP. q∈[1,4].

[0086] In other words, for any column of multi-bit weight generation unit, there are 4 C-XACNs and 4 C-XACPs.

[0087] Multiple bit weight generation units located in the same row share the same RWLAN, the same RWLBN, the same RWLAP, and the same RWLBP.

[0088] In the N multi-bit weight generation units in the same row, the qth single-bit weight generation unit of each multi-bit weight generation unit shares the same R-XACN and the same R-XACP.

[0089] In other words, for any row of multi-bit weight generation unit, there are 4 R-XACNs and 4 R-XACPs.

[0090] Based on the working principle of the single-bit weight generation unit in Embodiment 1, during forward propagation, among the N multi-bit weight generation units in the same column, the qth single-bit weight generation unit of each multi-bit weight generation unit will accumulate the XOR result onto the qth C-XACN.

[0091] During backpropagation, in the N multi-bit weight generation units in the same row, the qth single-bit weight generation unit of each multi-bit weight generation unit will accumulate the XOR result onto the qth R-XACN and the qth R-XACP respectively.

[0092] Example 3

[0093] See Figure 6 This is the in-memory computation macro provided in Embodiment 3. The in-memory computation macro includes: an array group, word line driver, backward channel input driver, forward bit line input, forward channel input driver, backward bit line input, in-memory computation controller, flash memory analog-to-digital converter, successive approximation analog-to-digital converter, and timing controller as disclosed in Embodiment 2.

[0094] The array group is used for forward or backward propagation. The word line driver controls the WL switch. The backward channel input driver controls the CWLAN, CWLAP, CWLBN, and CWLBP switches. The forward bit line input precharges C-XACN to VDD / 2 during forward propagation and connects C-XACN to VSS and C-XACP to VDD during backward propagation. The forward channel input driver controller controls the RWLAN, RWLAP, RWLBN, and RWLBP switches. The backward bit line input controller precharges R-XACN to VDD and R-XACP to VSS during backward propagation. The backward bit line input precharges R-XACN to VDD and R-XACP to VSS during backward propagation. The in-memory compute controller switches the array group's functions. The flash analog-to-digital converter provides a 4-bit output during forward propagation. A successive approximation analog-to-digital converter is used to obtain an 8-bit output during backward propagation. A timing controller is used to control the clock pulses of each signal.

[0095] Among them, the word line driver controller, the backward channel input driver controller, and the forward channel input driver controller are equivalent to switches: only when they are turned on can the corresponding signals be input from the timing controller into the array group.

[0096] The forward bit line input controller and the backward bit line input controller are equivalent to switches, used to precharge or connect the relevant signal lines to the corresponding level for forward or backward propagation.

[0097] The in-memory computing controller can switch between array group storage mode and computing mode: when the array group is in storage mode, weight values ​​can be written into the storage cells first; when the array group is in computing mode, the aforementioned forward propagation or backward propagation is performed.

[0098] Additionally, see Figure 7 For any column of multi-bit weight generation unit, there are 4 C-XACNs. The flash analog-to-digital converter is configured with 4 C-XACNs and connected one-to-one with each of the 4 C-XACNs.

[0099] Thus, during forward propagation, for any column of multi-bit weight generation unit, RWLAN, RWLBN, RWLAP, and RWLBP each carry a 4-bit pulse width signal, which forms a 4-bit input and is XORed with the weight to accumulate the result, which is reflected on the 4 C-XACNs. After being quantized by the flash memory analog-to-digital converter, 4 sets of 4-bit outputs are obtained.

[0100] See Figure 8 The following is a waveform diagram showing a partial case of forward propagation using a column of multi-bit weight generation units:

[0101] Suppose a column of multi-bit weight generation units contains 16 multi-bit weight generation units. The identifier corresponding to the j-th multi-bit weight generation unit is... <j-1>Then the j-th multi-bit weight generation unit contains WL <j-1>、LBL <j-1>,LBLB <j-1>、RWLAN <j-1>、RWLBN <j-1>、RWLAP <j-1>、RWLBP <j-1>、X-ACN <j-1>wait.

[0102] Using the corresponding WL, select and enable the first standard 6T-SRAM cell of each single-bit weight generation unit. Then, set the weight value stored in these 16 standard 6T-SRAM cells to 1. Figure 8 The example only shows the case of the first multi-bit weight generation unit – WL <0> Set to high level, LBL <0> High level, LBLB <0> When the level is low, Weight = +1.

[0103] Configure RWLAN <0> ~RWLAN <15> RWLBN <0> ~RWLBN <15> RWLAP <0> ~RWLAP <15> RWLBP <0> ~RWLBP <15> The input has a 4-bit pulse width signal.

[0104] If RWLAN <0> =0, RWLBN <0> 0, RWLAP <0> =1, RWLBP <0> When XNOR = 0, C-XACN is 1. <0> There is no change, therefore the XAC value is 0.

[0105] So, when RWLAN <0> ~RWLAN <15> All zeros, RWLBN <0> ~RWLBN <15> All zeros, RWLAP <0> ~RWLAP <15> All are 1, RWLBP <0> ~RWLBP <15> When all values ​​are 1, accumulate to C-XACN <0> There was still no change; the XAC value was 0.

[0106] If RWLAN <0> 1, RWLBN <0> 0, RWLAP <0> 0, RWLBP <0> =1, XNOR = +1, C - XACN <0> The value decreases, and after binary conversion, the XAC value is 15.

[0107] So, when RWLAN <0> ~RWLAN <15> All 1s, RWLBN <0> ~RWLBN <15> All zeros, RWLAP <0> ~RWLAP <15> All zeros, RWLBP <0> ~RWLBP <15> When all values ​​are 1, accumulate to C-XACN <0> There were 16 decreases, with an XAC value of 240.

[0108] See Figure 9 For any row of multi-bit weight generation units, there are 4 R-XACNs and 4 R-XACPs. Eight successive approximation analog-to-digital converters are set, of which 4 are connected one-to-one with the 4 R-XACNs and the other 4 are connected one-to-one with the 4 R-XACPs.

[0109] Thus, during backpropagation, for any row of multi-bit weight generation units, CWLAN, CWLBN, CWLAP, and CWLBP each carry a 4-bit pulse width signal, which forms a 4-bit input and is XORed with the weights to accumulate the result, which is reflected on 4 R-XACNs and 4 R-XACPs. After successive approximation analog-to-digital converter quantization, 8 groups of 8-bit outputs are obtained.

[0110] See Figure 10 The diagram shows waveforms for some cases of backpropagation using a single-row multi-bit weight generation unit:

[0111] Suppose a row of multi-bit weight generation units contains 16 multi-bit weight generation units. The identifier corresponding to the j-th multi-bit weight generation unit is... <j-1>Then the j-th multi-bit weight generation unit contains WL <j-1>、LBL <j-1>,LBLB <j-1>、CWLAN <j-1>、CWLBN <j-1>、CLAP <j-1>、CWLBP <j-1>、R-XACN <j-1>、R-XACP <j-1>.

[0112] Using the corresponding WL, select and enable the first standard 6T-SRAM cell of each single-bit weight generation unit. Then, set the weight value stored in these 16 standard 6T-SRAM cells to 1. Figure 10 The example only shows the case of the first multi-bit weight generation unit – WL <0> Set to high level, LBL <0> High level, LBLB <0> When the level is low, Weight = +1.

[0113] Configure CWLAN <0> ~CWLAN <15> CWLBN <0> ~CWLBN <15> CWLAP <0> ~CWLAP <15> CWLBP <0> ~CWLBP <15> The input has a 4-bit pulse width signal.

[0114] If CWLAN <0> =0, CWLBN <0> 0, CWLAP <0> 1. CWLBP <0> When XNOR = 0, R - XACN is 1. <0> R-XACP <0> No changes, R-XACN <0> The corresponding XAC1 is 0, R-XACP <0> The corresponding XAC2 is 0, therefore the overall XAC = XAC1 + XAC2 is also 0.

[0115] So, when CWLAN <0> ~CWLAN <15> All zeros, CWLBN <0> ~CWLBN <15> All zeros, CWLAP <0> ~CWLAP <15> All are 1, CWLBP <0> ~CWLBP <15> When all values ​​are 1, accumulate to R-XACN <0> R-XACP <0> There is still no change, so the overall XAC is still 0.

[0116] If CWLAN <0> 1. CWLBN <0> 0, CWLAP <0> 0, CWLBP <0> The value is 1, at which point R-XACN <0> Decrease, R-XACP <0> Unchanged, R-XACN <0> The corresponding XAC1 is 15, R-XACP <0> The corresponding XAC2 is 0, therefore the overall XAC = XAC1 + XAC2 is 15.

[0117] So, when CWLAN <0> ~CWLAN <15> All 1, CWLBN <0> ~CWLBN <15> All zeros, CWLAP <0> ~CWLAP <15> All zeros, CWLBP <0> ~CWLBP <15> When all values ​​are 1, accumulate to R-XACN <0> There were 16 decreases, accumulated to R-XACP <0> No changes, R-XACN <0> The corresponding XAC1 is 240, R-XACP <0> The corresponding XAC2 is still 0, so the overall XAC is 240.

[0118] It should be noted that for in-memory computation macros, inference operations involve forward propagation; while training operations involve both forward propagation and backward propagation. Based on the above structural design and working method, the integration of inference and training is achieved.

[0119] Finally, the inventors simulated the present invention with two existing arithmetic circuits, comparing the accuracy with 1-bit and 2-bit input weights. The first existing arithmetic circuit has 4-bit input and 4-bit output, while the second has 8-bit input and 8-bit output.

[0120] See the accuracy comparison results. Figure 11 As can be seen, compared with the existing 4-bit input and 4-bit output circuit structure, the accuracy of the present invention is significantly improved and reaches a level similar to that of the existing 8-bit input and 8-bit output circuit structure.

[0121] The inventors also conducted simulation experiments on the energy efficiency of the present invention—during forward propagation, the energy efficiency of the present invention was 106.85 Tops / W; during backward propagation, the energy efficiency was 18.25 Tops / W. This is because a flash memory analog-to-digital converter was used for quantization during forward propagation, thereby improving the energy efficiency.

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A single-bit weight generation unit, characterized in that, include: n standard 6T-SRAM cells are used as storage units; the reading of any standard 6T-SRAM cell is controlled by its word line WL, and the weight value stored therein is reflected in its BL and BLB. The bit lines BL of n standard 6T-SRAM cells are connected to the local bit line LBL, and the bit lines BLB of n standard 6T-SRAM cells are connected to the local bit line LBLB; n≥1; as well as One transposed XNOR accumulator unit serves as the computation unit; the transposed XNOR accumulator unit includes: NMOS transistors N1 to N6 and PMOS transistors P1 to P6; wherein... The drain of N1 is connected to the signal line C-XACN, and the gate is connected to the control signal RWLAN. The drain of N2 is connected to C-XACN, and the gate is connected to the control signal RWLBN; The drain of N3 is connected to the source of N1, the gate is connected to the control signal CWLAN, and the source is connected to C-XACN. The drain of N4 is connected to the source of N2, the gate is connected to the control signal CWLBN, and the source is connected to C-XACN; The drain of N5 is connected to the source of N1 and the drain of N3, the gate is connected to LBL, and the source is connected to the signal line R-XACN. The drain of N6 is connected to the source of N2 and the drain of N4, the gate is connected to LBLB, and the source is connected to R-XACN. The gate of P1 is connected to the control signal CWLAP, and the source is connected to the signal line C-XACP; The gate of P2 is connected to the control signal CWLBP, and the source is connected to C-XACP; The drain of P3 is connected to C-XACN, the gate is connected to the control signal RWLAP, and the source is connected to the drain of P1. The drain of P4 is connected to C-XACN, the gate is connected to the control signal RWLBP, and the source is connected to the drain of P2. The drain of P5 is connected to the drain of P1 and the source of P3, the gate is connected to LBL, and the source is connected to the signal line R-XACP. The drain of P6 is connected to the drain of P2 and the source of P4, the gate is connected to LBLB, and the source is connected to R-XACP.

2. The single-bit weight generation unit according to claim 1, characterized in that, n=8。 3. The single-bit weight generation unit according to claim 1, characterized in that, During forward propagation, C-XACN is precharged to VDD / 2; If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VDD, RWLBN is connected to VSS, N1 is on, and N2 is off; RWLAP is connected to VSS, RWLBP is connected to VDD, P3 is on, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN discharges to R-XACN through N1 and N5; If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VSS, RWLBN is connected to VDD, N2 is on, and N1 is off; RWLAP is connected to VDD, RWLBP is connected to VSS, P4 is on, and P3 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; R-XACP discharges to C-XACN through P4 and P6; If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, N1 is off, and N2 is off; RWLAP is connected to VDD, RWLBP is connected to VDD, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN is kept at VDD / 2.

4. The single-bit weight generation unit according to claim 1, characterized in that, During forward propagation, C-XACN is precharged to VDD / 2; If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VDD, RWLBN is connected to VSS, N1 is on, and N2 is off; RWLAP is connected to VSS, RWLBP is connected to VDD, P3 is on, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; R-XACP discharges to C-XACN through P5 and P3; If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, N1 is off, and N2 is off; RWLAP is connected to VDD, RWLBP is connected to VDD, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN discharges to R-XACN through N2 and N6. If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on; RWLAN is connected to VSS, RWLBN is connected to VSS, N1 is off, and N2 is off; RWLAP is connected to VDD, RWLBP is connected to VDD, P3 is off, and P4 is off; R-XACN is connected to VSS, and R-XACP is connected to VDD; C-XACN is kept at VDD / 2.

5. The single-bit weight generation unit according to claim 1, characterized in that, During backward propagation, R-XACN is precharged to VDD, and R-XACP is precharged to VSS; C-XACN is connected to VSS, and C-XACP is connected to VDD. If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on. CWLAP is connected to VSS, CWLBP is connected to VDD, P1 is on, P2 is off, CWLAN is connected to VDD, CWLBN is connected to VSS, N3 is on, N4 is off, and R-XACN discharges to C-XACN through N5 and N3. If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on. CWLAP is connected to VDD, CWLBP is connected to VSS, P2 is on, P1 is off, CWLAN is connected to VSS, CWLBN is connected to VDD, N4 is on, N3 is off, and C-XACP discharges to R-XACP through P2 and P6. If LBL is high and LBLB is low, N5 is on, P5 is off, N6 is off, and P6 is on; CWLAP is connected to VDD, CWLBP is connected to VDD, P1 is off, P2 is off, CWLAN is connected to VSS, CWLBN is connected to VSS, N3 is off, N4 is off, R-XACN remains at VDD, and R-XACP remains at VSS.

6. The single-bit weight generation unit according to claim 1, characterized in that, During backward propagation, R-XACN is precharged to VDD, and R-XACP is precharged to VSS; C-XACN is connected to VSS, and C-XACP is connected to VDD. If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on. CWLAP is connected to VSS, CWLBP is connected to VDD, P1 is on, P2 is off, CWLAN is connected to VDD, CWLBN is connected to VSS, N3 is on, N4 is off, and C-XACP discharges to R-XACP through P1 and P5. If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on. CWLAP is connected to VDD, CWLBP is connected to VSS, P2 is on, P1 is off, CWLAN is connected to VSS, CWLBN is connected to VDD, N4 is on, N3 is off, and R-XACN discharges to C-XACN through N6 and N4. If LBL is low and LBLB is high, P5 is on, N5 is off, P6 is off, and N6 is on. CWLAP is connected to VDD, CWLBP is connected to VDD, P1 is off, P2 is off, CWLAN is connected to VSS, CWLBN is connected to VSS, N3 is off, N4 is off, R-XACN is kept as VDD, and R-XACP is kept as VSS.

7. A multi-bit weight generation unit, characterized in that, The multi-bit weight generation unit includes four single-bit weight generation units as described in any one of claims 1-6; The four single-bit weight generation units are located in the same row and share the same RWLAN, the same RWLBN, the same RWLAP, the same RWLBP, the same CWLAP, the same CWLBP, the same CWLAN, and the same CWLBN. In the four single-bit weight generation units in the same row, the m-th standard 6T-SRAM cell of each single-bit weight generation unit shares the same WL, m∈[1,n].

8. An array group, characterized in that, including N x N, arrayed, multi-bit weight generating units as claimed in claim 7; N = 2 i , i > 0; Among them, multiple bit weight generation units located in the same column share the same CWLAP, the same CWLBP, the same CWLAN, and the same CWLBN; In the N multi-bit weight generation units in the same column, the qth single-bit weight generation unit of each multi-bit weight generation unit shares the same C-XACN and the same C-XACP; q∈[1,4]; Multiple bit weight generation units located in the same row share the same RWLAN, the same RWLBN, the same RWLAP, and the same RWLBP; In the N multi-bit weight generation units in the same row, the qth single-bit weight generation unit of each multi-bit weight generation unit shares the same R-XACN and the same R-XACP.

9. The array group according to claim 8, characterized in that, During forward propagation, in the N multi-bit weight generation units in the same column, the qth single-bit weight generation unit of each multi-bit weight generation unit will add the XOR result to the qth C-XACN. During backpropagation, in the N multi-bit weight generation units in the same row, the qth single-bit weight generation unit of each multi-bit weight generation unit will accumulate the XOR result onto the qth R-XACN and the qth R-XACP respectively.

10. An in-memory computation macro, characterized in that, include: The array group as described in claim 8 or 9 is used for forward propagation or backward propagation; Word line drive controller, which is used to control the WL switch; The backward channel input drive controller is used to control the CWLAN, CWLAP, CWLBN, and CWLBP switches; A forward bit line input controller is used to precharge C-XACN to VDD / 2 during forward propagation and connect C-XACN to VSS and C-XACP to VDD during backward propagation. Forward channel input drive controller, which is used to control the RWLAN, RWLAP, RWLBN, and RWLBP switches; Backward bit line input controller, which is used to precharge R-XACN to VDD and R-XACP to VSS during backward propagation; In-memory computing control circuit, which is used to switch array groups; A flash memory analog-to-digital converter used to obtain a 4-bit output during forward propagation; A successive approximation analog-to-digital converter is used to obtain an 8-bit output during backpropagation; as well as A timing controller is used to control the clock pulses of various signals.