MRAM (Magnetic Random Access Memory) self-decryption unit, storage and calculation array and working method thereof
By using MRAM self-decryption unit and hybrid voltage gated spin-orbit torque magnetic tunnel junction in the MRAM memory and computing integrated architecture, combined with the surrounding gate carbon nanotube field effect transistor, the security vulnerability problems in the memory and computing integrated architecture in nonvolatile memory are solved, and efficient and secure AI model decryption and multiplication operations are achieved.
Patent Information
- Application Number
- CN202510623958.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing integrated storage and computing architecture has security vulnerabilities in nonvolatile memory, especially when the fixed weight of the AI model may pose a threat to the inference engine, resulting in the risk of model leakage and reverse engineering.
Using MRAM self-decryption unit, combined with a surround gate carbon nanotube field effect transistor (GAA-CNTFET) and a hybrid voltage gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ), a 4T2M structure is built to realize a voltage divider network, supporting data access, decryption and full-precision multiplication operations.
It realizes that different voltages are generated through keys and encryption weights for decryption without additional decryption logic in the AI model, which significantly improves area efficiency and reduces calculation delay, and enhances the security of the AI model.
Smart Images

Figure CN120220752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing-in-memory, and particularly to an MRAM self-decryption unit, a computing-in-memory array, and a working method thereof. Background Art
[0002] With the development of artificial intelligence (AI) technology, the scale of models has been continuously expanding, leading to an increase in computing power and storage requirements. However, the von Neumann bottleneck limits the energy efficiency of AI chips, mainly due to frequent memory accesses. The computing-in-memory architecture can directly perform calculations in memory, thereby reducing the frequent access to storage during the calculation process. Therefore, it is a powerful AI computing architecture that can break through the von Neumann bottleneck. However, the rapid development of the computing-in-memory architecture has also brought new security vulnerabilities, especially in non-volatile memories (such as MRAM and RRAM), because the fixed weights of AI models stored in memory may pose a threat to the inference engine, including model leakage and reverse engineering. Generally, a well-trained AI model has extremely high value due to the large amount of resources and time cost required for its training. Therefore, it is crucial to explore a new CIM macro-architecture that can protect AI models.
[0003] With the continuous scaling down of the CMOS process, the short-channel effect in silicon-based transistors has become increasingly significant. This makes the traditional computing-in-memory architecture based on silicon transistors approach its limit in terms of energy efficiency. In contrast, the gate-all-around carbon nanotube field-effect transistor (GAA-CNTFET) has stronger gate control ability, which can reduce leakage current and increase the on-off ratio. In addition, both n-type and p-type CNTFETs exhibit symmetric transfer characteristic curves, and their threshold voltages can be adjusted by the flat-band voltage. And the device based on voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ) has lower write energy and a write-read separated structure, so it has higher durability than RRAM. In addition, the combination of these technologies can also achieve multi-layer monolithic 3D integration (low-temperature, back-end compatible process), and has excellent electrical properties such as superior radiation resistance, trinary system implementation ability, high speed, and low power consumption. Therefore, the combination of CNTFET and MTJ technologies provides a promising path for constructing an ultra-high-efficiency CIM macro-architecture in the post-Moore era.
[0004] Currently, the CIM research on the hybrid design of CNTFET and MTJ mainly focuses on two aspects:
[0005] Unit level: Current research on unit-level exploration includes ternary non-volatile memory cells for high-capacity storage, low-power true random number generators, near-memory logic computing units, approximate full adders, designs based on spintronic synapses and carbon nanotube transistors as neurons, and XNOR / XOR operation units for accelerating binary neural networks (BNNs). However, these studies are limited to unit-level exploration and have limitations in throughput and energy efficiency.
[0006] Array level: In 2024, Tong et al. proposed a 16kb MRAM based on GAA-CNTFET for full-array Boolean logic operations and in-situ storage. The array can be fully activated to perform half-add operations, achieving high throughput. In the same year, Tong implemented an 8kb highly robust BNN accelerator using a complementary sensing time-domain readout circuit. However, both of these studies are limited to single-bit operations, restricting their applicability in high-precision neural network models. Summary of the Invention
[0007] In view of this, the present invention provides an MRAM self-decryption unit, a memory-computation array, and its working method to solve at least one of the above-mentioned problems.
[0008] To achieve the above object, the present invention adopts the following solutions:
[0009] According to the first aspect of the present invention, there is provided an MRAM self-decryption unit, which includes: a first carbon-based transistor, a second carbon-based transistor, a third carbon-based transistor, a fourth carbon-based transistor, a first magnetic tunnel junction, and a second magnetic tunnel junction; the gate of the first carbon-based transistor is connected to a first write word line, the first end is connected to a bit line, and the second end is connected to the first end of the second magnetic tunnel junction; the gate of the second carbon-based transistor is connected to a first read word line, the first end is connected to an anti-bit line, and the second end is connected to the third end of the first magnetic tunnel junction; the gate of the third carbon-based transistor is connected to a second write word line, the first end is connected to the second end of the second magnetic tunnel junction, and the second end is connected to a source line; the gate of the fourth carbon-based transistor is connected to a second read word line, the first end is connected to the third end of the second magnetic tunnel junction, and the second end is connected to a read source line; the first end of the first magnetic tunnel junction is connected to the third end of the second magnetic tunnel junction, and the second end is connected to an anti-source line; the second and third ends of the first magnetic tunnel junction and the second magnetic tunnel junction are located at the bottom electrode of the magnetic tunnel junction, and the first ends of the first magnetic tunnel junction and the second magnetic tunnel junction are located at the top electrode of the magnetic tunnel junction.
[0010] As an embodiment of the present invention, the first magnetic tunnel junction and the second magnetic tunnel junction are used to store encrypted weight data.
[0011] As an embodiment of the present invention, the bit line and the anti-bit line are used to support the read and write operations of the MRAM self-decryption unit, and are also used to load key data during the decryption operation.
[0012] As an embodiment of the present invention, the first carbon-based transistor, the second carbon-based transistor, the third carbon-based transistor, and the fourth carbon-based transistor are surrounding-gate carbon nanotube field-effect transistors.
[0013] As an embodiment of the present invention, the first magnetic tunnel junction and the second magnetic tunnel junction are hybrid voltage-gated spin-orbit torque magnetic tunnel junctions.
[0014] According to a second aspect of the present invention, there is provided an MRAM computing-in-memory array, the MRAM computing-in-memory array including: a plurality of self-decryption modules, a plurality of transistor logic multiplication modules corresponding to the plurality of self-decryption modules one by one, the self-decryption module being composed of a plurality of self-decryption units as described above, the self-decryption module being configured to complete the decryption operation of the encrypted weight, obtain the original weight after decryption, and input the original weight into the corresponding transistor logic multiplication module for multiplication operation.
[0015] As an embodiment of the present invention, the self-decryption module includes a first decryption sub-module, a second decryption sub-module, a fifth carbon-based transistor, a sixth carbon-based transistor, a first low-threshold inverter, and a second low-threshold inverter; both the first decryption sub-module and the second decryption sub-module include at least one self-decryption unit as described above; the source lines, anti-source lines, bit lines, and anti-bit lines of the self-decryption units in the first decryption sub-module and the second decryption sub-module are connected in rows; the read source lines of the self-decryption units in the first decryption sub-module are connected in rows and are connected to the first end of the fifth carbon-based transistor and the input end of the first low-threshold inverter; the read source lines of the self-decryption units in the second decryption sub-module are connected in rows and are connected to the first end of the sixth carbon-based transistor and the input end of the second low-threshold inverter; the gate of the fifth carbon-based transistor is connected to the first global word line, and the second end is connected to the even data input line; the gate of the sixth carbon-based transistor is connected to the second global word line, and the second end is connected to the odd data input line.
[0016] As an embodiment of the present invention, the transistor logic multiplication module includes a dot product sub-module and two half adders, the dot product sub-module includes four NOR gates composed of eight transistors, the dot product sub-module is configured to perform a dot product operation on the decryption data output by the first low-threshold inverter and the second low-threshold inverter and the data input by the even data input line and the odd data input line; the half adder is configured to accumulate the data output by the dot product sub-module to obtain a partial product sum, realizing a full-precision 2b×2b local multiplication operation.
[0017] As an embodiment of the present invention, the half adder is composed of six transistors and has two input terminals and two output terminals.
[0018] According to the third aspect of the present invention, there is provided a working method of the MRAM computing-in-memory array as described above, the method comprising: pre-storing encrypted weight data in the magnetic tunnel junctions of the MRAM computing-in-memory array; loading a key onto the bit lines and anti-bit lines of the MRAM computing-in-memory array; selecting the MRAM self-decrypting unit to be involved in the operation by controlling the first write word line, the second write word line, the first read word line, and the second read word line, and performing data access and decryption operations; feeding the decrypted weight data and input data into a transistor logic multiplication module to complete a local multiplication operation.
[0019] The MRAM self-decrypting unit, the computing-in-memory array, and the working method thereof provided by the present invention adopt a hybrid voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ) and a gate-all-around carbon nanotube field-effect transistor (GAA-CNTFET), and support simultaneous data access, decryption, and full-precision multiplication operations. The present invention proposes that the MRAM self-decrypting unit uses a 4T2M structure to form a voltage-dividing network, which can generate different voltages according to the key and the encrypted weight, thereby realizing the decryption function without additional decryption logic. Finally, the present invention proposes a transistor logic multiplication module based on transmission gate logic, which can achieve a full-precision 2b-IN × 2b-W local operation with only 20 transistors (20T), significantly improving the area efficiency and reducing the calculation delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. In the drawings:
[0021] Figure 1 is a schematic structural diagram of an MRAM self-decrypting unit provided by an embodiment of the present application;
[0022] Figure 2 is a schematic diagram of writing encrypted weight data to the MRAM self-decrypting unit provided by an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of an MRAM array composed of self-decrypting units provided by an embodiment of the present application;
[0024] Figure 4It is the decryption operation mapping and decryption result diagram provided by the embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of an MRAM memory and computing array provided by the embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of the principle of a self-decryption multiplication operation provided by the embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of a half adder and a basic multiplication operation based on transmission gate logic provided by the embodiment of the present application;
[0028] Figure 8 It is a truth table of the local 2b-IN×2b-W multiplication result provided by the embodiment of the present application;
[0029] Figure 9 It is a schematic flow chart of a working method of an MRAM memory and computing array provided by the embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer and more understandable, the following further describes the embodiments of the present invention in detail with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0031] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0032] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values described in these embodiments do not limit the scope of the present application. At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that: similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0033] In the description of the present application, it should be understood that the orientation or positional relationship indicated by orientation words such as "front, back, top, bottom, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom", etc. is usually based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description. Without contrary explanation, these orientation words do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the protection scope of the present application; the orientation words "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0034] For convenience of description, spatial relative terms such as "above", "over", "on the upper surface", "upper" etc. may be used herein to describe the spatial positional relationship of one device or feature to another device or feature as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures of the device. For example, if the device in the figures is inverted, the device described as "above" or "over" another device or structure will then be positioned "below" or "under" another device or structure. Thus, the exemplary term "above" can include both orientations of "above" and "below". The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and the corresponding explanations are made to the spatial relative descriptions used herein.
[0035] In addition, it should be noted that the use of words such as "first", "second" etc. to limit components is only for the convenience of distinguishing the corresponding components. Without additional statement, the above words have no special meaning. Therefore, it should not be construed as a limitation on the protection scope of the present application. In addition, although the terms used in the present application are selected from well-known and commonly used terms, some of the terms mentioned in the specification of the present application may be selected by the applicant according to his or her judgment, and their detailed meanings are described in the relevant parts of the description herein. In addition, it is required to understand the present application not only through the actual terms used, but also through the meaning implied by each term.
[0036] It should be understood that when a component is referred to as "on another component", "connected to another component", "coupled to another component", or "in contact with another component", it can be directly on, connected to, or coupled to, or in contact with the other component, or there may be an intervening component. In contrast, when a component is referred to as "directly on another component", "directly connected to", "directly coupled to", or "directly in contact with" another component, there is no intervening component. Similarly, when a first component is referred to as "electrically in contact with" or "electrically coupled to" a second component, there is an electrical path allowing current flow between the first component and the second component. The electrical path can include capacitors, coupled inductors, and / or other components allowing current flow, even without direct contact between the conductive components.
[0037] As Figure 1 Shown is a schematic structural diagram of an MRAM self-decryption unit provided by an embodiment of the present application. The MRAM self-decryption unit of this embodiment has a 4T2M structure, that is, it is composed of 4 carbon-based transistors (4T) and 2 magnetic tunnel junctions (2M).
[0038] From Figure 1 It can be seen that the MRAM self-decryption unit includes a first carbon-based transistor M1, a second carbon-based transistor M2, a third carbon-based transistor M3, a fourth carbon-based transistor M4, a first magnetic tunnel junction MTJ0, and a second magnetic tunnel junction MTJ1.
[0039] The gate of the first carbon-based transistor M1 is connected to the first write word line WLR, the first end is connected to the bit line BL, and the second end is connected to the T1 end of the second magnetic tunnel junction MTJ1. Depending on the direction of current flow, the first end of the first carbon-based transistor M1 here can be either the drain or the source, and similarly the second end can also be either the drain or the source. The descriptions of the first end and the second end of the subsequent carbon-based transistors are the same as those of the first carbon-based transistor M1.
[0040] The gate of the second carbon-based transistor M2 is connected to the first read word line RWR, the first end is connected to the anti-bit line BLB, and the second end is connected to the T3 end of the first magnetic tunnel junction MTJ0.
[0041] The gate of the third carbon-based transistor M3 is connected to the second write word line WWL, the first end is connected to the T2 end of the second magnetic tunnel junction MTJ1, and the second end is connected to the source line SL.
[0042] The gate of the fourth carbon-based transistor M4 is connected to the second read word line RWL, the first end is connected to the T3 end of the second magnetic tunnel junction MTJ1, and the second end is connected to the read source line RSL.
[0043] The T1 terminal of the first magnetic tunnel junction MTJ0 is connected to the T3 terminal of the second magnetic tunnel junction, and the T2 terminal is connected to the anti-source line SLB.
[0044] Among them, the T2 and T3 terminals of the first magnetic tunnel junction MTJ0 and the T2 and T3 terminals of the second magnetic tunnel junction MTJ1 are located at the bottom electrode of the magnetic tunnel junction, and the T1 terminal of the first magnetic tunnel junction MTJ0 and the T1 terminal of the second magnetic tunnel junction MTJ1 are located at the top electrode of the magnetic tunnel junction.
[0045] Preferably, the above-mentioned first carbon-based transistor M1, second carbon-based transistor M2, third carbon-based transistor M3, and fourth carbon-based transistor M4 are gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs). More preferably, the above-mentioned first magnetic tunnel junction MTJ0 and second magnetic tunnel junction MTJ1 are hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJs). That is, the MRAM self-decryption unit of the present application is a voltage-divider self-decryption cell (VSDC) composed of four GAA-CNTFETs and two VGSOT-MTJs.
[0046] Preferably, in this embodiment, the first magnetic tunnel junction MTJ0 and the second magnetic tunnel junction MTJ1 are used to store encrypted weight data. In the above-mentioned first magnetic tunnel junction MTJ0 and second magnetic tunnel junction MTJ1, the spin-orbit torque current (I SOT ) is generated by the spin Hall effect current and applied between the T2 and T3 terminals. The MTJ can switch between a high-resistance state (R H ) and a low-resistance state (R L ) without an external magnetic field. In addition, by using the voltage-controlled magnetic anisotropy effect, the energy barrier can be adjusted, thereby significantly reducing the write energy consumption.
[0047] Before performing the decryption operation, through the write operation of the MRAM, the encrypted weight data is pre-written into the above-mentioned first magnetic tunnel junction MTJ0 and second magnetic tunnel junction MTJ1. Figure 2 FIG. is a schematic diagram of writing encrypted weight data into the MRAM self-decryption unit provided by the embodiment of the present application, where part (a) shows the write operation of the second magnetic tunnel junction MTJ1, and part (b) shows the write operation of the first magnetic tunnel junction MTJ0.
[0048] As can be seen from Figure 2 part (a) in, when performing the write operation on the second magnetic tunnel junction MTJ1, the first magnetic tunnel junction MTJ0 remains idle. The write operation on the second magnetic tunnel junction MTJ1 includes the following two stages:
[0049] Phase 1: First, activate the first write word line WLR, the second write word line WWL, and the second read word line RWL. When the source line SL is set to the write voltage (about 200 mV) and the read source line RSL is set to the negative power supply voltage (VSS), the current ISOT flows from the T2 end to the T3 end of the second magnetic tunnel junction MTJ1, thereby writing R L to the second magnetic tunnel junction MTJ1. Conversely, R H can also be written to the second magnetic tunnel junction MTJ1. In this stage, the BL is set to the positive power supply voltage (VDD), thereby enabling the bias voltage V b of the second magnetic tunnel junction MTJ1 to be > 0, reducing the minimum write current.
[0050] Phase 2: Subsequently, turn off the second read word line RWL, while the first write word line WLR and the second write word line WWL remain enabled. The source line SL and the bit line BL are respectively set to VDD and VSS, thereby achieving Vb < 0 and further stabilizing the write operation.
[0051] Similarly, the write operation of the first magnetic tunnel junction MTJ0 is as shown in part (b) below and will not be elaborated here. Figure 2 as shown in part (b) below and will not be elaborated here.
[0052] As can be seen from the above, when writing to the MTJ, this embodiment can optimize the writing process by controlling the voltage of the bit line BL and / or the source line SL. In Phase 1, the bit line BL is set to VDD, thereby enabling the bias voltage V b of the second magnetic tunnel junction MTJ1 to be > 0, and the positive bias voltage V b helps reduce the minimum current required to write to the second magnetic tunnel junction MTJ1; in Phase 2, the source line SL and the bit line BL are respectively set to VDD and VSS, thereby achieving Vb < 0, because the negative bias voltage V b can stabilize the magnetization reversal process of the second magnetic tunnel junction MTJ1, reduce the randomness of the write operation, and improve the reliability and stability.
[0053] Preferably, in this embodiment, the bit line BL and the anti-bit line BLB are used to support the read and write operations of VSDC and are also used to load key data during the decryption operation.
[0054] The decryption process of the MRAM self-decryption unit is described below. As shown in Figure 3 is a schematic diagram of an MRAM array composed of self-decryption units provided by an embodiment of the present application. Figure 3It can be seen that each column of the MRAM array is composed of multiple MRAM self-decryption units connected in series. During the decryption operation, the first write word line WLR, the first read word line RWR, and the second read word line RWL are all set to VDD. Then, the key Key n is loaded onto BL n / BLB n , so that the MRAM array can achieve column-by-column decryption.
[0055] As Figure 3 shown, the decryption result of VSDC is sensed by a low-threshold inverter (LVT-INV), which is composed of a p-CNTFET and an n-CNTFET. The flat-band voltages (V fb ) of the p-CNTFET and the n-CNTFET are -0.4 V and 0.2 V respectively, so that the low-threshold inverter obtains a threshold voltage of 0.3 V. The encrypted weight (W e =W⊕Key) is stored in MTJ1 and MTJ0, where W is the original weight data, Key is the key data, and W e is the encrypted weight data.
[0056] Figure 4 Shows the mapping table and decryption results of the decryption operation (XOR), where BL / BLB = VDD / VSS (VSS / VDD) represents Key = 0 (1) respectively, and MTJ1 / MTJ0 = R L / R H (R H / R L ) represents W e =0 (1) respectively. The result of the decryption operation is reflected in the voltage (V RSL ) of the read source line RSL. For example, when Key and W e are the same, V RSL is as follows:
[0057]
[0058] Among them, V H represents that the calculation result is greater than V LVT , V LVT is Figure 3 the threshold voltage of the LVT-INV in TH , V M1 and R M2 are the on-resistances (R ON ) of the first carbon-based transistor M1 and the second carbon-based transistor M2. At this time, the LVT-INV outputs logic 0. It should be noted that the R of VGSOT-MTJH and R L are approximately 975 kΩ and 332 kΩ respectively, the tunneling magnetoresistance ratio (TMR) is 200%, while R M1 and R M2 are approximately 3.3 kΩ. For simplicity, R ON can be ignored. When Key and W e are not synchronized, V RSL becomes:
[0059]
[0060] where V L represents that the calculation result is less than V LVT . Therefore, the decryption operation result (W e ⊕ Key) between Key and W e can be read through LVT-INV. At this time, LVT-INV outputs logic 1.
[0061] As Figure 5 shown is a schematic diagram of an MRAM arithmetic storage array provided by an embodiment of the present application. This MRAM arithmetic storage array can simultaneously complete decryption and multiplication operations. As Figure 5 can be seen, this MRAM arithmetic storage array includes multiple self-decryption modules 510, and multiple transistor logic multiplication modules 520 corresponding one-to-one to the multiple self-decryption modules 510. Among them, the self-decryption module 510 is composed of multiple VSDCs corresponding to Figure 1 . This self-decryption module 510 is used to complete the decryption operation of the encrypted weight, obtain the decrypted original weight, and input the original weight into the corresponding transistor logic multiplication module 510 for multiplication operation.
[0062] Next, one column of the above MRAM arithmetic storage array is selected to describe in detail how the present application realizes the self-decryption multiplication operation. As Figure 6 shown is a schematic diagram of a self-decryption multiplication operation provided by an embodiment of the present application. As Figure 6 can be seen, the above self-decryption module 510 includes a first decryption sub-module 511, a second decryption sub-module 512, a fifth carbon-based transistor M5, a sixth carbon-based transistor M6, a first low-threshold inverter 513, and a second low-threshold inverter 514.
[0063] Both the first decryption sub-module 511 and the second decryption sub-module 512 include at least one self-decryption unit (VSDC) corresponding to Figure 1 . Figure 6The description is given by taking the example that each of the first decryption sub-module 511 and the second decryption sub-module 512 includes 8 VSDCs (VSDC#0-7 and VSDC#8-15). From the foregoing description, it can be seen that the first decryption sub-module 511 and the second decryption sub-module 512 exist as two weight storage units in the decryption multiplication operation.
[0064] The source line SL, the anti-source line SLB, the bit line BL, and the anti-bit line BLB of the self-decryption units in the first decryption sub-module 511 and the second decryption sub-module 512 are connected row by row. That is, in each row of the column MRAM computing array, the source line SL, the anti-source line SLB, the bit line BL, and the anti-bit line BLB of each self-decryption unit are respectively connected together.
[0065] The read source line RSL of the self-decryption units in the first decryption sub-module 511 is connected row by row. That is, in the first decryption sub-module 511, the read source lines RSL of each self-decryption unit are connected together. And the read source line RSL is also connected to one end drain of the fifth carbon-based transistor M5 and the input end of the first low-threshold inverter 513.
[0066] The read source line RSL of the self-decryption units in the second decryption sub-module 512 is also connected row by row, and is also connected to one end drain of the sixth carbon-based transistor M6 and the input end of the second low-threshold inverter 514.
[0067] The gate of the fifth carbon-based transistor M5 is connected to the first global word line (GWL[1]), and the other end drain is connected to the even data input line GSL_E.
[0068] The gate of the sixth carbon-based transistor M6 is connected to the second global word line (GWL[0]), and the other end drain is connected to the odd data input line GSL_O.
[0069] As Figure 6 shown, the output ends of the first low-threshold inverter 513 and the second low-threshold inverter 514, as well as the even data input line GSL_E and the odd data input line GSL_O, are also connected to the transistor logic multiplication module 520 (PTL-MC). The transistor logic multiplication module 520 includes a dot product sub-module and two half adders (6T-HA). The dot product sub-module is composed of four NOR gates formed by eight transistors. The dot product sub-module is used to perform a dot product operation on the decryption data output by the first low-threshold inverter 513 and the second low-threshold inverter 514, as well as the data input by the even data input line GSL_E and the odd data input line GSL_O. The half adder is used to accumulate the data output by the dot product sub-module to obtain a partial product sum, realizing a full-precision 2b×2b local multiplication operation.
[0070] Preferably, the above half adder is composed of six transistors and has two input terminals (A, B), a sum output terminal (S), and a carry output terminal (CO).
[0071] From Figure 6 It can be seen that the output terminal of the first low-threshold inverter 513 is respectively connected to the first NOR gate and the third NOR gate, and the output terminal of the second low-threshold inverter 514 is respectively connected to the second NOR gate and the fourth NOR gate. GSL_E is respectively connected to the third NOR gate and the fourth NOR gate, and GSL_O is respectively connected to the first NOR gate and the second NOR gate. The output terminal of the second NOR gate is connected to the A terminal of the first half adder, the output terminal of the third NOR gate is connected to the B terminal of the first half adder, the CO terminal of the first half adder is connected to the B terminal of the second half adder, and the output terminal of the fourth NOR gate is connected to the A terminal of the second half adder. For the specific structure of the half adder, reference can be made to Figure 7 As shown, it is implemented using CMOS gate circuits inside.
[0072] Take Figure 6 as an example. When performing self-decrypting multiplication, set GWL[1:0] to VSS to isolate GSL_E / GSL_O from its adjacent RSL_E / RSL_O bit lines. Through the MRAM write operation, the encrypted weights are pre-stored in the MRAM computing-in-memory array. To improve throughput, two weight storage units (VSDC#0-7 and VSDC#8-15) can be accessed simultaneously. The key non is loaded onto BL and BLB. According to , the low-threshold inverter (LVT-INV) outputs the decrypted weights , and inputs them into the transistor logic multiplication module 520 for multiplication calculation.
[0073] Meanwhile, the 2-bit input is configured into GSL_E and GSL_O. Subsequently, through the processing of the NOR gate based on PTL-MC, the dot product operation is completed to generate a 4-bit partial product (P[3:0]). Finally, use Figure 7 the two transmission-gate-logic-based half adders (6T-HFs) shown to accumulate P[3:0] (where P[3]=W[1]&IN[1], P[2]=W[0]&IN[1], P[1]=W[1]&IN[0], P[0]=W[0]&IN[0]) to obtain the partial product sum (S[3:0] = IN[1:0]×W[1:0]).
[0074] As described above, the PTL-MC unit proposed in the embodiment of the present application includes 4 transmission-gate based NOR gates and two 6T-HFs, and only 20 transistors (20T) are required in total to achieve full-precision 2b-IN×2b-W local multiplication operation, greatly improving the area efficiency. Here, b represents bit, 2b-IN represents 2-bit input data, W represents the decrypted weight data, and 2b-W represents 2-bit decrypted weight data.
[0075] The following further illustrates the above decryption multiplication operation through a specific example:
[0076] Taking W e [1:0] = “10”, IN[1:0]= “10” and Key=“0” as an example. In VSDC #8, MTJ1 and MTJ0 are respectively stored as R H and R L , indicating that W e [1] = “1”; in VSDC #0, MTJ1 and MTJ0 are respectively stored as R L and R H , indicating that W e [0]= “0”. BL and BLB are respectively set to VSS and VDD, indicating =“1”. GSL_E and GSL_O are respectively set to “1” and “0”, corresponding to . At the same time, activate WLR[8&0] / RWL[8&0] and RWR[8&0], and set WWL[0] and WWL[8] to VSS to turn off M3. Under this condition, V RSL_O >V LVT , while V RSL_E < V LVT . Thus, can be obtained from D[1:0]. Then, the dot product operation of is completed through 4 transmission-gate logic based NOR gates. Finally, P[3:0] =“1000” is accumulated through two 6T-HFs to obtain S[3:0] = “0100”. Therefore, the decryption and multiplication operations can be completed simultaneously without additional decryption logic delay. Other cases can be referred to Figure 8 , which is the truth table of the local 2b-IN×2b-W multiplication result provided by the embodiment of the present application.
[0077] As Figure 9 shown is a schematic flow chart of a working method of an MRAM computing-in-memory array provided by the embodiment of the present application. The execution subject of this method is the Figure 5 corresponding MRAM computing-in-memory array. This method includes the following steps:
[0078] Step S901: Pre-store the encrypted weight data in the magnetic tunnel junctions of the MRAM memory and computing array.
[0079] Step S902: Load the key onto the bit lines and complementary bit lines of the MRAM memory and computing array.
[0080] Step S903: Select the MRAM self-decrypting units to be involved in the operation by controlling the first write word line, the second write word line, the first read word line, and the second read word line, and perform data access and decryption operations.
[0081] Step S904: Feed the decrypted weight data and the input data into the transistor logic multiplication module to complete the local multiplication operation.
[0082] Each step of the above method has been mentioned in the foregoing description and will not be elaborated further here.
[0083] As can be seen from the above, the MRAM self-decrypting unit, the memory and computing array, and their working methods provided by the present invention adopt a hybrid voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ) and a gate-all-around carbon nanotube field-effect transistor (GAA-CNTFET), and support simultaneous data access, decryption, and full-precision multiplication operations. The MRAM self-decrypting unit proposed by the present invention uses a 4T2M structure to form a voltage division network, which can generate different voltages according to the key and the encrypted weight, thereby realizing the decryption function without additional decryption logic. Finally, the present invention proposes a transistor logic multiplication module based on transmission gate logic, which can achieve a full-precision 2b-IN × 2b-W local operation with only 20 transistors (20T), significantly improving the area efficiency and reducing the calculation delay.
[0084] In the present invention, specific embodiments are used to elaborate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An MRAM self-decryption unit, characterized in that: The MRAM self-decryption unit includes: a first carbon-based transistor, a second carbon-based transistor, a third carbon-based transistor, a fourth carbon-based transistor, a first magnetic tunnel junction, and a second magnetic tunnel junction; The gate of the first carbon-based transistor is connected to the first write word line, the first end is connected to the bit line, and the second end is connected to the first end of the second magnetic tunnel junction; The gate of the second carbon-based transistor is connected to the first read word line, the first end is connected to the inverting bit line, and the second end is connected to the third end of the first magnetic tunnel junction; The gate of the third carbon-based transistor is connected to the second write word line, the first end is connected to the second end of the second magnetic tunnel junction, and the second end is connected to the source line; The gate of the fourth carbon-based transistor is connected to the second read word line, the first end is connected to the third end of the second magnetic tunnel junction, and the second end is connected to the read source line; The first end of the first magnetic tunnel junction is connected to the third end of the second magnetic tunnel junction, and the second end is connected to the anti-source line; The second end and the third end of the first magnetic tunnel junction and the second magnetic tunnel junction are located at the bottom electrode of the magnetic tunnel junction, and the first end of the first magnetic tunnel junction and the second magnetic tunnel junction are located at the top electrode of the magnetic tunnel junction.
2. The MRAM self-decryption unit according to claim 1, wherein: The first magnetic tunnel junction and the second magnetic tunnel junction are used to store encrypted weight data.
3. The MRAM self-decryption unit according to claim 1, characterized in that: The bit line and the inverted bit line are used to support the read and write operations of the MRAM self-decryption unit and are also used to load key data during a decryption operation.
4. The MRAM self-decryption unit according to claim 1, wherein: The first carbon-based transistor, the second carbon-based transistor, the third carbon-based transistor and the fourth carbon-based transistor are all-around gate carbon nanotube field effect transistors.
5. The MRAM self-decryption unit according to claim 1, wherein: The first magnetic tunnel junction and the second magnetic tunnel junction are hybrid voltage-gated spin-orbit torque magnetic tunnel junctions.
6. An MRAM storage and computing array, characterized in that: The MRAM storage and computing array includes: multiple self-decryption modules, and multiple transistor logic multiplication modules corresponding to the multiple self-decryption modules one by one, wherein the self-decryption modules are composed of multiple self-decryption units according to any one of claims 1 to 5, and the self-decryption modules are used to complete the decryption operation of the encrypted weights, obtain the decrypted original weights, and input the original weights into the corresponding transistor logic multiplication modules for multiplication operations.
7. The MRAM storage and computing array according to claim 6, characterized in that: The self-decryption module includes a first decryption submodule, a second decryption submodule, a fifth carbon-based transistor, a sixth carbon-based transistor, a first low-threshold inverter, and a second low-threshold inverter; The first decryption submodule and the second decryption submodule each include at least one self-decryption unit according to any one of claims 1 to 5; The source lines, anti-source lines, bit lines and anti-bit lines of the self-decryption units in the first decryption submodule and the second decryption submodule are connected in rows; The read source lines of the self-decryption units in the first decryption submodule are connected in rows and connected to the first end of the fifth carbon-based transistor and the input end of the first low threshold inverter; The read source lines of the self-decryption units in the second decryption submodule are connected in rows and connected to the first end of the sixth carbon-based transistor and the input end of the second low threshold inverter; The gate of the fifth carbon-based transistor is connected to the first global word line, and the second end is connected to the even data input line; The gate of the sixth carbon-based transistor is connected to the second global word line, and the second terminal is connected to the odd data input line.
8. The MRAM storage and computing array according to claim 7, characterized in that: The transistor logic multiplication module includes a dot product submodule and two half adders. The dot product submodule includes four NOR gates composed of eight transistors. The dot product submodule is used to perform a dot product operation on the decrypted data output by the first low threshold inverter and the second low threshold inverter and the data input by the even data input line and the odd data input line; the two half adders are used to accumulate the data output by the dot product submodule to obtain partial product sums, thereby realizing a full-precision 2b×2b local multiplication operation.
9. The MRAM storage and computing array according to claim 8, characterized in that: The half adder is composed of six transistors and has two input terminals and two output terminals.
10. A method for operating the MRAM storage and computing array according to any one of claims 6 to 9, characterized in that: The method comprises: Pre-storing the encrypted weight data in the magnetic tunnel junction of the MRAM storage array; Loading a key onto the bit lines and inverse bit lines of the MRAM storage array; The MRAM self-decryption unit to be involved in the operation is selected by controlling the first write word line, the second write word line, the first read word line and the second read word line, and performing data access and decryption operations; The decrypted weight data and input data are sent to the transistor logic multiplication module to complete the local multiplication operation.
Citation Information
Patent Citations
2T-2MTJ memory calculation unit and MRAM memory calculation circuit
CN117807021A