Three-dimensional storage array, manufacturing method thereof, and storage and calculation method
By setting up independent sub-storage blocks for parallel operation in a three-dimensional storage array, the performance bottleneck of NAND storage arrays in large-scale data processing is solved, achieving more efficient data access and computing capabilities to adapt to complex task requirements.
Patent Information
- Application Number
- CN202510846876.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When dealing with scenarios where large-scale data is frequently accessed, the existing three-dimensional architecture of NAND storage arrays faces problems such as limited data access rate, insufficient throughput, inflexible update mechanism, and excessively frequent erase and write operations, which lead to accelerated storage cell aging and increased system latency.
By setting multiple independent upper selection lines in the three-dimensional storage array, it is divided into multiple sub-storage blocks that can operate independently. Each sub-storage block operates in parallel or independently with other sub-storage blocks, using outer product to write data, and reading and writing data in parallel to optimize the storage and calculation process.
It improves the operational flexibility and utilization of three-dimensional storage arrays, reduces redundant waste, enhances parallel processing capabilities, improves computing speed and efficiency, reduces waiting time for storage or computing tasks, and adapts to complex task requirements.
Smart Images

Figure CN120379268B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of integrated circuit technology, and in particular to a three-dimensional storage array and a manufacturing method and storage and computing method thereof. Background Art
[0002] NAND Flash (NAND) is a type of non-volatile memory. NAND Flash offers advantages such as high storage density, low cost, long life, and low power consumption. NAND Flash increases storage density by vertically stacking memory cells, connecting them in series. This achieves higher storage density and lower costs. However, NAND Flash still has many issues. Summary of the Invention
[0003] Based on this, it is necessary to provide a three-dimensional storage array and its manufacturing method and storage and computing method to address the problems in the existing technology.
[0004] In a first aspect, the present application provides a three-dimensional storage array, comprising:
[0005] A plurality of memory strings, each comprising a plurality of memory transistors and an upper gate transistor sequentially connected in series along a first direction;
[0006] a plurality of word lines, the word lines being arranged at intervals along the first direction, each of the word lines being connected to a control gate of the storage transistor disposed in the same layer;
[0007] a plurality of bit lines, the bit lines extending along a second direction and arranged at intervals along a third direction, the bit lines being connected to the upper gate transistors arranged along the second direction; the second direction intersecting the third direction, the first direction being perpendicular to a plane containing the second direction and the third direction;
[0008] A plurality of upper gate lines are arranged at intervals, each of the upper gate lines is connected to the control gate of at least one of the upper gate tubes, the storage strings connected to each of the upper gate lines constitute a sub-storage block, and the plurality of upper gate lines divide the plurality of storage strings into a plurality of independent sub-storage blocks, each of the sub-storage blocks operates in parallel or independently with the other sub-storage blocks.
[0009] Optionally, at least two sub-memory blocks capable of operating in parallel are included, and the at least two sub-memory blocks capable of operating in parallel are connected to different upper gate lines and different bit lines.
[0010] Optionally, the upper gate lines extend along the third direction and are arranged at intervals along the second direction and the third direction.
[0011] Optionally, the upper gate line is led out from a side of the three-dimensional memory array and connected to a peripheral circuit; or, the upper gate line is led out from a top of the three-dimensional memory array and connected to the peripheral circuit.
[0012] Optionally, the three-dimensional storage array includes:
[0013] substrate;
[0014] a source line, disposed on the substrate;
[0015] a stacked structure, provided on a side of the source line away from the substrate, the stacked structure comprising a conductive layer and an insulating layer stacked in sequence;
[0016] a plurality of storage structures passing through the stacked structure along the first direction, wherein the bottom of the storage structure is connected to the source line; and a direction pointing from the stacked structure to the center of the storage structure, wherein the storage structure includes a gate dielectric layer, a charge trapping layer, a tunneling layer, a channel layer, and an isolation layer arranged in sequence;
[0017] The top conductive layer of the stacked structure is divided into a plurality of independently arranged upper gate lines.
[0018] In a second aspect, the present application provides a method for manufacturing a three-dimensional memory array, comprising:
[0019] providing a substrate, and forming a source line on one side of the substrate;
[0020] Alternatingly forming an insulating layer and a sacrificial layer on a side of the source line away from the substrate;
[0021] forming a first hole, wherein the first hole penetrates the alternating insulating layers and the sacrificial layers along a first direction perpendicular to the substrate, and the first hole exposes a portion of a top surface of the source line;
[0022] forming a storage structure in the first hole, the storage structure comprising a gate dielectric layer, a charge trapping layer, a tunneling layer, a channel layer, and an isolation layer sequentially covering a hole wall of the first hole;
[0023] The sacrificial layer is removed, and conductive layers arranged at intervals along the first direction are formed in the area where the sacrificial layer is removed; the multiple layers of the conductive layers are used to form a lower gate line, a plurality of word lines, and an upper gate line arranged along the first direction;
[0024] Etching the top conductive layer to form a plurality of upper gate lines spaced apart from each other;
[0025] A plurality of bit lines are formed, wherein the bit lines extend along the second direction and are arranged at intervals along the third direction, and the bit lines cover top surfaces of the memory structures arranged along the second direction.
[0026] Optionally, etching the top conductive layer to form a plurality of upper gate lines spaced apart from each other comprises:
[0027] Etching the top conductive layer to form a first trench extending along the second direction, wherein the first trench divides the top conductive layer into a plurality of upper gate lines arranged at intervals along the second direction;
[0028] Alternatively, the top conductive layer is etched to form a second trench extending along the third direction, wherein the second trench divides the top conductive layer into a plurality of upper gate lines arranged at intervals along the third direction.
[0029] Optionally, the top conductive layer is etched to form the first trench extending along the second direction and the second trench extending along the third direction, and the upper gate lines extend along the third direction and are arranged at intervals along the second and third directions.
[0030] In a third aspect, the present application provides a storage and computing method based on a three-dimensional storage array, wherein the three-dimensional storage array includes at least two sub-storage blocks capable of operating in parallel, wherein the at least two sub-storage blocks capable of operating in parallel are connected to different upper gate lines and different bit lines, and the storage and computing method includes:
[0031] Inputting a first vector and a second vector in parallel to an upper gate line and a bit line, and writing an outer product of the first vector and the second vector into a first storage transistor of a sub-memory block;
[0032] The third vector and the fourth vector are input in parallel to another upper gate line and another bit line, and the outer product of the third vector and the fourth vector is written into the second storage transistor of another sub-memory block.
[0033] Optionally, when reading:
[0034] precharging each of the bit lines to a preset voltage;
[0035] Applying a turn-on voltage to the upper gate line connected to the first storage transistor, applying a zero bias voltage to the other upper gate lines, and reading data stored in the first storage transistor according to discharge of the bit line connected to the first storage transistor;
[0036] An on-voltage is applied to the upper gate line connected to the second storage transistor, and a zero bias voltage is applied to the other upper gate lines. Data stored in the second storage transistor is read out according to discharge of the bit line connected to the second storage transistor.
[0037] The three-dimensional memory array and its manufacturing method and storage and computing method of the present application divide the three-dimensional memory array into multiple independently operable sub-memory blocks by providing multiple independent upper selection lines. Each sub-memory block operates in parallel or independently with the other sub-memory blocks, and storage and / or computing tasks can be assigned to each sub-memory block according to application requirements, thereby improving the operational flexibility and utilization of the three-dimensional memory array. Each sub-memory block operates in parallel or independently with the other sub-memory blocks, which alleviates the problem of a large number of transistors being idle during read operations of the three-dimensional memory array, thereby improving the utilization of the three-dimensional memory array and reducing redundant waste. At the same time, while operating one sub-memory block, the other sub-memory blocks can simultaneously perform storage or computing operations, thereby improving the parallel processing capability of the three-dimensional memory array and significantly improving the computing speed and efficiency of the three-dimensional memory array. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 is a circuit diagram of a three-dimensional memory array provided in one embodiment;
[0040] Figure 2 is a structural diagram of a three-dimensional storage array provided in one embodiment;
[0041] Figure 3 A projection diagram of a memory string on a substrate provided in one embodiment;
[0042] Figure 4 A projection diagram of a sub-memory block on a substrate provided in an embodiment;
[0043] Figure 5 is a process flow chart of a method for manufacturing a three-dimensional memory array provided in one embodiment;
[0044] Figure 6 A flowchart of a storage and computing method based on a three-dimensional storage array provided in one embodiment. DETAILED DESCRIPTION
[0045] To facilitate understanding of the present application, a more comprehensive description of the present application will be provided below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the present application. However, the present application may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present disclosure.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.
[0047] In recent years, deep learning models, especially large language models, have achieved remarkable results in many fields. These models often have parameters in the tens or even hundreds of billions, requiring large-capacity memory to store and load model weights. At the same time, to improve the efficiency of model fine-tuning and inference, these weight data must be able to be updated and read quickly. However, existing storage solutions struggle to strike a balance between capacity and speed, especially when dealing with models with large parameters. Traditional storage structures and data read and write methods face bottlenecks in storage density and read / write bandwidth. The three-dimensional architecture of NAND storage arrays can significantly increase storage capacity and reduce storage costs. Deep learning models (such as large language models) use NAND storage arrays to store model weights.
[0048] The three-dimensional architecture of NAND storage arrays failed to fully consider the special needs of emerging applications such as large models in its design. When dealing with scenarios where large-scale data is frequently called, its performance bottlenecks are becoming increasingly prominent, mainly manifested in problems such as limited data access rates, insufficient throughput, inflexible update mechanisms, and excessively frequent erase and write operations. These technical shortcomings not only accelerate the aging process of storage cells but also significantly increase system latency. Therefore, how to optimize the three-dimensional architecture of NAND storage arrays to optimize the storage mechanism, update strategy, and read process of large model weights to improve inference efficiency, reduce system latency, and extend storage life has become a key technical challenge that needs to be overcome.
[0049] According to an exemplary embodiment, this embodiment provides a three-dimensional storage array, referring to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4As shown, the three-dimensional memory array includes multiple memory strings Str, multiple word lines WL, multiple bit lines BL, and multiple upper selection lines TSG. The memory strings Str are arranged at intervals, and each memory string Str includes multiple memory transistors T and upper selection transistors TM connected in series along a first direction Z. The channels of the multiple memory transistors T and the upper selection transistors TM are connected in series. The source end of the memory string Str is connected to the source line SL; multiple word lines WL are arranged at intervals along a first direction Z, and each word line WL is connected to the control gate of a memory transistor T arranged in the same layer; bit lines BL extend along a second direction Y and are arranged at intervals along a third direction X, and each bit line BL is connected to an upper gate transistor TM arranged along the second direction Y, and each bit line BL controls a column of memory strings Str connected thereto; the second direction Y intersects with the third direction X, and the first direction Z is perpendicular to the plane containing the second direction Y and the third direction X; multiple upper gate lines TSG are arranged at intervals, each upper gate line TSG is connected to the control gate of at least one upper gate transistor TM, and each upper gate line TSG controls the memory string Str connected thereto, and the memory strings Str connected to each upper gate line TSG constitute a sub-memory block 100. The multiple upper gate lines TSG divide the multiple memory strings Str into multiple independent sub-memory blocks 100, and each sub-memory block 100 operates in parallel or independently with the other sub-memory blocks 100.
[0050] like Figure 1 As shown, the three-dimensional memory array includes at least n word lines WL1 to WLn, where n is the number of word lines and is a positive integer greater than or equal to 2. The three-dimensional memory array includes at least m upper gate lines TSG, where n is a positive integer greater than or equal to 2. The three-dimensional memory array may further include m+1 upper gate lines TSG, or m+2 upper gate lines TSG.
[0051] The three-dimensional memory array of this embodiment is divided into multiple independently operable sub-memory blocks 100 by providing multiple independent top gate lines TSG. Each sub-memory block 100 operates in parallel or independently of the other sub-memory blocks 100. Storage tasks and / or computing tasks can be assigned to each sub-memory block 100 according to application requirements, thereby improving the operational flexibility and utilization of the three-dimensional memory array. Each sub-memory block 100 operates in parallel or independently with the other sub-memory blocks 100, thereby improving the problem of a large number of transistors being idle during read operations in the three-dimensional memory array, thereby improving the utilization of the three-dimensional memory array and reducing redundancy. At the same time, while one sub-memory block 100 is operating, the other sub-memory blocks 100 can simultaneously perform storage or computing operations, thereby improving the parallel processing capability of the three-dimensional memory array and significantly improving the computing speed and efficiency of the three-dimensional memory array.
[0052] The traditional three-dimensional architecture of NAND memory arrays is limited by the order of operations and serial processing, resulting in long wait times for storage or computing tasks. In the three-dimensional memory array of this embodiment, multiple sub-memory blocks 100 can perform storage or computing operations simultaneously, reducing the wait time for storage or computing tasks and improving the response speed of the three-dimensional memory array. Furthermore, the three-dimensional memory array is divided into independent sub-memory blocks 100, allowing the three-dimensional memory array to adapt to more complex storage and computing tasks. As task requirements increase, the storage capacity of the three-dimensional memory array can be expanded by increasing the number of sub-memory blocks 100 in the three-dimensional memory array or adjusting the configuration of the sub-memory blocks 100, thereby meeting the increased storage and / or computing needs.
[0053] In some embodiments, reference Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, the three-dimensional memory array includes at least two sub-memory blocks 100 capable of operating in parallel. These at least two sub-memory blocks 100 are connected to different upper gate lines TSG and different bit lines BL. This allows the at least two sub-memory blocks 100 to operate simultaneously. By independently controlling the upper gate lines TSG and bit lines BL, parallel writing and reading of the at least two sub-memory blocks 100 can be achieved, significantly increasing the parallelism of the three-dimensional memory array and improving the throughput of the three-dimensional memory array.
[0054] In some embodiments, reference Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, a plurality of upper gate lines TSG are arranged at intervals along the second direction Y; and / or, a plurality of upper gate lines TSG are arranged at intervals along the third direction X. For example, the upper gate lines TSG are arranged at intervals along the second direction Y; or, the upper gate lines TSG are arranged at intervals along the third direction X; or, the upper gate lines TSG are arranged at intervals along both the second direction Y and the third direction X.
[0055] In some embodiments, reference Figure 1 、 Figure 2 、 Figure 3 、 Figure 4As shown, the upper gate lines TSG extend along the third direction X and are arranged at intervals along the second direction Y and the third direction X. Each memory string Str connected to each upper gate line TSG serves as a sub-memory block 100, and the upper gate line TSG controls the sub-memory block 100 to which it is connected. Multiple upper gate lines TSG are arranged at intervals along the second direction Y and the third direction X, dividing the three-dimensional memory array into multiple sub-memory blocks 100 arranged along the second direction Y and the third direction X. Each sub-memory block 100 can be independently controlled by the upper gate lines TSG and the bit lines BL, so that each sub-memory block 100 can operate in parallel with or independently of other sub-memory blocks 100.
[0056] In this embodiment, the upper gate line TSG is a linear structure extending along the third direction X. However, this embodiment does not limit the three-dimensional array structure of the present application. In other embodiments, the upper gate line TSG can be designed into other shapes. For example, the upper gate line TSG can connect several adjacent memory strings Str, and the upper gate line TSG can be square, circular, or other shapes.
[0057] In some other embodiments, the number of upper selection lines TSG can be the same as the number of storage strings Str, and the upper selection lines TSG are set corresponding to the storage strings Str. In this way, each storage string Str constitutes a sub-storage block 100, which is equivalent to that the storage transistor T of each storage string Str can be operated independently, and the operation flexibility and utilization rate of the three-dimensional storage array are higher.
[0058] In some embodiments, the upper gate line TSG is drawn from a side of the three-dimensional memory array and connected to a peripheral circuit; or, the upper gate line TSG is drawn from a top of the three-dimensional memory array and connected to a peripheral circuit.
[0059] In some examples, multiple upper gate lines (TSGs) are arranged in two rows, spaced apart along a third direction (X). Each row of upper gate lines (TSGs) includes multiple upper gate lines (TSGs) spaced apart along a second direction (Y). Alternatively, a three-dimensional memory array includes two upper gate lines (TSGs), spaced apart along the third direction (X). The upper gate lines (TSGs) are extended from two sidewalls of the three-dimensional memory array along the third direction (X). The upper gate lines (TSGs) are connected to peripheral circuits via metal lead wires. This makes the extension of the upper gate lines (TSGs) compatible with complementary metal oxide semiconductor (CMOS) processes, saving process costs and reducing manufacturing complexity and yield loss. Furthermore, the upper gate lines (TSGs) are extended from the sides of the three-dimensional memory array, and the lead wires assist in heat dissipation, thereby reducing heat accumulation within the three-dimensional memory array.
[0060] In other examples, multiple upper gate lines TSG are arranged in at least three rows along a third direction X, or the three-dimensional memory array includes at least three upper gate lines TSG spaced apart along the third direction X. The upper gate lines TSG are extended from the top of the three-dimensional memory array and connected to peripheral circuits. For example, the upper gate lines TSG are extended from the top of the three-dimensional memory array to the peripheral circuits via a vertical interconnect structure. For example, the upper gate lines TSG can be extended to the peripheral circuits via a through-silicon via (TSV) structure. In this way, more sub-memory blocks 100 can be constructed within the three-dimensional memory array, thereby increasing the utilization of the three-dimensional memory array. Furthermore, the vertical interconnect has a short transmission path, which can reduce signal transmission delay.
[0061] In some embodiments, reference Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, the three-dimensional memory array further includes a substrate 10, a source line SL, a stacked structure 20, and multiple memory structures 30. The source line SL is disposed on the substrate 10; the stacked structure 20 is disposed on the side of the source line SL away from the substrate 10. The stacked structure 20 includes a conductive layer 22 and an insulating layer 21 stacked in sequence. The conductive layer 22 of the stacked structure 20 forms a bottom gate line BSG, multiple word lines WL, and a top gate line TSG arranged along a first direction Z. Multiple memory structures 30 extend through the stacked structure 20 along the first direction Z, with the bottoms of the memory structures 30 connected to the source line SL. Along the stacked structure 20, toward the center of the memory structures 30, the memory structures 30 include a gate dielectric layer 31, a charge trapping layer 32, a tunneling layer 33, a channel layer 34, and an isolation layer 35. The top conductive layer 22 of the stacked structure 20 is divided into multiple independently arranged top gate lines TSG.
[0062] The stacked structure 20 is provided with multiple first holes (not numbered in the figure). These first holes extend through each conductive layer 22 and each insulating layer 21 of the stacked structure 20 along a first direction Z, with the bottom surface of each first hole extending to the source line SL. A storage structure 30 is correspondingly disposed within the first hole. The gate dielectric layer 31, charge trapping layer 32, tunneling layer 33, and channel layer 34 of the storage structure 30 sequentially cover the walls of the first hole. An isolation layer 35 covers the inner sidewalls of the channel layer 34 and fills the unfilled area of the first hole. The intersection of the storage structure 30 and each conductive layer 22 of the stacked structure 20 forms a transistor, comprising the multiple storage transistors T and the upper gate transistor TM of the storage string Str. The source terminal of the storage string Str is connected to the source line SL to facilitate data read and write operations. The storage transistor T is located at the intersection of the storage structure 30 and the word line WL. The storage transistor T is connected to the word line WL, which is provided on the same layer (physical layer). The storage transistor T is used to store data, and programming and erasing operations can be performed on the connected storage transistor T by controlling the voltage of the word line WL. The upper gate transistor TM is located at the intersection of the storage structure 30 and the upper gate line TSG. The control gate of the upper gate transistor TM is connected to the upper gate line TSG, and the drain end of the upper gate transistor TM is connected to the bit line BL. The upper gate transistor TM is used to control the connection between the storage string Str constructed by the storage structure 30 and the bit line BL to realize data transmission.
[0063] The top conductive layer 22 of the stacked structure 20 is divided into a plurality of independently set upper selection lines TSG, thereby dividing the plurality of storage strings Str into a plurality of independent sub-storage blocks 100 through the plurality of upper selection lines TSG, so that each sub-storage block 100 can operate in parallel or independently with other sub-storage blocks 100, thereby improving the problem of a large number of transistors being idle when the three-dimensional storage array performs a read operation, thereby improving the utilization rate of the three-dimensional storage array and reducing redundancy waste.
[0064] According to an exemplary embodiment, this embodiment provides a method for manufacturing a three-dimensional memory array, such as Figure 5 As shown, and combined with reference Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, the method for manufacturing a three-dimensional storage array includes the following steps:
[0065] Step S101 : providing a substrate 10 , and forming a source line SL on one side of the substrate 10 .
[0066] Step S102 : alternately forming an insulating layer 21 and a sacrificial layer on a side of the source line SL away from the substrate 10 .
[0067] Step S103 : forming a first hole, wherein the first hole penetrates the alternating insulating layers 21 and the sacrificial layer along a first direction Z perpendicular to the substrate 10 , and the first hole exposes a portion of the top surface of the source line SL.
[0068] Step S104 : forming a storage structure 30 in the first hole, the storage structure 30 including a gate dielectric layer 31 , a charge trapping layer 32 , a tunneling layer 33 , a channel layer 34 and an isolation layer 35 sequentially covering the hole wall of the first hole.
[0069] Step S105: removing the sacrificial layer and forming a conductive layer 22 arranged in a first direction Z in the area where the sacrificial layer is removed; the multi-layer conductive layer 22 is used to form a lower gate line BSG, a plurality of word lines WL and an upper gate line TSG arranged in the first direction Z.
[0070] Step S106 : etching the top conductive layer 22 to form a plurality of upper gate lines TSG spaced apart from each other.
[0071] Step S107 : forming a plurality of bit lines BL, wherein the bit lines BL extend along the second direction Y and are arranged at intervals along the third direction X. The bit lines BL cover the top surfaces of the memory structures 30 arranged along the second direction Y.
[0072] In the method for manufacturing a three-dimensional memory array of this embodiment, the top conductive layer 22 is etched to form multiple independent upper gate lines TSG. The memory strings Str connected to each upper gate line TSG constitute a sub-memory block 100. The multiple upper gate lines TSG divide the multiple memory strings Str into multiple independent sub-memory blocks 100. Each sub-memory block 100 can operate in parallel with or independently of other sub-memory blocks 100. Storage tasks and / or computing tasks can be assigned to each sub-memory block 100 according to application requirements. This can improve the operational flexibility and utilization of the three-dimensional memory array, enhance the parallel processing capability of the three-dimensional memory array, and significantly improve the computing speed and efficiency of the three-dimensional memory array.
[0073] In step S101, substrate 10 may be a semiconductor substrate, such as a silicon substrate, a germanium (Ge) substrate, a silicon-germanium (SiGe) substrate, a silicon-on-insulator (SOI) substrate, or a germanium-on-insulator (GOI) substrate. The semiconductor substrate may be doped with ions. For example, the semiconductor substrate may be a P-type doped substrate or an N-type doped substrate. In this embodiment, the substrate is a silicon crystal substrate.
[0074] The source line SL may be formed by depositing the source line SL using any one of chemical vapor deposition (CVD), physical vapor deposition (PVD), atomic layer deposition (ALD), plasma enhanced chemical vapor deposition (PECVD), or sputtering processes, and the source line SL covers the top surface of the substrate 10 .
[0075] For example, the material of the source line SL may include silicon, germanium, silicon germanium, etc. The source line SL may be doped with ions, for example, P-type doping or N-type doping.
[0076] In step S102, an insulating layer 21 and a sacrificial layer (not shown) are alternately deposited on the source line SL using any of the deposition processes, including chemical vapor deposition, physical vapor deposition, atomic layer deposition, or sputtering. The sacrificial layer has a high etching ratio relative to the insulating layer 21. For example, the insulating layer 21 may be made of silicon oxide, and the sacrificial layer may be made of silicon nitride. In this embodiment, both the bottom and top layers of the deposited stack are insulating layers 21.
[0077] It can be understood that the number of stacked layers of the insulating layer 21 and the sacrificial layer can be flexibly adjusted according to manufacturing requirements, and this embodiment does not impose any limitation on this.
[0078] In step S103, a pattern layer is formed on the alternating insulating layer 21 and sacrificial layer. The stacked structure 20 is etched according to the pattern layer until a portion of the top surface of the source line SL is exposed. A plurality of first holes are formed through the insulating layer 21 and the sacrificial layer. The bottom surfaces of the first holes expose a portion of the top surface of the source line SL. The plurality of first holes are arranged in a spaced relationship along the second direction Y and the third direction X.
[0079] In step S104, refer to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4As shown, a gate dielectric layer 31, a charge trapping layer 32, and a tunneling layer 33 can be sequentially deposited by chemical vapor deposition or atomic layer deposition. The gate dielectric layer 31, the charge trapping layer 32, and the tunneling layer 33 on the bottom surface of the first hole are then etched away. A channel layer 34 can then be formed by plasma-enhanced chemical vapor deposition. The channel layer 34 covers the source line SL exposed at the bottom surface of the first hole and the inner sidewalls of the tunneling layer 33. Subsequently, an isolation layer 35 can be deposited by chemical vapor deposition or plasma-enhanced chemical vapor deposition. The isolation layer 35 covers the inner sidewalls of the channel layer 34 and fills the unfilled area of the first hole.
[0080] The gate dielectric layer 31 is made of at least one of silicon oxide, silicon nitride, or silicon oxynitride. The charge trapping layer 32 is made of silicon nitride and is used to trap charges. The tunneling layer 33 is made of silicon oxide or silicon oxynitride. The isolation layer 35 is made of silicon oxide. The channel layer 34 is made of a semiconductor material (such as polycrystalline silicon) or a metal oxide material (such as indium zinc oxide (IGZO)).
[0081] IGZO is an amorphous or microcrystalline material with a relatively uniform molecular structure, a low processing temperature, and no obvious grain boundaries. Using IGZO as the channel layer 34 can reduce processing temperature, simplify production processes, and reduce costs. IGZO has no obvious grain boundaries, which can increase the carrier mobility of the channel layer 34 and help improve the switching speed and reliability of the transistor.
[0082] For example, the material of the metal oxide may be indium gallium zinc oxide (IGZO). The material of the metal oxide may also be ITO, IWO, ZnO x 、InO x 、In2O3、InWO、SnO2、TiO x 、InSnO x 、Zn x O y N z Mg x Zn y O z 、In x Zn y O z 、In x Ga y Zn z O a 、Zr x In y Zn z O a , Hf x Iny Zn z O a 、Sn x In y Zn z O a 、Al x Sn y In z Zn a O d 、Si x In y Zn z O a 、Zn x Sn y O z 、Al x Zn y Sn z O a 、Ga x Zn y Sn z O a 、Zr x Zn y Sn z O a , InGaSiO, IAZO, IGO, IZO (indium-zinc-oxide), IZO x and other materials.
[0083] In step S105, the insulating layer 21 and the sacrificial layer are etched layer by layer to form a second hole (not shown) penetrating the insulating layer 21 and the sacrificial layer. The second hole is staggered with the first hole. An etchant is injected into the second hole to dissolve and remove the entire sacrificial layer. In this embodiment, the etchant can be hot phosphoric acid.
[0084] Then, a conductive layer 22 can be deposited using physical vapor deposition, plasma-enhanced chemical vapor deposition, or atomic layer deposition. The conductive layer 22 covers the exposed surface of the insulating layer 21 and the exposed surface of the gate dielectric layer 31 and fills the area where the sacrificial layer was removed. The alternating conductive layers 22 and insulating layers 21 form a stacked structure 20. Next, the conductive material in the second hole is etched away, and an insulating material is deposited into the second hole to form an insulating column. The material of the insulating column can include at least one of silicon oxide, silicon nitride, or silicon oxynitride.
[0085] For example, the material of the conductive layer 22 may include one or more selected from metal vanadium (V), metal niobium (Nb), metal ruthenium (Ru), metal tungsten (W), metal tantalum (Ta), tantalum nitride (TaN), metal titanium (Ti), titanium nitride (TiN), metal hafnium (Hf), metal iridium (Ir), metal manganese (Mn), metal zinc (Zn), metal platinum (Pt), metal palladium (Pd), metal copper (Cu), or alloys thereof.
[0086] In step S106, in some embodiments, the top conductive layer 22 is etched to form a plurality of upper gate lines TSG spaced apart from each other. The following embodiments may be employed: etching the top conductive layer 22 to form first trenches 41 extending along the second direction Y, the first trenches 41 dividing the top conductive layer 22 into a plurality of upper gate lines TSG spaced apart from each other along the second direction Y; and / or etching the top conductive layer 22 to form second trenches 42 extending along the third direction X, the second trenches 42 dividing the top conductive layer 22 into a plurality of upper gate lines TSG spaced apart from each other along the third direction X. The second trenches 42 may be straight trenches extending along the third direction X, or may be zigzag trenches extending along the third direction X, which are not limited in this embodiment.
[0087] For example, a mask layer is defined on the top surface of the stacked structure 20, and the top insulating layer 21 and the top conductive layer 22 of the stacked structure 20 are etched according to the mask layer to form a first trench 41 and / or a second trench 42. The top conductive layer 22 is divided into a plurality of independently provided upper gate lines TSG. Each upper gate line TSG is connected to at least one storage structure 30 to form a sub-memory block 100 controlled by the upper gate line TSG. This divides the three-dimensional storage array into multiple independently operable sub-memory blocks 100. Each sub-memory block 100 can operate in parallel or independently with other sub-memory blocks 100. Storage tasks and / or computing tasks can be assigned to each sub-memory block 100 according to application requirements, thereby improving the operational flexibility and utilization of the three-dimensional storage array.
[0088] The arrangement and number of the first trench 41 and / or the second trench 42 can be set according to application requirements, thereby controlling the number and arrangement of the upper strobe lines TSG, so as to flexibly control the number and arrangement of the sub-memory blocks 100 of the three-dimensional memory array. For example, the upper strobe lines TSG can be configured into a linear, square, circular, etc. For another example, each upper strobe line TSG can be connected to one or more memory strings Str to adjust the number of sub-memory blocks 100. For another example, the number of memory strings Str connected to each upper strobe line TSG can be adjusted so that the number of memory strings Str in each sub-memory block 100 is the same or different.
[0089] In this embodiment, referring to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, the top conductive layer 22 is etched to form a first trench 41 extending along the second direction Y and a second trench 42 extending along the third direction X. The upper gate lines TSG extend along the third direction X and are arranged alternately along the second direction Y and the third direction X. In this way, the three-dimensional memory array is divided into a plurality of sub-memory blocks 100 arranged along the second direction Y and the third direction X. Each sub-memory block 100 can be independently controlled by the upper gate lines TSG and the bit lines BL, so that each sub-memory block 100 can operate in parallel with or independently of other sub-memory blocks 100.
[0090] In this embodiment, after etching the top conductive layer 22 to form a plurality of independently disposed upper gate lines TSG, a chemical vapor deposition insulating material may be used to fill the first trench 41 and the second trench 42 to electrically isolate the upper gate lines TSG.
[0091] In step S107, the isolation layer 35 is etched back so that the top surface of the isolation layer 35 is lower than the top surface of the first hole, exposing a portion of the inner sidewalls of the channel layer 34. A bitline conductive layer (not shown) can be deposited using physical vapor deposition, plasma-enhanced chemical vapor deposition, or atomic layer deposition processes. The bitline conductive layer covers the inner sidewalls of the top of the channel layer 34 and fills the unfilled area of the first hole and the top surface of the stacked structure 20. The bitline conductive layer is then patterned and etched to form bitlines BL. The bitlines BL extend along the second direction Y and are spaced apart along the third direction X. For example, the material of the bitlines BL may include tungsten, titanium, or the like.
[0092] In some embodiments, after forming the bit lines BL, the process further includes leading out the upper gate lines TSG. Depending on the number and arrangement of the upper gate lines TSG, metal lead lines can be formed on the sides of the three-dimensional memory array to lead out the upper gate lines TSG, or a vertical interconnect structure can be formed on the top of the three-dimensional memory array to lead out the upper gate lines TSG.
[0093] According to an exemplary embodiment, this embodiment provides a storage and calculation method based on a three-dimensional storage array, which is implemented based on the three-dimensional storage array of the above embodiment. Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 As shown, the three-dimensional memory array includes at least two sub-memory blocks 100 capable of operating in parallel, and the at least two sub-memory blocks 100 capable of operating in parallel are connected to different upper gate lines TSG and different bit lines BL. Figure 6 As shown, the storage calculation method includes steps S201 and S202.
[0094] Step S201: inputting a first vector and a second vector in parallel to an upper gate line and a bit line, and writing the outer product of the first vector and the second vector into a first storage transistor of a sub-memory block.
[0095] Step S202: inputting the third vector and the fourth vector in parallel to another upper selection line and another bit line, and writing the outer product of the third vector and the fourth vector into the second storage transistor of another sub-memory block.
[0096] In this embodiment, data is written into the storage transistor T in the form of the outer product of two vectors.
[0097] When writing data to the first storage transistor, a programming voltage is applied to the word line WL connected to the first storage transistor, a conduction voltage is applied to the other word lines WL, and the source line SL is grounded. The first vector and the second vector are input in parallel to the upper gate line TSG and the bit line BL connected to the sub-memory block 100 where the first storage transistor is located, so that the outer product of the first vector and the second vector is written into the first storage transistor.
[0098] Similarly, when writing data to the second storage transistor, a programming voltage is applied to the word line WL connected to the second storage transistor, and a turn-on voltage is applied to other word lines WL, and the third vector and the fourth vector are input in parallel to the upper selection line TSG and the bit line BL connected to the sub-storage block 100 where the second storage transistor is located, so as to write the outer product of the third vector and the fourth vector into the second storage transistor.
[0099] In this embodiment, data can be written to the first storage transistor and the second storage transistor in parallel at the same time, which significantly shortens the data writing time and can meet the storage requirements of the development of large language models.
[0100] In some embodiments, when reading data, steps S203 to S205 are executed.
[0101] Step S203: pre-charge each bit line BL to a preset voltage.
[0102] In this embodiment, before the read operation begins, the bit lines BL are connected to a voltage source through a precharge circuit for a certain period of time until all bit lines BL are precharged to a preset voltage Vpre, so that all bit lines BL are at the same initial potential, providing a unified reference point for the subsequent discharge process.
[0103] Step S204: applying a turn-on voltage to the upper gate line TSG connected to the first storage transistor, applying a zero bias voltage to other upper gate lines TSG, and reading data stored in the first storage transistor according to the discharge of the bit line BL connected to the first storage transistor.
[0104] In this embodiment, a turn-on voltage is applied to the upper gate line TSG connected to the first storage transistor to activate the storage transistor T in the sub-memory block 100 connected to the upper gate line TSG. A zero bias voltage is applied to the other upper gate lines TSG to keep the storage transistors T in the other sub-memory blocks 100 in an inactive state. The bit line BL connected to the first storage transistor begins to discharge, and the final discharge voltage of the bit line BL connected to the first storage transistor is detected to obtain the data stored in the first storage transistor (the outer product of the first vector and the second vector).
[0105] Step S205: applying a turn-on voltage to the upper gate line TSG connected to the second storage transistor, applying a zero bias voltage to other upper gate lines TSG, and reading data stored in the second storage transistor according to the discharge of the bit line BL connected to the second storage transistor.
[0106] In this embodiment, a turn-on voltage is applied to the upper gate line TSG connected to the second storage transistor, and a zero bias voltage is applied to the other upper gate lines TSG. The bit line BL connected to the second storage transistor begins to discharge. The discharged voltage of the bit line BL connected to the second storage transistor is detected to obtain the data stored in the second storage transistor (the outer product of the third and fourth vectors).
[0107] In this embodiment, data in the first storage transistor and the second storage transistor can be read in parallel, further improving the reading efficiency.
[0108] In this embodiment, the three-dimensional storage array sub-storage block 100 is used to read the data in the first storage transistor and the second storage transistor in parallel, which improves the reading efficiency and can increase the throughput of the three-dimensional storage array for model weight processing, so that the storage capacity of the three-dimensional storage array can meet the speed of model weight update, accelerate the speed of model weight update, reduce update delay, and improve the overall performance and cost-effectiveness of model weight storage, update and reading.
[0109] The technical features of the above-mentioned embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features of the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0110] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A three-dimensional storage array, characterized in that: include: A plurality of memory strings, each comprising a plurality of memory transistors and an upper gate transistor sequentially connected in series along a first direction; a plurality of word lines, the word lines being arranged at intervals along the first direction, each of the word lines being connected to a control gate of the storage transistor disposed in the same layer; a plurality of bit lines, the bit lines extending along a second direction and arranged at intervals along a third direction, the bit lines being connected to the upper gate transistors arranged along the second direction; the second direction intersecting the third direction, the first direction being perpendicular to a plane containing the second direction and the third direction; a plurality of upper gate lines, the upper gate lines being arranged at intervals, each of the upper gate lines being connected to a control gate of at least one of the upper gate transistors, the memory strings connected to each of the upper gate lines constituting a sub-memory block, the plurality of upper gate lines dividing the plurality of memory strings into a plurality of independent sub-memory blocks, the sub-memory blocks being arranged along the second direction and the third direction, and each of the sub-memory blocks operating in parallel or independently of the other sub-memory blocks; The plurality of sub-memory blocks include at least two sub-memory blocks capable of operating in parallel, and the at least two sub-memory blocks capable of operating in parallel are connected to different upper gate lines and different bit lines.
2. The three-dimensional storage array according to claim 1, wherein: The upper gate lines extend along the third direction and are arranged at intervals along the second direction and the third direction.
3. The three-dimensional storage array according to claim 1, wherein: The upper gate line is led out from the side of the three-dimensional memory array and connected to the peripheral circuit; or, the upper gate line is led out from the top of the three-dimensional memory array and connected to the peripheral circuit.
4. The three-dimensional storage array according to any one of claims 1 to 3, characterized in that: The three-dimensional storage array comprises: substrate; a source line, disposed on the substrate; a stacked structure, provided on a side of the source line away from the substrate, the stacked structure comprising a conductive layer and an insulating layer stacked in sequence; a plurality of storage structures passing through the stacked structure along the first direction, the bottom of each storage structure being connected to the source line; and a direction pointing from the stacked structure to the center of the storage structure, the storage structure comprising a gate dielectric layer, a charge trapping layer, a tunneling layer, a channel layer, and an isolation layer arranged in sequence; The top conductive layer of the stacked structure is divided into a plurality of independently arranged upper gate lines.
5. A method for manufacturing a three-dimensional storage array, for manufacturing the three-dimensional storage array according to any one of claims 1 to 4, characterized in that: include: providing a substrate, and forming a source line on one side of the substrate; Alternatingly forming an insulating layer and a sacrificial layer on a side of the source line away from the substrate; forming a first hole, wherein the first hole penetrates the alternating insulating layers and the sacrificial layers along a first direction perpendicular to the substrate, and the first hole exposes a portion of a top surface of the source line; forming a storage structure in the first hole, the storage structure comprising a gate dielectric layer, a charge trapping layer, a tunneling layer, a channel layer, and an isolation layer sequentially covering a hole wall of the first hole; The sacrificial layer is removed, and conductive layers arranged at intervals along the first direction are formed in the area where the sacrificial layer is removed; the multiple layers of the conductive layers are used to form a lower gate line, a plurality of word lines, and an upper gate line arranged along the first direction; Etching the top conductive layer to form a plurality of upper gate lines spaced apart from each other, each upper gate line corresponding to a sub-memory block; A plurality of bit lines are formed, wherein the bit lines extend along the second direction and are arranged at intervals along the third direction, and the bit lines cover top surfaces of the memory structures arranged along the second direction.
6. The method for manufacturing a three-dimensional storage array according to claim 5, wherein: Etching the top conductive layer to form a plurality of upper gate lines spaced apart from each other, comprising: Etching the top conductive layer to form a first trench extending along the second direction, wherein the first trench divides the top conductive layer into a plurality of upper gate lines arranged at intervals along the second direction; Alternatively, the top conductive layer is etched to form a second trench extending along the third direction, wherein the second trench divides the top conductive layer into a plurality of upper gate lines arranged at intervals along the third direction.
7. The method for manufacturing a three-dimensional storage array according to claim 6, wherein: The top conductive layer is etched to form the first trench extending along the second direction and the second trench extending along the third direction. The upper gate lines extend along the third direction and are arranged at intervals along the second and third directions.
8. A storage and calculation method based on a three-dimensional storage array, characterized in that: Applied to a three-dimensional memory array according to any one of claims 1 to 4, the three-dimensional memory array comprising at least two sub-memory blocks capable of operating in parallel, the at least two sub-memory blocks capable of operating in parallel being connected to different upper gate lines and different bit lines, the storage and calculation method comprising: Inputting a first vector and a second vector in parallel to an upper gate line and a bit line, and writing an outer product of the first vector and the second vector into a first storage transistor of a sub-memory block; The third vector and the fourth vector are input in parallel to another upper gate line and another bit line, and the outer product of the third vector and the fourth vector is written into the second storage transistor of another sub-memory block.
9. The storage and computing method based on a three-dimensional storage array according to claim 8, characterized in that: When reading: precharging each of the bit lines to a preset voltage; Applying a turn-on voltage to the upper gate line connected to the first storage transistor, applying a zero bias voltage to the other upper gate lines, and reading data stored in the first storage transistor according to discharge of the bit line connected to the first storage transistor; An on-voltage is applied to the upper gate line connected to the second storage transistor, and a zero bias voltage is applied to the other upper gate lines. Data stored in the second storage transistor is read out according to discharge of the bit line connected to the second storage transistor.
Citation Information
Patent Citations
Memory device and manufacturing method thereof
CN119855156A