Method, apparatus, and storage system for performing in-memory computing
By using a memory transistor array in flash memory as the carrier of the neural network weight matrix, and by adjusting the threshold voltage of the transistor dynamically configure the weight matrix, the problem of difficulty in mapping the memory to the neural network and dynamically adjusting the weight matrix in the prior art is solved, and efficient and energy-efficient neural network computing is achieved.
Patent Information
- Application Number
- CN202410001792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-01-02
AI Technical Summary
The prior art is difficult to effectively map memory to neural networks and dynamically adjust the weight matrix of neural networks, especially how to deal with the weight offset during backpropagation.
In flash memory, the memory transistor array is used as the weight matrix carrier of the neural network, and the weight matrix is dynamically configured by adjusting the threshold voltage of the transistor, and the weight size at various positions in the memory transistor array is reset during the backpropagation phase.
The dynamic configuration and update of the neural network weight matrix is realized, the weight offset is avoided, the computing efficiency and energy efficiency ratio of the neural network is improved, and the compilation wall, storage wall, and power consumption wall problems under the traditional von Neumann architecture are solved.
Smart Images

Figure CN119207518B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of data storage. Specifically, it relates to a method for performing in-memory computing in the field of data storage, an in-memory computing device capable of performing in-memory computing, and a corresponding storage system. Background Art
[0002] In previous electronic devices, computing and storage were separated. The computing part, such as a processor, and the storage part, such as a hard disk or a memory like a memory, each focused on their respective tasks, and computing and storage basically did not interfere with each other. There are also manufacturers who integrate the processor and the memory closer physically in order to improve the computing speed and computing power. For example, in the layout design of the circuit board, they are close to each other, with short circuits, and are directly stacked on top of each other during the packaging stage.
[0003] The improvement of computing power in the industry has been mainly advanced step by step according to the rhythm of Moore's Law, but the mainstream computing paradigm has always followed the von Neumann architecture. As a new type of computing power, in-memory computing in the industry mainly focuses on narrow in-memory computing, that is, using storage media for computing, but the computing medium methods have not been fully unified. As a new type of computing power, in-memory computing is expected to solve the problems of compilation wall, memory wall, and power consumption wall under the traditional von Neumann architecture. In the industry's consensus, in-memory computing has basically been determined as one of the key technologies to solve the computing power problem.
[0004] In-memory computing organically integrates storage and computing. Based on its huge potential for improving energy efficiency ratio, it is expected to become an advanced computing architecture solution in the computing power era. However, there are doubts: how to map the memory to the neural network and how to configure a dynamically adjustable weight matrix for the neural network to adapt to the actual applications of the neural network, such as backpropagation, etc. This is a key problem that must be addressed when using in-memory computing. In particular: how to address the weight offset and other problems that need to be solved urgently when updating the neural network weights inside the memory. For the memory, this belongs to a hidden problem. Summary of the Invention
[0005] The present application relates to a method for performing in-memory computing, which is characterized by including:
[0006] Using a storage transistor array in a flash memory as a carrier of the weight matrix of a neural network, setting the input data of the neural network to a series of bit lines of the storage transistor array, and setting the output data of the neural network to be retrieved from a series of source lines of the storage transistor array; wherein, the magnitude of the weight of any transistor in the storage transistor array is adjusted by using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network.
[0007] The above method, wherein the storage transistor array for the neural network is located on the same physical block of the flash memory; and the storage spaces where the input data and output data of the neural network are located respectively belong to different physical blocks of the flash memory from the storage space where the storage transistor array is located.
[0008] The above method, wherein the original input data of the neural network is converted into an input quantity in the form of a voltage value, and is transmitted to a series of bit lines of the storage transistor array based on the input quantities sorted into a vector.
[0009] The above method, wherein the output data of the neural network is characterized by the output current of the storage transistor array, and the output currents of a series of source lines of the storage transistor array are regarded as the output quantities sorted into a vector of the storage transistor array.
[0010] The above method, wherein the way to adjust the weight of any transistor includes changing the charge injected into its floating gate, and adjusting the weight magnitude at the position where any transistor is located in the storage transistor array (i.e., at an element position corresponding to the any transistor in the weight matrix) by changing the charge injected into the floating gate of the transistor.
[0011] The above method, in the backpropagation stage of the neural network, dynamically configures the weight matrix of the neural network by resetting the weight magnitudes at each position in the storage transistor array.
[0012] The above method, at each bit line of the storage transistor array, a digital-to-analog converter is configured, which is used to convert the input data of the neural network into a voltage value matching the input signal pattern of the bit line.
[0013] The above method, at each source line of the storage transistor array, an analog-to-digital converter is configured, which is used to convert the output signal in the format of the output current at the source line into output data adapted to the digital output quantity of the neural network.
[0014] This application relates to a device for performing in-memory computing, which is characterized by including:
[0015] A flash memory, with its storage transistor array as the carrier of the weight matrix of the neural network, setting the input data of the neural network to be transmitted to a series of bit lines of the storage transistor array, and setting the output data of the neural network to be retrieved from a series of source lines of the storage transistor array;
[0016] An input unit, which converts the original input data of the neural network into an input quantity in the form of a voltage value, and transmits it to a series of bit lines of the storage transistor array based on the input quantities sorted into a vector;
[0017] An output unit, which is configured to convert an output signal in the format of an output current at a source line into output data adapted to the digital output quantity of a neural network. The output currents of a series of source lines of a memory transistor array are regarded as the output quantities of the memory transistor array sorted into a vector.
[0018] The above-mentioned device for performing in-memory computing, wherein the memory transistor array for the neural network is located on the same physical block of a flash memory; and the storage spaces where the input data and output data of the neural network are located respectively belong to different physical blocks of the flash memory from the storage space where the memory transistor array is located.
[0019] The above-mentioned device for performing in-memory computing, wherein the original input data of the neural network is converted into an input quantity in the form of a voltage value and is transmitted to a series of bit lines of the memory transistor array based on the input quantity sorted into a vector.
[0020] The above-mentioned device for performing in-memory computing, wherein the output data of the neural network is characterized by the output current of the memory transistor array. The output currents of a series of source lines of the memory transistor arrays are regarded as the output quantities of the memory transistor array sorted into a vector.
[0021] The above-mentioned device for performing in-memory computing, wherein the way to adjust the weight of any transistor includes changing the charge injected into its floating gate, and the weight size at the position where the any transistor is located in the memory transistor array (i.e., at an element position corresponding to the any transistor in the weight matrix) is adjusted by changing the charge injected into the floating gate of the transistor.
[0022] The above-mentioned device for performing in-memory computing, in the backpropagation stage of the neural network, the weight matrix of the neural network can be dynamically configured by resetting the weight sizes at various positions in the memory transistor array.
[0023] The above-mentioned device for performing in-memory computing, the input unit at least includes a digital-to-analog converter. A digital-to-analog converter is configured at each bit line of the memory transistor array, and it is used to convert the input data of the neural network into a voltage value matching the input signal pattern of the bit line.
[0024] The above-mentioned device for performing in-memory computing, the output unit at least includes an analog-to-digital converter. An analog-to-digital converter is configured at each source line of the memory transistor array, and it is used to convert an output signal in the format of an output current at the source line into output data adapted to the digital output quantity of the neural network.
[0025] This application relates to a memory system for performing in-memory computing, which is characterized by including:
[0026] A flash memory-based storage chip utilizes a storage transistor array in the flash memory as a carrier for the weight matrix of a neural network. The magnitude of the weight of any transistor is adjusted using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network. The input data of the neural network is set to be delivered to a series of bit lines of the storage transistor array, and the output data of the neural network is also set to be retrieved from a series of source lines of the storage transistor array.
[0027] A digital-to-analog converter built into the storage chip, where a digital-to-analog converter is configured at each bit line of the storage transistor array to convert the input data of the neural network into a voltage value that matches the input signal pattern of the bit line.
[0028] An analog-to-digital converter built into the storage chip, where an analog-to-digital converter is configured at each source line of the storage transistor array to convert the output signal in the output current format at the source line into output data that is adapted to the digital output quantity of the neural network.
[0029] The above storage system, where the storage transistor array for the neural network is located on the same physical block of the flash memory; and the storage spaces where the input data and output data of the neural network are located respectively belong to different physical blocks of the flash memory from the storage space where the storage transistor array is located.
[0030] In the above storage system, the original input data of the neural network is converted into an input quantity in the form of a voltage value and is delivered to a series of bit lines of the storage transistor array based on the input quantity sorted into a vector.
[0031] In the above storage system, the output data of the neural network is characterized by the output current of the storage transistor array, and the output current of a series of source lines of the storage transistor array is regarded as the output quantity of the storage transistor array sorted into a vector.
[0032] In the above storage system, the method of adjusting the weight of any transistor includes changing the charge injected into its floating gate, and by changing the charge injected into the floating gate of the transistor, the weight magnitude at the position where any transistor is located in the storage transistor array (i.e., at the position of an element corresponding to any transistor in the weight matrix) is adjusted.
[0033] In the above storage system, during the backpropagation stage of the neural network, the weight matrix of the neural network is dynamically configured by resetting the weight magnitudes at various positions in the storage transistor array.
[0034] In the above storage system, a digital-to-analog converter is configured at each bit line of the storage transistor array, which is used to convert the input data of the neural network into a voltage value that matches the input signal pattern of the bit line.
[0035] In the above storage system, an analog-to-digital converter is arranged at each source line of the storage transistor array, and is configured to convert an output signal in the format of an output current at the source line into output data adapted to the digital output quantity of the neural network.
[0036] This application relates to a chip for performing in-memory computing, which is characterized by including:
[0037] A storage transistor array, which serves as a carrier of the weight matrix of the neural network. The magnitude of the weight of any transistor is adjusted by using its threshold voltage, so as to dynamically configure the weight matrix of the neural network; setting the input data of the neural network to be transmitted to a series of bit lines of the storage transistor array, and setting the output data of the neural network to be retrieved from a series of source lines of the storage transistor array;
[0038] A digital-to-analog converter integrated in the chip, wherein a digital-to-analog converter is arranged at each bit line of the storage transistor array to convert the input data of the neural network into a voltage value matching the input signal pattern of the bit line;
[0039] An analog-to-digital converter integrated in the chip, wherein an analog-to-digital converter is arranged at each source line of the storage transistor array to convert an output signal in the format of an output current at the source line into output data adapted to the digital output quantity of the neural network.
[0040] Based on the foregoing, one of the greatest advantages of this application is that the weight matrix of the neural network is directly mapped to the storage transistor array and the weight matrix of the neural network can be dynamically configured by adjusting the transistor threshold voltage, so as to adapt to the update of the weight matrix of the neural network (weight matrix), such as matching backpropagation. And thereby, it can also cope with relatively difficult problems such as weight offset faced when the neural network weights are updated inside the memory, and prevent one or more elements in the weight matrix mapped by the memory from having an unexpected hidden offset.
[0041] Based on the above, the second biggest advantage of this application is that it adopts a non-von Neumann architecture, which can solve the problems of compilation wall, storage wall, and power consumption wall under the traditional von Neumann architecture. One convenient aspect of compilation is to realize the programming of core weights by adjusting the threshold voltage of transistors. This solution only needs to configure the voltage or current of the three terminals of the transistor at the transistor to realize the compilation of core weights. Because it is a compilation compatible with information such as voltage and current on a traditional memory, there is no obstacle in designing a part of the compilation behavior. Moreover, unlike the traditional solution, which requires data exchange between the memory and the processor and the subsequent storage wall, the data storage and data calculation of this application, especially matrix operations, occur inside the memory, solving the problem of storage wall. Since the specific implementation of the neural network is based on the premise that there is no need to implement a large amount of data transfer between the memory and the processor, and since most of the energy consumption of the power consumption wall is generated during the data transfer process, for example, the data transfer power consumption is more than a thousand times the computing power consumption, it also solves the problem of high energy consumption in the application scenarios of neural networks and their artificial intelligence. In addition, AI computing requires 1PB / s, while data transfer speeds such as DRAM 40GB-1TB / s are far from sufficient to meet the needs of AI computing. If internal memory calculations as in this application are used, the concern that the data transfer speed cannot keep up with the AI computing speed can be better resolved.
[0042] Based on the foregoing, the third biggest advantage of this application is to promote a significant increase in computing power. By using storage-computing integrated technology, the weight part of a large number of multiplication and addition calculations in AI calculations can be stored in the storage unit, and modifications can be made to the existing storage circuit part of the storage unit, so that data input and calculation processing can be performed while reading, and neural networks such as convolution operations are completed in the storage array. Since convolution operations with a large number of multiplications and additions are the core components of deep learning, in-memory computing and in-memory logic are very suitable for deep neural network applications of artificial intelligence and big data technologies based on AI. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to make the above purposes, features and advantages more understandable, the following text will explain the specific implementation methods in detail in conjunction with the drawings. After reading the following description and referring to the following drawings, the features and advantages of the present application will become apparent.
[0044] Figure 1 A part of the transistors inside the flash memory are selected as the weight carriers of the neural network.
[0045] Figure 2 The input data to the neural network comes from a series of bit lines in the storage transistor array.
[0046] Figure 3For example, pixel values or grayscale values are input into the storage transistor array in the form of pictures or videos.
[0047] Figure 4 The weight matrix of the neural network is mapped to the storage transistor array of the flash memory.
[0048] Figure 5 The weight value of the transistor in the transistor array is adjusted by the change of the charge injected into the floating gate.
[0049] Figure 6 It is to address concerns such as weight offset faced during the update of the neural network weights inside the memory. Detailed implementation manners
[0050] The technical solutions disclosed in this application will be clearly and completely described below in conjunction with specific embodiments. However, the described embodiments are only examples used for narrative illustration in this application and not all embodiments. Based on these embodiments, those skilled in the art should recognize that any solution obtained without creative efforts falls within the protection scope of this application.
[0051] Refer to Figure 1 , this is a circuit of a common NOR flash memory that already exists. Its working principle is that between the source S and the drain D of the storage transistor, that is, on the semiconductor where current conducts unidirectionally, a floating gate (Floating Gate) that can store charges such as electrons can be formed. The floating gate is wrapped with a silicon oxide film insulator. Above the floating gate is a select / control gate (Control Gate) that controls the conduction current between the source and the drain. The charges such as electrons in the floating gate will not disappear due to power-off, so the flash memory is a non-volatile memory. For write operations, a positive voltage is applied to the control electrode to allow charges such as electrons to enter the floating gate through the insulating layer. The erasure operation is exactly the opposite. For example, a positive voltage is applied to the substrate to extract electrons from the floating gate to achieve erasure. The data operations of the flash memory belong to the prior art and will not be elaborated separately herein. The read, write, or erase operations of the flash memory in the prior art can be applied to the flash memory of this application.
[0052] Refer to Figure 1 , describe the flash memory in the form of an array: the gates of each row of transistors are connected to a common word line, and the drains of each column of transistors are connected to a common bit line, and the sources of each column of transistors are connected to a common source line.
[0053] Refer to Figure 1 , regarding the row word line example, the gates of the transistors in the first row are respectively connected to the word line G1, the gates of the transistors in the second row are respectively connected to the word line G2, and the gates of the transistors in the Mth row are respectively connected to the word line GM. More omitted rows are not shown repeatedly in the figure. A total of M rows of transistors are regarded as the row elements of the matrix M×N.
[0054] See Figure 1 , regarding the example of row source lines, the source electrodes of the transistors in the first row are respectively connected to the source line S1, the source electrodes of the transistors in the second row are respectively connected to the source line S2, and the source electrodes of the transistors in the Mth row are respectively connected to the source line SM. More omitted rows are not repeatedly shown in the figure.
[0055] See Figure 1 , regarding the example of column bit lines, the drain electrodes of the transistors in the first column are respectively connected to the bit line B1, the drain electrodes of the transistors in the second column are respectively connected to the bit line B2, and the drain electrodes of the transistors in the Nth column are respectively connected to the bit line BN. More omitted columns are not repeatedly shown in the figure. The total N columns of transistors are regarded as the column elements of the matrix M×N.
[0056] See Figure 1 , the original function of the illustrated flash memory is to store storage symbols 1 or 0. In other words, in a traditional computer, server, or other similar electronic device, the flash memory has no special use. However, in this application, the flash memory is no longer regarded as just a storage space for storing symbols 1 or 0, but it is proposed to change the circuit architecture of the flash memory into a computing unit similar to that which can participate in artificial intelligence, and the underlying physical structure of the flash memory remains almost unchanged.
[0057] See Figure 2 , using a computing unit similar to that based on a flash memory to replace a traditional computer is because there are several major drawbacks when a computer executes the running stage of a neural network: reading and writing data from the memory outside the processor, the data transfer time is hundreds or thousands of times the operation time, and the non-functional power consumption of the whole process is about between 60% - 90%, the energy efficiency is very low, and the accompanying memory wall problem has become a major obstacle to data computing applications. In particular, the biggest challenge in accelerating deep learning in neural networks is the frequent movement of data between the computing unit and the storage unit. Although, for example, multi-core CPUs or, for example, many-core GPUs parallel acceleration technologies are used to improve computing power, in the post-Moore era, the memory bandwidth restricts the effective bandwidth of the computing system and makes the growth of the overall computing power difficult. Running the weight calculation of a neural network or running the inference calculation of a neural network on a traditional computer, the time-consuming and power-consuming are one of the most prominent manifestations of this drawback.
[0058] See Figure 2, the introduction of a large number of parameters in neural networks has encountered bottlenecks in its development. Each output neuron in the fully connected layer is connected to all input neurons. According to the assumption in the figure: if the number of input neurons is N and the number of output neurons is M, then there are at least approximately N×M weights or M biases that need to be trained and calculated. For each output neuron in the convolutional layer, there are usually also input neurons within the coverage of the convolutional kernel connected to it, and all outputs use the same convolutional kernel. Although local connection and weight sharing reduce the parameters of the convolutional layer, there is still a large amount of computation that needs to be performed under the conditions of large-size image input and deep neural networks. The reading and calculation of a large amount of data face the memory bottleneck and computing bottleneck of the network, and the power consumption and latency of the neural network increase sharply, restricting the application of the neural network.
[0059] See Figure 2 , to solve the aforementioned bottlenecks, the illustrated non-Von Neumann architecture can provide greater computing power and higher energy efficiency in specific fields and can exceed the computing power of existing ASIC chips. In-memory computing and in-memory logic, that is, the computing-in-memory technology directly uses the memory for data processing or calculation, thereby integrating data storage and data calculation in the same chip or the same area of the same memory, which can completely eliminate the bottleneck of the Von Neumann computing architecture, especially being more suitable for application scenarios with large amounts of data and large-scale parallelism such as deep learning neural networks.
[0060] See Figure 2 , in the design, the bit lines B1 - BN are used as inputs and the source lines S1 - SM are used as outputs. Figure 1 The flash memory is regarded as the physical carrier that can implement the neural network NET.
[0061] See Figure 2 , to illustrate the specific implementation process, first design a matrix Ma_wei as follows.
[0062]
[0063] See Figure 2 , the inputs of the bit lines B1 - BN are represented by the vector Ma_in = [B1, B2... BN]. For the convenience of mathematical expression, the input quantity of bit line B1 needs to be mapped to X1, the input quantity of bit line B2 needs to be mapped to X2, the input quantity of bit line B3 needs to be mapped to X3,..., and the input quantity of bit line BN needs to be mapped to X N .
[0064] See Figure 2 , the outputs of the source lines S1 - SM are represented by the vector Ma_out = [S1, S2... SM]. For the convenience of mathematical expression, the output quantity of source line S1 needs to be mapped to Y1, the output quantity of source line S2 needs to be mapped to Y2, the output quantity of source line S3 needs to be mapped to Y3,..., and the output quantity of source line SM needs to be mapped to YM 。
[0065] See Figure 2 , the purpose of the flash memory is to perform the arithmetic functions of the neural network NET. The specific arithmetic process is expressed by the following formula. Substantially, this process is actually a detailed example of simulating the flash memory as a computing unit or a computing-like unit. Taking Y1 and Y2 and Y M etc. as representatives to illustrate this spirit.
[0066] Y1 = X1 * W 11 + X2 * W 12 + X3 * W 13 +…X N * W 1N
[0067] Y2 = X1 * W 21 + X2 * W 22 + X3 * W 23 +…X N * W 2N
[0068] Y M = X1 * W M1 + X2 * W M2 + X3 * W M3 +…X N * W MN
[0069] See Figure 2 , a more intuitive expression is to use vector matrix multiplication to show the aforementioned complex operations. The vector matrix shows the local data processing of the flash memory executing the neural network NET as follows.
[0070]
[0071] See Figure 2 , if X1 to X N are defined as the input layer of the neural network NET, then Y1 to Y N correspondingly are the output layer of the neural network NET. So far, it can be considered that Figure 1 the flash memory has completed an entire vector operation from the input layer to the output layer, and this operation is very different from that of a traditional von Neumann computer, because its operation only needs to measure the respective signal amounts from the source lines S1 - SM, and individual multiplications and additions do not need to be implemented one by one. The final operation result can be fully reflected from the signal amounts of the source lines S1 - SM at one time. Therefore, it is called computing-like. The operation of a traditional von Neumann computer must implement all multiplications and additions one by one, which is the biggest difference between the two.
[0072] See Figure 2, the flash memory-based neural network NET is applicable to the field of image processing. Suppose data such as pictures or videos (artificial intelligence, autonomous driving, face recognition, object detection, etc.) with pixel values or grayscale values are input into the flash memory-based neural network NET. Hereinafter, the image Gr_in is taken as an example.
[0073] See Figure 3 , the pixel points in the image Gr_in are divided into multiple rows and columns. Define the pixel corresponding to the address of the 0th pixel as g1, the pixel corresponding to the address of the 1st pixel as g2, the pixel corresponding to the address of the 2nd pixel as g3, and so on. The pixel corresponding to the last pixel address is gn. According to the arrangement rule of the pixel matrix and the address regulation, the grayscale value of any pixel point can be obtained through the captured image. It is allowed that the grayscale values or pixel values of n pixel points (belonging to the image Gr_in) are coupled to X1 to X of the flash memory N etc. Of course, the grayscale values or pixel values can also be normalized or similar image data processing can be implemented.
[0074] See Figure 3 , an issue worthy of attention is that the input quantity format of the bit lines B1 - BN should conform to the inherent characteristics of the flash memory for the flash memory. That is, the input values of the bit lines B1 - BN and the pixel values g1 - gn are not directly equivalent quantities. Therefore, in an optional embodiment, the pixel values g1 - gn are subjected to digital-to-analog conversion processing by a DAC device, such as a digital-to-analog converter packaged together with the flash memory. The analog quantities corresponding to the pixel values g1 - gn are then input to the bit lines B1 - BN. The DAC (Digital to Analog Converter) device can also be directly integrated onto the same semiconductor wafer with the flash memory during the wafer manufacturing stage. Thus, the data contained in the image Gr_in can be smoothly coupled with the flash memory. Another important factor is that: the final storage location or space of the image Gr_in is usually various memories. Then, the prominent advantage of the intrinsic characteristic that the storage location of the image Gr_in and the location for performing matrix operations are the same flash memory can be manifested. The processor CPU or GPU will no longer have to read data from the memory to itself to perform matrix operations. The storage operation and the calculation operation are all implemented in the memory, and the concerns such as the read / write data and data transfer time cost between the processor and the memory mentioned in the previous text will no longer exist.
[0075] See Figure 3 , an important problem has not been solved yet: In the previous text, the matrix Ma_wei participating in network calculation was specially designed to explain the specific implementation process of the neural network. However, how to dualize the M*N vector matrix Ma_wei into the transistors in the flash memory such asFigure 1 There are still urgent doubts about the M×N elements shown.
[0076] See Figure 4 , regarding the relationship between pixel values g1 - gn and bit lines B1 - BN, pixel value g1 is converted into voltage value V1 through DAC digital - to - analog conversion processing, pixel value g2 is converted into voltage value V2 through digital - to - analog conversion, pixel value g3 is converted into voltage value V3 through digital - to - analog conversion, and so on. Pixel value gn is converted into voltage value VN through digital - to - analog conversion. According to the corresponding relationship of the interfaces shown in the figure, voltage value V1 is connected to bit line B1, voltage value V2 is connected to bit line B2, voltage value V3 is connected and input to bit line B3, and so on. Voltage value VN is connected and input to bit line BN. According to the rule, the larger the pixel value, the larger the converted bit - line voltage value, and the smaller the pixel value, the lower the converted bit - line voltage value. Moreover, the layout quantity of transistors, that is, the size of the matrix M×N, should have a corresponding relationship with the pixel size of the image Gr_in.
[0077] See Figure 4 , regarding the matrix Ma_wei, the gates of the transistors in the first row are respectively connected to word line G1, the gates of the transistors in the second row are respectively connected to word line G2, and the gates of the transistors in the M - th row are respectively connected to word line GM. The weights generated by the transistors in the first row in the first row of the weight matrix are denoted as W 11 , W 12 , W 13 , ……W 1N . The weights generated by the transistors in the second row in the second row of the weight matrix are denoted as W 21 , W 22 , W 23 , ……W 2N . The weights generated by the transistors in the M - th row in the M - th row of the weight matrix are denoted as W M1 , W M2 , W M3 , ……W MN . So far, each transistor is configured with a corresponding weight value contributing to the matrix Ma_wei participating in network calculations. Traditional transistors themselves do not have characteristics associated with weights. Especially, flash memory has no direct relevance to neural networks. Therefore, how to map weights to transistors is another key problem to be solved.
[0078] See Figure 4, a total of M rows and N columns of transistors are regarded as matrix elements of M×N. The definition of each weight value in the matrix Ma_wei is the basis for implementing the neural network NET. In the figure, it shows that the source current of each row of transistors flows to the common source line of the transistors in that row, and it shows that the drain of each column of transistors is connected to the common bit line of the transistors in that column. This layout of transistors will bring a feasible solution for how to define the weights of transistors. For example, the source current convergence characteristic of each row of transistors corresponds to the multiplication and accumulation calculation algorithm of the vector matrix multiplication mentioned above.
[0079] See Figure 4 , to explain how to map weights to transistors, first introduce the relationship between the current, threshold, and applied voltage of a storage transistor: the threshold voltage of the gate is V TH ; the process parameter K related to the process of the storage transistor is K = μCW / L, where W / L is the aspect ratio size of the transistor. For a storage transistor with a determined manufacturing process and size, the process parameter K is almost a constant value; the voltage V GS between the gate and the source and the voltage V DS between the drain and the source of the transistor can satisfy K(V GS - V TH )V DS = I. When the current I flowing through the transistor is in the linear region or the triode region, if the current I of a single transistor performs multiplication operations and the output currents I of each transistor in a row perform accumulation, this is equivalent to the multiplication and accumulation calculations of matrix multiplication.
[0080] See Figure 4 , for the first row of transistors (the gate is coupled to G1): the current I1 of the first transistor is approximately calculated as K(V GS1 - V TH1 )V DS1 = I1. According to the same current formula, the current I2 of the second transistor is approximately calculated as K(V GS2 - V TH2 )V DS2 = I2, and the current of the Nth transistor K(V GSN - V THN )V DSN = IN. The voltage between the gate and the source of the first transistor is V GS1 and the voltage between the drain and the source is V DS1 , and the threshold voltage is V TH1 . The voltage between the gate and the source of the second transistor is V GS2 and the voltage between the drain and the source is V DS2 , and the threshold voltage is V TH2 . The voltage between the gate and the source of the Nth transistor is V GSNand the voltage between the drain and source is V DSN and the threshold voltage is V THN .
[0081] Refer to Figure 4 . For the first row of transistors (gate coupled to G1): for the first transistor to the Nth transistor, their respective source currents all flow to the source line S1 common to the transistors in this row, and the current of the source line S1 is denoted as IS1.
[0082] IS1 = I1 + I2 + I3 + … + IN
[0083] Refer to Figure 4 . For the first row of transistors (gate coupled to G1): if W is defined at the position of the first transistor 11 is proportional to or equal to K(V GS1 - V TH1 ) and V1 is proportional to or equal to V DS1 , then the turnable expression of the current I1 flowing through the first transistor mentioned above can be expressed as I1 = W 11 × V1. If W is defined at the position of the second transistor 12 is proportional to or equal to K(V GS2 - V TH2 ) and V2 is proportional to or equal to V DS2 , then the turnable expression of the current I2 flowing through the second transistor mentioned above can be expressed as I2 = W 12 × V2. If W is defined at the position of the third transistor 13 is proportional to or equal to K(V GS3 - V TH3 ) and V3 is proportional to or equal to V DS3 , then the turnable expression of the current I3 flowing through the third transistor mentioned above can be expressed as I3 = W 13 × V3. If W is defined at the position of the Nth transistor 1N is proportional to or equal to K(V GSN - V THN ) and VN is proportional to or equal to V DSN , then the turnable expression of the current IN flowing through the Nth transistor mentioned above can be expressed as IN = W 1N × VN.
[0084] Refer to Figure 4, for the first row of transistors (gate coupled to G1): The current of the source line S1, denoted as IS1, can be expressed in the following way after transformation by the formula (it can also be defined that Y1 = IS1). The prerequisite is the source current convergence characteristic of the first row of transistors, which corresponds to the multiplication and accumulation calculation algorithm of the vector matrix multiplication. Therefore, it is one of the main factors for the flash memory to be able to replace some functions of the traditional CPU or GPU.
[0085] IS1 = Y1 = W 11 *V1 + W 12 *V2 + W 13 *V3 + … W 1N *VN
[0086] See Figure 4 , for the second row of transistors (gate coupled to G2): The current of the source line S2, denoted as IS2, can be expressed in the following way after transformation by the formula (it can also be defined that Y2 = IS2).
[0087] IS2 = Y2 = W 21 *V1 + W 22 *V2 + W 23 *V3 + … W 2N *VN
[0088] See Figure 4 , for the Mth row of transistors (gate coupled to GM): The current of the source line SM, denoted as ISM, can be expressed in the following way after transformation by the formula (it can also be defined that Y M = ISM).
[0089] ISM = Y M = W M1 *V1 + W M2 *V2 + W M3 *V3 + … W MN *VN
[0090] See Figure 4 , as defined above, X1 to X N are the input layer of the neural network NET, and Y1 to Y N correspondingly are the output layer of the neural network NET. So far, the first row of transistors in the flash memory has completed a vector operation Y1 from the input layer to the output layer, and the result of Y1 can be directly calculated at once by sensing the current of the source line S1 of the first row of transistors, without separately calculating W 11 × V1 or W 12 × V2 or W 1N×VN, etc., belong to the category of computing under the non - von Neumann architecture. The calculation result is not obtained through numerical operations but by current sensing. Similarly, the second - row transistors of the flash memory complete a vector operation Y2 from the input layer to the output layer, the third - row transistors of the flash memory complete a vector operation Y3 from the input layer to the output layer, and the M - th row transistors of the flash memory complete a vector operation Y from the input layer to the output layer. M In an alternative embodiment, the currents IS1 - ISM undergo analog - to - digital conversion processing by an ADC device, such as an analog - to - digital converter encapsulated together with the flash memory. The digital quantities corresponding to the current values are applicable to various electronic devices. The ADC (Analog to Digital Converter) device can also be directly integrated onto the same semiconductor wafer with the flash memory during the wafer manufacturing stage.
[0091] See Figure 4 , one task completed by the transistor array of the flash memory is to construct the matrix Ma_wei, and each element in the matrix corresponds to the weight matrix of the neural network NET.
[0092] See Figure 4 , taking the input layer Lay1 / hidden layer Lay2 of the neural network NET as an example for the matrix Ma_wei. When the input data is input to the input layer Lay1, there is a matrix Ma_wei between the input layer Lay1 and the hidden layer Lay2. And the input data is converted into voltages V1 - VN and given to the input layer Lay1. Then the output obtained by the hidden layer Lay2 is equivalent to the operation result obtained by performing matrix operations on the input data by the weight matrix Ma_wei. The type of input is a voltage signal and the output data is a current - value signal. By the action of the weight matrix Ma_wei on the input voltage signal, the amplitude of the output current value is affected. This is roughly the process of implementing artificial intelligence operations through the NOR Flash memory.
[0093] See Figure 4 , in addition to being used for the above - mentioned AI artificial intelligence calculations, the memory - in - computing technology can also be used in other types of computing applications such as neuromorphic chips, representing a type of future mainstream big - data computing chip architecture, etc. At the same time, it can also be used as the core computing module of a low - energy - consumption supercomputer center. For example, NOR FLASH in - memory computing can support approximately 300M or so deep - learning weight parameters in a single chip and can perform calculations without additional memory.
[0094] See Figure 4 , the calculation of the flash memory for the neural network NET has several important characteristics described below compared with other types of general computing, which also belong to the advantages it exhibits.
[0095] SeeFigure 4 , in the first aspect, the neural network NET based on flash memory has a simple calculation process and allows a large amount of parallelism. The calculation of the neural network NET no longer requires mechanisms such as branch prediction and speculative execution commonly used in general-purpose computing by the CPU to operate efficiently. This feature enables the flash memory hardware accelerator for the neural network NET to minimize or eliminate the use of the CPU or GPU as much as possible, thereby performing low-cost and high-efficiency calculations.
[0096] See Figure 4 , in the second aspect, the calculations performed by the neural network NET consist of only a few single operations, such as matrix multiplication and vector operations, the application of convolutional kernels, and other linear algebra operations. This enables the neural network NET accelerator to be similar to the GPU, using a large number of simple computing units to handle the complex data operations required by the industry.
[0097] See Figure 4 , in the third aspect, the neural network NET itself exhibits a large sparsity characteristic, where a large number of its parameters and calculation results are very small, even approaching zero. Utilizing the sparsity characteristic can enable the operation to reasonably skip the calculation of multiplying by zero, thereby improving the overall calculation speed, especially for matrix operations.
[0098] See Figure 4 , in the fourth aspect, the neural network NET allows low-precision calculations such as low-bit precision. For example, for most artificial intelligence applications, 8-bit or 4-bit (meaning small transistor arrays) fixed-point numbers can be used for quantization with almost no loss of recognition accuracy. Low-precision fixed-point numbers can be used for calculation according to the artificial intelligence application scenario to improve the inference speed and reduce the power consumption. Power loss is a point worth emphasizing. Traditional CPUs / GPUs generate a large amount of heat during the model training and inference stages, and most of the power is lost in the heat dissipation of the processor. The calculation of this application completely avoids the process of single-term multiplication and the accumulation of multiplication results by the processor, and reduces unnecessary data transfer links between the processor and the memory, reducing the energy consumption to 1 / 10 to 1 / 100 of the original.
[0099] See Figure 4 , for example, the size of the intelligent voice model is usually in the range of several hundred K to several megabytes, and the size of the image inference model on the edge side is usually between several megabytes and dozens of megabytes. The storage space of conventional NOR FLASH on the market can easily handle this scale of model volume. Therefore, even the NOR FLASH in-memory computing chip of a handheld terminal is sufficient to meet the needs of most AI scenarios. On the premise of optimizing the NOR FLASH structure, such as the series-parallel or cascade connection of numerous flash memories, it can almost rival all artificial intelligence application scenarios that can be achieved by the GPU.
[0100] See Figure 5 , as described above for weights such as W 11 and transistor parameters such as K(V GS1 -V TH1 ), a proportional or equal mapping relationship was explained. In the neural network NET, weights are information that needs to be adjusted and updated at any time. Therefore, the part that is worth elaborating in detail also includes how to adaptively modify transistor parameters (or weights). In particular, neural networks include backpropagation, and the core of backpropagation is the update and iteration of weights. Obviously, the local calculation of flash memory is different from the weight iteration of traditional processors.
[0101] See Figure 5 , the threshold voltage of the transistor is V TH . If the threshold voltage is changed, it is equivalent to changing K(V GS -V TH ), which is equivalent to changing the weight value proportional to K(V GS -V TH ). Therefore, changing the threshold voltage V TH is an optional method. The writing and erasing of flash memory data are carried out by manipulating the charge inside the floating gate related to the control gate, such as charge release and tunneling. For example, in the writing method, NOR Flash uses the hot electron injection method. By increasing the voltage of the control gate during writing, for example, injecting charge from the source to the floating gate, the change in the amount of charge in the floating gate can change K(V GS -V TH ). Another example is in the erasing method. NOR Flash uses the tunneling effect to remove the charge in the floating gate. For example, a negative voltage is applied to the control gate.
[0102] See Figure 5 , taking the transistor in the Mth row (the gate is coupled to the word line GM) as an example. In the erase operation of programming PG0, the floating gate charge of the transistor determined by the bit line BN (weight is W MN ) can be cleared. Therefore, its threshold voltage is reduced and the weight W MN is increased. The erase of programming PG0 typically applies a negative voltage at the word line GM to cause the floating gate charge of the transistor to undergo quantum tunneling and enter the substrate where the transistor is located. According to the requirements for weight value modification described above, conversely, in the write operation of programming PG1, the floating gate charge of the transistor determined by the bit line BN (weight is W MN ) can be injected. In this case, the threshold voltage of the transistor is increased and the weight W MN is reduced. Furthermore, in the write operation of programming PG2, continue to inject the floating gate charge of the transistor determined by the bit line BN (weight is WMN ) The floating - gate charge of, programming PG2 is a fine - tuning relative to programming PG1, so the threshold voltage of the transistor continues to increase and the weight W MN is slightly reduced. Among them, the voltage of charge - injection programming PG1 is much higher than that of the fine - tuned programming PG1, and the small voltage scale of programming PG2 is to avoid over - adjustment of the weight. In essence, charge - injection programming can adopt more injection times than those shown in the figure, so that the weight W MN can not only meet the multiple - value requirements of the weight matrix Ma_wei but also meet the high - precision requirements of the weight value. The erase - programming operation and the charge - injection programming operation shown in the figure execute multiple rounds of cycles, representing the iteration of the weight.
[0103] See Figure 4 , in an alternative embodiment, the method of performing in - memory computing includes: using a storage - transistor array in a flash memory as a carrier of the weight matrix of a neural network NET, setting the input data of the neural network to be delivered to a series of bit - lines B1 - BN of the storage - transistor array, and also setting the output data of the neural network to be retrieved from a series of source - lines S1 - SM of the storage - transistor array. The magnitude of the weight of any transistor in the storage - transistor array is adjusted using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network. For example, for a transistor T determined by bit - line B1 and source - line S1 11 : The weight of transistor T 11 is adjusted using its threshold voltage. For a transistor T determined by bit - line B1 and source - line S2 21 : The weight of transistor T 21 is adjusted using its threshold voltage. For a transistor T determined by bit - line B1 and source - line S3 31 : The weight of transistor T 31 is adjusted using its threshold voltage. For a transistor T determined by bit - line B1 and source - line SM M1 : The weight of transistor T M1 is adjusted using its threshold voltage. Therefore, the magnitude of the weight of any transistor is adjusted using its threshold voltage, thereby dynamically configuring the weight matrix (weightmatrix) of the neural network. N and M are positive integers. Again, for a transistor T determined by bit - line BN and source - line SM MN : The weight of transistor T MN is adjusted by injecting floating - gate charges and the like using its threshold voltage.
[0104] See Figure 4, in an alternative embodiment, the transistor arrays for the neural network determined by bit lines B1 - BN and source lines S1 - SM are located on the same physical block of the flash memory; and the spaces where the input data (e.g., vector Ma_in) and output data (e.g., vector Ma_out) of the neural network are located belong to different physical blocks from the storage space where the transistor arrays are located. For example, the storage spaces where the input data and the output data are located belong to the same physical block, or the storage spaces where the input data and the output data are located belong to the same physical blocks respectively. However, the storage spaces where the input data and the output data are located respectively belong to different physical blocks from the storage space where the transistor arrays are located. This separation of storage spaces is mainly considered to allow different erasure and write operations for the storage - like space and the computing - like space. That is, the weight configuration operations such as PG0, PG1, PG2, etc. of using the threshold voltage Figure 5 are performed in different storage spaces from the data read - write operations of storing data such as the input and output data of the neural network. The weight configuration operations and the data read - write operations of storing data do not interfere with each other. Because the programming voltages (or programming currents) used for the weight configuration operations and the data read - write operations of storing data require different voltage levels (or current levels).
[0105] See Figure 4 , in an alternative embodiment, the original input data of the neural network NET is converted into an input quantity in the form of a voltage value, and is sent to a series of bit lines of the transistor array based on the input quantity sorted into a vector. The original input data includes, for example, data types such as pixel values or grayscale values of images or videos. The original input data is converted into the input quantities V1, V2, V3... VN in the form of the illustrated voltage values. The voltage value V1 is input to the bit line B1, the voltage value V2 is similarly input to the bit line B2, the voltage value V3 is input to the bit line B3, and the voltage value VN is input to the bit line BN.
[0106] See Figure 4 , the input is represented as vector Ma_in = [B1 = V1, B2 = V2... BN = VN]. Based on the input quantity [V1, V2... VN] sorted into a vector, it is sent to the series of bit lines B1 - BN of the array (transistor arrays).
[0107] See Figure 4, in an alternative embodiment, the output data of the neural network NET is characterized by the output current of the memory transistor array, and the output currents of a series of source lines of the memory transistor array are regarded as the output quantities sorted into a vector of the memory transistor array. For example, for the series of source lines S1-SM of the memory transistor array (transistor arrays), their respective output currents are denoted as IS1, IS2, … ISM. The output current IS1 of the source line S1 can be converted into a digital quantity through ADC analog-to-digital conversion, and similarly, the output current ISM of the source line SM can be converted into a digital quantity through ADC analog-to-digital conversion. The output data of the network NET is characterized by the output currents IS1, IS2, … ISM of the memory transistor array.
[0108] See Figure 4 , in an alternative embodiment, the output currents of the series of source lines S1-SM of the memory transistor array are regarded as the output quantities sorted into a vector Ma_out [IS1, IS2, … ISM] of the memory transistor array. It can also be considered that the output of the neural network is represented by the vector Ma_out = [S1 = IS1, S2 = IS2, … SM = ISM].
[0109] See Figure 5 , in an alternative embodiment, the way to adjust the weights of the adjustment transistors (such as T determined by BN / SM) MN includes changing the charge injected into the floating gate (such as changing the floating gate charge of T) MN , and by changing the charge injected into the floating gate of the transistor, the weight magnitude at any position (such as the position determined by BN / SM) of the transistor in the memory transistor array is adjusted. For example, the weight W of T MN is adjusted. MN of its magnitude.
[0110] See Figure 5 , in an alternative embodiment, in the backpropagation stage of the neural network NET, the weight matrix of the neural network is dynamically configured by resetting the weight magnitudes at each position in the memory transistor array. Each position in the array mentioned here includes each position where the bit lines B1-BN and the source lines S1-SM intersect, and in fact, each position mentioned also represents the row and column indices of each element in the weight matrix. For example, the row and column indices of the element W MN in the weight matrix are determined or corresponding to the position determined by the bit line BN / source line SM. The neural network needs to repeatedly iterate the magnitude values of each element in the weight matrix in the backpropagation algorithm until the requirements are met.
[0111] See Figure 4, in an alternative embodiment, the input data of the neural network is delivered to a series of bit lines of the memory transistor array and it is set to retrieve the output data of the neural network from a series of source lines of the memory transistor array: the sources of the memory transistors in each row of the transistor array are connected together and regarded as the source line output terminal of one path of the output data of the neural network, and the drains of the memory transistors in each column of the transistor array are connected together and regarded as the bit line input terminal of one path of the input data of the neural network. Additionally, the gates of the memory transistors in each row of the memory transistor array are connected together and coupled to one path of word lines.
[0112] See Figure 3 , in an alternative embodiment, a digital-to-analog converter DAC is arranged at each of the bit lines B1 - BN of the memory transistor array, and the digital-to-analog converter DAC is used to convert the input data of the neural network NET into a voltage value that matches the input signal pattern of the bit line (for example, the bit line requires a voltage signal as the input quantity).
[0113] See Figure 3 , in an alternative embodiment, an analog-to-digital converter ADC is arranged at each of the source lines S1 - SM of the memory transistor array, and the analog-to-digital converter ADC is used to convert the output signal in the format of the output current at the source line into output data that is adapted to the digital output quantity of the neural network NET. For example, the analog-to-digital converter arranged at the source line SM is used to convert the output signal in the format of the output current ISM at the source line SM into output data that is adapted to the digital output quantity of the neural network. The digital quantity of the current is different from the analog quantity. The analog-to-digital converter arranged at the source line S1 is used to convert the output signal in the format of the output current IS1 at the source line S1 into output data that is adapted to the digital output quantity of the neural network. The same applies to the other S2 to SM - 1.
[0114] See Figure 4 , in an alternative embodiment, the device for performing in-memory computing includes: a flash memory, an input unit, an output unit, etc. Regarding the flash memory part in the device, its memory transistor array can be used as the carrier of the weight matrix of the neural network, and it is set that the input data of the neural network is delivered to a series of bit lines of the memory transistor array and it is set to retrieve the output data of the neural network from a series of source lines of the memory transistor array. And regarding the input unit part in the device, it can convert the original input data of the neural network into an input quantity in the form of a voltage value, and based on the input quantities sorted into a vector, deliver them to a series of bit lines of the memory transistor array. Regarding the output unit of the device, it is used to convert the output signal in the format of the output current of the source line into output data that is adapted to the digital output quantity of the neural network, and the output current of a series of source lines of the memory transistor array is regarded as the output quantity sorted into a vector of the memory transistor array. The input unit and the output unit respectively include a digital-to-analog converter and an analog-to-digital converter.
[0115] See Figure 4 In an alternative embodiment, a memory system performing in-memory computing includes: a flash memory-based memory chip, a digital-to-analog converter built into the memory chip, and an analog-to-digital converter built into the memory chip. The flash memory-based memory chip, wherein a memory transistor array is utilized in the flash memory as a carrier of a weight matrix of a neural network and the magnitude of the weight of any one transistor is adjusted by using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network; setting input data of the neural network to be delivered to a series of bit lines of the memory transistor array and also setting output data of the neural network to be retrieved from a series of source lines of the memory transistor array. Regarding the digital-to-analog converter built into the memory chip, wherein a digital-to-analog converter is configured at each bit line of the memory transistor array, so that the input data of the neural network can be converted into a voltage value matching the input signal pattern of the bit line. Regarding the analog-to-digital converter built into the memory chip, wherein an analog-to-digital converter is configured at each source line of the memory transistor array to convert an output signal in the format of an output current at the source line into output data adapted to the digital output quantity of the neural network.
[0116] See Figure 5 In conventional electronic devices, computing and storage are separate. Computing, such as a processor, and storage, such as a hard disk or memory, etc., each focus on their respective tasks, and computing and storage basically do not interfere with each other. There are also manufacturers who integrate the processor and the memory closer in physical location in order to improve computing speed and computing power. For example, in the layout design of the circuit board, they are close to each other, have short circuits, and are directly stacked on top of each other during the packaging stage.
[0117] See Figure 5 The improvement of computing power has always been gradually advanced according to the rhythm of Moore's Law, but the mainstream computing paradigm has always followed the von Neumann architecture. As a new type of computing power, in-memory computing in the industry mainly focuses on narrow in-memory computing, that is, using storage media for computing, but the computing media methods have not been fully unified. As a new type of computing power, in-memory computing is expected to solve the problems of compilation wall, memory wall, and power consumption wall under the traditional von Neumann architecture. The consensus in the industry is that in-memory computing has basically been determined as one of the key technologies to solve the computing power problem.
[0118] See Figure 5, the input data or output data of the neural network mapped to the storage transistor array in this application can be the data of a certain layer (input layer / hidden layer / output layer, etc.) of the neural network. In other words, it is allowed that the storage transistor array is only a local sub-network structure of the neural network. Or, the neural network can use multiple storage transistor arrays as shown in the figure for serial-parallel cascade combination of front and back stages, and multiple storage transistor arrays are combined into a local sub-network structure of the neural network or the entire network structure.
[0119] See Figure 5 , the integration of storage and computing organically combines storage and computing. Based on its huge potential for improving energy efficiency ratio, it is expected to become an advanced computing architecture solution in the era of computing power. However, there are doubts: how to address the weight shift (such as the shift of W MN occurring) and other doubts that need to be solved urgently. This is an insidious problem that is difficult to detect for the memory, and will be elaborated in detail below.
[0120] See Figure 6 , it is known that the transistor arrays for the neural network NET can be determined by the bit lines B1 - BN and the source lines S1 - SM. Taking this array as an example to illustrate the problem of weight shift. The size of the storage transistor array should fit the actual application scenario of the neural network. For example, neural networks often need to perform convolution calculations on images of various lengths and widths to obtain convolution results. For example: even a simple handwritten recognition image with a length and width of 28 * 28 may require about 784 input ports. Obviously, in this case, the size of the storage transistor array is quite large. At this time, a more difficult problem is that the length of the word line will become very long as the size of the storage transistor array increases, because the adjustment of the weight depends on programming the charge writing or erasing of the transistor floating gate through the word line. The too-long word line will cause a large word line parasitic capacitance C P on the transistor array, resulting in errors in weight programming. The parasitic capacitance includes at least the gate-source parasitic capacitance and the gate-drain parasitic capacitance of each transistor coupled to the word line and the capacitance of the word line to the substrate. The storage transistor array is fabricated on a semiconductor substrate. Programming errors also mean deviations in weight configuration. It should be noted that in many neural networks, such as convolutional neural networks in the intelligent driving domain video perception, the input image resolution related to the number of input ports of the transistor array is commonly 1628 * 1236, 1920 * 1080, 2048 * 1024, 3840 * 2160, etc. Whether the length and width of the image are scaled or not, it will be accompanied by a huge array size and the resulting parasitic capacitance C PComplication, because the amount of parasitic capacitance does not have a completely corresponding relationship with the size of the array. Parasitic capacitance includes both parallel capacitance and series capacitance, such as the body capacitance at the transistor and also coupling capacitance, such as the parasitic coupling capacitance between different word lines. This brings a great error rate to the weight iteration, such as the backpropagation algorithm, and there are great irrationalities in both training and inference. In the conventional application scenarios of the memory, this problem is not worried about, because when the memory faces ultra-large or massive data storage applications, it will use segmented physical blocks, planes or LUNs to meet this storage demand. However, in the neural network scenario, if the weight matrix is segmented in different physical areas, the inconsistency of the programming conditions of the weight elements will tend to be serious.
[0121] See Figure 6 , in an optional embodiment, for any transistor on any word line of the storage transistor array during the stage of dynamically configuring weights (dynamically configuring weights because the parameters of the neural network are continuously iterating), before any weight adjustment is performed on this any transistor, first transiently mount the reverse voltage of the programming voltage used by other transistors that were weight-adjusted adjacent to this any transistor on the previous time on this any word line, and then apply the programming voltage set for this any transistor to this any word line to perform this any weight adjustment. One of the purposes is to suppress the failure of the weight update iteration caused by the parasitic charge of the parasitic capacitance existing at this any word line of the storage transistor array. Parasitic charge is usually affected by the residue of the programming voltage of other transistors that perform weight adjustment on this any word line. In the case of high image resolution, assuming that the expected number of training epochs of the neural network is epoch = 5000, with the superposition of such a large array size and such a large number of iteration times, the parasitic charge of any word line will quickly get out of control in a hidden manner. It is very difficult to master the true value of the actual parasitic charge of the word line, and even if it is known, it will soon be refreshed by the next round of iteration.
[0122] See Figure 6 , in an optional embodiment, regarding how to better address the urgent doubts such as weight offset faced during the weight update of the internal neural network of the memory: The basic solution is to adjust the size of the weight of any transistor in the storage transistor array using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network. Further optimization is that for any transistor on any word line of the storage transistor array during the stage of dynamically configuring weights, before any weight adjustment is performed on this any transistor, first transiently mount the reverse voltage of the programming voltage used by other transistors that were weight-adjusted adjacent to this any transistor on the previous time on this any word line, and then apply the programming voltage set for this any transistor to this any word line to perform this any weight adjustment.
[0123] See Figure 6, for example, in transistor arrays, transistor T MN has its weight adjusted by its threshold voltage V THMN . Transistor T MN is a transistor determined by bit line BN and source line SM, and the word line it belongs to is GM. Let positive integer M = 1, 2, 3..., positive integer N = 1, 2, 3.... Thus, the weight matrix Ma_wei of neural network NET can be dynamically configured, and V THMN in V TH is the threshold, and the subscript MN is the index, which represents the threshold voltage and weight W MN of any certain transistor, and both can be adjusted by the index position.
[0124] See Figure 6 , for example, any transistor such as T on any word line such as GM in transistor arrays (transistor arrays) MN is in the stage of dynamically configuring the corresponding weight W MN . Before any weight adjustment is performed on this any transistor such as T MN , for example, before the weight adjustment with the execution number of epoch = 1455, it is better to first transiently mount the reverse voltage of the programming voltage used by the transistor adjacent to the previous weight adjustment on the same word line GM (assuming that transistor T M3 on GM is T MN that performs the previous weight adjustment adjacent to the current weight adjustment with epoch = 1455 times) once onto the same word line GM. After that, the programming voltage such as V PGM3 set for this any transistor such as transistor T on GM MN is applied to, for example, word line GM to perform this any weight adjustment, such as the adjustment action of weight W PGMN with the execution number of epoch = 1455. Inhibit the weight (such as W MN ) caused by the parasitic charge of the parasitic capacitance existing at the any word line GM of the memory transistor array MN) The update iteration fails and it is worth emphasizing that: this scheme is not suitable for data storage applications in the memory, otherwise it will cause the logical value of the stored data bit to flip unexpectedly with a high probability. Note that the parasitic charge of the word line GM has a volatile characteristic, but the rate of this volatile power loss may be lower than the speed of weight matrix iteration, which is related to the transistor deployment density of the storage transistor array and the selection of components such as wiring and the critical size of the transistor. The above scheme can make up for and save the error rate of the weight as much as possible, and it complements the improvement and training of neural network parameters such as weights. An important goal of the above scheme is also to compress the step difference between the programming voltage set by any transistor to perform any weight adjustment and the programming voltage used by other transistors adjacent to the previous weight adjustment on any word line, so as to avoid the programming voltage set by any transistor to perform any weight adjustment to be coupled to the adjacent word line of any word line. Thereby preventing the weight values of some transistors on adjacent word lines from being changed in unexpected iterations.
[0125] See also Figure 6 , for example, the above scheme tries to compress any transistor such as T MN The programming voltage set for any weight adjustment is V PGMN Other transistors adjacent to any word line such as GM in the previous weight adjustment, such as transistor T M3 The programming voltage used is V PGM3 The voltage step between them is to avoid any transistor, such as transistor T MN Execute any one time (for example, the execution number is epoch=1455) to adjust the weight setting programming voltage, such as voltage V PGMN Coupled to the adjacent word line GM-1 or GM+1 of any word line GM. In the memory transistor array, if the voltage V PGMN When coupled to positions such as GM-1 or GM+1, the weight values of some transistors coupled to word lines GM-1 or GM+1 (such as the intersection of GM-1 / GM+1 and BN) will be inadvertently changed in a non-iterative manner and the weight matrix will be modified locally rather than globally. In this case, the neural network will produce wrong results whether it is training or reasoning. The above-mentioned solution designed by the present application cleverly resolves the problem of undesired iterative changes in the weight values of some transistors on adjacent word lines.
[0126] See also Figure 6, Based on the foregoing, the convolutional neural network is trained with training samples, and the trained parameters such as the weight matrix have good generalization performance for test samples. It has considerable robustness and can recover non-accidental offsets in resisting the incorrect weights caused by the parasitic capacitance on any word line in the memory transistor array. It belongs to a convolutional neural network or sub-network with good performance. Neural networks are typically applicable to intelligent speech models, image inference models on the edge side, and most other AI artificial intelligence scenarios.
[0127] Through the above description and drawings, typical embodiments of specific structures of the specific implementation manners are given. The above invention proposes existing preferred embodiments, but these contents are not restrictive. To those skilled in the art, various changes and modifications will undoubtedly be obvious after reading the above description. Therefore, the appended claims should be regarded as covering all changes and modifications that cover the true intention and scope of the present invention. Any and all equivalent scopes and contents within the scope of the claims should be considered to still fall within the intention and scope of the present invention.
Claims
1. A method for performing in-memory computing, characterized in that: include: In a flash memory, a storage transistor array is used as a carrier of a weight matrix of a neural network, and input data of the neural network is set to be transmitted to a series of bit lines of the storage transistor array, and output data of the neural network is set to be captured from a series of source lines of the storage transistor array; wherein the weight of any transistor in the storage transistor array is adjusted by using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network; During the stage of dynamically configuring weights for any transistor on any word line of the storage transistor array, before executing any weight adjustment, the reverse voltage of the programming voltage used by other transistors on the word line that were adjacent to the previous weight adjustment is transiently mounted on the word line once, and then the programming voltage set for the transistor is applied to the word line to execute the weight adjustment, thereby suppressing the weight update iteration failure caused by the parasitic charge of the parasitic capacitance existing at the word line of the storage transistor array.
2. The method according to claim 1, characterized in that: The memory transistor array for the neural network is located on the same physical block of flash memory; and The storage space where the input data and output data of the neural network are located and the storage space where the storage transistor array is located belong to different physical blocks of the flash memory respectively.
3. The method according to claim 1, characterized in that: The raw input data of the neural network is converted into input quantities in the form of voltage values, which are transmitted to a series of bit lines of the memory transistor array based on the input quantities sorted into vectors.
4. The method according to claim 1, characterized in that: The output data of the neural network is represented by the output current of the storage transistor array, and the output current of a series of source lines of the storage transistor array is regarded as the output quantity of the storage transistor array sorted into a vector.
5. The method according to claim 1, characterized in that: The method of adjusting the weight of any transistor includes changing the floating gate injection charge, and adjusting the weight of any transistor at its position in the memory transistor array by changing the floating gate injection charge of the transistor.
6. The method according to claim 5, characterized in that: During the back-propagation phase of the neural network, the weight matrix of the neural network is dynamically configured by resetting the weights at various locations in the storage transistor array.
7. The method according to claim 3, characterized in that: A digital-to-analog converter is configured at each bit line of the storage transistor array, and is used to convert the input data of the neural network into a voltage value that matches the input signal pattern of the bit line.
8. The method according to claim 4, characterized in that: An analog-to-digital converter is configured at each source line of the storage transistor array, and is used to convert an output signal in an output current format at the source line into output data that is compatible with the digital output quantity of the neural network.
9. A device for performing in-memory computing, characterized in that: include: A flash memory, wherein the storage transistor array is used as a carrier of a weight matrix of a neural network, and the input data of the neural network is set to be transmitted to a series of bit lines of the storage transistor array, and the output data of the neural network is set to be captured from a series of source lines of the storage transistor array, wherein the weight of any transistor in the storage transistor array is adjusted by using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network, and for any transistor on any word line of the storage transistor array, in the stage of dynamically configuring the weight, before performing any weight adjustment, the reverse voltage of the programming voltage used by other transistors on the word line adjacent to the previous weight adjustment is first transiently mounted on the word line, and then the programming voltage set for the transistor is applied to the word line to perform the weight adjustment, thereby suppressing the weight update iteration failure caused by the parasitic charge of the parasitic capacitor existing at the word line of the storage transistor array; An input unit that converts raw input data of the neural network into an input quantity in the form of a voltage value to be transmitted to a series of bit lines of the storage transistor array based on the input quantity sorted into a vector; An output unit is used to convert an output signal in the output current format at the source line into output data compatible with the digital output quantity of the neural network, and the output current of a series of source lines of the storage transistor array is regarded as the output quantity of the storage transistor array sorted into a vector.
10. A storage system for performing in-memory computing, characterized in that: include: A memory chip based on a flash memory uses a memory transistor array in the flash memory as a carrier of a weight matrix of a neural network, wherein the size of the weight of any transistor is adjusted using its threshold voltage, thereby dynamically configuring the weight matrix of the neural network; setting the input data of the neural network to be transmitted to a series of bit lines of the memory transistor array and also setting the output data of the neural network to be captured from a series of source lines of the memory transistor array, and for any transistor on any word line of the memory transistor array in the stage of dynamically configuring the weight, before any weight adjustment is performed, the reverse voltage of the programming voltage used by other transistors on the word line adjacent to the previous weight adjustment is first transiently mounted on the word line once, and then the programming voltage set for the transistor is applied to the word line to perform the weight adjustment, thereby suppressing the weight update iteration failure caused by the parasitic charge of the parasitic capacitor existing at the word line of the memory transistor array; A digital-to-analog converter built into the memory chip, wherein a digital-to-analog converter is configured at each bit line of the memory transistor array to convert the input data of the neural network into a voltage value that matches the input signal pattern of the bit line; An analog-to-digital converter is built into the memory chip, wherein an analog-to-digital converter is configured at each source line of the memory transistor array to convert an output signal in the output current format at the source line into output data that is compatible with the digital output quantity of the neural network.
Citation Information
Patent Citations
Neural network, operation method thereof, and neural network information processing system
CN110543937A