A high-low weight splicing method based on a memory-computing integrated architecture

CN120929425BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511096349.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2026-09-25
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

[0005]为了解决现有大规模存算一体架构中,因器件非理想特性和电导编程限制导致的计算精度下降问题,本发明提出了一种基于高低位权重拼接技术的解决方案

Benefits of technology

[0015]S4,矩阵向量乘法计算与高低位权重计算结果拼接:在存算一体架构中,高低位权重阵列并行完成矩阵向量乘法运算,最终通过数字移位相加模块合成输出结果,从而达到高精度计算的目标。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929425B_ABST
    Figure CN120929425B_ABST
Patent Text Reader

Abstract

The application discloses a high-low weight splicing method based on a memory-computing integrated architecture, which is used for solving the problem of calculation precision decline caused by non-ideal characteristics of a memory. The neural network weight trained offline is stored by a multi-bit memristor device, and the high-low weight splicing method is combined with the block storage of high-low weight and the error compensation mechanism. The high weight is used for storing main information, and the low weight is used for error compensation, so that high-precision matrix-vector multiplication operation is realized. The memory-computing integrated architecture comprises a memory-computing array, an analog input-output module, a digital shift-adding module, an arithmetic logic unit, a control unit and a register. Through hardware optimization and module cooperation, high-precision calculation under low-power consumption is realized. The application is suitable for various memory devices, such as flash memory, resistive random access memory (RRAM) and phase change memory (PCM), and has wide application scenarios, including the fields of speech recognition, image processing and deep learning inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of novel storage and computing, specifically relating to a high-low weight concatenation method based on an integrated storage and computing architecture, which can be used in edge computing scenarios such as speech recognition and image processing. Background Technology

[0002] Currently, the efficient implementation of neural network computing faces the "memory wall" problem caused by the separation of storage and processor (von Neumann architecture). In-memory computing architectures, by integrating storage and computing functions into a single array, can significantly improve computational energy efficiency. Media such as flash memory and resistive switching memory can store multiple bits of data in a single cell. However, non-ideal factors of the devices (such as programming errors and non-linear operating characteristics) significantly affect computational accuracy.

[0003] Although existing technologies attempt to mitigate the aforementioned problems using error compensation or array calibration methods, they still suffer from the following drawbacks: programming errors are difficult to completely eliminate, and accumulated errors lead to significant deviations in calculation results; the nonlinear characteristics of flash memory devices limit the storage and computational capabilities of high-precision weights; and the accuracy degradation of large-scale in-memory arrays is caused by programming limitations. To address these issues, this invention proposes a solution based on high-low bit weight splicing technology, which can significantly improve the adaptability of in-memory arrays while maintaining computational accuracy.

[0004] Programmable Linear Random Access Memory (PLRAM) is a novel memory device based on an improved NOR-Flash structure. It employs an Etox structure, consisting of a stack of dual floating gates and a control gate, and uses a Selective Eraser Gate (SEG) for lateral erasure. PLRAM utilizes a bidirectional Fowler-Nordheim tunneling mechanism instead of traditional hot carrier injection (HCI), significantly reducing stress and trap generation in the gate oxide layer, thereby improving the stability and repeatability of the stored state. The separation of the tunneling oxide layer from the gate oxide layer effectively reduces electron trapping and releasing effects, while optimizing the transistor's operating range, enabling the device to maintain a longer linear operating interval during forward inference. Through self-calibrating programming and erasing mechanisms, the device can achieve programming accuracy up to 7 bits / cell, with good linearity and reliability, and can efficiently perform operations such as matrix-vector multiplication in in-memory computing architectures. Summary of the Invention

[0005] To address the issue of decreased computational accuracy in existing large-scale in-memory computing architectures due to non-ideal device characteristics and limitations in conductivity programming, this invention proposes a solution based on high-low bit weight concatenation technology. Offline-trained neural network weights are stored in a programmable linear random access memory (RAM), and matrix-vector multiplication operations are performed in the in-memory array using Ohm's law and Kirchhoff's laws. Furthermore, by storing and computing high- and low-bit weights in blocks, and through two parallel hardware forward propagation calculations, the computational accuracy of the in-memory array is significantly improved. This solution is also compatible with other multi-bit memories, including but not limited to non-volatile memories based on resistive random-access memory (RRAM), phase-change memory (PCM), and magnetoresistive random access memory (MRAM), exhibiting high versatility and adaptability. Compared to traditional computing methods, both time costs and circuit power consumption are significantly reduced.

[0006] The in-memory computing architecture of this invention is based on an improved PLRAM device that employs a self-calibrating programming / erasing scheme to achieve precise control over the device's conductivity. Before programming, the conductivity of each memory cell is measured, and a feedback control loop dynamically adjusts the programming voltage based on the actual measured value, thereby ensuring that each memory cell accurately reaches the target conductivity value.

[0007] The in-memory computing architecture of this invention consists of an in-memory array, an analog input module (including I / O interfaces and a DAC), an analog output module (including an ADC), a digital shift-and-add module, an arithmetic logic unit (ALU), a control unit (CU), and registers (including on-chip SRAM and on-chip Flash). The high-low weight concatenation technique is implemented at the hardware level through independent high-weight arrays (MSB arrays) and low-weight arrays (LSB arrays). This high-low weight concatenation technique uses high-weights to store key information and low-weights to compensate for errors, significantly improving computational accuracy and is suitable for neural network accelerators with in-memory computing architectures. Signals are loaded through the analog input module and fed into the in-memory array for computation; the output current is converted into a digital signal through the analog output module. The digital shift-and-add module concatenates the computation results from the high- and low-weight arrays and outputs them to the ALU for subsequent operations, achieving high-precision computation. The control unit (CU) coordinates the operation of each module, and registers are used to store intermediate data and instructions, improving system computational efficiency. Specifically, it includes the following core steps:

[0008] S1, Weight Quantization and Hardware Mapping: The trained neural network weights W are mapped to the in-memory array, with the following mapping relationship:

[0009]

[0010] Where G represents the target conductivity value, W min and W max G represents the minimum and maximum values ​​of the neural network weights, respectively. min and G max These are the minimum and maximum conductance values ​​of the memory cell, respectively. By applying a fixed voltage (e.g., 0.4V) to each memory cell, the output current of the memory cell is measured, and the actual conductance value G is calculated using the formula G = I / V. a .

[0011] S2, Programming Error Calculation: For each memory cell, the actual conductance value G a There is a deviation from the target conductance value G. The programming error e = G is calculated. a -G, this error conforms to a normal distribution e~N(0,σ) 2 This will serve as the basis for subsequent compensation.

[0012] S3, Conductivity Error Compensation: The programming error e is amplified according to a preset scaling factor s to generate compensation weights, which are then stored in a low-order weight array (LSB array) to reduce the impact of the error on the calculation results. The compensated conductivity value G. c The expression is as follows:

[0013]

[0014] Compensation for low-order weights can significantly reduce the cumulative impact of errors on the results of matrix-vector multiplication, reducing the error e by a factor of s.

[0015] S4, Matrix-vector multiplication calculation and high-low weight calculation results concatenation: In the in-memory computing architecture, the high-low weight arrays complete the matrix-vector multiplication operation in parallel, and finally synthesize the output result through the digital shift and add module, thereby achieving the goal of high-precision calculation.

[0016] The high-low weighting method of this invention has the following outstanding advantages: by storing the main information with high weights and using low weights to compensate for errors, it solves the problem of insufficient computational accuracy caused by the nonlinear characteristics of flash memory devices in large-scale in-memory arrays; in addition to flash memory, this technology can be adapted to multi-bit memories such as RRAM and PCM, and has high device adaptability; by adopting a hierarchical storage structure and a self-calibration programming mechanism, it significantly reduces energy consumption and latency while maintaining high computational accuracy; this invention can be widely applied to edge computing scenarios that require high precision, such as speech recognition and image processing. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the device structure of the programmable linear random access memory (PLRAM) in this invention;

[0018] Figure 2 This is a system architecture diagram of the PLRAM-based in-memory computing array chip in this invention;

[0019] Figure 3 This is a diagram showing the accuracy distribution of PLRAM devices under different array sizes in this invention.

[0020] Figure 4 This is a diagram illustrating the simulation error compensation scheme for the in-memory computing system in this invention.

[0021] Figure 5 This is a schematic diagram of the architecture of the high- and low-weight splicing storage and computing system in this invention. Detailed Implementation

[0022] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This is intended to provide a basic understanding of the present invention and is not intended to confirm the key or decisive elements of the invention or the scope of protection sought. It is readily understood that, without altering the essential spirit of the invention, those skilled in the art can make various substitutions, changes, and modifications without departing from the spirit and scope of the invention and the appended claims. Therefore, the following specific embodiments and accompanying drawings are merely illustrative examples of the technical solution of the present invention and should not be considered as the entirety of the invention or as a limitation or restriction of the technical solution of the present invention.

[0023] Figure 1 The device structure of the programmable linear random access memory (PLRAM) of this invention includes a silicon channel, a source, a drain, a floating gate, a control gate, and a select gate. The floating gate is isolated from the silicon channel by a gate oxide layer and is used to store charge information; the control gate is located above the floating gate and is used to adjust the potential of the floating gate; the select gate is located between the floating gates and is used to control programming and erasing operations. By applying voltage to the select gate and the control gate, bidirectional FN tunneling operation of electrons between the floating gate and the silicon channel can be realized, thereby completing the programming and erasing of the memory cell. This structure can significantly reduce the programming voltage while improving programming accuracy and reliability.

[0024] Figure 2This invention presents the system architecture of a compute-in-memory (CIM) array chip based on programmable linear random access memory (PLRAM). The system architecture includes an I / O interface, on-chip SRAM, a digital-to-analog converter (DAC), a compute array using PLRAM as storage units, an analog-to-digital converter (ADC), an arithmetic logic unit (ALU), and a control unit (CU). Input signals enter the on-chip SRAM through the I / O interface, are converted to analog signals by the DAC, and then applied to the word lines (WL) of the compute array. In the compute array, weights are stored in a differential manner, allocated to adjacent bit lines (BL), and the input signals are weighted and summed using the storage units. The output current is processed by a differential amplifier, converted to a digital signal by the ADC, and then sent to the arithmetic logic unit for subsequent calculations. This architecture, through differential weight storage, significantly improves the dynamic range and calculation accuracy of the weights. Simultaneously, it utilizes the high-precision storage characteristics of PLRAM devices to achieve low-power, high-efficiency neural network computation, making it suitable for edge computing and artificial intelligence inference tasks.

[0025] Figure 3 To determine the accuracy distribution of PLRAM devices under different array sizes in this invention, the cumulative distribution function (CDF) is used to statistically analyze the current distribution of each target conductance state across multiple cells. Figure 3 (a) in the diagram corresponds to a 128×128 small array. All 16384 cells are programmed into 128 conductance states. The CDF curves are closely arranged, indicating that the error is small and the linearity is high. It can guarantee 7-bit precision within a 1 sigma range. Figure 3 (b) corresponds to a large-scale array of 1024×4096. Although the number of programmed states is the same, the CDF curve distribution becomes wider and the conductivity accuracy decreases to about 6 bits due to process drift and block programming errors. This shows that expanding the array size has a significant impact on device consistency and memory accuracy.

[0026] Figure 4This invention presents a method for compensating for system and random errors in in-memory computing chips, aiming to improve the accuracy of in-memory computing operations. In the example shown in the figure, the target weight vector is [8,8,9], corresponding to an ideal output of 126. However, due to device fluctuations and programming limitations in the PLRAM array, the actual stored weights become [8.1,7.6,9.3]. The error [-0.1,0.4,-0.3] is amplified using a scaling factor S=16, resulting in a theoretically compensated weight vector of [-1.6,6.4,-4.8]. Assuming the hardware error offset of the lower-order weight array is the same as that of the higher-order weight array, the actual compensated weight vector is [-1.5,6,-4.5], which is stored in the corresponding position of the lower-order weight array. The matrix-vector multiplication result of the first column of the lower-order weight array is -33. After being converted into a digital signal by an analog-to-digital converter, it is reduced to 1 / 16 by a digital shift and adder module and added to the output of the higher-order weight array to obtain the corrected output. The compensation scheme reduced the output error from the original 2.2 to 0.1, verifying the effectiveness of the simulation compensation method in dealing with the non-idealities of in-memory computing chip devices.

[0027] Figure 5 The high- and low-weight concatenation storage and computation system of this invention uses a high-weight array primarily responsible for performing weighted summation operations in neural network computation tasks, while a low-weight array stores amplified error compensation weights to correct computational errors in the high-weight array. The red overlapping area indicates the overlap between the high and low weights during the bit-width concatenation process; its number of bits can be dynamically configured according to the network weight accuracy requirements, equivalent to adjusting the scaling factor S of the compensation weights. This mechanism enables the invention to perform accuracy-performance co-optimization for different tasks or network models. The specific operation steps are as follows:

[0028] 1. Weight Programming and Conductivity Measurement: The trained neural network weights are first programmed into the memory cells of a maximum weights array (MSB array). These memory cells store the main computational weights of the neural network. For each memory cell in the MSB array, a fixed voltage (e.g., 0.4V) is applied, and the output current of the memory cell is measured to determine the corresponding actual conductivity value G. a .

[0029] 2. Programming error calculation: Based on the target conductivity value G and the actual conductivity value G a Calculate the programming error e = G for each memory cell. a -G.

[0030] 3. Error Scaling and Compensation Weight Generation: The calculated programming error e is amplified according to a preset scaling factor s to generate compensation weights. The amplified compensation weights are used to correct the errors of the high-order weights. The amplified compensation weights are programmed into the storage cells of the low-order weight array (LSB array) to form the final high- and low-order weight concatenation structure.

[0031] 4. Analog Error Compensation: During system operation, the calculation results are corrected through two parallel hardware forward propagation calculations, combined with the main output of the high-weight array and the compensation output of the low-weight array. Since the weights of the low-weight array are amplified by a factor of s, the error of the result after digital shifting and addition is correspondingly reduced by a factor of s, thereby significantly improving the calculation accuracy.

[0032] The above description is merely a preferred embodiment of the present invention. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.

Claims

1. A method for concatenating high and low weights based on a memory-computing architecture, characterized in that, The in-memory computing architecture includes an in-memory array, an analog input module, an analog output module, a digital shift-and-add module, an arithmetic logic unit (ALU), a control unit, and registers. The in-memory array comprises a high-weight array and a low-weight array, both implemented at the hardware level. Both the high-weight array and the low-weight array are based on a programmable linear random access memory (PLM). Signals are loaded through the analog input module and fed into the in-memory array for computation; the output current is converted into a digital signal by the analog output module. The digital shift-and-add module concatenates the computation results from the high-weight array and the low-weight array and outputs them to the ALU for subsequent computation. The control unit coordinates the operation of each module, and the registers store intermediate data and instructions. Includes the following steps: S1, Weight Programming and Conductivity Measurement: The trained neural network weights are programmed into the storage cells of the high-order weight array; A fixed voltage is applied to each memory cell of the high-weighted array, and the output current is measured to determine the actual conductance value G. a ; S2, Programming Error Calculation: Based on the target conductivity value G and the actual conductivity value G a Calculate the programming error e=G for each memory cell. a -G; S3, Error scaling and compensation weight generation: The programming error e is amplified according to a preset scaling factor s to generate compensation weights; The compensation weights are programmed into the storage unit of the low-order weight array to form a high-low order weight splicing structure. S4, Matrix-vector multiplication calculation and high-low weight calculation results concatenation: In the in-memory computing architecture, the high-weight array and the low-weight array perform matrix-vector multiplication operations in parallel, and finally synthesize the output results through the digital shift and add module.

2. The high-low weight concatenation method based on in-memory computing architecture according to claim 1, characterized in that, The method can be adapted to non-volatile memories based on resistive random access memory, phase-change memory, and magnetoresistive random access memory.

Citation Information

Patent Citations

  • Accuracy compensation device and method for temperature excursion characteristics of storage and calculation integrated device

    CN116384456A

  • Three-dimensional stacked storage and calculation integrated SRAM and CPU integrated storage and calculation integrated architecture and implementation method

    CN118445089A