In-memory array testing device capable of flexibly pulsing and control method of in-memory array testing device
By designing a flexible in-memory array testing device, using daughterboard arrays and FPGAs to realize matrix computing and testing of in-memory chips, it solves the difficulty of existing test systems to meet the performance evaluation requirements in large-scale integration and high-speed computing environments, and achieves high flexibility and adaptability testing results.
Patent Information
- Application Number
- CN202510022844.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
AI Technical Summary
Existing test systems are difficult to meet the performance evaluation needs of memristor arrays in large-scale integration and high-speed computing environments, and lack flexibility and are difficult to adapt to changing technological development and application needs.
A flexible in-memory array testing device is designed, including a daughterboard array composed of multiple daughterboards. Each daughterboard includes a chip base, a digital-to-analog conversion module, a driver module, a differential reading module, an analog-to-digital conversion module, a shift accumulation module and a bidirectional interface. These components realize matrix computing and testing of in-memory chips, and realize data scheduling and control through FPGA.
It realizes flexible testing and verification of in-memory chips, can calculate different bit numbers according to needs, improves the flexibility and adaptability of the device, is suitable for operations with different accuracy and multi-bit wide fusion, further promoting the practicality of memristor technology.
Smart Images

Figure CN120072016A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to microelectronics, and more specifically, relates to a test device for an in-memory array capable of flexible pulsation and a control method therefor. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the wide application of deep learning algorithms in fields such as autonomous driving, industrial control, and intelligent monitoring, higher requirements are put forward for the performance and reliability of edge computing hardware. These devices need to achieve efficient data processing and rapid intelligent decision-making under limited energy and space conditions. Among them, the computing-in-memory (CIM) technology represented by memristors has become one of the key technologies to solve the AI computing power demand due to its significant advantages in providing high computing power, high energy efficiency, and low latency.
[0003] A memristor, as a non-volatile memory device, has the dual functions of storage and computing. Its unique physical properties make it show great potential in computing-in-memory applications. By placing data storage and computing in the same location, the memristor can significantly reduce the data transmission distance, lower power consumption, and improve computing efficiency. Therefore, the memristor is usually used as the core component of CIM technology, and its unique data storage and processing capabilities make it possible to achieve high computing power, high energy efficiency, and low latency AI hardware.
[0004] Although memristor technology has obvious advantages in theory, there are still many challenges in designing and fabricating a full-chip SOC chip based on in-memory computing in practical applications. For example, in order to detect the circuit characteristics of the memristor array, such as conductivity and durability, and to implement the application of memristors in deep learning algorithms, a large amount of preliminary exploration and verification of the memristor array are required. At present, the test and verification of the memristor array mainly rely on simulation models, which usually cannot meet the performance evaluation requirements of the memristor array under actual working conditions, especially in large-scale integration and high-speed operation environments. In addition, existing test systems often lack flexibility and are difficult to adapt to the changing technological development and application requirements.
[0005] Currently, there is a lack of an effective circuit system for testing and architecture verification of non-volatile memory array chips represented by memristors. Therefore, how to design a circuit test device based on FPGA to test the memristor array chip, so as to achieve digital logic verification of the chip architecture and algorithm deployment, has become a key issue in promoting the practical application of memristor technology.
[0006] Therefore, designing a flexible test device to test the in-memory array chip, so as to achieve digital logic verification of the chip architecture and algorithm deployment, is the key issue to promote the practical application of memristor technology. Summary of the Invention
[0007] In view of the above defects or improvement requirements of the prior art, the present invention provides a flexible pulsating in-memory array test device and its control method, aiming to realize the test of in-memory chips.
[0008] To achieve the above object, the present invention provides a flexible pulsating in-memory array test device, which includes a sub-board array composed of a plurality of sub-boards. Each sub-board includes a chip socket, a digital-to-analog conversion module, a driving module, a differential reading module, an analog-to-digital conversion module, a shift accumulation module, and first to fourth bidirectional interfaces;
[0009] The chip socket is used to place the in-memory array chip for matrix operation with the input vector;
[0010] The first bidirectional interface can perform bidirectional transmission of the input vector between the second bidirectional interface of the previous sub-board in the same row and the second bidirectional interface of its own sub-board;
[0011] The digital-to-analog conversion module is used to obtain the input vector of the first bidirectional interface and convert it into a voltage signal, and the driving module is used to connect the voltage signal to the chip and provide a driving current;
[0012] The third bidirectional interface can perform bidirectional transmission of the result vector between the fourth bidirectional interface of the previous sub-board in the same column and its own shift accumulation module, and the fourth bidirectional interface can perform bidirectional transmission of the result vector with the shift accumulation module;
[0013] The differential reading module includes two inverting amplifier circuits and an addition and subtraction circuit. The non-inverting input terminals of the two inverting amplifier circuits respectively sample the output currents of two adjacent columns in the chip, the inverting input terminals are connected to the clamping voltage, and the output terminals are connected to the addition and subtraction circuit for differential calculation and then output to the analog-to-digital conversion module for analog-to-digital conversion. The shift accumulation module is used to perform shift accumulation on the analog-to-digital conversion result and the digital signal of the third bidirectional interface or the fourth bidirectional interface to achieve aggregation and obtain the result vector.
[0014] Optionally, the resistor connected between the inverting input terminal and the output terminal of the inverting amplifier circuit is an adjustable resistor.
[0015] Optionally, the value range of the clamping voltage connected to the inverting amplifier circuit is 0.4V to 0.6V.
[0016] Optionally, the driving module is a voltage follower formed by an operational amplifier.
[0017] Optionally, the driving current provided by the driving module is 40 mA to 60 mA.
[0018] Optionally, the in-memory array chip is a non-volatile memory array chip.
[0019] Optionally, two of the chips can be placed in one daughter board, and matrix expansion is achieved through the two chips.
[0020] Optionally, the device further includes an FPGA programmable logic array for implementing data scheduling and control.
[0021] The present invention also provides a control method for a test device of an in-memory array capable of flexible pulsation. The test device is the test device described in any one of the above, and the control method includes:
[0022] Obtain N groups of input vectors in the to-be-input daughter board array. One group of input vectors corresponds to one row of the daughter board, and N is the number of rows of the daughter board array;
[0023] Perform data scheduling and control according to a set pulsation direction. The pulsation direction includes a row pulsation direction of input vectors between daughter boards in the same row and a column pulsation direction of result vectors between daughter boards in the same column. The performing data scheduling and control includes:
[0024] Control the input vectors to be transmitted from the previous daughter board to the next daughter board in a pipeline mode along the row pulsation direction through the first bidirectional interface and the second bidirectional interface, and control the result vectors to be transmitted from the previous daughter board to the next daughter board in a pipeline mode along the column pulsation direction through the third bidirectional interface and the fourth bidirectional interface. The transmission time of the input vectors between adjacent daughter boards and the processing time for the daughter board to receive the input vectors and obtain the result vectors are both one pulsation period. The daughter board obtains new input vectors within each pulsation period to complete matrix operations and obtains result vectors of other daughter boards for shift accumulation to obtain new result vectors.
[0025] Optionally, the pulsation direction is set as any one of the following:
[0026] Setting one: The input vectors pulsate to the right, and the result vectors pulsate downwards;
[0027] Setting two: The input vectors pulsate to the right, and the result vectors pulsate upwards;
[0028] Setting three: The input vectors pulsate to the left, and the result vectors pulsate upwards;
[0029] Setting four: The input vectors pulsate to the left, and the result vectors pulsate downwards;
[0030] Among them, left and right are respectively two opposite directions in which the rows extend in the daughter board array; up and down are respectively two opposite directions in which the columns extend in the daughter board array.
[0031] Generally speaking, compared with the prior art by the above technical solutions conceived in the present invention, the present invention mainly has the following beneficial effects:
[0032] 1. In the present invention, the test device is provided with a daughter board array. Each daughter board is provided with a chip socket, a digital-to-analog conversion module, a driving module, a differential readout module, an analog-to-digital conversion module, a shift and accumulation module, and first to fourth bidirectional interfaces. After connecting the memory chip to be tested, the matrix operation between the in-memory chip and the input vector can be realized, and the test of the memory chip can be realized.
[0033] 2. In the present invention, a daughter board array is provided. Signals are transmitted between the daughter board arrays through interfaces. The input vector can be transmitted in the row direction of the daughter board array, and the obtained result vectors of each daughter board can be transmitted in the column direction of the daughter board array. The result vectors of each daughter board can be gradually shifted and accumulated along the column direction. Therefore, calculations with different bit numbers can be performed as needed, and the flexibility is relatively strong. Moreover, the interfaces are set as bidirectional interfaces, and bidirectional data passing is set inside and between the daughter boards. The pulsation in different directions can be selected as needed, further improving the flexibility of the device. Moreover, based on the test device in the present invention, calculations with different precisions can be realized, and multi-bit width fusion can also be realized. Therefore, it can be flexibly applied to perform different operations. In addition, the differential readout module realizes the storage and operation of the positive and negative weights of the neural network by designing an inverting amplifier to read out and adding and subtracting operation circuits, which is fully adapted to the neural network calculation.
[0034] 3. Optionally, the resistor connected between the inverting input terminal and the output terminal of the inverting amplifier circuit is an adjustable resistor. In array chips with different resistance state distributions, the output current ranges vary greatly. By configuring different feedback resistors, current sampling under different resistance state distributions can be realized, further enhancing the flexibility of the test device. Description of the Drawings
[0035] Figure 1 is a schematic structural diagram of an in-memory array test device in an embodiment of the present invention;
[0036] Figure 2 is a schematic structural diagram of a daughter board in an embodiment of the present invention;
[0037] Figure 3 is a 1T1R array structure diagram of a chip crossbar in an embodiment of the present invention;
[0038] Figure 4It is a schematic structural diagram of a differential readout module in an embodiment of the present invention. Among them, (a) is a composition diagram of the differential readout module, and (b) is a schematic diagram of its adjustable resistor;
[0039] Figure 5 It is a schematic structural diagram of the part participating in matrix operation in a daughter board in an embodiment of the present invention;
[0040] Figure 6 It is a schematic structural diagram of a digital-to-analog conversion module and a driving module in an embodiment of the present invention;
[0041] Figure 7 It is a schematic diagram of a systolic principle in an embodiment of the present invention;
[0042] Figure 8 It is a schematic diagram of performing operations with different precisions in an embodiment of the present invention;
[0043] Figure 9 It is a schematic diagram of performing multi-bit width fusion in an embodiment of the present invention;
[0044] Figure 10 It is a schematic diagram of the computing acceleration of a neural network in an embodiment of the present invention. Detailed implementation manners
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0046] The present invention provides a test device for an in-memory array that can be flexibly systolic, as Figure 1 shown is a schematic structural diagram of a test device for an in-memory array in an embodiment of the present invention, as Figure 2 shown is a schematic structural diagram of a daughter board in an embodiment of the present invention.
[0047] Referring to Figure 1 and Figure 2 shown, the test device for the in-memory array includes a mother board, and the mother board is a daughter board array composed of a plurality of daughter board PEs. Each of the daughter board PEs includes a chip socket, a digital-to-analog conversion module DAC, a driving module DRV, a differential readout module Readout, an analog-to-digital conversion module ADC, a shift and accumulation module Shift&Add, a first bidirectional interface WS-port, a second bidirectional interface ES-port, a third bidirectional interface NS-port, and a fourth bidirectional interface SS-port.
[0048] The chip socket is used to place the in-memory array chip crossbar for matrix operations with the input vector. Each node in the in-memory array chip crossbar is used to store weight information. Define the input vector as vector A, the weight information stored in the in-memory array chip crossbar as matrix W, and the operation performed by the daughter board as: C = A * W, where C is the result vector. Specifically, the in-memory array chip is a non-volatile memory array chip. When the in-memory array is used in a deep learning model, it is configured to store the weights of a certain layer in the deep learning model. These weights are solidified after the model training is completed and will not change after the model is deployed to the hardware, and are tested through this test device.
[0049] The first bidirectional interface WS-port can perform bidirectional transmission of the input vector A between the second bidirectional interface ES-port of the previous daughter board in the same row and the second bidirectional interface ES-port of its own daughter board.
[0050] The digital-to-analog conversion module DAC is used to obtain the input vector A of the first bidirectional interface WS-port and convert it into a voltage signal.
[0051] The drive module DRV is used to connect the voltage signal to the chip crossbar and provide a drive current.
[0052] The third bidirectional interface NS-port can perform bidirectional transmission of the result vector C between the fourth bidirectional interface SS-port of the previous daughter board in the same column and its own shift-and-add module Shift&Add. The fourth bidirectional interface SS-port can perform bidirectional transmission of the result vector C with the shift-and-add module Shift&Add.
[0053] The differential readout module Read out is used to read the current information of every two adjacent columns of the chip crossbar and convert it into a differential voltage for storage.
[0054] The analog-to-digital conversion module ADC performs analog-to-digital conversion on the differential voltage.
[0055] The shift-and-add module Shift&Add is used to perform shift-and-add on the analog-to-digital conversion result and the digital signal of the third bidirectional interface NS-port or the fourth bidirectional interface SS-port to achieve aggregation and obtain the result vector C. For example, when performing matrix multiplication, each PE calculates a partial product, and the shifter and adder add these partial products in the correct order and position to obtain the final matrix multiplication result. This operation is performed in parallel, and the output of each PE is immediately involved in the final accumulation, thus achieving efficient matrix multiplication calculation.
[0056] Since the WS-port and ES-port are bidirectional interfaces, this design can achieve the left and right pulsations of the input vector A. Moreover, the NS-port and SS-port are also bidirectional interfaces, and this design can achieve the up and down pulsations of the result vector C, making the pulsating array architecture more flexible.
[0057] Therefore, the test device of the present invention has a total of four pulsation directions. During specific tests, any one of the following pulsation settings can be selected according to needs:
[0058] Setting 1: The input vector A pulsates to the right, and the result vector C pulsates downwards;
[0059] Setting 2: The input vector A pulsates to the right, and the result vector C pulsates upwards;
[0060] Setting 3: The input vector A pulsates to the left, and the result vector C pulsates upwards;
[0061] Setting 4: The input vector A pulsates to the left, and the result vector C pulsates downwards;
[0062] Among them, left and right are respectively two opposite directions of row extension in the daughter board array; up and down are respectively two opposite directions of column extension in the daughter board array.
[0063] Taking the case where the input vector A pulsates to the right and the result vector C pulsates downwards as an example: The first bidirectional interface WS-port transmits the input vector A to the DAC, and after being loaded onto the crossbar of the array chip through the DRV, the differential readout module Read out collects the current operation result and converts it into a differential voltage, and then conveys it to the ADC for quantization encoding. The ADC aggregates the encoded data with the result vector of the previous daughter board in the column direction through the Shift&Add module and then transmits the obtained new result vector to the fourth bidirectional interface SS-port, and is transmitted by the fourth bidirectional interface SS-port to the next daughter board in the column direction for aggregation to achieve the downward pulsation of the result vector; and the input vector A is transmitted to the next daughter board in the row direction through the first bidirectional interface WS-port and the second bidirectional interface ES-port to achieve the left pulsation of the input vector. The whole process completes the process of converting data from the digital domain to the analog domain, performing matrix operations in the analog domain, and then converting it back to the digital domain.
[0064] If the input vector A pulsates to the left, the input vector A is directly transmitted into the WS-port through the ES-port and enters the matrix operation circuit of the left daughter board.
[0065] If the result vector C pulsates upwards, the result vector C is directly transmitted to the NS-port and enters the shift and accumulation calculation of the upper daughter board.
[0066] In the present invention, a daughter board array is provided. Signals are transmitted between the daughter board arrays through interfaces. The input vector can be transmitted in the row direction of the daughter board array, and the resulting vectors obtained by each daughter board can be transmitted in the column direction of the daughter board array. The resulting vector of each daughter board can be gradually shifted and accumulated along the column direction. Therefore, calculations with different bit numbers can be performed as needed, and the flexibility is relatively high. Moreover, the interface is set as a bidirectional interface, and bidirectional data passing is set inside and between the daughter boards, and pulsation in different directions can be selected as needed, further improving the flexibility of the device.
[0067] As Figure 3 shown is the 1T1R array structure diagram of the chip crossbar in an embodiment of the present invention. Among them, the word line WL is parallel to the signal line SL, and the bit line BL is perpendicular to WL and SL. A voltage is loaded onto BL, and then a pulse is sent through the control module to turn on WL for matrix operations, and the output current of SL is sampled by the differential readout module.
[0068] As Figure 4 shown is the structural schematic diagram of the differential readout module in an embodiment of the present invention. The differential readout module includes two inverting amplifier circuits and an addition and subtraction circuit. The non-inverting input terminals of the two inverting amplifier circuits respectively sample the output currents of two adjacent columns in the chip, the inverting input terminals are connected to the clamping voltage, and the output terminals are connected to the addition and subtraction circuit for differential calculation and then output to the analog-to-digital conversion module.
[0069] Specifically, the value range of the clamping voltage connected to the inverting amplifier circuit is 0.4V to 0.6V. For example, it can be 0.5V.
[0070] With the above design of the differential readout module, the output terminal is voltage-clamped and current-sampled through the inverting amplifier circuit, and then the adjacent two columns of the in-memory array are differentially read out through an addition and subtraction circuit. The above readout circuit design scheme is fully adapted to neural network calculations. Specifically as follows: The weights of the neural network have positive and negative distributions, but based on the memristors in the memristive array, only conductance values can be stored. Two adjacent memristors are used as a differential pair, and the storage of positive and negative weights can be realized. For example: The memristor can store a low conductance state and a high conductance state, representing 0 and 1 respectively. The positive-phase memristor in the differential pair stores the high conductance state 1, and the negative-phase memristor stores the low conductance state 0, and the neural network weight represented is 1; the positive-phase memristor in the differential pair stores the low conductance state 0, and the negative-phase memristor stores the low conductance state 0, and the neural network weight represented is 0; the positive-phase memristor in the differential pair stores the low conductance state 0, and the negative-phase memristor stores the high conductance state 1, and the neural network weight represented is -1. The inverting amplifier converts the readout result of each column into a voltage and then inputs it into the addition and subtraction circuit. Among them, the positive column readout voltage Vp is input to the addition end of the addition and subtraction circuit, and the negative column readout voltage Vn is input to the subtraction end of the addition and subtraction circuit, and the result Vout is:
[0071] Vout = Vp - Vn
[0072] Therefore, through the inverting amplifier readout and the addition and subtraction operation circuit, the storage and operation of the positive and negative weights of the neural network are realized.
[0073] In a specific embodiment, the resistor connected between the inverting input terminal and the output terminal of the inverting amplifier circuit is an adjustable resistor R adj , for example, R adj There are 8 selection switches, which can control and select different sampling resistor gears according to the resistance state of the array chip, namely 100Ω, 200Ω, 400Ω, 800Ω, 1kΩ, 1.5kΩ, 2kΩ, 3kΩ. The adjustable resistor R adj can be implemented by a programmable resistor string and can be controlled and adjusted by an FPGA. In array chips with different resistance state distributions, the output current ranges vary greatly. By configuring different feedback resistors, current sampling under different resistance state distributions can be achieved. Adjacent two columns store information as a differential pair, and the voltages read out by two inverting amplifier circuits are input into the addition and subtraction circuit for differential readout. Finally, the addition and subtraction circuit subtracts the two readout voltages and outputs them to the ADC for sampling and quantization.
[0074] As Figure 5 shown is the structural schematic diagram of the part participating in matrix operation in the daughter board in an embodiment of the present invention. The input vector A with a data type of 8bit is converted into a voltage signal of 0V to 0.3V by the digital-to-analog conversion module and loaded into the memristor array. According to Kirchhoff's law I = V·G, matrix multiplication operation is performed, and the calculated analog current vector result is input into the differential readout module and the ADC. The ADC converts the current vector into a signed 8bit result vector C. The specific matrix operation formula is: C = A·W.
[0075] As Figure 6 shown is the structural schematic diagram of the digital-to-analog conversion module and the driving module in an embodiment of the present invention. The digital-to-analog conversion module DAC converts the input vector A into a voltage vector, and then loads the input vector A onto the memristor array chip through the driving module. Among them, the driving module is a voltage follower composed of an operational amplifier. The operational amplifier adopts class AB output, and the provided driving current is 40mA to 60mA. For example, it can be 50mA, which is sufficient for the testing of various non-volatile memory array chips. In the present invention, the driving module is set to provide a driving current to drive the in-memory array for testing.
[0076] In one embodiment, the test device has an FPGA circuit, and the scheduling and control of data are implemented based on the FPGA circuit. When different deep learning models need to be deployed or the parameters of existing models need to be adjusted, the configuration of the FPGA can be modified to adapt to new requirements without replacing the hardware. This flexibility enables the test device to adapt to changing computing tasks and technological developments. Specifically, the first to fourth bidirectional interfaces can all be bidirectional data transfer interfaces implemented by the FPGA. The FPGA can implement the left or right pulsation of the input vector A, and the FPGA can implement the up or down pulsation of the result vector C. Specifically, the DAC and ADC can also be directly connected to the FPGA, and the FPGA is used for the scheduling and control of the input vector A and the result vector C. The gates of the same column of select tubes in the in-memory array are connected together as the control terminal of the word line WL, and all WLs are directly connected to the FPGA for control. After the voltage vector in the memristor array is loaded, the WL is opened through the control pulse of the FPGA for analog matrix multiplication. Moreover, the RISC-V, address control, data cache, etc. in the test device can all be implemented and verified by the FPGA.
[0077] Correspondingly, the present invention also provides a control method for a test device of a flexible pulsating in-memory array, including:
[0078] Obtain N groups of input vectors in the to-be-input daughter board array, where one group of input vectors corresponds to one row of daughter boards, and N is the number of rows of the daughter board array;
[0079] Perform the scheduling and control of data according to the set pulsation direction, where the pulsation direction includes the row pulsation direction of the input vectors between the same row of daughter boards and the column pulsation direction of the result vectors between the same column of daughter boards: The performing of the scheduling and control of data includes:
[0080] Control the input vectors to be sequentially transmitted from the previous daughter board to the next daughter board in a pipeline mode along the row pulsation direction through the first and second bidirectional interfaces, and control the result vectors to be sequentially transmitted from the previous daughter board to the next daughter board in a pipeline mode along the column pulsation direction through the third and fourth bidirectional interfaces. The transmission time of the input vectors between adjacent daughter boards and the processing time for the daughter board to receive the input vectors and obtain the result vectors are both one pulsation period. The daughter board obtains new input vectors within each pulsation period to complete matrix operations and obtains the result vectors of other daughter boards for shift accumulation to obtain new result vectors.
[0081] Generally speaking, the flow of data in the pulsating array is pipeline-like, which means that each PE can immediately start processing the next set of data after processing the current input data. This immediate transfer mechanism significantly reduces the stagnation time of data between PEs because each PE is constantly receiving new inputs and generating outputs instead of waiting for the entire array to complete all calculations.
[0082] Specifically, when conducting the test, the test steps include:
[0083] Place the memristor array chip to be tested on the chip socket in the test device;
[0084] Load the input vector A into the BL terminal of the memristor chip through the FPGA;
[0085] Send a control pulse through the FPGA control pin to turn on all WLs, and the memristor array chip performs analog matrix operations;
[0086] The readout circuit reads out the current of the array and converts it into a voltage, which is then sampled and converted by the ADC;
[0087] The FPGA control interface receives the result vector from the previous daughter board and performs shift accumulation to obtain a new result vector.
[0088] As Figure 7 shown is the systolic schematic diagram in an embodiment of the present invention. Suppose there are 2 groups of input vectors, namely A00 and A01, and the mother board is a 2*2 daughter board array, and the stored weights are W00, W01, W10, and W11 respectively. A00 is input to the first row of daughter boards, and A01 is input to the second row of daughter boards. According to the data scheduling rules:
[0089] At T1: A00 is input to daughter board W00, and daughter board W00 calculates the result vector A00W00, and daughter boards W10 and W11 have no output;
[0090] At T2: A00 is transmitted from daughter board W00 to daughter board W01, and daughter board W01 calculates the result vector A00W01; A01 is input to daughter board W10, and daughter board W10 outputs the result vector A00W00 + A01W10, and daughter board W11 has no output;
[0091] At T3: A01 is transmitted from daughter board W10 to daughter board W11, and daughter board W11 outputs the result vector A00W01 + A01W11, and daughter board W10 has no output.
[0092] Based on this test device, operations with different precisions can be achieved. As Figure 8The figure shows a schematic diagram of performing operations with different precisions in an embodiment of the present invention. The input vector A has three bit widths: 1 bit, 4 bits, and 8 bits. The PEs in the architecture can store data of four bit width types: 1 bit, 2 bits, 3 bits, and 4 bits. The ADC in the PE has working modes of 1 bit, 4 bits, and 8 bits, so it can directly output data of three bit width types. By configuring the data types in the PE, matrix operations with different precisions can be achieved. For example, in the high-precision mode, the input vector A is configured with an 8-bit width, the memristive array in the PE stores weight data with a 4-bit width, and the ADC samples with 8 bits. In the low-precision mode, the input vector A is configured with a 1-bit width, the memristive array in the PE stores weight data with a 1-bit width, and the ADC samples with 1 bit. Lower power consumption can be achieved in the low-precision mode.
[0093] Based on this test device, multi-bit width fusion can also be achieved. For example, Figure 9 The figure shows a schematic diagram of multi-bit width fusion in an embodiment of the present invention. Assume that the highest precision of a single original PE is only 8 bits. Through multiple PEs for bit width expansion and fusion, a computing precision far exceeding 16 bits can be achieved. For example, the 8-bit data in the input vector A is differentially divided into a lower 4-bit data stream dataflow0 L01 and a higher 4-bit data stream dataflow1 H23, and are input into the systolic array as two input vectors. The 8-bit weight data is split into a lower 4-bit L01 and a higher 4-bit H23 and stored in four PEs respectively. Through the shift and accumulation module in the PE, the matrix operation results of LL+HL and LH+HH can be obtained in two output channels. Finally, the shift and accumulation of LL+HL and LH+HH are executed in the FPGA on the motherboard tile of the systolic array architecture, achieving a high-precision operation result.
[0094] Based on this test device, the acceleration of neural networks can also be achieved. For example, Figure 10 The figure shows a schematic diagram of the computational acceleration of a neural network in an embodiment of the present invention. The feature data is input from the FPGA into the PE. After being accelerated by matrix operations, the data results are transmitted back to the FPGA to perform various activations, and the activation functions can be implemented by the FPGA. The final processing results are transmitted from the FPGA back to the host computer. The above entire process realizes the acceleration verification of the systolic array architecture based on the non-volatile memory chip, which can be used for the verification and development of neural network chips.
[0095] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. It should be noted that the "in one embodiment", "for example", "for another example", etc. of the present invention are intended to illustrate the present invention, rather than to limit the present invention.
[0096] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A memory array test device with flexible pulsation, characterized in that: A sub-board array is formed by a plurality of sub-boards, each of which includes a chip seat, a digital-to-analog conversion module, a driving module, a differential readout module, an analog-to-digital conversion module, a shift-accumulation module and first to fourth bidirectional interfaces; The chip seat is used to place an in-memory array chip that performs matrix operations with input vectors; The first bidirectional interface can bidirectionally transmit input vectors with the second bidirectional interface of the previous daughter board in the same row and with the second bidirectional interface of the own daughter board; The digital-to-analog conversion module is used to obtain the input vector of the first bidirectional interface and convert it into a voltage signal, and the driving module is used to connect the voltage signal to the chip and provide a driving current; The third bidirectional interface can bidirectionally transmit the result vector with the fourth bidirectional interface of the previous sub-board in the same column and with the shift-and-accumulate module thereof, and the fourth bidirectional interface can bidirectionally transmit the result vector with the shift-and-accumulate module; The differential readout module includes two inverting amplifier circuits and an addition and subtraction circuit. The in-phase input terminals of the two inverting amplifier circuits respectively sample the output currents of two adjacent columns in the chip, the inverting input terminals are connected to the clamping voltage, and the output terminals are connected to the addition and subtraction circuits for differential calculation and then output to the analog-to-digital conversion module for analog-to-digital conversion. The shift and accumulation module is used to perform shift and accumulation on the analog-to-digital conversion result and the digital signal of the third bidirectional interface or the fourth bidirectional interface to achieve aggregation and obtain a result vector.
2. The flexible pulsation in-memory array test device according to claim 1, characterized in that: The resistor connected between the inverting input terminal and the output terminal of the inverting amplifier circuit is an adjustable resistor.
3. The flexible pulsation in-memory array test device according to claim 1, characterized in that: The clamping voltage connected to the inverting amplifier circuit has a value range of 0.4V to 0.6V.
4. The flexible pulsation in-memory array test device according to claim 1, characterized in that: The driving module is a voltage follower formed by an operational amplifier.
5. The flexible pulsation in-memory array test device according to claim 1, characterized in that: The driving module provides a driving current of 40 mA to 60 mA.
6. The flexible pulsation in-memory array testing device according to claim 1, characterized in that: The memory array chip is a non-volatile memory array chip.
7. The flexible pulsation in-memory array test device according to claim 1, characterized in that: Two chips can be placed in one daughter board, and matrix expansion can be achieved through the two chips.
8. The flexible pulsation in-memory array testing device according to any one of claims 1 to 7, characterized in that: The device also includes an FPGA programmable logic array for realizing data scheduling and control.
9. A method for controlling a memory array test device capable of flexible pulsation, characterized in that: The testing device is a testing device according to any one of claims 1 to 8, and the control method comprises: Obtain N groups of input vectors to be input into the sub-board array, where one group of input vectors corresponds to one row of sub-boards, and N is the number of rows of the sub-board array; The data is dispatched and controlled according to the set pulsation direction, wherein the pulsation direction includes the row pulsation direction of the input vector between the sub-boards in the same row and the column pulsation direction of the result vector between the sub-boards in the same column. The data dispatching and controlling includes: The control input vector is transmitted from the previous sub-board to the next sub-board in a pipeline mode in the row pulsation direction through the first bidirectional interface and the second bidirectional interface. The control result vector is transmitted from the previous sub-board to the next sub-board in a pipeline mode in the column pulsation direction through the third bidirectional interface and the fourth bidirectional interface. The transmission time of the input vector between adjacent sub-boards and the processing time of the sub-board receiving the input vector and obtaining the result vector are both one pulsation cycle. The sub-board obtains a new input vector in each pulsation cycle to complete the matrix operation and obtains the result vectors of other sub-boards for shift accumulation to obtain a new result vector.
10. The control method according to claim 9, characterized in that: The pulsation direction is any of the following settings: Setting 1: The input vector pulsates to the right, and the result vector pulsates downward; Setting 2: The input vector pulsates to the right, and the result vector pulsates upward; Setting 3: The input vector pulsates to the left, and the result vector pulsates upward; Setting 4: The input vector pulsates to the left, and the result vector pulsates downward; Wherein, leftward and rightward are two opposite directions in which rows in the daughter board array extend; upward and downward are two opposite directions in which columns in the daughter board array extend.
Citation Information
Cited By
Read-write multiplexing circuit applied to RRAM storage and calculation array and RRAM storage and calculation system
CN120612986A
Memristor differential double-subarray structure, convolution acceleration system and application
CN121306215A