A Configurable and Expandable Vector Matrix Multiplication Device and Working Method

By adopting the integrated storage and computing design and digital logic analog current convergence in neural network accelerators, efficient matrix multiplication operation is achieved, solving the lack of scalability and flexibility in existing accelerator designs, and improving performance and energy efficiency ratio.

CN115130058BActive Publication Date: 2025-06-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210672628.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-06-27
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The existing neural network accelerator design lacks scalability and flexibility, and it is difficult to meet the needs of complex neural network computing. Especially under the separate architecture of storage and computing, there are problems with large data handling and storage overhead.

Method used

Using a integrated storage and computing design, the matrix multiplication operation is realized through digital logic analog current convergence and digital-to-analog conversion, reducing data handling and reducing memory area, and a configurable and extended vector matrix multiplication device is designed, including data reception, depackaging, matrix storage, verification, multiplication, packaging, transmission and monitoring modules.

Benefits of technology

It realizes the efficiency and flexibility of matrix multiplication operations, reduces the demand for memory fetching and computing resources, improves performance and energy efficiency ratio, and supports the scalability of neural network computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130058B_ABST
    Figure CN115130058B_ABST
Patent Text Reader

Abstract

The present invention provides a configurable and extensible vector matrix multiplication device and working method. The device includes a data receiving module, a data unpacking module, a matrix storage module, a matrix input module, a matrix verification module, a matrix multiplication module, a data packing module, a data sending module, and a data monitoring module. The data receiving module is successively connected to the data unpacking module, the matrix input module, the matrix storage module, the data packing module, and the data sending module; the data unpacking module is also respectively connected to the matrix verification module and the matrix multiplication module; the matrix verification module and the matrix multiplication module are also respectively connected to the matrix storage module; the data monitoring module is respectively connected to the matrix input module, the matrix storage module, the matrix verification module, and the matrix multiplication module. The present invention adopts a design of integrating storage and computing, uses digital logic to simulate the convergence of current and digital-to-analog conversion to implement multiply-accumulate operations, reduces data transfer and shrinks the memory area, and can significantly reduce the area cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a configurable and extensible vector matrix multiplication device and working method, belonging to the field of digital signal processing of very large scale integrated circuits. Background Art

[0002] As neural networks become larger in scale and more diverse in type, the scale of neural networks is getting increasingly huge, and the resulting computational load cannot be underestimated. Therefore, various neural network accelerators emerge in an endless stream. However, the architecture design with separate storage and computing can never fundamentally eliminate the storage and data transfer overhead brought by such an astonishing amount of data. Along with the renewed popularity of the in-memory computing architecture and the proposal of related in-memory computing device designs, these problems may be solved by the in-memory computing architecture design. However, most current dedicated accelerator designs are targeted at specific networks and lack some scalability. The industry's requirements for neural network computing accelerators are not only limited to the acceleration of specific networks and low power consumption, but also have more requirements for flexibility, efficiency, configurability and other characteristics. Summary of the Invention

[0003] In order to meet the increasing requirements related to neural network computing devices, the present invention provides a configurable and extensible vector matrix multiplication device and working method.

[0004] The technical solution adopted by the present invention is as follows:

[0005] A configurable and extensible vector matrix multiplication device, the device includes a data receiving module, a data unpacking module, a matrix storage module, a matrix input module, a matrix verification module, a matrix multiplication module, a data packing module, a data sending module and a data monitoring module, characterized in that the data receiving module is sequentially connected to the data unpacking module, the matrix input module, the matrix storage module, the data packing module and the data sending module; the data unpacking module is also respectively connected to the matrix verification module and the matrix multiplication module; the matrix verification module and the matrix multiplication module are also respectively connected to the matrix storage module; the data monitoring module is respectively connected to the matrix input module, the matrix storage module, the matrix verification module and the matrix multiplication module;

[0006] The data receiving module is used to receive input data and send it to the data unpacking module;

[0007] The data unpacking module is used to decode the received data and send it to the matrix input module, the matrix verification module and the matrix multiplication module;

[0008] The matrix storage module is used to simulate an in-memory computing matrix array and store matrix weight information;

[0009] The matrix input module is used to write the received weights into the matrix storage module according to the specified rows and columns;

[0010] The matrix verification module is used to read out the weight information stored in the matrix storage module according to the specified rows and columns;

[0011] The matrix multiplication module is used to split the input excitation bit by bit and splice it in the order of the input excitation, and then send it to the matrix storage module;

[0012] The data packing module is used to encode the weights sent by the matrix verification module or the calculation results of the matrix multiplication module according to the belonging rows and columns;

[0013] The data sending module is used to send the data packet generated by the data packing module to the outside;

[0014] The data monitoring module is used to store the important information of each of the above modules.

[0015] The working method of the above-mentioned configurable and extensible vector matrix multiplication device of the present invention includes the following steps: the data receiving module receives the input mode data packet, the data unpacking module reads the input mode data packet, and after unpacking, sends the control signal for matrix input to the matrix input module, and the matrix input module performs the writing operation and writes the weights into the matrix storage module;

[0016] After all the weights are written, the data receiving module receives the verification mode data packet, the data unpacking module reads the verification mode data packet, and after unpacking, sends the row and column selection signals to the matrix verification module. The matrix verification module performs the reading operation, reads the weights from the matrix storage module, and then sends the weights to the data packing module; the data packing module packs and sends the received weights and their belonging rows and columns to the data sending module, and the data sending module sends the received data packet to the outside;

[0017] After all the weights or the specified weights are verified, the data receiving module receives the calculation mode data packet, the data unpacking module reads the calculation mode data packet, and after unpacking, sends the excitation and quantization coefficients to the matrix multiplication module. The matrix multiplication module receives the excitation, splits it by bit, and performs the column shift and accumulation operation in the matrix storage module. The obtained multiplication and accumulation result is sent to the data packing module. The data packing module packs and sends the received operation result and its belonging column to the data sending module, and the data sending module sends the received data packet to the outside.

[0018] Further, the data packet received by the data receiving module contains a pattern recognition code: two consecutive 10000000 represent that the data packet is an input mode data packet, two consecutive 10000001 represent that the data packet is a check mode data packet, and two consecutive 11111111 represent that the data packet is a calculation mode data packet; the pattern recognition code only appears in the packet header of a group of data packets.

[0019] Further, the data packet received by the data receiving module contains a checksum. The last 8-bit data of each group of data packets is the sum of all data except the two pattern recognition code packet headers and the checksum itself. If the bit width of this sum overflows 8 bits, the lower 8 bits are intercepted as the checksum. The checksum only appears at the end of a group of data packets.

[0020] The present invention adopts a design of computing in memory, uses digital logic to simulate the convergence of current and digital-to-analog conversion to implement multiply-accumulate operations, reduces data transfer and shrinks the memory area, and can significantly reduce the area cost. The present invention can parallelly implement the operation of the entire array matrix, and through in-memory computing, can achieve large-scale budgeting. The device and method of the present invention can reduce the memory access and computing resources required in actual operation when applied to a neural network, and improve the performance and energy efficiency ratio. Description of the Drawings

[0021] Figure 1 It is a structural block diagram of the device of the present invention.

[0022] Figure 2 It is a schematic diagram of the matrix input module. Among them, the global input control selects the input mode, controls the input loading module to write data into the column register respectively, and then controls the input to write this column into the matrix. I_sys_clk is the system clock, and I_rst_n is the system reset. I_lit_data_valid is that there is a signal at the input, I_lit_mode is used to specify whether the current operation is to store data in the register or write to the matrix, I_lit_weight is the weight data at the input, and I_lit_addr is the row where the current data is written or the column where data is written to the matrix; O_lit_vg, O_lit_vs, and O_lit_vd are control signals output to the matrix for writing data into the matrix; O_lit_ready is a handshake signal for sending whether the current input module is working to the previous-level module data unpacking module.

[0023] Figure 3Schematic diagram of the receiving module and the unpacking module. Among them, the receiving module receives an external input data packet and transmits it to the unpacking module for data extraction. I_rd_clk is the system clock, the same as I_sys_clk. I_rd_rst_n is the reset signal, the same as I_rst_n. I_wr_clk is the off-chip slow clock, and I_wr_rst_n is the reset signal. I_ctrl_mode is the system working mode, for external control. I_data is the data input connected to the host computer. wr_en is the input valid signal, and full is the full signal of the internal FIFO. When the full signal is pulled high, no more data packets are sent to the chip. rd_en is the read signal sent by the unpacking module to the receiving module, used to read data packets from the upper-level fifo. valid is the handshake signal sent by the receiving module to the unpacking module, meaning that the data sent to the unpacking module is valid. O_data is the data sent by the receiving module to the unpacking module. O_lit_data_valid is the handshake signal sent by the unpacking module to the input module, meaning that O_lit_mode, O_lit_addr, and O_lit_weight are valid. O_lit_weight is the data sent to the input module. O_lit_mode is the two input forms of the input module, which are writing data into the column register and mapping the full register to the matrix respectively. O_lit_addr is the positioning, which represents the specified row in a certain column to which the sent data belongs or the specified column in the matrix to which the data in the full register needs to be written in different data packets. O_lpf_data_valid is the handshake signal sent by the unpacking module to the verification module, meaning that O_lpf_col and O_lpf_row are valid. O_lpf_cow and O_lpf_col are the row and column signals sent to the matrix verification module respectively. O_data_valid is the handshake signal sent by the unpacking module to the matrix multiplication module, meaning that O_data and O_quan are valid. O_data is the data sent to the matrix multiplication module, and O_quan is the quantization coefficient sent to the matrix multiplication module. O_header_cnt is the number of data packets recognized by the unpacking module and sent to the detection module. O_success_cnt is the number of packets recognized successfully and with valid data packets by the unpacking module and sent to the detection module.

[0024] Figure 4There are a packing module and a sending module. Among them, after the packing module completes data encoding, it sends the data to the sending module, and the sending module transmits it outside the chip. I_ad_data_valid is a signal that can be selected from the verification module and the matrix multiplication module. When the on-chip module is the verification module, the valid signal of the verification module is selected; when the on-chip mode is the calculation mode, the valid signal of the matrix multiplication module is selected; I_data is the result read from the ADC; I_ctrl_mode is the current on-chip working mode; I_proof_valid is the current on-chip verification start signal, and I_cal_data_valid is the current on-chip calculation start signal; I_proof_row and I_proof_col are the row and column to which the read weight belongs; O_pack_ready is a signal for handshaking with the previous stage, used to indicate whether there is an unfinished process currently. Full is the fifo full signal of the sending module. When this signal is high, the packing module no longer sends data packets to the sending module; O_data2send is the data sent to the fifo, and O_data_valid is the valid signal of this data.

[0025] Figure 5It is a schematic diagram of the matrix verification module. Among them, I_proof_data_valid is the current input valid signal, used to determine whether I_lpf_col (column) and I_proof_row (row) are valid. O_lpf_vg, O_lpf_cd, and O_lpf_vs respectively correspond to row and column switches, used to control the read and write operations of the matrix; O_data_valid is the valid signal of the above signals; O_ready is the handshake signal sent to the previous stage, that is, the handshake signal with the unpacking module, used to indicate whether the current verification module is idle; I_pack_ready is the handshake signal of the subsequent module, that is, the data packing module, used to indicate whether the current data packing module can receive the output of the matrix verification module. IN_RO_IMIRR_CTR, IN_RO_BLP_CTR, IN_RO_BLN_CTR, and IN_CORR_IMIRR_CTR shown in the figure are dummy signals, always 0, IN_RO_MAC_CTR is the enable signal, used to control the matrix to read data, IN_RO_CAP0_CTR, IN_RO_CAP1_CTR, IN_RO_CAP2_CTR, IN_RO_CAP3_CTR, IN_RO_CAP4_CTR, IN_RO_CAP5_CTR, IN_RO_CAP6_CTR, IN_RO_CAP7_CTR are accumulation signals, which are dummy signals and always 0, IN_RO_RELU_CTR and IN_RO_RELU_COMP_CTR are activation function signals, used for quantization output, IN_RO_ADC_CS_CTR is the read signal, used to control the read time delay of the matrix data, and signals such as IN_RO_CVREF_CTR, IN_RO_VREF_CTR, and IN_RO_IREF_CTR are external reference input signals, always 0.

[0026] Figure 6 It is a schematic diagram of array row and column selection. Among them, WLi is the i-th row, BLj represents the j-th column, and W(i, j) represents the weight value of the i-th row and j-th column of the array.

[0027] Figure 7 It is a schematic diagram of the accumulation pipeline of the matrix storage module. Among them, cnt represents the n-th clock cycle, buffer_pos represents the positive column weight accumulation register, ad_pos_pipline1[i] is the weight value of the i-th row of the positive column, and it is also the storage register of the first stage of the positive column accumulation pipeline, ad_pos_pipline2_i is the storage register of the second stage of the positive column accumulation pipeline, ad_pos_pipline3_i is the storage register of the third stage of the positive column accumulation pipeline, and ad_pos_pipline4 is the storage register of the fourth stage of the positive column accumulation pipeline.

[0028] Figure 8 Schematic diagram for implementing matrix multiplication by the matrix storage module. Specific implementation mode

[0029] As Figure 1 As shown, a configurable and extensible vector matrix multiplication device of the present invention includes a data receiving module, a data unpacking module, a matrix storage module, a matrix input module, a matrix verification module, a matrix multiplication module, a data packing module, a data sending module, and a data monitoring module. The data receiving module is sequentially connected to the data unpacking module, the matrix input module, the matrix storage module, the data packing module, and the data sending module. The data unpacking module is also respectively connected to the matrix verification module and the matrix multiplication module; the matrix verification module and the matrix multiplication module are also respectively connected to the matrix storage module; the data monitoring module is respectively connected to the matrix input module, the matrix storage module, the matrix verification module, and the matrix multiplication module.

[0030] Among them, the data receiving module is used to receive the input data of different mode data packets encoded specifically sent by the off-chip slow clock and send it to the on-chip fast clock data unpacking module; the input end of the data receiving module is connected to the clock signal, the reset signal, the write enable signal, the read enable signal, and the write data, and the output end of the data receiving module is connected to the data unpacking module; further, the write enable signal and the write data of the data receiving module are connected to the system input port, and the read enable signal is connected to the output end of the data unpacking module.

[0031] The data unpacking module is used to decode the received input data and send the corresponding signals to the matrix input module, the matrix verification module, or the matrix multiplication module. The data unpacking module receives a group of 8-bit data packets. Among them, the first 2 data of a group of data packets are used for mode recognition, and the last data of a group of data packets is the checksum. The checksum is the sum of all data in a group of data packets except the 2 identification codes in the packet header and the checksum itself. If it overflows the 8-bit width, the lower 8 bits are intercepted. The input end of the data unpacking module is connected to the clock signal, the reset signal, the input data, the input data valid signal, the mode control signal, the matrix verification module handshake signal, the matrix multiplication module handshake signal, and the matrix input module handshake signal. The output ports of the data unpacking module are connected to the matrix input module, the matrix verification module, and the matrix multiplication module; the input data and the input data valid signal of the data unpacking are connected to the output end of the data receiving module, the mode control valid signal of the data unpacking module is connected to the system input port, and the input handshake signals of the data unpacking module are respectively connected to the output ends of the corresponding matrix input module, matrix verification module, and matrix multiplication module.

[0032] The matrix storage module simulates all functions of the matrix array, including input verification and calculation. The matrix storage module realizes matrix-vector multiplication or data reading by controlling the multiple shift accumulations of a column of data in the module through row and column switches; the matrix storage module calls four static random access memories with a bit width of 128 bits and a depth of 128 bits to splice into a static random access memory with a bit width of 512 bits and a depth of 128 bits.

[0033] The matrix input module is used to write a column of input data into the column register and write this column of data into the specified column of the storage matrix; its writing is a variable change process. For the data stored in each row of a column in the matrix, only one unit jumps in each cycle, specifically decreasing from 255 in sequence until the specified size is reached. The input end of the matrix input module is connected to the clock signal, reset signal, input valid signal, input data, and data address. The output end of the matrix input module is connected to the matrix storage module, and the output signals include row control switches and column control switches.

[0034] The matrix verification module is used to read the data of the specified row and column from the matrix storage to verify the writing validity and send it to the data packing module. The input end of the matrix verification module is connected to the clock signal, reset signal, handshake signal of the data packing module, row, column, and input valid signal. The output end of the matrix verification module is connected to the matrix storage module and the data packing module.

[0035] The matrix multiplication module splits the input excitation bit by bit and splices it in the order of the input excitation and then sends it to the matrix storage module. The matrix multiplication module includes a multiplication module and a multiplication delay module called by it. The multiplication delay module can control the time from the data input to this module to the output.

[0036] The data packing module is used to encode the weights sent by the matrix verification module or the calculation results of the matrix multiplication module according to the belonging rows and columns and send them to the data sending module. The input end of the data packing module is connected to the clock signal, reset signal, mode control signal, input data, input data valid signal, row, column, handshake signal of the data sending module, and matrix calculation start signal. The output end of the data packing module is connected to the data sending module. The output handshake signal of the data packing module is connected to the matrix verification module and the matrix multiplication module, and this data is pulled high when the data packing module is not working.

[0037] The data sending module is used to send the data packet generated by the on-chip fast clock domain data packing module to the external slow clock. The input end of the data sending module is connected to the clock signal, reset signal, write enable signal, and write data. The write enable signal and write data of the data sending module are connected to the output end of the data packing module.

[0038] The data monitoring module stores some important information of the above - mentioned modules in this module. The data monitoring module includes a monitoring module and a data monitoring memory module called by it. The data monitoring memory consists of a static random - access memory with a bit width of 16 bits and a depth of 512 bits, which is used to store globally important information.

[0039] In this embodiment, the data receiving module includes a slow - in - fast - out FIFO (First Input First Output). The FIFO input - output ports include a write clock, a read clock, a reset signal, a write enable signal, a data write port, a read enable signal, a data read port, a read data valid signal, and a full signal. The FIFO stores the written data using registers. Further, when the FIFO receives the write enable signal, it stores the written data in the registers in order. When the FIFO receives the read enable signal, it reads the data in the registers in order. If the FIFO is not empty, it simultaneously raises the output valid signal. If the FIFO is empty, the output valid signal is pulled low and invalid. Further, the write clock of the FIFO is a slow clock, the data write end is an external interface input, the write enable signal is an external interface input, the read clock of the FIFO is an on - chip fast clock, the data read end is connected to the data unpacking module, the data output valid signal end is connected to the data unpacking module, and the read enable signal is connected to the unpacking module and sent by the unpacking module. When the FIFO is full, the data receiving module no longer receives external writes.

[0040] The matrix input module has two working modes: 1. For each group of inputs received, it writes the only data in this group of inputs to the row corresponding to this data in the column register; 2. When a group of column registers is full, it writes this column of data to the matrix storage.

[0041] The matrix storage module receives the excitation split by bit and performs multiple shift - accumulation operations on the values in the matrix to achieve in - memory computing. The excitation received by the matrix storage module corresponds to the row - selection signal of the matrix storage module. As Figure 8 shown, taking a group of 8 - bit binary excitations as an example, the matrix multiplication module sends this excitation to the matrix storage module starting from the low - order bit. If it receives a '1' at the Nth bit, the weight value stored in the corresponding row of the matrix is left - shifted by N bits, and an accumulation operation is performed on each column. If the excitation received by the matrix storage module is '0', no shift - accumulation operation is performed. After all 8 - bit excitations are input, the value obtained after several shift - accumulation operations is the calculation result. The function of the matrix storage module is to simulate analog operations using digital logic. When using an in - memory computing device array as the matrix storage module, this process of shift - accumulation is actually a process of capacitor current convergence and then analog - to - digital conversion.

[0042] When performing column-wise accumulation in the matrix storage module, a pipelined structure is adopted to reduce the computational overhead and critical path. Taking the positive column of a certain calculation as an example, the structure is as Figure 7 shown. The accumulation time for each bit has 4 clock cycles: In the first clock cycle, the weights in the weight register are stored in the first-level pipeline register through the row selection signal; in the second clock cycle, every 16 first-level pipeline registers are added together, and the result is stored in the second-level pipeline register; in the third clock cycle, every two second-level pipeline registers are added together, and the result is stored in the third-level pipeline register; in the fourth clock cycle, the two third-level pipeline registers are added together, and the result is stored in the fourth-level pipeline register. The result stored in the fourth-level pipeline register is the final column accumulation value sum.

[0043] The usage method of the above configurable and extensible vector matrix multiplication device is as follows. The specific steps are:

[0044] (1) The host computer encodes the input. The encoding format of the data packet is described in detail below. The data packet should be set as multiple groups of data with a width of 8 bits. Among them, the first two 8 bits are specified as the packet header of the data packet for mode judgment. The host computer should send the data packet without interruption and only stop sending when the FIFO is full;

[0045] (2) Connect an external slow clock, and the on-chip required clock is generated by the internal phase-locked loop. Therefore, when the host computer sends the data packet, it should consider the power-on stabilization time of the phase-locked loop, that is, data can be sent only after the power-on is greater than 2 μs;

[0046] (3) Externally receive the result packaged by the chip. The host computer should receive the data packet on the chip without interruption, that is, any data packet sent by the on-chip sending module should be received.

[0047] Among them, the encoding format of the configurable and extensible vector matrix multiplication device is:

[0048] 1) Input: The matrix input module receives a set of data packets from the unpacking module at a time. The data packets include two 8-bit "10000000" packet headers used to determine the packet's belonging mode. The data packets also contain an 8-bit data value, an 8-bit position information with configurable bit width from 0 to 127, and a 2-bit input mode. The last data of the data packet is the checksum, which is the sum of all the above data excluding the packet headers. If the sum overflows 8 bits, the lower 8 bits are taken. The matrix input module has two input modes: 1. Store the value in the column register; 2. Write the information stored in the column register into the matrix storage. The specific steps are as follows: Continuously receive multiple rows of information of a specified column, that is, multiple data packets. In the state of storing in the column register, the position information contained in the data packet is the specified row; Receive a set of write data packets specifying which column of the column register in the previous step is to be written into the matrix. In this state, the position information contained in the data packet is the specified column of the matrix. The specific steps of writing the data stored in the register into the matrix storage are as follows: The reset value of the matrix storage module is 255. In each specific cycle, each element in the matrix only toggles once. In the state of writing the column register information into the matrix storage, the matrix input module controls the toggle switch of the values stored in the rows and columns of the matrix storage module, that is, the value toggle is specified by a set of coordinates; For a column of data to be written into the matrix, the corresponding bit of this column of the column switch is pulled high to 1, and the rest of the bits are 0. The row switches that have not reached the required stored value are always 1, that is, they will continue to toggle in the next cycle until the data writing is completed.

[0049] 2) Verification: The matrix verification module receives a set of data packets from the unpacking module at a time. The data packets include two 8-bit "10000001" packet headers used to determine the packet's belonging mode. The data packets also contain an 8-bit configurable row from 0 to 127 and an 8-bit configurable column from 0 to 127. The last data of the data packet is the checksum, which is the sum of all the above data excluding the packet headers. If the sum overflows 8 bits, the lower 8 bits are taken. When the matrix verification module receives the data packet for verifying the specified row and column, it opens the position switch, and reads the data at this position by the matrix storage module. There is a certain delay for data reading.

[0050] 3) Calculation: When calculating, the matrix storage module selects the calculation mode and receives the row and column selection signals from the matrix multiplication module. Each bit of the row and column selection signals is 1 or 0. The row selection signal has the number of row bits, and the column selection signal has the number of column bits. 1 means the corresponding row or column is selected, and 0 means not selected. Select the corresponding rows and columns according to the row and column selection signals, and accumulate by column to get the column number of accumulated sums. Since each excitation has 8 bits, after each accumulation is completed, a shift accumulation is required. For example, the accumulated sum of the second bit should be shifted left by one bit and accumulated with the accumulated sum of the first bit, and the same applies to the subsequent bits. After the eight-bit accumulation is completed, perform subsequent activation function operations, quantization, and rounding operations.

[0051] 4) Setting: When setting, the matrix storage module selects the setting mode, and the weight is incremented by 1 every 50,000 clock cycles (50M clock) until it accumulates to 255.

[0052] Example 1

[0053] In this specific implementation of this example, when the device is in the input mode, the system receives the mode control signal as "input", that is, I_ctrl_mode is the input mode "00". The system receives several groups of input mode data packets sent by the host computer, including a packet header (mode recognition code), weight, the address where the weight is to be stored, write mode (should first be the weight write register mode), checksum, etc. The data receiving module stores the above data in the FIFO. The data unpacking module sequentially reads the data packets from the data receiving module, accumulates the data except the packet header mode recognition code and the packet tail checksum, and compares it with the checksum. If they are equal and the mode recognition code is 10000000 continuously, that is, the input mode, the corresponding address, write mode, and weight are sent to the data input module. The data input module stores the received weight in the specified row of the column register according to the address match. After the system receives a complete column of weights, it receives a data packet including a packet header, address, write mode (at this time, it is to write the weights in the column register into the matrix), checksum, etc. The data receiving module stores the above data in the FIFO. After the data unpacking module reads the above information from the data receiving module, it checks the checksum of the data packet. If the checksum match is successful, the address and write mode are sent to the data input module, and the data input module writes the weight matrix in the column register to the address column. The specific steps are as follows: After the column is selected, the corresponding column switch is turned on, and the on time of the row switch for each row is the number of cycles of the input data size. As Figure 2 shown, the row and column switches O_lit_vg and O_lit_vd are represented by 0 and 1. 1 means that this row or column is selected, and the matrix storage data of this row and column is about to change. 0 means that this row and column are not selected. Each row and each column correspond to a 0 or 1 selection, and the data writing of each point of the matrix is controlled by the row and column combination. Each time the matrix performs a write operation, it writes a whole column. Then the switch of this column is turned on, that is, in the row and column selection signals where the matrix input module is connected to the matrix storage module, the column selection signal of this column is 1, and the column selection signals of the other columns are 0. According to the input value I_lit_weight of each row in this column, the duration of the row selection signal being 1 is controlled. For example, if the input of a certain row is 5 and the input of another row is 37, then after the selection signal in the first example lasts for 5 cycles of 1, it jumps back to 0. At this time, the writing of this row is completed, while the writing of the second row is still in progress, and the selection signal will not jump to 0 until 32 cycles later.

[0054] Example 2

[0055] In the verification mode, this embodiment reads out the weights configured in Embodiment 1. The specific implementation is as follows: When the system receives the mode control signal as "verification", that is, I_ctrl_mode is the verification mode "01". The system receives several groups of verification mode data packets sent by the host computer, which contain information such as a packet header (mode identification code), row, column, checksum, etc. The data reception module stores the above data in the FIFO. The data unpacking module sequentially reads out the data packets from the data reception module. If the mode identification code is the corresponding verification mode, that is, two consecutive 10000001, and the checksum is verified and the verification is successful, the corresponding row and column signals are sent to the data verification module and the data packing module. After the data verification module receives the specified row and column signals, as Figure 5 shown, the row-column switches O_lit_vg and O_lit_vd are represented by 0 and 1, where 1 means that this row or column is selected. At the same time, IN_RO_MAC_CTR is pulled high. After 35 clock cycles, IN_RO_ADC_CS_CTR is pulled high, and the matrix storage module reads out the weight stored at the specified position. In the verification mode, the row and column are uniquely specified. The matrix storage module receives the read instruction IN_RO_MAC_CTR and the read positions O_lit_vg and O_lit_vd, and reads out the weight at the specified position. The reading process is the process of assigning the weight itself to the output of the matrix storage module. This process is the process of capacitor current convergence in the analog memory-computation integrated device. After the above process is completed, the IN_RO_ADC_CS_CTR signal is pulled high, and the matrix storage module sends the read weight to the data packing module. After the data packing module receives the weight read from the matrix, it sends the weight and its corresponding row and column in the matrix to the data sending module. This operation is only performed when the FIFO in the data sending module is not full. The data sending module always sends the data stored in the FIFO to the outside of the system when the FIFO is not empty.

[0056] Embodiment 3

[0057] In the calculation mode, the weights completed using the configuration in Embodiment 1 are used for calculation in this embodiment. The specific implementation is as follows: When the system receives the mode control signal as "calculation", that is, I_ctrl_mode is the calculation mode "10". The system receives several groups of calculation mode data packets sent by the host computer, which contain information such as packet headers (pattern recognition codes), excitations, and quantization coefficients. The data receiving module stores the above data in the FIFO. The data unpacking module sequentially reads the data packets from the data receiving module. If the packet header is two consecutive 11111111, all of a column of excitations and quantization coefficients are sent to the calculation module. The matrix calculation module splits the received excitation into 8 bits and sequentially sends them to the matrix storage module. Specifically, when the matrix calculation module receives several 8-bit excitations, it first sends the least significant bit of these excitations to the matrix storage module. The signal composed of these least significant bits corresponds to the row signal of the matrix storage module. After receiving the row signal, the matrix storage module selects the corresponding row and then sequentially selects the columns. Each bit of the row and column selection signals is 1 or 0. The row selection signal has the number of row bits, and the column selection signal has the number of column bits. 1 means selecting the corresponding row or column, and 0 means not selecting. According to the row and column selection signals, the corresponding rows and columns are selected and accumulated column by column to obtain the column number of accumulated sums. Since each excitation has 8 bits, after each accumulation, a shift accumulation is required. As Figure 8 shown, for example, the accumulated sum of the second bit should be shifted left by one bit and accumulated with the accumulated sum of the first bit, and the same applies to the subsequent bits. To reduce the calculation overhead and critical path, taking the positive column of a certain calculation as an example, the structure is as Figure 7 shown. Each bit of the accumulation time has 4 clock cycles: In the first clock cycle, the weights in the weight register are stored in the first-level pipeline register through the row selection signal; in the second clock cycle, every 16 first-level pipeline registers are added together, and the result is stored in the second-level pipeline register; in the third clock cycle, every two second-level pipeline registers are added together, and the result is stored in the third-level pipeline register; in the fourth clock cycle, two third-level pipeline registers are added together, and the result is stored in the fourth-level pipeline register. The result stored in the fourth-level pipeline register is the final column accumulated value sum. After the 8-bit accumulation is completed, subsequent activation function operations, quantization, and rounding operations are performed. For the activation function operation, only need to judge whether the most significant bit of the accumulated sum is 1. If it is 1, set the accumulated sum to zero; during quantization, only need to shift the accumulated sum after the activation function operation to the right by the corresponding number of bits; during rounding, only need to consider whether the first bit overflowed during the right shift during quantization is 1. If it is 1, add 1 after the right shift. All of the above three functions can be selected to be turned on or off through the corresponding switches, and the quantization coefficient size can be selected for quantization.

[0058] Embodiment 4

[0059] In order to achieve the expansion of multiple devices in this embodiment, a downward-compatible network mode is adopted, that is, the network structure that is strictly smaller than (both the number of rows and columns are smaller than) the default network structure of the device specified in the present invention is compatible. The specific implementation method is as follows: control the number of bits of the row and column selection signals to the number of bits corresponding to the number of rows and columns, and set the extra bits to zero, that is, do not select them.

Claims

1. A configurable and extensible vector matrix multiplication device, which includes a data receiving module, a data unpacking module, a matrix storage module, a matrix input module, a matrix verification module, a matrix multiplication module, a data packing module, a data sending module, and a data monitoring module, characterized in that, The data receiving module is successively connected to the data unpacking module, the matrix input module, the matrix storage module, the data packing module, and the data sending module; the data unpacking module is also respectively connected to the matrix verification module and the matrix multiplication module; the matrix verification module and the matrix multiplication module are also respectively connected to the matrix storage module; the data monitoring module is respectively connected to the matrix input module, the matrix storage module, the matrix verification module, and the matrix multiplication module; The data receiving module is used to receive input data and send it to the data unpacking module; The data unpacking module is used to decode the received data and send it to the matrix input module, the matrix verification module, and the matrix multiplication module; The matrix storage module is used to simulate a memory-computation integrated matrix array and store matrix weight information; The matrix input module is used to write the received weights into the matrix storage module according to the specified rows and columns; The matrix verification module is used to read out the weight information stored in the matrix storage module according to the specified rows and columns; The matrix multiplication module is used to split the input excitation bit by bit and splice it in the order of the input excitation and then send it to the matrix storage module; The data packing module is used to encode the weights sent by the matrix verification module or the calculation results of the matrix multiplication module according to the belonging rows and columns; The data sending module is used to send the data packet generated by the data packing module to the outside; The data monitoring module is used to store the important information of each of the above modules.

2. The configurable and extensible vector matrix multiplication device according to claim 1, characterized in that The matrix storage module realizes matrix-vector multiplication or data reading by controlling the multiple shift accumulations of a column of data in the module through row-column switches; the matrix storage module calls four static random access memories with a bit width of 128 bits and a depth of 128 bits to be spliced into a static random access memory with a bit width of 512 bits and a depth of 128 bits.

3. The configurable and extensible vector matrix multiplication device according to claim 1, characterized in that, The matrix input module is used to write a column of input data into the column register and write this column of data into the specified column of the matrix storage module; the process of writing it into the matrix storage module is a quantitative change process, and for the data stored in each row of a column in the matrix, only one unit jumps in each cycle.

4. The configurable and extensible vector matrix multiplication device according to claim 1, wherein The matrix multiplication module includes a multiplication module and a multiplication delay module called by it. The multiplication delay module is used to control the time from when the data enters the matrix multiplication module to when it is output, and this time is the analog-to-digital conversion time in analog operations.

5. The configurable and extensible vector matrix multiplication device according to claim 1, characterized in that, The matrix multiplication module adopts the method of excitation decomposition, splits the excitation bit by bit and splices it in the order of the input excitation and then sends it to the matrix storage module.

6. The configurable and extensible vector matrix multiplication device according to claim 1, characterized in that The data monitoring module includes a monitoring module and a data monitoring memory module called by it. The data monitoring memory module consists of a static random access memory with a bit width of 16 bits and a depth of 512 bits.

7. The working method of a configurable and extensible vector matrix multiplication device as described in claim 1, characterized in that, The steps of this method include: the data receiving module receives an input mode data packet, the data unpacking module reads the input mode data packet, after unpacking, sends a control signal for matrix input to the matrix input module, and the matrix input module performs a write operation to write weights into the matrix storage module; After all weights are written, the data receiving module receives a check mode data packet. The data unpacking module reads the check mode data packet, unpacks it, and sends row and column selection signals to the matrix check module. The matrix check module performs a read operation, reads the weights from the matrix storage module, and then sends the weights to the data packing module. The data packing module packs the received weights and their corresponding rows and columns and sends them to the data sending module. The data sending module sends the received data packet to the outside. After all weights or specified weights are verified, the data receiving module receives a calculation mode data packet. The data unpacking module reads the calculation mode data packet, unpacks it, and sends excitation and quantization coefficients to the matrix multiplication module. The matrix multiplication module receives the excitation, splits it by bit, and performs a column shift and accumulation operation in the matrix storage module. The resulting multiply-accumulate result is sent to the data packing module. The data packing module packs the received operation result and its corresponding column and sends them to the data sending module. The data sending module sends the received data packet to the outside.

8. The working method according to claim 7, characterized in that, The data packets received by the data receiving module contain a pattern recognition code: two consecutive 10000000 represent that the data packet is an input mode data packet, two consecutive 10000001 represent that the data packet is a check mode data packet, and two consecutive 11111111 represent that the data packet is a calculation mode data packet. The pattern recognition code only appears at the header of a group of data packets.

9. The working method according to claim 8, characterized in that, The data packets received by the data receiving module contain a checksum. The last 8-bit data of each group of data packets is the sum of all data except the two pattern recognition code headers and the checksum itself. If the width of this sum overflows 8 bits, the lower 8 bits are intercepted as the checksum. The checksum only appears at the end of a group of data packets.

Citation Information

Patent Citations

  • Extensible fixed-point number matrix multiply-add operation in-memory calculation structure and method

    CN110427171A

  • Self-supervised learning acceleration system and method based on storage and calculation integrated device array

    CN110647983A