Chip architecture for AI computing based on NVM

By combining NPU and NVM for digital domain calculations in the AI chip, the problems of inflexible analog signal transmission and noise error are solved, and efficient, flexible and reliable AI computing is achieved, reducing external storage power consumption.

CN113127407BActive Publication Date: 2025-08-01NANJING UCUN TECHNOLOGY INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110541351.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-18
Publication Date
2025-08-01
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

In the existing NVM-based AI computing solution, analog signals are not flexible when transmitted between neural network layers, analog computing array structure is rigid, noise and error affect model reliability and calculation accuracy, and external storage speed bottlenecks and power consumption are high.

Method used

Using a digital domain computing architecture combining NPU and NVM, the neural network weight parameters are stored in the internal NVM array of chip. The NPU and NVM array are controlled by MCU for AI calculation, and combined with high-speed data reading channels and data conversion units, it supports flexible neural network structure.

Benefits of technology

Improves computing flexibility and reliability, breaks through external storage speed bottlenecks, reduces power consumption, and improves read accuracy and implementability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113127407B_ABST
    Figure CN113127407B_ABST
Patent Text Reader

Abstract

A chip architecture for AI computing based on NVM provided by the present invention includes an NVM array, an external interface module, an NPU, and an MCU that are communicatively connected via a bus; the NPU and NVM are combined to perform AI neural network computing. The weight parameters of the neural network are digitally stored in the NVM array. The MCU receives external AI operation instructions to control the NPU and the NVM array to implement neural network computing. The MCU controls the NVM array to load the weight parameters of the neural network stored therein, and performs AI computing through the running program and the neural network model. Compared with various existing memory-computation schemes that use NVM for analog computing, the digital storage and operation method has a flexible operation structure, good reliability, high precision, and high reading accuracy. Therefore, while breaking through the bottleneck of the storage speed of off-chip NVM and reducing the external input power consumption, the present invention also has high implementability, flexibility, and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI (Artificial Intelligence) technology, and particularly to a chip architecture for AI computing based on NVM (non-volatile memory). Background Art

[0002] The algorithms of AI are inspired by the structure of the human brain. The human brain is a network of complex connections of a large number of neurons. Each neuron receives information by connecting a large number of other neurons through a large number of dendrites. Each connection point is called a synapse. When the external stimuli accumulate to a certain extent, a stimulus signal is generated and transmitted through the axon. The axon has a large number of terminals, which are connected to the dendrites of a large number of other neurons through synapses. It is such a network composed of neurons with simple functions that realizes all human intelligent activities. Human memory and intelligence are generally considered to be stored in the different coupling strengths of each synapse.

[0003] Since the 1960s, the neural network algorithm has been used to imitate the function of neurons with a function. The function accepts multiple inputs from other neurons, each input has a different weight, and the output is the sum of the products of each input and the corresponding neuron connection weight. The function output is then input to other neurons in the next layer to form a neural network.

[0004] Common AI chips optimize matrix parallel computing for network computing in algorithms. However, because AI computing requires extremely high storage read bandwidth, the architecture that separates the processor from memory and storage encounters a bottleneck in read speed and is also limited by the external storage read power consumption. The industry has begun to widely research the architecture of in-memory-computing.

[0005] Currently, the in-memory-computing solutions using NVM all use NVM to store the weights in a neural network in the form of analog signals, and implement the calculation of the neural network through the method of analog signal addition and multiplication. For specific examples, see the Chinese patent application with the publication number CN109086249A. Although there have been many research results in such solutions, there are still difficulties in practical applications. Because practical neural networks basically have many layers and very complex connection structures, it is very inconvenient for analog signals to be transmitted between layers and for various signal processing during the implementation of neural network computing. The analog computing array structure is rigid and is not conducive to supporting flexible neural network structures. In addition, various noises and errors in the storage, reading, writing, and calculation of analog signals will affect the reliability of the stored neural network model and the accuracy of the calculation is limited. Summary of the Invention

[0006] The object of the present invention is to provide a chip architecture for AI computing based on NVM, which overcomes the deficiencies of existing in-memory computing solutions that use NVM to store the weights in a neural network in the form of analog signals. When the analog signals are used for the transfer between layers and various signal processing during the implementation of neural network computing, it is very inconvenient. The analog computing array structure is rigid and not conducive to supporting flexible neural network structures. In addition, various noises and errors in the storage, reading, and writing of analog signals and during computing will affect the reliability of the stored neural network model and the accuracy of computing. The present invention provides a chip architecture for AI computing based on NVM, which has better flexibility, high implementability, and reliability while breaking through the speed bottleneck of external storage and reducing the power consumption of external inputs.

[0007] To achieve the above object, the present invention provides a chip architecture for AI computing based on NVM, which includes an NVM array, an external interface module, an NPU (embedded neural network processor), and an MCU (Microcontroller Unit) communicatively connected via a bus;

[0008] The NVM array is used for on-chip storage of the weight parameters of a digitalized neural network, the program run by the MCU, and the neural network model;

[0009] The NPU is used for digital domain acceleration computing of the neural network;

[0010] The external interface module is used for receiving external AI operation instructions, input data, and outputting the results of AI computing outward;

[0011] The MCU is used for executing the program based on the AI operation instructions to control the NVM array and the NPU to perform AI computing on the input data to obtain the results of the AI computing.

[0012] In the chip architecture provided by this solution, the NPU and the NVM are combined to perform AI neural network computing. Among them, the weight parameters of the neural network are digitally stored in the NVM array inside the chip, and the neural network computing is also digital domain computing. Specifically, it is realized by the MCU controlling the NPU and the NVM array based on external AI operation instructions. The MCU controls the NVM array to include loading the weight parameters of the neural network, the program run by the MCU, and the neural network model stored therein to perform AI computing. Compared with existing memory-computation solutions that use NVM for analog computing, the digital storage and operation method has a flexible operation structure. The information stored in the NVM has better reliability, higher precision, and higher reading accuracy compared with the multi-level storage of analog signals. Therefore, this solution has high implementability, flexibility, and reliability while breaking through the speed bottleneck of off-chip NVM storage and reducing the power consumption of external inputs.

[0013] Furthermore, the chip architecture further includes a high-speed data reading channel, and the NPU reads the weight parameters from the NVM array through the high-speed data reading channel.

[0014] In addition to the on-chip bus, this solution also sets up a high-speed data reading channel between the NPU and the NVM array to support the bandwidth requirement for the high-speed reading of the weight parameters (i.e., weight data) of the neural network when the NPU performs digital domain operations.

[0015] Furthermore, the NVM array is provided with N read channels, where N is a positive integer. In one read cycle, the read channels read N bits of data in total, and the NPU is used to read the weight parameters from the NVM array through the read channels via the high-speed data reading channel.

[0016] This solution sets up N read channels. Preferably, N ranges from 128 to 512. In one read cycle (usually 30 - 40 nanoseconds), N bits of data can be read. The NPU reads the weight parameters of the neural network from the NVM array through the read channels via the high-speed data reading channel. This bandwidth is much higher than the reading speed supported by off-chip NVM and can support the parameter reading speed requirements for common neural network inference calculations.

[0017] Furthermore, the bit width of the high-speed data reading channel is m bits, where m is a positive integer; the chip architecture further includes a data conversion unit, and the data conversion unit includes a cache module and a sequential reading module. The cache module is used to cache the weight parameters output through the read channels in sequence by cycle. The capacity of the cache module is N * k bits, where k represents the number of cycles; the sequential reading module is used to convert the cached data in the cache module into m-bit width and then output it to the NPU through the high-speed data reading channel, where N * k is an integer multiple of m.

[0018] This solution also includes a data conversion unit. In the case where the number of read channels is inconsistent with the bit width of the high-speed data reading channel and / or the frequencies are asynchronous, the data conversion unit is used to convert the data into a combination of data with the same bit width as the high-speed data reading channel, usually a combination of words with a small width (such as 32 bits). The NPU reads data from the data conversion unit through the high-speed data reading channel at its own clock frequency (which can reach above 1 GHz).

[0019] The data conversion unit provided by this solution includes a cache containing N*k bits and a sequential reader that outputs m bits at a time, where N*k is an integer multiple of m; the read channel is connected to the NVM array and can output N bits in each cycle, and the cache can store data for k cycles; the width of the high-speed data read channel is m bits. Among them, the high-speed data read channel can include read / write instructions (CMD) and reply (ACK) signals and is connected to the NVM array read control circuit. After the read operation is completed, the ACK signal notifies the high-speed data read channel and can also notify the on-chip bus at the same time. The high-speed data read channel inputs the data in the cache into the NPU asynchronously in multiple times through the sequential read module.

[0020] Furthermore, the chip architecture also includes SRAM (Static Random-Access Memory), and the SRAM is communicatively connected to the NVM array, the external interface module, the NPU, and the MCU through the bus; the SRAM is used to cache the data during the execution of the program by the MCU, the data during the operation of the NPU, and the input / output data of the neural network model operation.

[0021] The chip architecture provided by this solution includes an embedded SRAM, which is used as a cache required for the operation and calculation of the internal system of the chip and is used to store input / output data, intermediate data generated by calculations, etc. Specifically, it can include the cache during the execution of the program by the MCU, storing the executable program during the operation of the MCU, system configuration parameters, calculation network structure configuration parameters, etc.; the operation cache of the NPU, and storing the input / output data during the operation of the neural network model.

[0022] Furthermore, multiple neural network models are stored in the NVM array, and the AI operation instructions include algorithm selection instructions, and the algorithm selection instructions are used to select one of the multiple neural network models as the algorithm for performing AI calculations.

[0023] The neural network models in this solution are digitally stored in the NVM array and can be multiple according to the number of application scenarios. For the case where multiple neural network models correspond to multiple application scenarios, the MCU can flexibly select any one of the pre-stored neural network models for AI calculation according to the externally input algorithm selection instructions, overcoming the problem that the analog computing array structure in the existing memory-computation integrated solution is rigid and not conducive to supporting flexible neural network structures.

[0024] Further, the NVM array adopts one of a flash memory process, an MRAM (Magnetoresistive Random Access Memory) process, a RRAM (resistive random access memory) process, an MTP (Multiple Time Programming) process, and an OTP (One Time Programming) process, and / or the interface standard of the external interface module is at least one of SPI (Serial Peripheral Interface), QPI (Quad SPI), and a parallel interface.

[0025] Further, the MCU is further configured to receive, through the external interface module, a data access instruction from the outside for operating the NVM array, and the MCU is further configured to complete logical control of basic operations on the NVM array based on the data access instruction.

[0026] Further, the NVM array adopts one of a SONOS (a flash memory process) flash memory process, a Floating Gate (a flash memory process) flash memory process, and a Split Gate (a flash memory process) flash memory process, and the interface standard of the external interface module is SPI and / or QPI;

[0027] The data access instruction is a standard flash memory operation instruction; the AI operation instruction and the data access instruction adopt the same instruction format and rules; the AI operation instruction includes an operation code, the AI operation instruction further includes an address part and / or a data part, and the operation code of the AI operation instruction is different from the operation code of the standard flash memory operation instruction.

[0028] The chip architecture provided by this solution is an improvement on the basis of the traditional flash memory chip architecture. Specifically, an MCU and an NPU are embedded inside the flash memory chip and are communicatively connected through an on-chip bus. The on-chip bus can be an AHB (Advanced High Performance Bus), or other communication buses that meet the requirements, which are not limited here. In this solution, the NPU and the NVM are combined, that is, both computing and storage are on-chip. Among them, the weight parameters of the neural network are digitally stored in the NVM array, and the neural network calculation is also digital domain calculation. Specifically, it is realized by the MCU controlling the NPU and the NVM array based on an external AI operation instruction. Thereby, while breaking through the storage speed bottleneck of off-chip NVM and reducing the external input power consumption, it also has high feasibility, flexibility, and reliability.

[0029] This solution realizes digital operations on the NVM array based on the MCU, which can specifically include basic operations such as reading, writing, and erasing of flash memory. The external data access instructions and external interfaces can adopt the standard flash memory chip format, making it easy for the chip to be flexibly and simply applied. The MCU embedded in this solution serves as the logical control unit of the NVM, replacing the logical state machine in the standard flash memory, simplifying the chip structure and saving chip area.

[0030] In addition to storing the neural network model, the weight parameters, and the programs for the internal system operation of the chip, the NVM array in this solution can also be used to store externally input data that is not limited to AI computing-related data, that is, it can also be used to store other externally input data related to AI computing and externally input data unrelated to AI computing. The unrelated data specifically includes information such as system parameters, configurations, and / or codes of external devices or systems; in addition to operations such as reading, writing, and erasing of the neural network model, the weight parameters, and the programs for the internal system operation, the basic operations also include operations such as direct storage reading, writing, and erasing of the stored externally input data in the NVM array.

[0031] The instructions for direct operation of the NVM and the instructions for AI computing processing in this solution adopt the same instruction format and rules. Taking the SPI and QPI interfaces as an example, based on the traditional SPI and QPI flash memory operation instructions op_code, select the op_code not used by the flash memory operation to express AI instructions, transmit more information in the address part, and implement AI data transfer during the data exchange cycle. Only need to expand the instruction decoder to realize the reuse of the interface, and add several status registers and configuration registers to realize AI computing.

[0032] Furthermore, the chip architecture also includes a DMA (Direct Memory Access) channel, and the DMA channel is used for external devices to directly read and write the SRAM.

[0033] The positive and progressive effects of the present invention are as follows:

[0034] A chip architecture for AI computing based on NVM provided by the present invention combines an NPU and NVM for AI neural network computing. The weight parameters of the neural network are digitally stored in the NVM array inside the chip, and the neural network computing is also digital-domain computing. Specifically, it is realized by an MCU controlling the NPU and the NVM array based on external AI operation instructions. The MCU controls the NVM array to load the weight parameters of the neural network stored therein, the program run by the MCU, and the neural network model for AI computing. Compared with various existing memory-computation solutions using NVM for analog operations, the digital storage and operation method has a flexible operation structure. The information stored in NVM has better reliability, higher precision, and higher reading accuracy compared to the multi-level storage of analog signals. Therefore, while breaking through the bottleneck of the storage speed of off-chip NVM and reducing the external input power consumption, the present invention also has high feasibility, flexibility, and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is a neuron diagram in the AI algorithm of the prior art;

[0037] Figure 2 It is a three-layer neural network diagram of the prior art;

[0038] Figure 3 It is a convolutional neural network diagram of the prior art;

[0039] Figure 4 It is a schematic diagram of the prior art for AI computing by adding circuits in a standard NVM array;

[0040] Figure 5 It is a schematic diagram of a chip architecture for AI computing based on NVM of the present application;

[0041] Figure 6 It is a schematic diagram of the data conversion unit of the chip architecture of the present application;

[0042] Figure 7 It is a flowchart of the instruction operation for the chip architecture of the present application to call NVM read and write operations;

[0043] Figure 8 It is a flowchart of running an AI operation instruction based on the chip architecture of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0045] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0047] This application first explains the basic architectures and relationships of neural networks, artificial intelligence, non-volatile storage, and in-memory computing.

[0048] As mentioned above: Artificial intelligence (AI) algorithms are generated by imitating the human brain structure. Through the synapses between neurons, they are connected to the dendrites of a large number of other neurons, forming a neural network with simple functions, realizing all human intelligent activities. Human memory and intelligence are generally considered to be stored in the different coupling strengths of each synapse.

[0049] The neural network algorithms that emerged in the 1960s use a function to imitate the function of neurons. The function accepts multiple inputs, each input has a different weight, and the output is the sum of the products of each input and its weight, as shown in the exemplary neuron diagram of the AI algorithm Figure 1 shown. The process of learning and training is to adjust each weight. The function output is sent to many other neurons, forming a network. Such algorithms have achieved rich results and have been widely used. Practical neural networks all have a hierarchical structure. Neurons within the same layer do not communicate with each other. The input of each neuron is connected to the outputs of multiple or all neurons in the previous layer, as shown in the three-layer neural network diagram Figure 2 shown, which includes an input layer, a hidden layer, and an output layer. The input layer and the hidden layer include 784 and 15 neurons respectively. There are different connection methods between different layers of the neural network, which is a fully connected network.

[0050] More commonly, as shown in Figure 3The shown convolutional neural network diagram. This convolutional neural network has a two-dimensional structure (image) at both the input and output, and only has connections between adjacent points.

[0051] However, a practical neural network often consists of many layers, including one or more of a convolutional layer, a layer for reducing the image size, and a fully connected layer to form a network structure.

[0052] Non-volatile storage:

[0053] Non-volatile memory (NVM) is a semiconductor storage medium that can retain its content after power-off. Common NVMs include flash memory, EEPROM (Electrically Erasable Programmable read only memory), MRAM, RRAM, FeRAM (Ferroelectric RAM), MTP, OTP, etc. The most widely used NVM currently is flash memory. Among them, the NOR flash memory structure has higher reliability and faster read speed compared to the NAND flash memory structure, and is often used for storing system code, parameters, algorithms, etc. In specific applications, the system can use an external stand-alone NVM or an embedded NVM embedded in the system. Embedded NVM is usually compatible with the CMOS (Complementary Metal-Oxide-Semiconductor) semiconductor process and can be integrated with the logic computing chip, and has a relatively fast read speed within the system.

[0054] Compared with other NVMs, current flash memory has advantages in both cost and capacity. Multiple flash memory technologies on the current market already have the ability to store multiple bits. The flash memory has a relatively slow erase and write speed (millisecond level), but the read speed is much faster (nanosecond level). The read speed of flash memory and other NVMs can support the high bandwidth required for neural network computing.

[0055] In-Memory-computing

[0056] Because AI computing requires extremely high storage bandwidth, the architecture that separates the processor from the memory / storage encounters a bottleneck in insufficient read speed. The industry has begun to widely study the architecture of in-memory computing, that is, computing within storage. For example, such as Figure 4The schematic diagram of the prior art that adds circuits in a standard NVM array for AI computing, that is, the architecture for neural network computing, which uses non-volatile memory to store the weights required for neural network computing, uses analog circuits for vector multiplication, and large-scale multiplication and addition can be performed in parallel, which can improve the operation speed and save power consumption.

[0057] For the in-memory computing of NVM adopted by the prior art, non-volatile memory is used to store the weights in the neural network, and the computing of the neural network is implemented through analog methods. However, since practical neural networks basically have many layers and very complex connection structures, it is very inconvenient for analog signals to be transmitted between layers and various processes to be performed, which is not conducive to supporting flexible neural network structures, resulting in considerable difficulties in the implementation and application of the overall neural network model. Moreover, the storage, reading and writing of analog signals and various noises and errors in computing significantly affect the reliability of the model and the accuracy of computing.

[0058] Figure 5 This is a chip architecture diagram for AI computing based on NVM of the present invention. As Figure 5 shown, a chip architecture for AI computing based on NVM of the present invention includes an NVM array 7, an external interface module 2, an SRAM 5, an NPU 6, and an MCU 1 that are communicatively connected through a bus 4. The MCU 1 reads and writes the SRAM 5 and the internal NVM array 7 through the bus 4 and communicates with the NPU 6. The NVM array 7 is used for on-chip storage of the weight parameters of the digitalized neural network, the programs run by the MCU 1, and the neural network model. The NPU 6 is used for digital domain acceleration computing of the neural network. The external interface module 2 is used to receive external AI operation instructions, input data, and output the results of AI computing outward. The MCU 1 is used to execute the program stored in the NVM array 7 based on the external AI operation instructions to control the NVM array 7 and the NPU 6 to perform AI computing on the input data to obtain the results of AI computing.

[0059] The SRAM 5 is used as a cache required for the operation and computing of the internal system of the chip, and is used to store input and output data, intermediate data generated by computing, etc. Specifically, it can include the cache during the execution of the program by the MCU 1, storing the executable program during the operation of the MCU 1, system configuration parameters, computing network structure configuration parameters, etc.; the operation cache of the NPU 6, and storing the input and output data during the operation of the neural network model.

[0060] In the chip architecture provided in this embodiment, the NPU 6 and the NVM are combined to perform AI neural network calculations. Among them, the weight parameters of the neural network are digitally stored in the NVM array 7 inside the chip. The neural network calculation is also a digital domain calculation. Specifically, the MCU 1 controls the NPU 6 and the NVM array 7 based on external AI operation instructions. The MCU 1 controls the NVM array 7 to include loading the weight parameters of the neural network stored therein, the program run by the MCU 1, and the neural network model to perform AI calculations. Compared with various existing memory-computation schemes that use NVM for analog operations, the digital storage and operation method has a flexible operation structure. The information stored in the NVM has better reliability, higher precision, and higher reading accuracy compared to the multi-level storage of analog signals. Therefore, the solution provided in this embodiment not only breaks through the bottleneck of the storage speed of off-chip NVM and reduces the external input power consumption, but also has high feasibility, flexibility, and reliability.

[0061] In one embodiment, the neural network model is stored in the NVM array 7 in a digitalized manner, and there can be multiple neural network models stored in the NVM array 7. The external AI operation instructions include algorithm selection instructions, and one of the multiple neural network models is selected as the algorithm for performing AI calculations through the algorithm selection instructions.

[0062] The neural network model in this embodiment is digitally stored in the NVM array 7 and can be multiple according to the number of application scenarios. For the case where multiple application scenarios correspond to multiple neural network models, the MCU 1 can flexibly select any one of the pre-stored neural network models for AI calculations according to the externally input algorithm selection instructions, overcoming the problem in the prior art that the analog computing array structure of memory-computation integration is rigid and not conducive to supporting flexible neural network structures.

[0063] In one embodiment, the NVM array 7 employs, but is not limited to, one of flash memory, MRAM, RRAM, MTP, and OTP. The interface standard of the external interface module 2 is at least one of SPI, QPI, and parallel interfaces.

[0064] In other embodiments, the NVM array 7 employs, but is not limited to, one of SONOS flash memory, Floating Gate flash memory, and Split Gate flash memory processes. The interface standard of the external interface module 2 is SPI and / or QPI.

[0065] The chip architecture provided in this embodiment is an improvement based on the traditional flash memory chip architecture. Specifically, an MCU1 and an NPU6 are embedded inside the flash memory chip and are communicatively connected through an on-chip bus 4. The on-chip bus 4 can be an AHB bus or other communication buses that meet the requirements, which is not limited herein. In the solution provided in this embodiment, the NPU6 is combined with the NVM, that is, both computing and storage are on-chip. Among them, the weight parameters of the neural network are digitally stored in the NVM array 7, and the neural network calculation is also digital domain calculation. Specifically, the MCU1 controls the NPU6 and the NVM array 7 based on external AI operation instructions, thereby breaking through the storage speed bottleneck of using off-chip NVM and reducing the external input power consumption, while also having high feasibility, flexibility, and reliability.

[0066] In one embodiment, in addition to the on-chip bus 4 communication, the chip architecture further includes a high-speed data reading channel; specifically, a high-speed data reading channel is established between the NPU6 and the NVM array 7, and the NPU6 is further configured to read the weight parameters from the NVM array 7 through the high-speed data reading channel. In this embodiment, the high-speed data reading channel is used to support the bandwidth requirement for the high-speed reading of the weight parameters, that is, the weight data, of the neural network when the NPU6 performs digital domain operations. The bit width of the high-speed data reading channel is m bits, and m is a positive integer.

[0067] In addition, the NVM array 7 is provided with N read channels, where N is a positive integer. In one read cycle, the read channels read a total of N bits of data. The NPU6 is configured to read the weight parameters from the NVM array 7 through the read channels via the high-speed data reading channel. Preferably, N ranges from 128 to 512. In one read cycle (usually 30 - 40 nanoseconds), the NPU6 reads the weight parameters of the neural network from the NVM array 7 through the read channels via the high-speed data reading channel with an m-bit width. Compared with the reading speed supported by off-chip NVM in the prior art, the bandwidth of the present invention is much higher and can support the parameter reading speed requirements for common neural network inference calculations.

[0068] In one embodiment, the chip architecture further includes a data conversion unit. In the case where the number of read channels is inconsistent with the bit width of the high-speed data reading channel and / or the frequencies are asynchronous, the data conversion unit is configured to convert the data into a combination of data with the same bit width as the high-speed data reading channel, usually a combination of words with a small width (such as 32 bits). The NPU6 reads the data from the data conversion unit through the high-speed data reading channel at its own clock frequency (up to more than 1 GHz).

[0069] Figure 6 It is a schematic diagram of the data conversion unit of the chip architecture of the present application. As Figure 6As shown, the data conversion unit includes a cache module and a sequential reading module. The cache module is used to sequentially cache N-bit data output from the NVM array 7 via the read channel in cycles. The capacity of the cache module is N*k bits, where k represents the number of cycles. The sequential reading module is used to convert the cached data in the cache module into an m-bit width and output it to the NPU 6 via the high-speed data reading channel, where N*k is an integer multiple of m.

[0070] Among them, the data conversion unit includes a cache module with N*k bits and a sequential reader that outputs m bits at a time, that is, the sequential reading module. N*k is an integer multiple of m. The read channel is connected to the NVM array 7 and can output N bits in each cycle. The cache can store data for k cycles. The width of the high-speed data reading channel is m bits. The high-speed data reading channel can include read / write instructions (CMD) and reply (ACK) signals and is connected to the read control circuit of the NVM array 7. After the read operation is completed, the ACK signal notifies the high-speed data reading channel and can also notify the on-chip bus at the same time. The high-speed data reading channel inputs the data in the cache into the NPU 6 asynchronously in multiple times through the sequential reading module.

[0071] In one embodiment, the MCU 1 is further configured to receive, through the external interface module 2, an external data access instruction for operating the NVM array 7. The MCU 1 is further configured to complete the logical control of the basic operations on the NVM array 7 based on the data access instruction. The data access instruction is a standard flash memory operation instruction. The AI operation instruction and the data access instruction adopt the same instruction format and rules. The AI operation instruction includes an operation code, and the AI operation instruction further includes an address part and / or a data part. The operation code of the AI operation instruction is different from the operation code of the standard flash memory operation instruction.

[0072] In this embodiment, the instructions for direct operation of the NVM and the instructions for AI computing processing adopt the same instruction format and rules. Taking the SPI and QPI interfaces as examples, based on the traditional SPI and QPI flash memory operation instruction op_code, the unused op_code for flash memory operation is selected to express the AI instruction, more information is passed in the address part, and AI data transfer is implemented in the data exchange cycle. Only by expanding the instruction decoder to realize the multiplexing of the interface and adding several status registers and configuration registers can AI computing be realized.

[0073] The MCU 1 realizes the digital operation of the NVM array 7, which can specifically include basic operations such as reading, writing, and erasing of flash memory. The external data access instruction and the external interface can adopt the standard flash memory chip format, which is easy to apply the chip flexibly and simply. The MCU 1 embedded in this chip, as the logical control unit of the NVM, replaces the logical state machine in the standard flash memory, simplifies the chip structure, and saves the chip area.

[0074] In this embodiment, in addition to storing neural network models, weight parameters, and programs for the internal system operation of the chip, the NVM array 7 can also be used to store data input from the outside that is not limited to AI computing-related data, that is, it can also be used to store other AI computing-related data input from the outside, as well as data input from the outside that is not related to AI computing. The unrelated data specifically includes information such as system parameters, configurations, and / or codes of external devices or systems; the basic operations include not only read, write, and erase operations on neural network models, weight parameters, and programs for the internal system operation, but also direct storage read, write, and erase operations on the stored data input from the outside in the NVM array 7.

[0075] During the specific implementation process, the MCU1 receives instructions for read, write, and other operations on the NVM array 7 from the outside and completes the logical control of the basic NVM operations. These basic operations include storing and reading the AI operation model algorithms and parameters, and can also be used for directly storing and reading system parameters, configurations, codes, etc. in the NVM array 7. The MCU1 also receives external AI operation instructions, controls the internal operation logic and input / output, and is also used for internal control of the AI algorithm logic.

[0076] Figure 7 It is a flowchart of the instruction operation for the chip architecture of the present application to call NVM read and write operations. As Figure 7 shown, the instruction operation process is as follows:

[0077] Step S101: The external device starts the chip where the NVM is located, and the MCU1 is powered on.

[0078] Step S102: Without external instructions, the MCU1 loads the required code and parameters from the NVM array 7 into the SRAM5, and the chip is in the standby state.

[0079] Step S103: The external device sends an NVM operation instruction, and the MCU1 receives and processes the instruction. The format and processing method of the NVM operation instruction are the same as those of the traditional standard NVM.

[0080] In one embodiment, the chip architecture further includes a DMA channel 3, and the DMA channel 3 is used for the external device to directly read and write the SRAM5. The external interface module 2 realizes the multiplexing of data and instructions, and through the DMA channel 3, it realizes the direct read and write operation of the external device on the SRAM5 inside the chip, improving the data transmission efficiency. The external device can also call the SRAM5 as a system memory resource through the DMA channel 3, increasing the flexibility of chip applications.

[0081] Figure 8 It is a flowchart for running AI operation instructions based on the chip architecture of the present application. As Figure 8 shown, the AI operation instruction operation process includes:

[0082] Step S201: The external device activates the chip where the NVM is located, and MCU1 is powered on.

[0083] Step S202: Without external instructions, the required code and parameters for MCU1 to run are loaded from the NVM array 7 into the SRAM5, and the chip is in the standby state.

[0084] Step S203: The external device sends an algorithm selection instruction to select a certain neural network model stored in the NVM array 7 of the chip.

[0085] Step S204: MCU1 processes this instruction, and the corresponding internal storage module is powered on and addressed.

[0086] Step S205: The external device sends an AI operation instruction and input data, and this data is cached in the SRAM5.

[0087] Step S206: MCU1 activates the NPU6 and performs recognition processing on the input data according to the AI operation instruction.

[0088] Step S207: The NPU6 reads the weight parameter data corresponding to the neural network model from the NVM array 7 for calculation.

[0089] Step S208: The external device reads the result of the AI calculation from the chip through the external interface module 2.

[0090] Among them, steps S205 to S208 can be reciprocally looped to continuously perform data input, calculation, and output.

[0091] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A chip architecture for AI computing based on NVM, characterized in that, It includes an NVM array, an external interface module, an NPU, and an MCU that are communicatively connected via a bus; The NVM array is used for on-chip storage of the weight parameters of a digitized neural network, the program run by the MCU, and the neural network model; The NPU is used for digital domain acceleration calculation of the neural network; The external interface module is used to receive external AI operation instructions, input data, and output the results of AI calculation outward; The MCU is used to execute the program based on the AI operation instructions. The MCU processes the instructions, and the corresponding internal storage module is powered on and addressed to control the NVM array and the NPU to perform AI calculation on the input data to obtain the results of the AI calculation; The NVM array is provided with a read channel. The read channel is N-way, where N is an integer greater than or equal to 128. In one read cycle, the read channel reads a total of N bits of data. The NPU is used to read the weight parameters from the NVM array through the high-speed data read channel via the read channel; The bit width of the high-speed data read channel is m bits, where m is a positive integer. The chip architecture further includes a data conversion unit, which includes a cache module and a sequential reading module. The cache module is used to sequentially cache the weight parameters output via the read channel by cycle. The capacity of the cache module is N*k bits, where k represents the number of cycles. The sequential reading module is used to convert the cached data in the cache module into an m-bit width and output it to the NPU via the high-speed data read channel, where N*k is an integer multiple of m.

2. The chip architecture for AI computing based on NVM according to claim 1, wherein The chip architecture further includes a high-speed data read channel. The NPU is also used to read the weight parameters from the NVM array through the high-speed data read channel; 3. The chip architecture for AI computing based on NVM according to claim 1, wherein The chip architecture further includes an SRAM, which is communicatively connected to the NVM array, the external interface module, the NPU, and the MCU via the bus. The SRAM is used to cache the data during the execution of the program by the MCU, the data during the operation of the NPU, and the input and output data of the neural network model operation; 4. The chip architecture for AI computing based on NVM according to claim 1, wherein Multiple neural network models are stored in the NVM array. The AI operation instructions include an algorithm selection instruction, which is used to select one of the multiple neural network models as the algorithm for performing AI calculation; 5. The chip architecture for AI computing based on NVM according to claim 1, wherein The NVM array adopts one of the flash memory process, MRAM process, RRAM process, MTP process, and OTP process, and / or the interface standard of the external interface module is at least one of SPI, QPI, and parallel interface; 6. The chip architecture for AI computing based on NVM as claimed in claim 5, wherein The NVM array adopts one of the SONOS flash memory process, Floating Gate flash memory process, and Split Gate flash memory process, and the interface standard of the external interface module is SPI and / or QPI; 7. The chip architecture for AI computing based on NVM according to claim 6, wherein The MCU is also used to receive external operation instructions for operating the NVM array through the external interface module, and the MCU is also used to complete the logical control of the basic operations of the NVM array based on the operation instructions; And / or, the operation instructions and the AI operation instructions adopt the same instruction format and rules.

8. The chip architecture for AI computing based on NVM according to claim 3, wherein, The chip architecture further includes a DMA channel, and the DMA channel is used for an external device to directly read and write the SRAM.

Citation Information

Patent Citations

  • Analog Vector-Matrix Multiplication Circuit

    CN109086249A

  • Chip architecture for carrying out AI calculation based on NVM

    CN214846708U

  • Processor with memory array operable as either cache memory or neural network unit memory

    US20180157970A1