In-memory computing low-power integrated image recognition method
Through a system combining a low-power camera and an integrated memory chip, a lightweight model is trained using a memory array for matrix multiplication and addition operations and knowledge distillation, the problem of large image recognition in the existing technology is solved, and a small-scale integrated image recognition with low power consumption and high energy efficiency is achieved.
Patent Information
- Application Number
- CN202111659219.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The existing image recognition systems are difficult to meet the low-power demand of small Always-On battery-powered low-power consumption in terms of power consumption and volume, especially systems based on GPU and FPGA, and the image recognition application of the integrated memory structure in actual scenarios is not integrated and low-power consumption.
A system composed of a low-power camera, a memory-based integrated chip, a low-power output module and a battery and power circuit is adopted, and a memory-based integrated architecture is used to perform matrix multiplication and addition operations, and a lightweight image recognition model is trained in combination with a knowledge distillation method, and communication is carried out through NB-IoT or LoRa.
It realizes low-power, high-energy-efficient small-scale integrated image recognition, suitable for Always-On-type applications, such as face recognition, face recognition and weather recognition, with a power consumption of less than 100mW, and is suitable for battery-powered terminal devices.
Smart Images

Figure CN114445607B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and more specifically, to an integrated in-memory computing and low-power integrated image recognition method.
Background Art
[0002] In recent years, image recognition methods based on deep neural networks have achieved a series of breakthroughs and have significant application effects, which are closely related to the growth of chip computing power following Moore's Law. Currently, there are mainly two types of implementation methods for image recognition: One is to collect images through a camera and push them to a cloud computer for image recognition. This mode has a long process flow, a large perception delay, great difficulty in real-time perception of complex time-varying environments, a large demand for computing resources, and high energy consumption. The other method is to use local computing to achieve image recognition, which has the advantages of strong real-time performance, privacy protection, saving network and cloud resources, and has a wider range of applicable scenarios. To achieve local image recognition, on the one hand, the deep neural network model has a large demand for computing resources, and on the other hand, for terminal systems, their energy, volume, and weight are usually limited (especially for ubiquitously deployed Internet of Things terminals). Therefore, there is an urgent need for low-power and high energy efficiency ratio methods and systems to achieve local image recognition.
[0003] To solve the problem that a large amount of network weight data is used in the inference of the deep neural network in the image recognition system, one form is composed of a CPU and GPU / DSP / FPGA / ASIC paired with DRAM and FLASH (flash memory). For example, the image system implemented by the cuda programming language based on the Jetson Nano GPU chip and the neural network model disclosed in Chinese Patent CN112308096A. The image recognition method, system, device, and storage medium based on the FPGA chip disclosed in Chinese Patent CN113011223A. Since the architectures of the above technical solutions belong to the traditional von Neumann architecture of separate storage and computing, there are still bottlenecks of the "memory wall" and "power wall". Another form is to use an in-memory computing architecture, which breaks through the bottleneck of the von Neumann computing architecture and directly uses the memory for data processing, thus integrating data storage and computing. Compared with the traditional computing system based on the von Neumann architecture, it can greatly reduce the time and energy consumption of data transfer. For example, the paper "Fully Hardware-implemented Memristor Convolutional Neural Network" published by Yao P, Wu H, Gao B, et al. implemented 0-9 handwritten digit recognition on the MNIST dataset (image resolution 28*28) based on a 2Mbit RRAM in-memory computing array.
[0004] However, for an image recognition terminal system implemented based on a GPU chip, although it has relatively strong computing power, its volume and power consumption are both relatively large, making it difficult to meet the requirements of small-sized, battery-powered, Always-On, low-power image recognition. For an image recognition terminal system implemented based on an FPGA chip, although its volume is controllable and its computing performance is also relatively strong, its power consumption is still relatively large, its energy efficiency ratio is not high, and there are still difficulties in meeting the requirements of small-sized, battery-powered, Always-On, low-power image recognition. Using a memory-computation integrated structure for image recognition is currently mainly in the laboratory exploration stage. Yao's paper only achieved recognition of static handwritten digital pictures in the MNIST dataset based on an RRAM memory-computation integrated array development board in a laboratory scenario. The development board has a large volume and does not implement a low-power, miniaturized integrated architecture and system including image data acquisition, preprocessing, and image recognition. The neural network implemented only includes 5 layers, and the system does not implement image recognition in a video stream in an actual scenario.
Summary of the Invention
[0005] In order to overcome the above-mentioned deficiencies of the prior art, especially to meet the requirements of Always-On image recognition tasks powered by batteries, and aiming at the problems of high power consumption, large volume, and low integration degree of the aforementioned prior art, the present invention proposes a low-power, small-sized integrated image recognition method.
[0006] The present invention proposes a memory-computation integrated image recognition method based on a low-power integrated image recognition system with memory-computation integration. The system includes:
[0007] A low-power camera, a memory-computation integrated chip, a low-power output module, and a battery and power supply circuit; wherein,
[0008] The low-power camera is used to collect image data,
[0009] The memory-computation integrated chip, which embeds an image recognition neural network model algorithm, is used for image analysis and recognition based on the neural network model;
[0010] The low-power output module is used to display the recognition result;
[0011] The battery and power supply circuit is used to provide the required power for each module of the system;
[0012] The memory-computation integrated chip further includes:
[0013] A memory-computation array, on-chip SRAM, on-chip FLASH, an arithmetic logic unit and a control unit, and an input / output I / O interface. Among them, the memory-computation array is composed of memory-computation units in the form of a Crossbar array, responsible for performing fast matrix multiplication and addition operations, and the memory-computation unit is implemented by a Flash or MRAM non-volatile storage device;
[0014] The on-chip SRAM is used for caching calculation instructions, input image data, neural network outputs, and intermediate data;
[0015] The on-chip FLASH is used for storing calculation instruction codes;
[0016] The arithmetic logic unit and the control unit are responsible for instruction execution and communicate with external chips or interfaces through the I / O interface;
[0017] The method includes a process for developing an in-memory computing image recognition model and an in-memory computing image recognition process;
[0018] The process for developing the in-memory computing image recognition model includes:
[0019] Steps for training data collection and production: The training dataset includes three parts. The first part is collected and labeled based on the in-memory computing low-power integrated image recognition system. The second part is labeled through an open-source face dataset. The third part of the data is produced by data augmentation means including image noise addition and affine transformation based on the first two parts of the datasets;
[0020] Steps for training a lightweight image recognition model for in-memory computing: Use a method for training a lightweight image recognition model for in-memory computing based on knowledge distillation to train the image recognition model;
[0021] Steps for embedding the in-memory computing chip model: Use an embedding method of a neural network model algorithm on the in-memory computing chip to transplant and embed the image recognition model into the in-memory computing chip.
[0022] Furthermore, the system further includes a low-power communication module for transmitting recognition results or other relevant data to a host computer or a cloud system and receiving response control and update instructions.
[0023] Furthermore, the low-power communication module uses NB-IoT or LoRa for communication.
[0024] Furthermore, the in-memory computing chip directly receives image data from the low-power camera and performs preprocessing, performs image analysis and recognition based on the image recognition neural network model algorithm embedded in the in-memory computing chip, and gives recognition results by integrating business logic, outputs through the low-power output module, and directly uploads data and results and receives and responds to control and update instructions through the low-power communication module.
[0025] Furthermore, the system further includes a low-power microprocessor MCU. The MCU is responsible for preprocessing the image data by cropping and conversion, and sending the preprocessed image data to the in-memory computing integrated chip for recognition;
[0026] The MCU also generates a recognition result based on the model recognition result of the in-memory computing integrated chip and combines the business logic decision, and sends the recognition result to the low-power output module;
[0027] The MCU is also responsible for calling the low-power communication module to establish communication with the host computer system, and transmitting the recognition result data according to the recognition result and the business logic.
[0028] Furthermore, the in-memory computing integrated image recognition process includes the following steps:
[0029] Image data acquisition step: Based on the low-power camera, according to the scene requirements, continuously acquire the images to be recognized with a specific resolution at a certain frame rate;
[0030] Image preprocessing step: Perform preprocessing on the input images to be recognized, including format conversion, cropping, size conversion, and filtering. The preprocessed images are input into the in-memory computing integrated chip for recognition;
[0031] In-memory computing integrated image recognition step: Based on the in-memory computing integrated chip and the embedded image recognition neural network model algorithm, analyze and recognize the preprocessed images;
[0032] Recognition result output step: Give the final recognition result by integrating the analysis and recognition result given by the in-memory computing integrated image recognition step and the business logic.
[0033] Furthermore, the method for training the in-memory computing integrated lightweight image recognition model based on knowledge distillation includes:
[0034] Teacher network training step: According to the image recognition application requirements, select the AlexNet, ResNet50, ResNet101, or VGG network architecture to pre-train a deep image recognition teacher network model;
[0035] Student network training step for in-memory computing integrated: According to the size and computing power of the in-memory computing array of the in-memory computing integrated chip, design a fully convolutional network as the student network. The fully convolutional network is stacked by convolutional modules, and the convolutional module is composed of a convolutional layer, a Pooling layer, and a Relu activation function in sequence; Using the pre-trained deep image recognition teacher network model as a guide and the cross-entropy loss as the error loss function for knowledge distillation, train the student network for in-memory computing integrated.
[0036] Furthermore, the embedding method includes:
[0037] Steps for disassembling the neural network model: Analyze the trained neural network model for image recognition to obtain the operator types and weight parameters of each layer of the model;
[0038] Steps for weight fixed-point quantization: After normalizing the weight parameters of the neural network layer by layer, perform 8-bit fixed-point quantization processing;
[0039] Steps for input and output fixed-point quantization: Adopt the strategy of adaptive floating fixed-point quantization to retrieve and scale the data of the input and output of each operator for each frame, and make the significant bits of the fixed-point data the most;
[0040] Steps for weight arrangement and correction: Adopt the implementation method of convolutional network memory and computing integration with parallel loops to realize the weight arrangement of the convolutional network layer; For the fixed-point weights, correct them according to the circuit and physical characteristics of the memory and computing array until the high-precision analog computing requirements of the memory and computing array are met;
[0041] Steps for computing pipeline arrangement: The neural network model performs calculations layer by layer. For each convolutional neural network layer, adopt the implementation method of convolutional network memory and computing integration with parallel loops for calculation; If the resolution of the image is not greater than 320×320, input the whole image into each layer of the neural network for calculation, otherwise, after dividing the image into blocks, input each block into each layer of the neural network for calculation;
[0042] Steps for generating assembly code: Generate the model algorithm assembly code on the memory and computing integrated chip according to the weight arrangement and computing pipeline arrangement;
[0043] Steps for burning weights: Write the weight parameters of the neural network after weight arrangement and correction into the memory and computing array of the memory and computing integrated chip;
[0044] Steps for generating and downloading machine code: Assemble and generate the machine code of the memory and computing integrated chip and download it to the on-chip FLASH of the memory and computing integrated chip;
[0045] Steps for testing and debugging: Conduct tests and optimizations on the quantization and fixed-point quantization effects, and evaluate the time consumption, weight storage consumption, and cache consumption required for each layer of calculation. Test the overall running performance indicators. If the application requirements are met, the embedding process ends, otherwise repeat all the above steps.
[0046] Furthermore, the final recognition result in the recognition result output step is displayed in the form of an LED light group or an LCD screen, or the recognition result is uploaded to the cloud through the low-power communication module.
[0047] Furthermore, the implementation method of convolutional network memory and computing integration with parallel loops includes the following steps in the case of a single input channel and a single convolutional kernel:
[0048] Let N be the width of the complete input image or the intermediate layer feature map, the convolution kernel size be m×m, and the stride be d, where N, d, and m are all positive integers, and d ≤ m.
[0049] Steps for weight arrangement: First, split the convolution kernel into m m-dimensional weight vectors by columns and connect them head-to-tail to form a weight vector WS(1) of m1 dimensions, where m1 is the result of multiplying m by m; then shift the last d1 dimensions of the weight vector WS(1) to the front d1 dimensions in a circular shift manner to form a weight vector WS(2), where d1 is the result of multiplying d by m; similar to the process of forming WS(2), continue the circular shift p - 2 times, and get a weight vector each time of the circular shift, a total of p weight vectors are formed. These p weight vectors are denoted as: WS(i), where 1 ≤ i ≤ p, and p = LCM(d, m), and LCM(d, m) represents the least common multiple of d and m; arrange the above p obtained weight vectors in parallel to form the final convolution kernel weight arrangement.
[0050] Steps for data input and calculation: First, arrange the input image data block or feature map data block covered by the convolution kernel into an m1-dimensional vector by columns as the input data S(1), input it into the memory and computing array one by one, and open the corresponding columns of the weight vector WS(1) to output the result of a matrix multiply-accumulate operation.
[0051] After the convolution kernel slides horizontally according to the stride, the newly added part of the input image data block or feature map data block it covers compared with before sliding is denoted as S+, and the reduced part compared with before sliding is denoted as S-. Replace the corresponding part of S- in the data S(1) with S+, that is, cyclically replace the data that the convolution kernel no longer covers with the newly covered data, and keep the rest unchanged to obtain the input data S(2). Input it into the memory and computing array one by one, and the memory and computing array open the corresponding columns of the corresponding weight vector WS(2) to give the matrix multiply-accumulate result.
[0052] Similar to obtaining the input data S(2), continue to slide the convolution kernel horizontally according to the stride. Each time it slides, a new input data is obtained, denoted as: S(q + 1), and it is correspondingly input into the memory and computing array. The memory and computing array open the corresponding columns of the weight vector WS(q + 1) corresponding to the new input data S(q + 1) to give the matrix multiply-accumulate operation result, where q is the number of times the convolution kernel slides horizontally, and 2 ≤ q ≤ N - m.
[0053] When a round of horizontal sliding ends, then slide vertically by one stride. At this time, slide the convolution kernel vertically according to the stride in the same way as obtaining the input data S(1), S(2), S(q + 1) above to generate the corresponding input data, open the corresponding columns of the corresponding weight vector and calculate to obtain the result until the convolution calculation of the complete input image or feature map is completed.
[0054] In this process, the updated amount of the input data only needs to be d / m of the full amount of the input data.
[0055] Furthermore, the steps of the method for implementing the convolutional network with in-memory computing in a parallel loop manner in the case of multiple input channels and multiple convolutional kernels include:
[0056] Let N be the width of the complete input image or the intermediate layer feature map, the number of input channels be s, the number of convolutional kernels be k, the size of each convolutional kernel be s×m×m, and the stride be d, where N, k, s, m, and d are all positive integers, and d ≤ m.
[0057] Weight arrangement step: For the first convolutional kernel, each of the s channels of the convolutional kernel is divided into m m-dimensional weight vectors by columns, and they are connected end to end to form a weight vector WS(1) of m1 dimensions, where m1 is the result of multiplying m by m; the vectors corresponding to the s channels are connected into a weight vector of m2 dimensions and arranged in a total column to obtain the arrangement of the first convolutional kernel, where m2 is the result of multiplying s by m1; for other convolutional kernels, they are each arranged in a column in the same way and arranged side by side with the first convolutional kernel to form a weight arrangement block WSB(1) of k columns.
[0058] For the m m1-dimensional weight vectors of each channel of each convolutional kernel, they are arranged into a weight vector of m2 dimensions in a cyclic shift manner, and the k m2-dimensional weight vectors corresponding to the k channels are arranged in parallel to form a new weight arrangement block of k columns; each time a cycle is completed, a weight arrangement block is obtained, and a total of p weight arrangement blocks are obtained. Each weight arrangement block is denoted as: WSB(i), where 1 ≤ i ≤ p, and p = LCM(d, m), and LCM(d, m) represents the least common multiple of d and m.
[0059] Data input and calculation step: First, each channel of the input image data block or the feature map data block of the convolutional kernel size is arranged into a vector of m1 dimensions by columns, and the m1-dimensional vectors of the s channels are connected into a vector of m2 dimensions to form the input data S(1), which is input into the in-memory computing array one by one, and the weight arrangement block WSB(1) is opened to output the matrix multiplication and addition results corresponding to the k convolutional kernels in parallel.
[0060] After the convolutional kernel slides horizontally by the stride, each input channel uses the newly covered part of the data of the convolutional kernel to cyclically update the part of the data that is no longer covered to form the input data S(2), and then the weight arrangement block WSB(2) is opened to output the matrix multiplication and addition results corresponding to the k convolutional kernels in parallel.
[0061] Continue to horizontally slide the convolution kernel by the said step size. Use a similar method as obtaining the input data S(2) to obtain new input data, denoted as S(q + 1). The memory - computing array opens the weight arrangement block WSB(q + 1) corresponding to the new input data S(q + 1), and parallelly gives the matrix multiply - add operation results of k convolution kernels, where q is the number of times the convolution kernel slides horizontally, and 2 ≤ q ≤ N - m;
[0062] When a round of horizontal sliding ends, then slide vertically by one step size. At this time, slide the convolution kernel vertically by the step size in the same way as obtaining the input data S(1), S(2), S(q + 1) above to generate corresponding input data, open the corresponding weight arrangement block to calculate the results until the convolution calculation of the complete input image or feature map is completed.
[0063] Further, m is 3 and d is 1.
[0064] Further, s, m, and k are all 3 and d is 1.
[0065] Based on the inventive concept of the present invention, the present invention can obtain the following beneficial technical effects:
[0066] Integrate the memory - computing integrated chip with components such as a low - power camera, a low - power MCU (optional), a power module, and a low - power communication module, and make full use of the low - power and high - energy - efficiency characteristics of the memory - computing integrated chip to perform inference calculations for image recognition deep neural networks. Since the main computing requirements of the image recognition system lie in the image recognition deep neural network calculations, therefore, combined with other low - power components, the memory - computing integrated image recognition system can achieve the characteristics of low power consumption and high energy efficiency. Based on this system and method, Always - On type low - power and small - sized integrated image recognition related applications can be realized, including but not limited to frontal face recognition, face recognition, weather recognition, gesture recognition, etc. For example, when running the frontal face recognition model inference, the power consumption of the memory - computing integrated chip is less than 30 mW, and the overall power consumption of the system can be less than 100 mW.
BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is an embodiment of the composition structure of the memory - computing integrated low - power integrated image recognition system proposed by the present invention.
[0068] Figure 2 This is another embodiment of the composition structure of the memory - computing integrated low - power integrated image recognition system proposed by the present invention.
[0069] Figure 3 This is an embodiment of the memory - computing integrated chip architecture proposed by the present invention.
[0070] Figure 4This is an embodiment of the development process of the in-memory computing image recognition model and the flowchart of the in-memory computing image recognition inference process proposed by the present invention.
[0071] Figure 5 This is an embodiment of the in-memory computing low-power integrated image recognition model training method based on knowledge distillation proposed by the present invention.
[0072] Figure 6 This is an embodiment of the image recognition network model for in-memory computing proposed by the present invention.
[0073] Figure 7 This is an embodiment of the neural network model algorithm embedding process on the in-memory computing chip proposed by the present invention.
[0074] Figure 8 It is a schematic diagram of the in-memory computing implementation method of a single-channel single-convolution kernel parallel loop convolution network.
[0075] Figure 9 It is a schematic diagram of the in-memory computing implementation method of a multi-channel multi-convolution kernel parallel loop convolution network.
Detailed Embodiment
[0076] For ease of understanding, this specific embodiment is a preferred embodiment of the in-memory computing image recognition method proposed by the present invention, to illustrate in detail the structure and inventive points of the present invention, but does not limit the protection scope of the claims of the present invention.
[0077] See Figure 1 , for the scenario with high image resolution and heavy preprocessing tasks, the composition structure of an in-memory computing low-power integrated image recognition system given by a preferred embodiment includes 6 basic modules, namely a low-power camera, a low-power MCU, an in-memory computing chip, a low-power output module, a low-power communication module, and a battery and power supply circuit. Among them:
[0078] (1) The low-power camera is responsible for collecting image data and transmitting it to the low-power MCU. The image acquisition resolution can be set to 720p / VGA / QVGA, etc. according to application requirements, and the whole picture or the picture automatically cropped based on the camera can be collected. The field of view angle is not less than 30°, and the focal length can be configured according to the scene requirements.
[0079] (2) Low-power MCU: (a) Responsible for preprocessing image data such as cropping and conversion, and sending the preprocessed image data to the in-memory computing chip for recognition; (b) At the same time, responsible for data splitting coordination and task scheduling between in-memory computing chips; (c) According to the model recognition results of the in-memory computing chip, combined with business logic decisions, generate recognition results and send them to the low-power output module; (d) Responsible for calling the low-power communication module to establish communication with the host computer system, and transmitting recognition result data according to the recognition results and business logic.
[0080] (3) In-memory computing chip, which embeds an image recognition neural network model algorithm and is responsible for image analysis and recognition based on the neural network model. The in-memory computing chip has the characteristics of low power consumption and high energy efficiency.
[0081] (4) Low-power communication module, responsible for transmitting data such as recognition results to the host computer or cloud system, and receiving response control and update instructions, and can communicate using NB-IoT or LoRa.
[0082] (5) Low-power output module, including a low-power LCD display and / or an LED indicator group. The low-power LCD display is used to display the original image and the recognition results, and the LED indicator group is used to indicate the recognition results. For example, when performing frontal face recognition based on this system, the red, yellow, and green indicator lights being on respectively indicate that a frontal face, a left face, and a right face are recognized. The output module can also output a specific level signal for connecting to external devices such as locks.
[0083] (6) Battery and power supply circuit, providing the required power for each module of the system, and can output multiple different voltages to meet the needs of different modules of the system, including 3.3V, 2.8V, 1.8V, 1.5V, etc.
[0084] See Figure 2 , as another preferred embodiment, the in-memory computing low-power integrated image recognition system, mainly for scenarios with relatively light image preprocessing tasks, includes: a low-power camera, an in-memory computing chip, a low-power output module, a battery and power supply circuit, 4 basic modules, and 1 optional low-power communication module. For image recognition scenarios that do not require networking, the low-power communication module can be not configured, thus having lower power consumption, smaller volume, and cost than Figure 1 the implementation mode 1 shown. Among them:
[0085] (1) The functions and roles of the low-power camera, low-power output module, and battery and power supply circuit are similar to the corresponding modules in the Figure 1 implementation mode shown.
[0086] (2) The in-memory computing integrated chip directly receives image data from a low-power camera and performs preprocessing. It conducts image analysis and recognition based on the image recognition neural network model algorithm embedded in the chip, and gives recognition results by integrating business logic and outputs them through a low-power output module. For a system configured with a low-power communication module, the in-memory computing integrated chip is directly responsible for data and result uploading, as well as receiving and responding to control and update instructions.
[0087] For a deep neural network, its main computation is matrix multiplication and addition operations. The present invention utilizes the in-memory computing architecture to break through the bottleneck of the von Neumann computing architecture, directly performing matrix multiplication and addition operations using the memory, thereby integrating data storage and computation. Compared with traditional computing systems based on the von Neumann architecture, it can greatly reduce the time and energy consumption of data transfer, featuring low power consumption and high energy efficiency, and is more suitable for deep neural network computing. See Figure 3 , which is a preferred embodiment of the architecture of the in-memory computing integrated chip proposed by the present invention and can be applied to Figure 1 、 Figure 2 In the embodiment shown, the in-memory computing integrated chip mainly consists of an in-memory computing array, on-chip SRAM (static random access memory), an arithmetic logic unit and a control unit, and an I / O interface, etc. Among them, the in-memory computing array is composed of in-memory computing units in the form of a Crossbar array, responsible for performing fast matrix multiplication and addition operations. The in-memory computing units can be implemented by non-volatile storage devices such as Flash (flash memory) / MRAM (magnetic random access memory); the on-chip SRAM is used for caching computing instructions, input image data, neural network outputs, and intermediate data; the arithmetic logic unit and the control unit are responsible for instruction execution and communicate with external chips or interfaces through the I / O interface.
[0088] On the other hand, the in-memory computing integrated image recognition method proposed by the present invention generally includes: the in-memory computing integrated image recognition model development process and the in-memory computing integrated image recognition process. The in-memory computing integrated image recognition model development process mainly trains the model based on the collected data samples and embeds it into the in-memory computing integrated chip, including three steps: training data collection and production, model design and training, and in-memory computing integrated chip model embedding. The in-memory computing integrated image recognition process includes four steps: image data collection, image preprocessing, in-memory computing integrated image recognition, and recognition result output, which is the recognition and inference process in actual applications. The in-memory computing integrated image recognition method is applicable to general image recognition scenarios, including but not limited to frontal face recognition, face recognition, weather recognition, gesture recognition, etc.
[0089] As Figure 4 shown, an embodiment of the in-memory computing integrated image recognition model development process specifically includes:
[0090] (a) Training data collection and production
[0091] The training dataset mainly consists of three parts. The first part is collected and labeled based on the in-memory computing and low-power integrated image recognition system of the present invention. The second part is labeled through an open-source face dataset. The third part of the data is produced by means of data augmentation such as image noise addition and affine transformation based on the datasets of the previous two parts.
[0092] (b) Training of a lightweight image recognition model for in-memory computing
[0093] Considering the limited resources of the in-memory computing array and on-chip SRAM of the in-memory computing chip, in order to reduce the model size and simplify the operator complexity, a training method for the in-memory computing lightweight image recognition model based on knowledge distillation is adopted to train the image recognition model.
[0094] (c) Embedding of the model in the in-memory computing chip
[0095] An embedding method of the neural network model algorithm on the in-memory computing chip is adopted to transplant and embed the image recognition model into the in-memory computing chip.
[0096] Figure 4 The in-memory computing image recognition process is also shown. The specific steps include:
[0097] (a) Image data acquisition. Based on a low-power camera, according to the scene requirements, continuously acquire images to be recognized with a specific resolution at a certain frame rate.
[0098] (b) Image preprocessing. Perform preprocessing such as format conversion, cropping, size conversion, and filtering on the input image. The preprocessed image is input into the in-memory computing chip for recognition.
[0099] (c) In-memory computing image recognition. Based on the in-memory computing chip and the embedded image recognition neural network model algorithm, perform image analysis and recognition.
[0100] (d) Output of the recognition result. Combining the in-memory computing image recognition result and the business logic, give the final recognition result, and display it in the form of an LED light group or an LCD screen, etc. It can also be set to upload the recognition result to the cloud through a low-power communication module according to the requirements.
[0101] Figure 5 A specific implementation manner of the training method for the in-memory computing lightweight image recognition model based on knowledge distillation is shown. First, according to the application requirements of image recognition, pre-train a larger image recognition teacher network model, such as ResNet50, etc.; then, based on the knowledge distillation architecture, train a lightweight image recognition network for in-memory computing. The lightweight image recognition network for in-memory computing can adopt a fully convolutional neural network or a neural network based on a lightweight architecture such as MobileNet according to the size and computing power of the in-memory computing array of the in-memory computing chip.
[0102] Taking frontal face recognition as an example (recognizing whether the image is a frontal face / left face / right face / no face), for the lightweight image recognition network model for in-memory computing, a convolutional neural network model structure is adopted, including an input layer, 3 convolutional layers, 3 pooling layers (downsampling layers), a Dropout layer, a fully connected layer, and an output layer, specifically as Figure 6 shown.
[0103] The input image resolution is 80×80, the parameter size of the pre-trained teacher network ResNet50 is 25.6MB, and the required computing power is 1.14G FLOPS; after knowledge distillation, the trained student network, as a frontal face detection model for in-memory computing, has a model size of only 436kB, and the required computing power is reduced to 25.8MFLOPS.
[0104] The method for embedding a neural network model algorithm on an in-memory computing chip includes the process of embedding a neural network model algorithm on an in-memory computing chip and a method for implementing an in-memory computing of a parallel loop convolutional network.
[0105] The process of the method for embedding a neural network model algorithm on an in-memory computing chip is as Figure 7 shown, mainly including steps such as neural network model disassembly, weight fixed-point quantization, weight arrangement and correction, weight programming, input / output fixed-point quantization, calculation pipeline arrangement, assembly code generation, machine code generation and download, and test debugging.
[0106] (a) Neural network model disassembly. Parse the trained neural network model (model in ONNX format) to obtain the operator types and weight parameters of each layer of the model.
[0107] (b) Weight fixed-point quantization. After normalizing the neural network weights layer by layer, perform 8-bit fixed-point quantization processing.
[0108] (c) Input / output fixed-point quantization. Adopt the strategy of adaptive floating fixed-point quantization to retrieve and scale the data of the input and output of each operator for each frame, and make the significant bits of the fixed-point data the most.
[0109] (d) Weight arrangement and correction. Adopt the method for implementing an in-memory computing of a parallel loop convolutional network to implement the weight arrangement of the convolutional network layer. For the fixed-point weights, correct them according to the circuit and physical characteristics of the in-memory computing array to meet the high-precision analog computing of the in-memory computing array.
[0110] (e) Compute the streaming layout. The neural network model performs computations layer by layer. For each layer of the convolutional neural network, a parallel loop-based in-memory computing implementation method for convolutional networks is used for computation; if the resolution of the image is no greater than 320×320, the entire image is input into each layer of the neural network for computation, otherwise, after the image is divided into blocks, each block is input into each layer of the neural network for computation. Among them, a parallel loop-based in-memory computing implementation method for convolutional networks is used to implement data input and convolutional computation for a single layer of convolutional neural network.
[0111] (f) Generate assembly code. Generate the model algorithm assembly code on the in-memory computing chip according to the computed streaming layout and the weight layout.
[0112] (g) Burn weights. Write the arranged and corrected neural network weights into the in-memory computing queue of the in-memory computing chip.
[0113] (h) Generate and download machine code. Assemble and generate the machine code of the in-memory computing chip and download it to the on-chip FLASH of the in-memory computing chip.
[0114] (i) Test and debug. Conduct quantization and fixed-point effect inspection tests and optimizations, and evaluate the time consumption, weight storage consumption, and cache consumption required for each layer of computation, and test the overall running performance metrics. If the application requirements are met, the embedding process ends, otherwise, repeat all the above steps.
[0115] The parallel loop-based in-memory computing implementation method for convolutional networks is as Figure 8 、 Figure 9 shown, where Figure 8 For the case of a single input channel and a single convolutional kernel, Figure 9 For the case of multiple input channels and multiple convolutional kernels.
[0116] See Figure 8 The shown in-memory computing implementation method for a single-channel single-convolutional-kernel parallel loop-based convolutional network. For the case of a single input channel and a single convolutional kernel (for ease of understanding, assume the convolutional kernel size is 3×3 and the stride is 1, but not limited to convolutional kernels and strides of the above sizes), the in-memory computing implementation method of the convolutional network is as follows:
[0117] ① Weight layout. Divide the 3*3 convolutional kernel into 3 three-dimensional vectors by column, and cycle and replace the 3 vectors 3 times to form a 3-column 9-dimensional vector layout, denoted as weights WS(1), WS(2), and WS(3).
[0118] ② Data input and calculation. First, arrange the input image data block (or feature map data block) with the size of the convolution kernel into a 9-dimensional vector column by column to obtain the input data S(1), which is input into the memory and computing array one by one, and open the corresponding columns of the weight WS(1). Output the result of a matrix multiplication and addition once. After the convolution kernel slides, only partial updates are needed for the input data S2 (the newly covered part of the convolution kernel, that is, In10, In11, and In12 update In1, In2, and In3 in the figure). Then, the memory and computing array opens the corresponding columns of the weight WS(2), that is, uses the data corresponding to the columns to give the result of matrix multiplication and addition. The data update amount in this process only needs d / k of the full-scale input data update (d is the step size, which is 1 here, and k is the side length of the convolution kernel, which is 3 here), which can effectively reduce the data preparation time and improve the computing efficiency. Repeat the above process of updating the input data, opening the corresponding columns of the memory and computing array, and giving the result of matrix multiplication and addition operations to quickly realize the convolution operation process.
[0119] See Figure 9 The memory and computing integrated implementation method of a multi-channel and multi-convolution kernel parallel loop convolution network shown in
[0120] ① Weight arrangement. For the first convolution kernel 1: Divide each channel (3×3) of the 3-channel convolution kernel into 3 3-dimensional vectors column by column, and arrange them into 1 column (9-dimensional vector) in the same order respectively; Connect the vectors corresponding to the 3 channels into one vector (27-dimensional) and arrange it into a total of 1 column to obtain the arrangement of convolution kernel 1. For convolution kernel 2, convolution kernel 3, …, convolution kernel n, use the same method to arrange each into 1 column and arrange them side by side with convolution kernel 1 to form the weight arrangement block WSB(1). For the 3 3-dimensional vectors of each channel of each convolution kernel, rearrange them in a cyclic replacement manner to re-form the weight arrangement block WSB(2) and the weight arrangement block WSB(3).
[0121] ② Data input and calculation. First, each channel of the input image data block (or feature map data block) with the size of the convolution kernel is arranged into a 9-dimensional vector column by column, and the 3 channels are connected into a 27-dimensional vector to form the input data S(1), which is input into the memory-computation array one by one, and the weight arrangement block WSB(1) is opened, that is, the data corresponding to the weight arrangement block is used to parallelly output the matrix multiplication and addition results of multiple convolution kernels; after the convolution kernel slides, each input channel correspondingly updates part of the data (the part newly covered by the convolution kernel) to form the input data S(2), and then the weight arrangement block WSB(2) is opened to parallelly output the matrix multiplication and addition results of multiple convolution kernels. Similarly, the data update amount in this process only needs d / k of the full-scale input data update (d is the step size, here d = 1, k is the side length of the convolution kernel, here k = 3), which can effectively reduce the data preparation time and improve the calculation efficiency; repeat the above process of updating the input data, open the corresponding weight arrangement block of the memory-computation array, and parallelly give the matrix multiplication and addition operation results of multiple convolution kernels, so as to parallelly and quickly implement the convolution operation in the case of multiple input channels and multiple convolution kernels.
[0122] Through the above method for realizing the integration of memory and computation in a convolutional network with a parallel loop structure, the compact arrangement of neural network weights and parallel and efficient convolutional network calculations can be achieved, so that the memory-computation integrated chip has the characteristics of low power consumption and high energy efficiency when running the image recognition neural network model.
[0123] The above are some specific implementation manners of the present invention, but the present invention is not limited to the above manners. All simple transformations of the technical features of the present invention, and any equivalent changes or modifications made according to the structure, features and principles described in the scope of the patent application of the present invention, will fall within the protection scope of the present invention.
Claims
1. An in-memory computing-based image recognition method for an in-memory computing low-power integrated image recognition system, characterized in that, The system includes: a low-power camera, a memory-computation integrated chip, a low-power output module, and a battery and power circuit; wherein, the low-power camera is used for collecting image data, the memory-computation integrated chip which embeds an image recognition neural network model algorithm is used for image analysis and recognition based on the neural network model; the low-power output module is used for displaying the recognition result; the battery and power circuit is used for providing the required power for each module of the system; the memory-computation integrated chip further includes: a memory-computation array, on-chip SRAM, on-chip FLASH, an arithmetic logic unit and a control unit, and an input / output I / O interface, wherein, the memory-computation array is composed of memory-computation units in the form of a Crossbar array, responsible for performing fast matrix multiplication and addition operations, and the memory-computation units are implemented by Flash or MRAM non-volatile storage devices; the on-chip SRAM is used for caching calculation instructions, input image data, neural network outputs, and intermediate data; the on-chip FLASH is used for storing calculation instruction codes; the arithmetic logic unit and the control unit are responsible for instruction execution and communicate with external chips or interfaces through the I / O interface; The method includes a memory-computation integrated image recognition model development process and a memory-computation integrated image recognition process; The memory-computation integrated image recognition model development process includes: A training data collection and production step: The training data set includes three parts. The first part is collected and labeled based on the memory-computation integrated low-power integrated image recognition system, the second part is labeled through an open-source face data set, and the third part of the data is produced by data augmentation means including image noise addition and affine transformation based on the first two parts of the data sets; A lightweight image recognition model training step for memory-computation integration: Adopt a memory-computation integrated lightweight image recognition model training method based on knowledge distillation to perform image recognition model training; A memory-computation integrated chip model embedding step: Adopt an embedding method of a neural network model algorithm on the memory-computation integrated chip to transplant and embed the image recognition model into the memory-computation integrated chip.
2. The image recognition method according to claim 1, wherein: The system further includes a low-power communication module, which is used for transmitting the recognition result or other relevant data to the host computer or cloud system, and receiving response control and update instructions.
3. The image recognition method according to claim 2, characterized in that: The low-power communication module uses NB-IoT or LoRa for communication.
4. The image recognition method according to claim 2 or 3, wherein: The memory-computation integrated chip directly receives image data from the low-power camera and performs preprocessing, performs image analysis and recognition based on the image recognition neural network model algorithm embedded in the memory-computation integrated chip, gives a recognition result based on comprehensive service logic, outputs it through the low-power output module, and through the low-power communication module, the memory-computation integrated chip directly uploads data and results and receives and responds to control and update instructions.
5. The image recognition method according to claim 2 or 3, characterized in that, The system further includes a low-power microprocessor MCU, wherein the MCU is responsible for cropping and converting the image data for preprocessing, and sending the preprocessed image data to the memory-computation integrated chip for recognition; The MCU also generates an identification result according to the model identification result of the memory-computation integrated chip and combines the service logic decision, and sends the identification result to the low-power output module; The MCU is also responsible for calling the low-power communication module to establish communication with the host computer system, and transmitting the identification result data according to the identification result and the service logic.
6. The image recognition method according to claim 2 or 3, characterized in that The memory-computation integrated image recognition process includes the following steps: Image data acquisition step: Based on the low-power camera, according to the scene requirements, continuously acquire the images to be recognized with a specific resolution at a certain frame rate; Image preprocessing step: Perform preprocessing on the input image to be recognized, including format conversion, cropping, size conversion, and filtering. The preprocessed image is input into the memory-computation integrated chip for recognition; Memory-computation integrated image recognition step: Based on the memory-computation integrated chip and the embedded image recognition neural network model algorithm, analyze and recognize the preprocessed image; Recognition result output step: Combine the analysis and recognition result given in the memory-computation integrated image recognition step and the service logic to give the final recognition result.
7. The image recognition method according to claim 6, wherein The method for training the memory-computation integrated lightweight image recognition model based on knowledge distillation includes: Teacher network training step: According to the image recognition application requirements, select the AlexNet, ResNet50, ResNet101, or VGG network architecture to pre-train a deep image recognition teacher network model; Student network training step for memory-computation integration: According to the size and computing power of the memory-computation array of the memory-computation integrated chip, design a fully convolutional network as the student network. The fully convolutional network is stacked by convolutional modules, and each convolutional module consists of a convolutional layer, a Pooling layer, and a Relu activation function in sequence; Using the pre-trained deep image recognition teacher network model as a guide and the cross-entropy loss as the error loss function for knowledge distillation, train the student network for memory-computation integration.
8. The image recognition method according to claim 7, wherein The embedding method includes: Neural network model disassembling step: Analyze the trained neural network model for image recognition to obtain the operator types and weight parameters of each layer of the model; Weight fixed-point quantization step: After normalizing the weight parameters of the neural network layer by layer, perform 8-bit fixed-point quantization processing; Input and output fixed-point quantization step: Adopt the strategy of adaptive floating-point quantization to retrieve and scale the data of the input and output of each operator for each frame, and make the effective bits of the fixed-point data the most; Weight arrangement and correction step: Adopt the parallel loop-based implementation method of the convolutional network for memory-computation integration to implement the weight arrangement of the convolutional network layer; Correct the fixed-point weights according to the circuit and physical characteristics of the memory-computation array until the high-precision analog computing requirements of the memory-computation array are met; Computing pipeline arrangement step: The neural network model is calculated layer by layer. For each convolutional neural network layer, adopt the parallel loop-based implementation method of the convolutional network for memory-computation integration for calculation; If the resolution of the image is not greater than 320×320, input the whole image into each layer of the neural network for calculation, otherwise, after dividing the image into blocks, input each block into each layer of the neural network for calculation; Steps for generating assembly code: Generate the model algorithm assembly code on the memory - computing integrated chip according to the weight arrangement and the computing pipeline arrangement; Steps for weight programming: Write the weight arrangement and the weight parameters of the corrected neural network into the memory - computing array of the memory - computing integrated chip; Steps for generating and downloading machine code: Assemble to generate the machine code of the memory - computing integrated chip and download it to the on - chip FLASH of the memory - computing integrated chip; Steps for testing and debugging: Conduct quantization, fixed - point effect inspection tests and optimizations, and evaluate the time consumption, weight storage consumption, and cache consumption required for each layer of calculation. Test the overall running performance indicators. If the application requirements are met, the embedding process ends; otherwise, repeat all the above steps.
9. The image recognition method according to claim 8, wherein The final recognition result in the recognition result output step is displayed in the form of an LED light group or an LCD screen, or the recognition result is uploaded to the cloud through the low - power communication module.
10. The image recognition method according to claim 9, characterized in that The method for realizing the convolutional network memory - computing integration in a parallel loop form, in the case of a single - input channel and a single convolutional kernel, the steps include: Let N be the width of the complete input image or intermediate - layer feature map, the convolutional kernel size be m×m, and the stride be d, where N, d, and m are all positive integers, and d ≤ m. Steps for weight arrangement: First, split the convolutional kernel into m m - dimensional weight vectors by columns and connect them head - to - tail to form a weight vector WS(1) of m1 dimensions, where m1 is the result of multiplying m by m; then shift the last d1 dimensions of the weight vector WS(1) to the front d1 dimensions in a circular - shift manner to form a weight vector WS(2), where d1 is the result of multiplying d by m; similar to the process of forming WS(2), continue to perform circular - shift p - 2 times, and get a weight vector each time a cycle is completed, a total of p weight vectors are formed. These p weight vectors are denoted as: WS(i), where 1 ≤ i ≤ p, and p = LCM(d,m), and LCM(d,m) represents the least common multiple of d and m; Arrange the above - obtained p weight vectors in parallel to form the final convolutional - kernel weight arrangement; Steps for data input and calculation: First, arrange the input - image data block or feature - map data block covered by the convolutional kernel into an m1 - dimensional vector by columns as the input data S(1), input it into the memory - computing array one - by - one, and open the corresponding columns of the weight vector WS(1) to output the result of a matrix multiply - add operation; After the convolutional kernel slides horizontally by the stride, the newly added part of the input - image data block or feature - map data block covered by it compared with before the slide is denoted as S +, and the reduced part compared with before the slide is denoted as S -. Replace the corresponding part of S - in the data S(1) with S +, that is, circularly replace the data that the convolutional kernel no longer covers with the newly covered data of the convolutional kernel, and keep the rest unchanged to obtain the input data S(2). Input S(2) into the memory - computing array one - by - one, and the memory - computing array opens the corresponding columns of the weight vector WS(2) to give the result of the matrix multiply - add operation; Similar to obtaining the input data S(2), continue to horizontally slide the convolution kernel by the said step length. Each time it slides, a new input data is obtained, denoted as: S(q + 1), which is correspondingly input into the memory - computing array. The memory - computing array opens the column corresponding to the weight vector WS(q + 1) corresponding to the new input data S(q + 1), and gives the result of matrix multiplication and addition. Here, q is the number of times the convolution kernel slides horizontally, and 2 ≤ q ≤ N - m; When a round of horizontal sliding ends, then vertically slide by one step length. At this time, slide the convolution kernel vertically by the step length in the same way as obtaining the input data S(1), S(2), S(q + 1) above, generate the corresponding input data, open the column corresponding to the corresponding weight vector and calculate the result until the convolution calculation of the complete input image or feature map is completed; In this process, the update amount of the input data only needs d / m of the full - volume input data.
11. The image recognition method according to claim 9, characterized in that The steps of the parallel - loop - type convolution network memory - computing integrated implementation method in the case of multiple input channels and multiple convolution kernels include: Let N be the width of the complete input image or the intermediate - layer feature map, the number of input channels be s, the number of convolution kernels be k, the size of each convolution kernel be s×m×m, and the step length be d. Here, N, k, s, m, d are all positive integers, and d ≤ m. Weight arrangement step: For the first convolution kernel, each of the s channels of the convolution kernel is divided into m m - dimensional weight vectors by columns and connected head - to - tail to form an m1 - dimensional weight vector WS(1), where m1 is the result of multiplying m by m; the vectors corresponding to the s channels are connected into an m2 - dimensional weight vector and arranged in a total column to obtain the arrangement of the first convolution kernel, where m2 is the result of multiplying s by m1; the other convolution kernels are arranged in a column in the same way and arranged side - by - side with the first convolution kernel to form a k - column weight arrangement block WSB(1); For the m m1 - dimensional weight vectors of each channel of each convolution kernel, they are arranged into an m2 - dimensional weight vector in a circular - shift manner. The k m2 - dimensional weight vectors corresponding to the k channels are arranged in parallel to form a new k - column weight arrangement block; each time it loops, a weight arrangement block is obtained, and a total of p weight arrangement blocks are obtained. Each weight arrangement block is denoted as: WSB(i), where 1 ≤ i ≤ p, and p = LCM(d,m), and LCM(d,m) represents the least common multiple of d and m; Data input and calculation step: First, each channel of the input - image data block or feature - map data block with the size of the convolution kernel is arranged into an m1 - dimensional vector by columns, and the m1 - dimensional vectors of the s channels are connected into an m2 - dimensional vector to form the input data S(1), which is correspondingly input into the memory - computing array, and the weight arrangement block WSB(1) is opened to parallel - output the matrix - multiplication - and - addition results with the k convolution kernels; After the convolution kernel slides horizontally by the step length, each input channel uses the newly covered part of the data of the convolution kernel to circularly update the non - covered part of the data to form the input data S(2), and then the weight arrangement block WSB(2) is opened to parallel - output the matrix - multiplication - and - addition results with the k convolution kernels; Continue to horizontally slide the convolution kernel by the said step size. Use a similar method to obtain the input data S(2) to obtain new input data, denoted as S(q + 1). The computing-in-memory array opens the weight arrangement block WSB(q + 1) corresponding to the new input data S(q + 1), and parallelly gives the matrix multiplication and addition operation results of k convolution kernels, where q is the number of times the convolution kernel slides horizontally, and 2 ≤ q ≤ N - m; When a round of horizontal sliding ends, then slide vertically by one step size. At this time, slide the convolution kernel vertically by the step size in the same way as obtaining the input data S(1), S(2), S(q + 1) above to generate corresponding input data, open the corresponding weight arrangement block to calculate the results until the convolution calculation of the complete input image or feature map is completed.
12. The image recognition method according to claim 11, wherein The said m is 3 and d is 1.
13. The image recognition method according to claim 12, characterized in that The said s, m, and k are all 3 and d is 1.
Citation Information
Patent Citations
Image recognition method and device, computer equipment and storage medium
CN112308096A
Image recognition method, system and device and storage medium
CN113011223A
Heterogeneous storage and calculation fusion system and method supporting deep neural network reasoning acceleration
CN112149816A
Data processing method and device, equipment and storage medium
CN113222107A