Neural network inference chip, neural network inference method and terminal
By using computing modules consisting of SSP, VSP, MMP, and PSUM, the hardware structure is optimized to reduce memory access power consumption, solving the problems of performance, power consumption, and scalability of neural network inference chips, and achieving high energy efficiency and flexibility.
Patent Information
- Application Number
- CN202111545989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-12-16
AI Technical Summary
Existing neural network inference chips are not ideal in terms of performance, power consumption and scalability, especially lacking flexibility and efficiency when handling different neural network algorithms.
The computing module consists of SSP, VSP, MMP and PSUM. By reusing weights, reusing some input data and distributing weight storage, the hardware structure is optimized to reduce memory access power consumption and supports a variety of convolution operators.
It improves the energy efficiency of neural network inference chips, enhances computing performance and flexibility, and meets the performance, cost and scalability requirements of terminal devices.
Smart Images

Figure CN114169512B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chip design, and more particularly to a neural network inference chip, a neural network inference method, and a terminal. Background Technology
[0002] With the deepening of research in the field of deep learning, convolutional neural networks have been applied to various fields of computer vision. Recent research and experiments have shown that convolutional neural networks exhibit absolute dominance over traditional computer vision algorithms in many image processing tasks such as object detection, face recognition, image classification, and semantic segmentation.
[0003] Current neural network inference chip architectures mainly include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs).
[0004] However, both GPUs and FPGAs use general-purpose computing modules to handle the computation of different operators in neural network algorithms. Therefore, their performance, power consumption, and scalability are not ideal for deep learning. Thus, how to more specifically optimize the hardware structure and layers to achieve optimal performance, power consumption, and scalability for deep learning algorithms has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a neural network inference chip, a neural network inference method, and a terminal, and proposes a high-efficiency neural network inference chip architecture for terminal devices, which meets the new requirements faced by terminals in terms of performance, cost, and scalability.
[0006] The technical solution of this application embodiment is implemented as follows:
[0007] In a first aspect, embodiments of this application provide a neural network inference chip, the neural network inference chip including an SSP and a computing module, wherein the computing module is composed of a VSP, an MMP and a PSUM;
[0008] The SSP is used to generate control commands based on the feature map and neural network structure parameters;
[0009] The VSP is used to generate convolution calculation instructions and convolution calculation data based on the control instructions;
[0010] The MMP stores weight data and is used to perform convolution calculations according to the convolution calculation instructions, the convolution calculation data, and the weight data to obtain a first calculation result.
[0011] The PSUM is used to determine the convolution result based on the first calculation result.
[0012] Secondly, embodiments of this application provide a neural network inference method, which is applied to a neural network inference chip. The neural network inference chip includes an SSP and a computing module, wherein the computing module is composed of a VSP, an MMP, and a PSUM; the method includes:
[0013] The SSP generates control commands based on the feature map and neural network structure parameters;
[0014] The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions;
[0015] The MMP performs convolution calculations according to the convolution calculation instructions, the convolution calculation data, and the stored weight data to obtain a first calculation result;
[0016] The PSUM determines the convolution result based on the first calculation result.
[0017] Thirdly, embodiments of this application provide a terminal, the terminal comprising: a generation unit, a calculation unit, and a determination unit.
[0018] The generation unit is used to generate control instructions based on the feature map and neural network structure parameters; and to generate convolution calculation instructions and convolution calculation data based on the control instructions.
[0019] The computing unit is configured to perform convolution calculation according to the convolution calculation instruction, the convolution calculation data, and the weight data to obtain a first calculation result.
[0020] The determining unit is used to determine the convolution result based on the first calculation result.
[0021] Fourthly, embodiments of this application provide a terminal, which includes a neural network inference chip, a processor, and a memory storing executable instructions. When the instructions are executed, the neural network inference method as described in the second aspect is implemented.
[0022] This application provides a neural network inference chip, a neural network inference method, and a terminal. The neural network inference chip includes an SSP (Signal Component Provider) and a computing module. The computing module consists of a VSP (Variable Component Provider), an MMP (Multi-Component Provider), and a PSUM (Power Component Provider). The SSP generates control instructions based on feature maps and neural network structure parameters. The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions. The MMP stores weight data and performs convolution calculations based on the convolution calculation instructions, convolution calculation data, and weight data to obtain a first calculation result. The PSUM determines the convolution result based on the first calculation result. In other words, in this application's embodiment, based on the neural network inference chip architecture, through weight reuse, partial input data reuse, and distributed weight storage, the energy consumption of the neural network inference chip for accessing off-chip data is greatly reduced, improving the energy efficiency of the neural network inference chip. It can support multiple types of convolution operators and has strong scalability and flexibility. Therefore, the high-efficiency neural network inference chip architecture for terminal devices proposed in this application meets the new requirements faced by terminals in terms of performance, cost, and scalability. Attached Figure Description
[0023] Figure 1 Schematic diagram of the composition structure of a neural network inference chip Figure 1 ;
[0024] Figure 2 Schematic diagram of the composition structure of a neural network inference chip Figure 2 ;
[0025] Figure 3 This is a schematic diagram of the VSP architecture;
[0026] Figure 4 This is a schematic diagram illustrating the storage method of feature maps;
[0027] Figure 5 This is a schematic diagram of the MMP architecture;
[0028] Figure 6 This is a schematic diagram of the mac tree structure;
[0029] Figure 7 A schematic diagram of the convolution of a 4x6 input feature block with a 3×3 kernel;
[0030] Figure 8 This is a schematic diagram of convolution based on a mac tree.
[0031] Figure 9 This is a schematic diagram of convolution calculation with a kernel size of 1×1;
[0032] Figure 10 This is a schematic diagram of the PSUM architecture;
[0033] Figure 11 This is a schematic diagram of the architecture of VSP, MMP, and PSUM;
[0034] Figure 12 Implementation flow diagram of neural network inference method Figure 1 ;
[0035] Figure 13 Implementation flow diagram of neural network inference method Figure 2 ;
[0036] Figure 14 Implementation flow diagram of neural network inference method Figure 3 ;
[0037] Figure 15 Implementation flow diagram of neural network inference method Figure 4 ;
[0038] Figure 16 Diagram of the terminal's structural composition Figure 1 ;
[0039] Figure 17 Diagram of the terminal's structural composition Figure 2 . Detailed Implementation
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant application and not for limiting the application. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant application are shown in the accompanying drawings.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.
[0042] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. It should also be noted that the terms "first, second, third" used in the embodiments of this application are merely for distinguishing similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0043] With the deepening of research in the field of deep learning, convolutional neural networks have been applied to various fields of computer vision. Recent research and experiments have shown that convolutional neural networks exhibit absolute dominance over traditional computer vision algorithms in many image processing tasks such as object detection, face recognition, image classification, and semantic segmentation.
[0044] As an important method in computer vision, neural networks take image data as input and, through a trained neural network, calculate the required semantic information, such as the category of objects. This process is called neural network inference.
[0045] Neural network inference chips have become a hot research topic in the field of artificial intelligence hardware in recent years. Neural networks are characterized by computational complexity and large data volumes, requiring extensive parallel computing and memory access, thus placing extremely high demands on computing power. When running neural network algorithms in terminal devices (such as mobile phones, smart homes, and autonomous vehicles), power consumption is strictly limited. Furthermore, different neural networks have significantly different structures, varying data sizes, and diverse operator types, thus also requiring a certain degree of chip flexibility.
[0046] Current neural network inference chip architectures mainly include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs).
[0047] While convolutional neural networks (CNNs) are continuously improving their performance in tasks such as classification and detection, the number of parameters and computational load of neural network models are also increasing significantly. This presents many challenges when applying neural network algorithms in practical applications. Specifically, both GPUs and FPGAs use general-purpose computing modules to handle the computation of different operators in neural network algorithms. Therefore, their performance and power consumption are not ideal when performing deep learning. GPUs, in particular, cannot fully leverage their parallel computing advantages in real-world inference scenarios, nor can they flexibly configure their hardware structure. FPGAs, on the other hand, have lower computational resource ratios, and their performance in terms of speed and power consumption is also not ideal.
[0048] With the development of artificial intelligence, the massive number of terminal applications require robust local hardware support, thus creating a significant demand for high-efficiency dedicated chips for neural network inference. How to more effectively optimize the hardware structure and layers to achieve optimal performance, power consumption, and area for deep learning algorithms has become a pressing issue.
[0049] To address the aforementioned issues, the embodiments of this application, based on a neural network inference chip architecture, significantly reduce the energy consumption of the neural network inference chip for accessing off-chip data through methods such as weight reuse, partial input data reuse, and distributed weight storage. This improves the energy efficiency of the neural network inference chip, supports various types of convolution operators, and exhibits strong scalability and flexibility. Therefore, the high-energy-efficiency neural network inference chip architecture for terminal devices proposed in this application meets the new requirements faced by terminals in terms of performance, cost, and scalability.
[0050] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0051] One embodiment of this application provides a neural network inference chip, which may include a scheduler sequence processor (SSP) and a computing module. The computing module consists of a vector scalar processor (VSP), a matrix multiply processor (MMP), and a partial sum accumulator (PSUM).
[0052] It should be noted that, in the embodiments of this application, the VSP may include N vsp_cores, and the MMP may include N 2 N mmp_cores, where N 2 Each mmp_core is arranged in N rows and N columns; PSUM can include N psum_cores.
[0053] It is understood that in this application, N can be an integer greater than 0.
[0054] It should be noted that, in the embodiments of this application, the N vsp_cores can work in parallel, and the N psum_cores can also work in parallel.
[0055] Furthermore, in the embodiments of this application, the neural network inference chip can be any form such as a terminal, chip, integrated circuit (IC), or application-specific integrated circuit (ASIC). For example, the neural network inference chip proposed in the embodiments of this application can be an embedded neural network processing unit (NPU).
[0056] The NPU employs a "data-driven parallel computing" architecture, making it particularly adept at processing massive amounts of multimedia data, such as video and images. The NPU works by simulating human neurons and synapses at the circuit layer and directly processing large-scale neurons and synapses using a deep learning instruction set; a single instruction completes the processing of a group of neurons. Compared to the Central Processing Unit (CPU) and GPU, the NPU improves operational efficiency by emphasizing weights to integrate storage and computation.
[0057] NPUs are built to mimic biological neural networks. While CPUs and GPUs require thousands of instructions to process neurons, NPUs can accomplish the same task with just one or a few instructions, giving them a significant advantage in processing efficiency for deep learning. Like GPUs, NPUs also require the assistance of the CPU to complete specific tasks.
[0058] Typically, a dedicated system-on-chip (SoC) for mobile devices integrates a CPU, GPU, and NPU. The CPU handles smooth switching between mobile applications, the GPU supports fast loading of game graphics, and the NPU is specifically responsible for artificial intelligence (AI) computations and the implementation of AI applications. In other words, the CPU is responsible for computation and overall coordination, the GPU handles image-related tasks, and the NPU handles AI-related tasks. The workflow is as follows: any task must first pass through the CPU, which then determines which component to assign it based on the nature of the task. Graphics-related calculations are assigned to the GPU, while AI-related calculations are assigned to the NPU.
[0059] Unlike GPU acceleration, NPUs do not output the computation results of each neuron to main memory. Instead, they are passed to the next neuron for further computation according to the connections of the neural network. Therefore, they have a significant improvement in both computing performance and power consumption.
[0060] NPUs simulate neurons at the circuit level, integrating storage and computation through synaptic weights. A single instruction completes the processing of a group of neurons, improving operational efficiency. They are primarily used in communications, big data, and image processing. As a type of dedicated ASIC (Application-Specific Integrated Circuit), the NPU is a chip customized to meet specific requirements. Aside from its lack of scalability, it offers advantages in power consumption, reliability, and size, especially in high-performance, low-power mobile devices.
[0061] Figure 1 Schematic diagram of the composition structure of a neural network inference chip Figure 1 ,like Figure 1As shown, the neural network inference chip 10 may include an SSP 11 and a computing module 12. The computing module 12 consists of a VSP 121, an MMP 122, and a PSUM 123. The SSP 11 is used to parse the information of each layer of the neural network and then generate corresponding control instructions according to the computational requirements, thereby enabling the computing module 12 to perform the corresponding calculations. The VSP 121 is used to store the input feature data of each layer of the neural network and perform operator calculations other than convolution. The MMP 122 is used to perform convolution operator calculations, and the PSUM 123 is used to perform partial sum accumulation.
[0062] Furthermore, in the embodiments of this application, SSP11 can be used to generate control instructions based on feature maps and neural network structure parameters.
[0063] After generating the control commands, SSP11 can also send the control commands to VSP121 of the computing module 12.
[0064] Furthermore, in the embodiments of this application, VSP121 can be used to generate convolution calculation instructions and convolution calculation data based on the control instructions after receiving the control instructions sent by SSP11.
[0065] After generating the convolution calculation instructions and convolution calculation data, VSP121 can also send the convolution calculation instructions and convolution calculation data to MMP122.
[0066] Furthermore, in the embodiments of this application, MMP122 stores weight data, wherein the weight data can be stored in mmp_core, and can then be used to perform convolution calculation based on the convolution calculation instruction and convolution calculation data sent by VSP121 after receiving the convolution calculation instruction and convolution calculation data, and the weight data stored in mmp_core to obtain a first calculation result.
[0067] After completing the first calculation, MMP122 can also send the first calculation result to PSUM123.
[0068] It should be noted that, in the embodiments of this application, the first calculation result represents the calculation result obtained after performing convolution calculation. The first calculation result can be the convolution calculation result transmitted from MMP122 to PSUM123, or it can be a partial sum of the convolution calculation results transmitted from MMP122 to PSUM123.
[0069] Furthermore, in the embodiments of this application, PSUM123 can be used to obtain the final convolution result based on the first calculation result after receiving the first calculation result sent by MMP122.
[0070] After determining the final convolution result, PSUM123 can also transmit the convolution result to VSP121 via MMP122. That is, in the embodiments of this application, MMP122 can also be used to transmit the convolution result determined by PSUM123 to VSP121, and correspondingly, VSP121 can also be used to output the convolution result.
[0071] It should be noted that, in the embodiments of this application, VSP121 can also be used to generate non-convolution calculation instructions and non-convolution calculation data based on the control instructions after receiving the control instructions sent by SSP11, and then perform non-convolution calculation according to the non-convolution calculation instructions and non-convolution calculation data to finally obtain the second calculation result.
[0072] It is understood that, in the embodiments of this application, the second calculation result represents the calculation result obtained after performing non-convolution calculation.
[0073] In other words, in the embodiments of this application, the operators of the neural network are very diverse, including convolutional operators and non-convolutional operators. Corresponding to the input feature map and neural network structure parameters, SSP11 can realize the inference of the neural network by sending control commands to the computing module 12. Specifically, VSP121 in the computing module 12 can generate convolutional calculation commands and convolutional calculation data, as well as non-convolutional calculation commands and non-convolutional calculation data, according to the control commands. The convolutional calculation commands and convolutional calculation data will be sent by VSP121 to MMP122 in the computing module 12 for corresponding convolutional calculation processing, while the non-convolutional calculation commands and non-convolutional calculation data will be directly processed by the internal logic processing part of VSP121 for corresponding non-convolutional calculation processing.
[0074] It should be noted that, in the embodiments of this application, whether it is the first calculation result obtained by MMP performing convolution calculation, the second calculation result obtained by the logic processing part in VSP performing non-convolution calculation, or the final convolution result, all can be stored in the memory in VSP, that is, stored in the data storage part of VSP.
[0075] Furthermore, in the embodiments of this application, the SSP in the neural network inference chip can be composed of a configuration information storage module CMEM and a ring buffer module.
[0076] It should be noted that, in the embodiments of this application, the ring cache module can be used to obtain feature maps and neural network structure parameters from external memory, and then store the feature maps and neural network structure parameters in CMEM.
[0077] Accordingly, in the embodiments of this application, SSP can also be used to read feature maps and neural network structure parameters from CMEM.
[0078] Figure 2 Schematic diagram of the composition structure of a neural network inference chip Figure 2 ,like Figure 2 As shown, the SSP11 in the neural network inference chip 10 can be composed of CMEM111 and Ring buffer112. The Ring buffer112 is a module for exchanging image data between the neural network inference chip 10 and external Double Data Rate (DDR) / system memory. After reading data from the external source, the Ring buffer112 can store the data in CMEM111. CMEM111 stores the configuration information of each layer of the network, such as operator type, input feature map size, and output feature map size. The SSP11 can read the configuration information of each layer stored in CMEM111, parse and process the configuration information of each layer, and finally generate control instructions for VSP121, MMP122, and PSUM123. These control instructions then schedule the computation module 12 to cooperate in implementing the corresponding operator operations.
[0079] Furthermore, in the embodiments of this application, the feature map and neural network structure parameters are input to N vsp_cores in the VSP through N channels respectively.
[0080] It should be noted that, in the embodiments of this application, corresponding to the N vsp_cores included in the VSP, the feature map and neural network structure parameters can be input to each of the N vsp_cores through N channels respectively.
[0081] Furthermore, in the embodiments of this application, the i-th vsp_core among the N vsp_cores includes an i-th logical processing part and an i-th data storage part. Here, i is an integer greater than or equal to 0 and less than N. That is, each vsp_core in the VSP is configured with one logical processing part and one data storage part.
[0082] Figure 3 A schematic diagram of the VSP architecture, as shown below. Figure 3As shown, a VSP can include 8 (N=8) vsp_cores, namely vsp_core0, vsp_core1, ..., vsp_core7. Each vsp_core includes a logical processing part and a data storage (Data Memory, DM) part, such as vsp_core0 being configured with DM_0 and vsp_core1 being configured with DM_1.
[0083] Furthermore, in the embodiments of this application, for the i-th vsp_core, the i-th logic processing part can be used to perform non-convolution calculations based on the data of the i-th channel; the i-th data storage part can be used to store the data of the i-th channel.
[0084] In other words, in the embodiments of this application, corresponding to N vsp_cores, the feature maps and neural network structure parameters are input to the N vsp_cores through N channels respectively, and the N vsp_cores process the data input from their respective channels. Each vsp_core can analyze and process the control commands sent by the SSP and the data input from the corresponding channel, generating corresponding convolution-related calculation commands and data, as well as non-convolution-related calculation commands and data. Then, it can send the convolution-related calculation commands and data to the MMP to perform convolution calculations, and simultaneously send the non-convolution-related calculation commands and data to the internally configured logic processing section to perform non-convolution calculations. The data storage section in each vsp_core can store the data input from the corresponding channel. For example, the data input from the first channel can be stored in DM_1 configured in vsp_core1, and the data input from the seventh channel can be stored in DM_7 configured in vsp_core7.
[0085] Figure 4 This is a schematic diagram illustrating the storage method of feature maps, such as... Figure 4As shown, a VSP can include 8 (N=8) vsp_cores: vsp_core0, vsp_core1, ..., vsp_core7, each configured with DM_0, DM_1, ..., DM_7. For each layer of feature maps, 8 channels can be allocated and stored in DM_0, DM_1, ..., DM_7 of the 8 vsp_cores. For the data storage portion DM of each vsp_core, assuming the width, height, and number of channels of the input feature map from the corresponding channel are IW, IH, and IC respectively, then: all input feature map data with IC = 8 × m input to vsp_core0 can be stored in DM_0; all input feature map data with IC = 8 × m + 1 input to vsp_core1 can be stored in DM_1; and all input feature map data with IC = 8 × m + 7 input to vsp_core7 can be stored in DM_7, where m is an integer greater than or equal to 0.
[0086] It is understood that, in the embodiments of this application, the data storage portion DM of each vsp_core can store the corresponding channel input feature map data in NHCW format. Here, N represents the number of convolutions, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map.
[0087] Furthermore, in embodiments of this application, the MMP may include an N row and N column N 2 There are N mmp_cores. 2 The j-th mmp_core can include M MACs and the j-th weight storage part. Here, j is greater than or equal to 0 and less than N. 2 The integer M is a positive integer. That is, each mmp_core in MMP is configured with M MACs and a weight storage section.
[0088] Figure 5 The following is a schematic diagram of the MMP architecture, such as... Figure 5 As shown, MMP can include 64 (N=8) mmp_cores in 8 rows and 8 columns, namely mmp_core00, mmp_core01, ..., mmp_core63. Each mmp_core includes a weight memory (WM) part, such as vsp_core00 being configured with WM_00 and mmp_core01 being configured with WM_01.
[0089] Furthermore, in the embodiments of this application, for the j-th mmp_core, the M MACs can be used to perform convolution calculations; the j-th weight storage part can be used to store the weight data of the convolution calculation.
[0090] It should be noted that, in the embodiments of this application, the weight data stored in the j-th weight storage part corresponding to the j-th mmp_core can be the weights used for convolution calculations by the M MACs in the j-th mmp_core. That is, the weight data for performing convolution calculations can be stored in a distributed manner, where N... 2 Each mmp_core only needs to store the weight data corresponding to the convolution calculation performed by that mmp_core.
[0091] It is understood that, in the embodiments of this application, N 2 Each mmp_core can be used to perform a convolution operation on a single input channel based on the input feature map and a convolution kernel to obtain the output point on the output channel plane (OC).
[0092] Furthermore, in the embodiments of this application, for any mmp_core in the MMP, the M MACs included can be configured in the form of at least one column of mac tree, wherein each column of mac tree can consist of at least one MAC.
[0093] Figure 6 Here is a schematic diagram of the mac tree structure, such as Figure 6 As shown, each mmp_core can consist of several columns of mactrees, with each column containing 9 MACs, enabling 3×3 convolution operations. For example, assuming an mmp_core contains 8 columns of mactrees, then in one round of convolution, 8 output points can be calculated. Each column of the mactree can perform convolution calculations within a 3×3 convolution window, resulting in 8 output parts.
[0094] It is understood that, in the embodiments of this application, the number of MAC trees inside each mmp_core can be customized according to the requirements of computational parallelism.
[0095] Furthermore, in embodiments of this application, the j-th mmp_core can be used to configure the input data format according to the structural parameters of at least one column of mac trees. The structural parameters of the at least one column of mac trees may include the number of mac trees and the number of MACs in each column of mac trees.
[0096] In other words, in the embodiments of this application, for any mmp_core in the MMP, the data format of the data input to the mmp_core can be restricted according to the number of MAC trees inside the mmp_core and the number of MACs in each column of the MAC tree. For example, if the number of MAC trees inside an mmp_core is 8 and the number of MACs in each column of the MAC tree is 9, then the input data can be restricted to a 9-row, 8-column data format.
[0097] Figure 7 This is a schematic diagram of the convolution of a 4x6 input feature block with a 3x3 kernel, as shown below. Figure 7 As shown, assuming the input feature map patch size passed from VSP to MMP is 4 rows and 6 columns, after multiplying with a 3×3 convolution kernel (with a stride of 1), an output feature map patch of 2 rows and 4 columns can be output.
[0098] Figure 8 This is a schematic diagram of convolution based on a mac tree, as shown below. Figure 8 As shown, based on the above Figure 7 Assuming the input feature map patch size from VSP to MMP is 4 rows and 6 columns, when performing convolution processing within an mmp_core, based on the structural parameters that the mmp_core contains 8 MAC trees and each column of the MAC tree contains 9 MACs, the 4 rows and 6 columns of feature map patches can first be rearranged into a 9 rows and 8 columns data format. Here, the 9 input data points in each column belong to the same 3×3 convolution window. These 9 data points are multiplied and accumulated with the 9 weight data points (k0 to k8) of the 3×3 convolution kernel to finally obtain the corresponding output points.
[0099] Furthermore, in the embodiments of this application, the j-th mmp_core can also be used to control at least one column of mac tree weight data to reuse during convolution calculation.
[0100] It should be noted that, in the embodiments of this application, since at least one column of the MAC tree in mmp_core uses the weighted repetition method, therefore, based on the above... Figure 8 After the 8 input data columns are distributed to 8 MAC trees, the 9 weight data obtained by the 8 MAC trees are completely identical when performing convolution calculation. That is, the 8 input data columns are convolved with the same convolution kernel, and each MAC tree outputs a partial sum of an output point.
[0101] It is understood that, in the embodiments of this application, for the same column of mmp_core, all mmp_cores in that column process the same IW (width) and IH (height) coordinates of the input feature map, only the input feature map data of different channels. The partial convolution calculation results of different channels calculated by the mmp_cores in the same column are accumulated sequentially from top to bottom and passed to the corresponding psum_core in PSUM by the mmp_core of the last row. The mmp_cores of different columns in MMP are used to process the convolution calculation of different convolution kernels with the input feature map. That is, the entire MMP can complete the accumulation of N input channels (IC) in the column direction each time, and can complete the partial convolution operation with N convolution kernels in the row direction each time, outputting the partial sum of the outputs on N OCs. Assuming N=8, and each mmp_core contains 8 columns of MAC trees, then an MMP's mmp_core can simultaneously process convolution calculations with a maximum IC=8, OC=8, OW=4, and OH=2, where OW represents the width of the output feature map and OH represents the height of the output feature map. When the values of IC, OC, OW, and OH exceed these upper limits, the input feature map and the number of kernels need to be divided multiple times, and convolution calculations are performed through cyclic block division.
[0102] It should be noted that, in the embodiments of this application, when the convolution kernel size is 3×3 and the stride is 1, the above mapping method can be used for convolution calculation, and the MAC utilization rate can reach 100%. Furthermore, the neural network inference chip proposed in the embodiments of this application can also support the calculation of convolution kernels of other sizes. For example, Figure 9 This is a schematic diagram of convolution calculation with a kernel size of 1×1, as shown below. Figure 9 As shown, for convolution computation with a 1×1 kernel, one MAC is used out of the eight MAC trees of each mmp_core. When performing a 1×1 convolution on the WH plane, the accumulation on the input channels is still achieved by passing the mmp_cores in the same column from top to bottom.
[0103] Furthermore, in the embodiments of this application, the i-th vsp_core of the VSP in the calculation module can be specifically used to control N based on the data of the i-th channel. 2 The i-th row of mmp_core performs convolution calculations.
[0104] It should be noted that, in the embodiments of this application, the N vsp_cores correspond to the N rows of mmp_cores, that is, each vsp_core in the VSP corresponds to N. 2A row of mmp_cores. Each vsp_core can control the convolution calculation of all mmp_cores in its corresponding row. Specifically, the convolution calculation instructions and data output by the i-th vsp_core can be passed sequentially to all mmp_cores in the i-th row.
[0105] In other words, in the embodiments of this application, the instructions and data obtained by the mmp_core in the same row for performing convolution calculations can be exactly the same. The difference is that there is a time delay in the time it takes for different mmp_cores in the same row to obtain the instructions and data for convolution calculations.
[0106] Furthermore, in the embodiments of this application, N 2 The mmp_core in the first row and i-th column of the mmp_core can be used to send the convolution calculation result corresponding to the mmp_core in the first row and i-th column to the mmp_core in the second row and i-th column.
[0107] In other words, in the embodiments of this application, for N rows and N columns N 2 For each mmp_core, after completing the convolution calculation based on the input instructions and data and obtaining the corresponding convolution calculation result, the mmp_core in the first row can send the convolution calculation result to the mmp_core in the same column of the next row, that is, to the second row of mmp_core in the same column.
[0108] Furthermore, in the embodiments of this application, N 2 The mmp_core in the kth row and ith column of the mmp_core can be used to send the sum of the k convolution calculation results corresponding to the mmp_core in the first k rows and ith column to the mmp_core in the (k+1)th row and ith column; where k is an integer greater than 1 and less than N.
[0109] In other words, in the embodiments of this application, for N rows and N columns N 2For each mmp_core, each mmp_core in the k-th row can receive the sum of the (k-1) convolution calculation results corresponding to the i-th column of the (k-1)-th row mmp_core sent by the mmp_core in the same column of the previous row. Then, after completing the convolution calculation according to the input instructions and data and obtaining the corresponding convolution calculation result, the convolution calculation result is accumulated with the sum of the (k-1) convolution calculation results corresponding to the i-th column of the (k-1)-th row mmp_core to determine the sum of the k convolution calculation results corresponding to the i-th column of the (k)-th row mmp_core. The sum of the k convolution calculation results is then sent to the mmp_core in the same column of the next row, that is, to the (k+1)-th row mmp_core in the same column.
[0110] Furthermore, in the embodiments of this application, N 2 The mmp_core in the Nth row and ith column of the mmp_core can be used to send a portion of the convolution calculation results of the N convolutions corresponding to all mmp_cores in the N rows and ith columns to the ith psum_core among the N psum_cores.
[0111] In other words, in the embodiments of this application, for N rows and N columns N 2 For each `mmp_core`, each `mmp_core` in the Nth row can receive the sum of the (N-1) convolutional calculation results corresponding to the `mmp_core` in the i-th column of the previous (N-1) rows, sent by the `mmp_core` in the same column of the previous row. Then, after completing the convolutional calculation according to the input instructions and data, and obtaining the corresponding convolutional result, this result is accumulated with the sum of the (N-1) convolutional calculation results corresponding to the `mmp_core` in the i-th column of the previous (N-1) rows. This determines the partial sum of the N convolutional calculation results corresponding to all `mmp_cores` in the i-th column of the N rows, and this partial sum is sent to `PSUM` for further accumulation calculation. Specifically, the `mmp_core` in the i-th column of the N rows can send this partial sum to the i-th `psum_core` among the N `psum_cores`.
[0112] It should be noted that, in the embodiments of this application, the N psum_cores correspond to the N columns of mmp_cores, that is, each psum_core in PSUM corresponds to N columns. 2 One column of mmp_cores. The partial sum of the N convolutional results corresponding to each mmp_core column can be transferred to a corresponding psum_core for further calculation and processing.
[0113] It is understood that, in the embodiments of this application, the first calculation result obtained by MMP performing convolution calculation may include N. 2 The partial sum of the N convolution calculation results corresponding to each column of mmp_core in mmp_core.
[0114] Furthermore, in the embodiments of this application, Figure 10 A schematic diagram of the PSUM architecture, as shown below. Figure 10 As shown, PSUM can include 8 (N=8) psum_cores, namely psum_core0, psum_core1, ..., psum_core7.
[0115] It should be noted that, in the embodiments of this application, the i-th psum_core among the N psum_cores can be used to process N. 2 The partial sums of the convolution calculation results of each round of the i-th column of mmp_core are stored, and the total partial sums are accumulated to obtain the final result.
[0116] It is understood that, in the embodiments of this application, the N psum_cores in PSUM can be used to implement the accumulation processing of partial convolution sums. If the number of channels IC of the input feature map is equal to N, N 2 The N columns of `mmp_core` perform convolution calculations on their respective channels, and the results are passed to the next row of `mmp_core` for sequential accumulation. The last row (the Nth row) of `mmp_core` stores the sum of all N convolution results, which is then passed to the corresponding `psum_core` in `PSUM`. If the number of channels (IC) of the input feature map is greater than N, then N... 2 The N columns of `mmp_core` cannot calculate all the input data from all channels at once, so they need to be calculated in multiple mappings. For example, assuming IC equals 2N, each column of `mmp_core` needs to first calculate the convolution of the first N channels, store the partial sums of the first N channels in the corresponding `psum_core` in `PSUM`, and then perform the convolution of the last N channels, storing the partial sums of the last N channels in the corresponding `psum_core` in `PSUM`. Finally, the corresponding `psum_core` can then be used to sum these two partial sums to obtain the accumulated result.
[0117] Furthermore, in the embodiments of this application, the i-th psum_core among the N psum_cores can be used to quantize the accumulated result to obtain the convolution result, and the convolution result is passed to the i-th vsp_core in sequence through N different mmp_cores.
[0118] It should be noted that, in the embodiments of this application, after the N psum_cores in PSUM have completed the storage and accumulation processing of partial convolution sums and obtained the accumulation result, the accumulation result can be further quantized to unify the format, precision and bit width of the accumulation result, and obtain the final convolution result. Then the convolution result can be written back to VSP.
[0119] It is understood that in the embodiments of this application, since PSUM and VSP are not directly connected, PSUM needs to pass the convolution result to VSP through MMP when returning the convolution result. Specifically, the i-th psum_core among the N psum_cores can sequentially pass the convolution result to the i-th vsp_core through N different mmp_cores.
[0120] In other words, in the embodiments of this application, when PSUM writes the convolution result back to VSP, the convolution result returned by each psum_core can be passed through N different mmp_cores before being written back to the corresponding vsp_core. For example, the convolution result returned by psum_core1 is passed to vsp_core1 after passing through mmp_core08, mmp_core09, mmp_core17, mmp_core25, mmp_core33, mmp_core41, mmp_core49, and mmp_core57 in sequence.
[0121] Furthermore, Figure 11 The following is a schematic diagram of the architecture of VSP, MMP, and PSUM, as shown below. Figure 11As shown, for the computation module in the neural network inference chip, the VSP includes 8 (N=8) vsp_cores, namely vsp_core0, vsp_core1, ..., vsp_core7, each configured with DM_0, DM_1, ..., DM_7. The MMP includes 64 mmp_cores arranged in 8 rows and 8 columns, namely mmp_core00, mmp_core01, ..., mmp_core63. Each mmp_core includes a weight storage part WM, such as vsp_core00 configured with WM_00, mmp_core01 configured with WM_01. The PSUM includes 8 psum_cores, namely psum_core0, psum_core1, ..., psum_core7. The 8 vsp_cores correspond to the 8 rows of mmp_cores, and each vsp_core can control the convolution calculation of all mmp_cores in its corresponding row. For example, the convolution calculation instructions and data output by vsp_core7 can be sequentially passed to all mmp_cores in row 8 (including mmp_core56 to mmp_core63), thereby controlling all mmp_cores in row 8 to perform the corresponding convolution calculations. The eight psum_cores correspond to the eight columns of mmp_cores, and the partial sum of the convolution calculation results for each column of mmp_cores can be transmitted to the corresponding psum_core for subsequent calculations and processing.
[0122] It should be noted that, in the embodiments of this application, each vsp_core can be responsible for controlling the convolution calculation of all mmp_cores in the corresponding row. For example, vsp_core0 can be responsible for controlling the convolution calculation of the mmp_cores (mmp_core00 to mmp_core07) in the first row of the MMP. A systolic array can be used to sequentially transmit the MMP-related control commands and input feature data output by the vsp_core to all mmp_cores in the corresponding row from left to right. That is, the input feature data obtained by the mmp_cores in the same row is exactly the same, only with a slight delay in the time it takes for each mmp_core to acquire data.
[0123] It should be noted that, in the embodiments of this application, based on the above... Figure 11When performing convolution calculations, time-division multiplexing can be used. Specifically, the neural network inference chip can compute a small block of output data (cube) using only a small block of input data in one cycle, and then compute the next output block (cube) in the next cycle, thus completing the mapping from the network algorithm to the fixed hardware structure. For example, each input block can be 4 rows and 6 columns of input data. The division of input blocks can be completed in the logic processing part of the VSP. Therefore, the input data passed from each vsp_core to the mmp_core in the same row is the divided input block.
[0124] Furthermore, in the embodiments of this application, based on the above... Figure 11 For mmp_core in the same column, it can process input feature map data with the same IW and IH coordinates but different channels. The partial convolution results of different channels calculated by mmp_core in the same column will be accumulated from top to bottom and passed to the corresponding psum_core by mmp_core in the last row.
[0125] It is understood that, in the embodiments of this application, based on the above... Figure 11 In MMP, different columns of `mmp_core` can be used to process convolution calculations between different convolution kernels and the input feature maps. That is, the entire MMP can perform the accumulation of 8 input channels (IC) at a time in the column direction, and perform partial convolution operations with 8 kernels at a time in the row direction, outputting partial sums of output channels (OC). Within each `mmp_core`, each `mmp_core` needs to perform convolution calculations between the input data from the corresponding `vsp_core` and the corresponding kernel.
[0126] Furthermore, in the embodiments of this application, based on the above... Figure 11 PSUM can be used to implement the accumulation of partial convolution sums. Each psum_core is responsible for the convolution and accumulation calculation of one column of mmp_cores. When the IC of the input feature map is greater than 8, the mmp_cores in each column of MMP cannot calculate all ICs at once, so multiple mappings and convolution calculations are required. For example, when the IC is equal to 16, each column of mmp_cores needs to first calculate the convolution of the first 8 channels, store the accumulation of the 8 channels in the corresponding psum_core, and then perform the convolution calculation of the last 8 channels. After the partial sum of the convolution of the last 8 channels is passed to the corresponding psum_core, the accumulation of these two partial sums is performed inside the corresponding psum_core to obtain the final accumulation result.
[0127] It should be noted that, in the embodiments of this application, in order to unify the format, precision and bit width of the output result, psum_core can further quantize the convolution accumulation result to obtain the final convolution result, and then write the final convolution result back to the data storage part DM configured by vsp_core.
[0128] Furthermore, in the embodiments of this application, based on the above... Figure 11 Since PSUM and VSP are not directly connected, the output of PSUM needs to be transmitted to VSP via MMP. The transmission path is as follows: Figure 11 As shown by the dashed lines, each psum_core needs to be passed through 8 mmp_cores before the data can be written back to the corresponding vsp_core. For example, psum_core0 passes the output convolution result through 8 mmp_cores to vsp_core0; psum_core1 passes the output convolution result through mmp_core08, mmp_core09, mmp_core17, mmp_core25, mmp_core33, mmp_core41, mmp_core49, and mmp_core57 in the MMP to vsp_core1.
[0129] It should be noted that the neural network inference chip proposed in this application can be a terminal or mobile chip architecture applicable to all applications of neural network inference. That is, this application proposes a dedicated neural network chip architecture design suitable for terminal devices. It can achieve a good mapping from algorithm to hardware through a variety of innovative design and architecture optimization methods, achieve very high energy efficiency, and the architecture has very strong scalability and flexibility, and can be applied to a variety of mobile application scenarios.
[0130] Furthermore, the neural network inference chip proposed in this application can significantly reduce the energy consumption of accessing off-chip data by means of weight reuse, partial input data reuse, and distributed weight storage, thereby improving the energy efficiency of the neural network inference chip. At the same time, the neural network inference chip proposed in this application can support multiple types of convolution operators, exhibiting extremely high flexibility.
[0131] It is understood that the number of computing units configured in the neural network inference chip proposed in this application can be modified as needed, thus exhibiting good scalability.
[0132] This application provides a neural network inference chip, which includes an SSP (Signal Component Provider) and a computing module. The computing module consists of a VSP (Variable Component Provider), an MMP (Multi-Component Provider), and a PSUM (Power Component Provider). The SSP generates control instructions based on feature maps and neural network structure parameters. The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions. The MMP stores weight data and performs convolution calculations based on the convolution calculation instructions, convolution calculation data, and weight data to obtain a first calculation result. The PSUM determines the convolution result based on the first calculation result. In other words, in this application's embodiment, based on the neural network inference chip architecture, through weight reuse, partial input data reuse, and distributed weight storage, the energy consumption for accessing off-chip data is greatly reduced, improving the energy efficiency of the neural network inference chip. It can support multiple types of convolution operators and has strong scalability and flexibility. Therefore, the high-energy-efficiency neural network inference chip architecture proposed in this application for terminal devices meets the new requirements faced by terminals in terms of performance, cost, and scalability.
[0133] Based on the above embodiments, another embodiment of this application provides a neural network inference method, which can be applied to a neural network inference chip. The neural network inference chip may include a control scheduler SSP and a computing module. The computing module consists of a vector processor VSP, a multiplication matrix processor MMP, and a partial sum accumulator PSUM.
[0134] It should be noted that, in the embodiments of this application, the VSP may include N vsp_cores, and the MMP may include N 2 N mmp_cores, where N 2 The mmp_cores are arranged in N rows and N columns; PSUM can include N psum_cores. N can be an integer greater than 0.
[0135] It should be noted that, in the embodiments of this application, the N vsp_cores can work in parallel, and the N psum_cores can also work in parallel.
[0136] Figure 12 Implementation flow diagram of neural network inference method Figure 1 ,like Figure 12 As shown in the embodiments of this application, the method for a neural network inference chip to perform neural network inference may include the following steps:
[0137] Step 101: SSP generates control commands based on feature maps and neural network structure parameters.
[0138] In the embodiments of this application, the SSP in the neural network inference chip can first generate control instructions based on the input feature map and neural network structure parameters, and then send the control instructions to the VSP in the neural network inference chip.
[0139] It should be noted that, in the embodiments of this application, the neural network inference chip can take any form, such as a terminal, a chip, an integrated circuit, or an application-specific integrated circuit. For example, the neural network inference chip proposed in the embodiments of this application can be a neural network processor (NPU).
[0140] Furthermore, in the embodiments of this application, the SSP in the neural network inference chip can be composed of a configuration information storage module CMEM and a ring buffer module.
[0141] It should be noted that, in the embodiments of this application, the circular cache module can obtain feature maps and neural network structure parameters from external memory, and then store the feature maps and neural network structure parameters in CMEM. Correspondingly, SSP can read feature maps and neural network structure parameters from CMEM.
[0142] Step 102: VSP generates convolution calculation instructions and convolution calculation data based on control instructions.
[0143] In the embodiments of this application, after the SSP generates control instructions based on the input feature map and neural network structure parameters and sends the control instructions to the VSP, the VSP can generate convolution calculation instructions and convolution calculation data based on the control instructions, and then send the convolution calculation instructions and convolution calculation data to the MMP.
[0144] Figure 13 Implementation flow diagram of neural network inference method Figure 2 ,like Figure 13 As shown in the embodiments of this application, after the SSP generates control commands based on the feature map and neural network structure parameters, and sends the control commands to the VSP, i.e., after step 101, the method for the neural network inference chip to perform neural network inference may include the following steps:
[0145] Step 105: The VSP generates non-convolution calculation instructions and non-convolution calculation data based on the control instructions, and performs non-convolution calculations according to the non-convolution calculation instructions and non-convolution calculation data to obtain the second calculation result.
[0146] In the embodiments of this application, after the SSP generates control instructions based on the input feature map and neural network structure parameters and sends the control instructions to the VSP, the VSP can also generate non-convolutional computation instructions and non-convolutional computation data based on the control instructions. Then, it can perform non-convolutional computation based on the non-convolutional computation instructions and non-convolutional computation data to obtain a second computation result.
[0147] In other words, in the embodiments of this application, the operators of the neural network are very diverse, including convolutional operators and non-convolutional operators. Corresponding to the input feature map and neural network structure parameters, the SSP can realize the inference of the neural network by sending control commands to the computing module. The VSP in the computing module can generate convolutional calculation commands and convolutional calculation data, as well as non-convolutional calculation commands and non-convolutional calculation data, according to the control commands. The convolutional calculation commands and convolutional calculation data will be sent by the VSP to the MMP in the computing module for corresponding convolutional calculation processing, while the non-convolutional calculation commands and non-convolutional calculation data will be directly processed by the internal logic processing part of the VSP for corresponding non-convolutional calculation processing.
[0148] Furthermore, in the embodiments of this application, the feature map and neural network structure parameters can be input into the N vsp_cores in the VSP through N channels respectively.
[0149] It should be noted that, in the embodiments of this application, corresponding to the N vsp_cores included in the VSP, the feature map and neural network structure parameters can be input to each of the N vsp_cores through N channels respectively.
[0150] Furthermore, in the embodiments of this application, the i-th vsp_core among the N vsp_cores includes an i-th logical processing part and an i-th data storage part. Here, i is an integer greater than or equal to 0 and less than N. That is, each vsp_core in the VSP is configured with a logical processing part and a data storage part. For example, the VSP may include 8 (N=8) vsp_cores, namely vsp_core0, vsp_core1, ..., vsp_core7, where each vsp_core includes a logical processing part and a data storage (Data Memory, DM) part, such as vsp_core0 being configured with DM_0 and vsp_core1 being configured with DM_1.
[0151] Furthermore, in the embodiments of this application, for the i-th vsp_core, the i-th logic processing part can perform non-convolutional computation based on the data of the i-th channel; the i-th data storage part can be used to store the data of the i-th channel.
[0152] In other words, in the embodiments of this application, corresponding to N vsp_cores, the feature maps and neural network structure parameters are input to the N vsp_cores through N channels respectively, and the N vsp_cores process the data input from their respective channels. Each vsp_core can analyze and process the control commands sent by the SSP and the data input from the corresponding channel, generating corresponding convolution-related calculation commands and data, as well as non-convolution-related calculation commands and data. Then, it can send the convolution-related calculation commands and data to the MMP to perform convolution calculations, and simultaneously send the non-convolution-related calculation commands and data to the internally configured logic processing section to perform non-convolution calculations. The data storage section in each vsp_core can store the data input from the corresponding channel. For example, the data input from the first channel can be stored in DM_1 configured in vsp_core1, and the data input from the seventh channel can be stored in DM_7 configured in vsp_core7.
[0153] It is understood that, in the embodiments of this application, the data storage portion DM of each vsp_core can store the corresponding channel input feature map data in NHCW format. Here, N represents the number of convolutions, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map.
[0154] Step 103: MMP performs convolution calculation according to the convolution calculation instructions, convolution calculation data and stored weight data to obtain the first calculation result.
[0155] In the embodiments of this application, after the VSP generates convolution calculation instructions and convolution calculation data based on control instructions and sends the convolution calculation instructions and convolution calculation data to the MMP, the MMP can perform convolution calculation according to the convolution calculation instructions, convolution calculation data and stored weight data to obtain a first calculation result, and then send the first calculation result to the PSUM.
[0156] It should be noted that, in the embodiments of this application, MMP stores weight data, which can be stored in mmp_core.
[0157] Furthermore, in embodiments of this application, the MMP may include an N row and N column N 2 There are N mmp_cores. 2 The j-th mmp_core can include M MACs and the j-th weight storage part. Here, j is greater than or equal to 0 and less than N. 2The integer M is a positive integer. That is, each mmp_core in the MMP is configured with M MACs and a weight storage part. For example, an MMP can include 64 (N=8) mmp_cores in 8 rows and 8 columns, namely mmp_core00, mmp_core01, ..., mmp_core63. Each mmp_core includes a weight storage part WM, such as vsp_core00 being configured with WM_00 and mmp_core01 being configured with WM_01.
[0158] Furthermore, in the embodiments of this application, for the j-th mmp_core, M MACs can perform convolution calculations; the j-th weight storage part can store the weight data for the convolution calculation. The weight data stored in the j-th weight storage part can be the weights used by the M MACs in the j-th mmp_core to perform convolution calculations. That is, the weight data for performing convolution calculations can be stored in a distributed manner, where N... 2 Each mmp_core only needs to store the weight data corresponding to the convolution calculation performed by that mmp_core.
[0159] It is understood that, in the embodiments of this application, for any mmp_core in an MMP, the M MACs included can be configured in the form of at least one column of MAC trees, where each column of MAC trees can consist of at least one MAC. The number of MAC trees within each mmp_core can be customized according to the computational parallelism requirements. For example, assuming an mmp_core includes 8 columns of MAC trees, and each column of MAC trees includes 9 MACs, then for one round of convolution operation, the computation of 8 output points can be completed. Each column of MAC trees can perform convolution computation with a 3×3 convolution window size, resulting in 8 output parts.
[0160] Furthermore, in embodiments of this application, the j-th mmp_core can be configured with the input data format according to the structural parameters of at least one column of mac trees. The structural parameters of at least one column of mac trees may include the number of mac trees and the number of MACs in each column of mac trees. For example, if an mmp_core has 8 mac trees and each column of mac trees contains 9 MACs, then the input data can be restricted to a 9x8 data format.
[0161] Furthermore, in the embodiments of this application, the j-th mmp_core can also be used to control at least one column of mac tree to reuse weight data when performing convolution calculation. Specifically, since at least one column of mac tree in mmp_core adopts the weight reuse method, the weight data obtained by at least one column of mac tree is exactly the same when performing convolution calculation.
[0162] Furthermore, in the embodiments of this application, the i-th vsp_core of the VSP in the calculation module can control N based on the data of the i-th channel. 2 The i-th row of mmp_core performs convolution calculations.
[0163] In other words, in the embodiments of this application, N vsp_cores correspond to N rows of mmp_cores, that is, each vsp_core in the VSP corresponds to N rows of mmp_cores. 2 A row of mmp_cores. Each vsp_core can control the convolution calculation of all mmp_cores in its corresponding row. Specifically, the convolution calculation instructions and data output by the i-th vsp_core can be sequentially passed to all mmp_cores in the i-th row. Therefore, the instructions and data obtained by mmp_cores in the same row for convolution calculation can be exactly the same; the difference lies in the time delay between different mmp_cores in the same row in obtaining the convolution calculation instructions and data.
[0164] Furthermore, in the embodiments of this application, N 2 The `mmp_core` in the first row and i-th column of an `mmp_core` can send the convolution calculation result corresponding to the `mmp_core` in the first row and i-th column to the `mmp_core` in the second row and i-th column. That is, after each `mmp_core` in the first row completes the convolution calculation according to the input instructions and data and obtains the corresponding convolution calculation result, it can send the convolution calculation result to the `mmp_core` in the second row of the same column.
[0165] Furthermore, in the embodiments of this application, N 2The `mmp_core` in the `k`th row and `i`th column of an `mmp_core` can send the sum of the k convolution results corresponding to the `mmp_core` in the `k`th row and `i`th column of the previous `mmp_core` to the `mmp_core` in the `(k+1)`th row and `i`th column of the previous `mmp_core`; where `k` is an integer greater than 1 and less than N. That is, each `mmp_core` in the `k`th row can receive the sum of the (k-1) convolution results corresponding to the `mmp_core` in the `(k-1)`th row and `i`th column of the previous `mmp_core`. Then, after completing the convolution calculation according to the input instructions and data and obtaining the corresponding convolution result, it accumulates this convolution result with the sum of the (k-1) convolution results corresponding to the `mmp_core` in the `(k-1)`th row and `i`th column of the previous `mmp_core` to determine the sum of the k convolution results corresponding to the `mmp_core` in the `k`th row and `i`th column of the previous `mmp_core`, and sends this sum of the k convolution results to the `mmp_core` in the `(k+1)`th row of the same column.
[0166] Furthermore, in the embodiments of this application, N 2 The `mmp_core` in the Nth row and i-th column of the `mmp_core` can send the partial sum of the N convolution calculation results corresponding to all `mmp_cores` in the N rows and i-th columns to the `ipsum_core` in the N `psum_cores`. That is, each `mmp_core` in the Nth row and i-th column can receive the sum of the (N-1) convolution calculation results corresponding to the `mmp_cores` in the previous (N-1) rows and i-th columns sent by the `mmp_core` in the same column of the previous row. Then, after completing the convolution calculation according to the input instructions and data and obtaining the corresponding convolution calculation result, it accumulates this convolution calculation result with the sum of the (N-1) convolution calculation results corresponding to the previous (N-1) rows and i-th columns of the `mmp_core`, thereby determining the partial sum of the N convolution calculation results corresponding to all `mmp_cores` in the N rows and i-th columns. This partial sum of the N convolution calculation results is then sent to `PSUM` for further accumulation calculation. Specifically, the `mmp_core` in the Nth row and i-th column can send this partial sum to the `i`-th `psum_core` in the N `psum_cores`.
[0167] It should be noted that, in the embodiments of this application, the N psum_cores correspond to the N columns of mmp_cores, that is, each psum_core in PSUM corresponds to N columns. 2 One column of mmp_cores. The partial sum of the N convolutional results corresponding to each mmp_core column can be transferred to a corresponding psum_core for further calculation and processing.
[0168] Step 104: PSUM determines the convolution result based on the first calculation result.
[0169] In the embodiments of this application, after the MMP performs convolution calculation according to the convolution calculation instructions, convolution calculation data and weight data stored in mmp_core, obtains the first calculation result and sends the first calculation result to PSUM, PSUM can further determine the convolution result according to the first calculation result and pass the convolution result to VSP through MMP.
[0170] In other words, in the embodiments of this application, MMP can also be used to transmit the convolution result determined by PSUM to VSP, and correspondingly, VSP can also be used to output the convolution result.
[0171] Furthermore, in embodiments of this application, PSUM may include N psum_cores, namely psum_core0, psum_core1, ..., psum_core(N-1). The i-th psum_core among the N psum_cores can be assigned to N... 2 The partial sums of the convolution calculation results of each round of the i-th column of mmp_core are stored, and the total partial sums are accumulated to obtain the final result.
[0172] It is understood that, in the embodiments of this application, the N psum_cores in PSUM can be used to implement the accumulation processing of partial convolution sums. If the number of channels IC of the input feature map is greater than N, N 2 The N columns of `mmp_core` cannot calculate all the input data from all channels at once, so they need to be calculated in multiple mappings. For example, assuming IC equals 2N, each column of `mmp_core` needs to first calculate the convolution of the first N channels, store the partial sums of the first N channels in the corresponding `psum_core` in `PSUM`, and then perform the convolution of the last N channels, storing the partial sums of the last N channels in the corresponding `psum_core` in `PSUM`. Finally, the corresponding `psum_core` can then be used to sum these two partial sums to obtain the accumulated result.
[0173] Furthermore, in the embodiments of this application, the i-th psum_core among the N psum_cores can quantize the accumulated result to obtain the convolution result, and then pass the convolution result to the i-th ivsp_core sequentially through N different mmp_cores. Specifically, after obtaining the accumulated result, the N psum_cores in PSUM can further quantize the accumulated result to unify the format, precision, and bit width of the accumulated result, thereby obtaining the final convolution result, which can then be written back to the VSP. The i-th psum_core among the N psum_cores can pass the convolution result to the i-th vsp_core sequentially through N different mmp_cores.
[0174] It should be noted that, in the embodiments of this application, whether it is the first calculation result obtained by MMP performing convolution calculation, the second calculation result obtained by the logic processing part in VSP performing non-convolution calculation, or the final convolution result, all can be stored in the memory in VSP, that is, stored in the data storage part of VSP.
[0175] Another embodiment of this application proposes a neural network inference method that is applied to a terminal equipped with a neural network inference chip. The neural network inference chip may include a control scheduler SSP and a computing module. The computing module consists of a vector processor VSP, a multiplication matrix processor MMP, and a partial sum accumulator PSUM.
[0176] It should be noted that, in the embodiments of this application, the VSP may include N vsp_cores, and the MMP may include N 2 N mmp_cores, where N 2 The mmp_cores are arranged in N rows and N columns; PSUM can include N psum_cores. N can be an integer greater than 0.
[0177] Furthermore, in the embodiments of this application, the SSP in the neural network inference chip can be composed of a configuration information storage module CMEM and a ring buffer module.
[0178] The neural network inference method proposed in this application can be applied to a terminal, wherein the terminal is equipped with a neural network inference chip. Figure 14 Implementation flow diagram of neural network inference method Figure 3 ,like Figure 14 As shown in the embodiments of this application, the method for a terminal to perform neural network inference may include the following steps:
[0179] Step 201: The SSP generates control commands based on the feature map and neural network structure parameters, and sends the control commands to the VSP.
[0180] Step 202: The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions, and sends the convolution calculation instructions and convolution calculation data to the MMP.
[0181] Step 203: MMP performs convolution calculation according to the convolution calculation instructions, convolution calculation data and weight data stored in mmp_core, obtains the first calculation result, and sends the first calculation result to PSUM.
[0182] Step 204: PSUM determines the convolution result based on the first calculation result and passes the convolution result to VSP through MMP.
[0183] In embodiments of this application, the neural network inference chip in the terminal may include a control scheduler (SSP) and a computation module (Tile). The Tile includes a vector processing module (VSP), a convolution processing module (MMP), and a partial sum processing module (PSUM). The SSP is responsible for parsing the information of each layer of the convolutional neural network and generating corresponding control instructions to control the Tile according to computational requirements. The VSP is mainly responsible for storing the input feature data of each layer of the network and performing operator computations other than convolution. The MMP mainly consists of a MAC array and is responsible for the computation of convolution operators. The PSUM is responsible for accumulating the convolution partial sums and writing the final convolution result back to the data storage module of the VSP.
[0184] The neural network inference method proposed in this application can be applied to a terminal, wherein the terminal is equipped with a neural network inference chip. Figure 15 Implementation flow diagram of neural network inference method Figure 4 ,like Figure 15 As shown in the embodiments of this application, the method for a terminal to perform neural network inference may include the following steps:
[0185] Step 301: Obtain feature maps and neural network structure parameters.
[0186] Step 302: Generate control instructions based on the feature map and neural network structure parameters.
[0187] Step 303: Determine whether to perform convolution calculation based on the control instructions. If yes, proceed to step 305; otherwise, proceed to step 306.
[0188] Step 304: Schedule MMP and PSUM to perform convolution calculations.
[0189] Step 305: Schedule the VSP to perform non-convolutional computation.
[0190] It should be noted that, in the embodiments of this application, the SSP includes a configuration information storage module (CMEM) and a ring buffer module. The CMEM stores configuration information for each layer of the network, such as operator type, input feature map size, and output feature map size. The ring buffer module is used for exchanging image data between the neural network inference chip and external DDR / system memory. The SSP reads the configuration information of each layer from the CMEM, parses and processes it, and generates control instructions for VSP, MMP, and PSUM, thereby scheduling these three parts in the computation module Tile to cooperate in implementing the corresponding operator operations.
[0191] Furthermore, in the embodiments of this application, the VSP comprises n vsp_cores, which operate in parallel, each processing and storing data input from different input channels. Each vsp_core may consist of a logic processing section and a data storage section. Each vsp_core receives control commands sent by the SSP, and after further analysis and processing, sends convolution-related calculation commands and data to MMP and PSUM to perform convolution calculations. Non-convolution-related commands and data are sent to the logic processing section within the vsp_core to perform non-convolution-related calculations. The storage section of each vsp_core stores the data of the corresponding channel of the input feature map for each layer. Specifically, assuming there are 8 vsp_cores, the channels are allocated to the memory of each vsp_core in an 8-aligned manner. The memory of each vsp_core stores the input feature map in NHCW format.
[0192] Furthermore, in the embodiments of this application, each mmp_core is mainly responsible for performing convolution operations between the input feature map and a kernel on a single input channel to obtain several output points on the output channel plane (OC). Each mmp_core internally consists of several columns of MAC trees, each column including several MACs. For example, assuming an mmp_core has 8 columns of MAC trees, one cycle can complete the calculation of 8 output points. If each column of MAC trees includes 9 MACs, then a convolution calculation with a 3×3 convolution window size can be achieved, obtaining 8 output partial sums. Specifically, since a weight reuse method can be used, the 8 columns of input data can be respectively assigned to the 8 columns of MAC trees, and the 9 weight data obtained from the 8 columns of MAC trees can be completely identical.
[0193] It should be noted that, in the embodiments of this application, the number of MAC trees inside each mmp_core can be configured according to the requirements of computational parallelism.
[0194] Furthermore, in the embodiments of this application, for all mmp_cores in the same column, they process input feature map data with the same IW and IH coordinates but different channels. The partial convolution results of different channels calculated by the mmp_cores in the same column are accumulated sequentially from top to bottom and passed to the corresponding psum_core in PSUM by the mmp_core in the last row. The mmp_cores of different columns in the MMP can be used to process convolution calculations between different kernels and the input feature map. That is, the entire MMP can complete the accumulation of N input channels in the column direction each time, and can complete partial convolution operations with N kernels in the row direction each time, outputting partial sums of N OCs. Assuming N=8, and each mmp_core has 8 columns of MAC trees, the MMP array can process convolution calculations with a maximum of IC=8, OC=8, OW=4, and OH=2 simultaneously. When the values of IC, OC, OW, and OH exceed these upper limits, the input feature map and the number of kernels need to be divided multiple times, and convolution calculations are performed through cyclic block division.
[0195] Furthermore, in the embodiments of this application, PSUM mainly implements the accumulation of partial convolution sums. Specifically, each column of mmp_core can correspond to a psum_core, that is, each kernel has a corresponding psum_core for storing and accumulating partial sums. For example, when IC=8, the 8 mmp_cores perform convolution calculations for their respective channels and pass the calculation results to the mmp_core in the next row for sequential accumulation. The last row of mmp_cores passes the convolution result accumulated from the 8 channels to the corresponding psum_core in PSUM for storage. When the number of channels in the input feature map is greater than 8, the mmp_core in each column of MMP cannot complete the calculation of all ICs at once, so it needs to be calculated in multiple mappings. For example, when IC=16, the mmp_core in each column needs to first calculate the convolution calculation of the first 8 channels, store the accumulated sum in the psum_core in PSUM, and then perform the calculation of the last 8 channels. After the convolutional parts of the last 8 channels are passed to psum_core in PSUM, psum_core can perform the accumulation of these two partial sums.
[0196] In summary, based on steps 101 to 105, 201 to 204, and 301 to 305, the neural network inference method proposed in this application is a solution applicable to all terminal or mobile chip architectures that apply neural network inference. It can achieve a good mapping from algorithm to hardware through various innovative design and architecture optimization methods, achieving very high energy efficiency. Furthermore, the architecture has very strong scalability and flexibility, making it suitable for a wide variety of mobile application scenarios.
[0197] Furthermore, the neural network inference method proposed in this application can significantly reduce the energy consumption of the neural network inference chip for accessing off-chip data by means of weight reuse, partial input data reuse, and distributed weight storage, thereby improving the energy efficiency of the neural network inference chip. At the same time, the neural network inference method proposed in this application can support multiple types of convolution operators, exhibiting extremely high flexibility.
[0198] This application provides a neural network inference method applied to a neural network inference chip and a terminal equipped with such a chip. The neural network inference chip includes an SSP (Signal Component Provider) and a computing module. The computing module consists of a VSP (Variable Component Provider), an MMP (Multi-Component Provider), and a PSUM (Power Component Provider). The SSP generates control instructions based on feature maps and neural network structure parameters. The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions. The MMP stores weight data and performs convolution calculations based on the convolution calculation instructions, convolution calculation data, and weight data to obtain a first calculation result. The PSUM determines the convolution result based on the first calculation result. In other words, in this application's embodiment, based on the neural network inference chip architecture, by using weight reuse, partial input data reuse, and distributed weight storage, the energy consumption of the neural network inference chip for accessing off-chip data is significantly reduced, improving the energy efficiency of the neural network inference chip. It can support multiple types of convolution operators and has strong scalability and flexibility. Therefore, the high-energy-efficiency neural network inference chip architecture proposed in this application for terminal devices meets the new requirements faced by terminals in terms of performance, cost, and scalability.
[0199] Based on the above embodiments, in another embodiment of this application... Figure 16 Diagram of the terminal's structural composition Figure 1 ,like Figure 16 As shown, the terminal 20 proposed in this embodiment may include a generation unit 21, a calculation unit 22, and a determination unit 23.
[0200] The generation unit 21 is used to generate control instructions based on the feature map and neural network structure parameters; and to generate convolution calculation instructions and convolution calculation data based on the control instructions.
[0201] The calculation unit 22 is used to perform convolution calculation according to the convolution calculation instruction, the convolution calculation data and the weight data to obtain a first calculation result;
[0202] The determining unit 23 is used to determine the convolution result based on the first calculation result.
[0203] In the embodiments of this application, further, Figure 17 Diagram of the terminal's structural composition Figure 2 ,like Figure 17 As shown, the terminal 20 proposed in this application embodiment may further include a memory 25 storing executable instructions of the processor 24. Furthermore, the neural network inference chip 10 may further include a communication interface 26, and a bus 27 and a neural network inference chip 28 for connecting the processor 24, the memory 25 and the communication interface 26.
[0204] In the embodiments of this application, the processor 24 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other types, and this application embodiment does not specifically limit this. The neural network inference chip 10 may also include a memory 25, which can be connected to the processor 24. The memory 25 is used to store executable program code, which includes computer operation instructions. The memory 25 may include high-speed RAM memory and may also include non-volatile memory, such as at least two disk drives.
[0205] In embodiments of this application, bus 27 is used to connect communication interface 26, processor 24, and memory 25, as well as the mutual communication between these devices.
[0206] In embodiments of this application, memory 25 is used to store instructions and data.
[0207] In practical applications, the aforementioned memory 25 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 24.
[0208] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0209] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0210] This application provides a terminal equipped with a neural network inference chip. The neural network inference chip includes an SSP (Surface Mount Technology SP) and a computing module. The computing module consists of a VSP (Variable Scale SP), an MMP (Multi-Scale Module), and a PSUM (Power Scale Unit). The SSP generates control instructions based on feature maps and neural network structure parameters. The VSP generates convolution calculation instructions and convolution calculation data based on the control instructions. The MMP stores weight data and performs convolution calculations based on the convolution calculation instructions, convolution calculation data, and weight data to obtain a first calculation result. The PSUM determines the convolution result based on the first calculation result. In other words, in this application's embodiment, based on the neural network inference chip architecture, through weight reuse, partial input data reuse, and distributed weight storage, the energy consumption of the neural network inference chip for accessing off-chip data is greatly reduced, improving the energy efficiency of the neural network inference chip. It can support multiple types of convolution operators and has strong scalability and flexibility. Therefore, the high-energy-efficiency neural network inference chip architecture for terminal devices proposed in this application meets the new requirements faced by terminals in terms of performance, cost, and scalability.
[0211] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0212] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the schematic and / or block diagrams can be implemented by computer program instructions, and combinations of blocks in the schematic and / or block diagrams can be implemented. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the schematic and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0213] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0214] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0215] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A neural network inference chip, comprising: The neural network inference chip comprises a control scheduler SSP and a calculation module, wherein the calculation module is composed of a vector processor VSP, a multiplication matrix processor MMP and a partial sum accumulator PSUM; The SSP is configured to generate a control instruction according to the feature map and the neural network structure parameter; The VSP is configured to generate a convolution calculation instruction and convolution calculation data based on the control instruction; The MMP is configured to store weight data and perform convolution calculation according to the convolution calculation instruction, the convolution calculation data and the weight data to obtain a first calculation result; The PSUM is configured to determine a convolution result according to the first calculation result; The MMP is further configured to transmit the convolution result determined by the PSUM to the VSP.
2. The neural network inference chip according to claim 1, wherein The VSP is further configured to output the convolution result.
3. The neural network inference chip according to claim 2, wherein The VSP is further configured to generate a non-convolution calculation instruction and non-convolution calculation data based on the control instruction, and perform non-convolution calculation according to the non-convolution calculation instruction and the non-convolution calculation data to obtain a second calculation result; The VSP is further configured to store the first calculation result, the second calculation result and the convolution result.
4. The neural network inference chip of claim 3, wherein, The VSP comprises N vsp_cores, and the feature map and the neural network structure parameter are input to the N vsp_cores in the VSP in N channels respectively; wherein N is an integer greater than 0.
5. The neural network inference chip according to claim 4, wherein The i-th vsp_core in the N vsp_cores comprises an i-th logical processing part and an i-th data storage part; the MMP comprises N 2 mmp_cores; the j-th mmp_core in the N 2 mmp_cores comprises M multiply-accumulate units (MACs) and a j-th weight storage part; wherein i is an integer greater than or equal to 0 and less than N; j is an integer greater than or equal to 0 and less than N 2 ; and M is an integer greater than 0.
6. The neural network inference chip according to claim 5, wherein The i-th logic processing part is configured to perform non-convolution calculation according to the data of the i-th channel; The i-th data storage part is configured to store the data of the i-th channel.
7. The neural network inference chip according to claim 5, wherein The M MACs are configured to perform convolution calculation; The j-th weight storage part is configured to store the weight data for convolution calculation.
8. The neural network inference chip according to claim 5, wherein The M MACs are configured in the form of at least one column of mac tree; wherein each column of mac tree is composed of at least one MAC.
9. The neural network inference chip according to claim 8, wherein The j-th mmp_core is configured to configure the input data form according to the structure parameter of the at least one column of mac tree; The j-th mmp_core is further configured to control the at least one column of mac tree to multiplex the weight data when performing convolution calculation.
10. The neural network inference chip according to claim 5, wherein The ith vs p_core is configured to control the ith row mmp_core in the N 2 mmp_cores to perform convolution calculation according to the data of the ith channel.
11. The neural network inference chip of claim 5, wherein, The PSUM comprises N psum_cores; The N 2 The first row i column mmp_core in the mmp_core is configured to send the convolution calculation result corresponding to the first row i column mmp_core to the second row i column mmp_core. The N 2 The mmp_core in the i-th column of the k-th row, for sending a sum value of k convolution calculation results corresponding to the i-th column mmp_core of the first k rows to the i-th column mmp_core of the (k+1)-th row; wherein k is an integer greater than 1 and less than N. The N 2 The Nth row i-th column mmp_core in the mmp_core is configured to send the partial sum of the N convolution calculation results corresponding to all N rows i-th column mmp_core to the i-th psum_core in the N psum_core.
12. The neural network inference chip according to claim 11, wherein The i-th psum_core among the N psum_cores is used to process the N 2 The partial sums of the convolution calculation results of each round of the i-th column of mmp_core are stored, and the total partial sums are accumulated to obtain the cumulative result.
13. The neural network inference chip according to claim 12, wherein An i-th psum_core in the N psum_cores is configured to quantize the accumulation result, obtain the convolution result, and sequentially transmit the convolution result to the i-th vsp_core through N different mmp_cores.
14. The neural network inference chip of claim 1, wherein, The SSP is composed of a configuration information storage module (CMEM) and a ring buffer module, The ring buffer module is configured to obtain the feature map and the neural network structure parameter from an external memory and store the feature map and the neural network structure parameter to the CMEM. The SSP is configured to read the feature map and the neural network structure parameter from the CMEM.
15. A neural network inference method, comprising: The neural network inference method is applied to a neural network inference chip, and the neural network inference chip includes a control scheduler (SSP) and a calculation module, wherein the calculation module is composed of a vector processor (VSP), a multiplication matrix processor (MMP), and a partial sum accumulator (PSUM); the method includes: The SSP generates a control instruction according to a feature map and a neural network structure parameter; The VSP generates a convolution calculation instruction and convolution calculation data based on the control instruction; The MMP performs convolution calculation according to the convolution calculation instruction, the convolution calculation data, and stored weight data to obtain a first calculation result; The PSUM determines a convolution result according to the first calculation result; The MMP transmits the convolution result determined by the PSUM to the VSP.
16. The method of claim 15, wherein The VSP outputs the convolution result.
17. The method of claim 16, wherein The VSP generates a non-convolution calculation instruction and non-convolution calculation data based on the control instruction, and performs non-convolution calculation according to the non-convolution calculation instruction and the non-convolution calculation data to obtain a second calculation result; The VSP stores the first calculation result, the second calculation result, and the convolution result.
18. The method of claim 17, wherein, The VSP includes N vsp_cores; wherein N is an integer greater than 0; The feature map and the neural network structure parameter are respectively input to the N vsp_cores in the VSP according to N channels.
19. The method of claim 18, wherein, The i-th vsp_core in the N vsp_cores comprises an i-th logical processing part and an i-th data storage part; the MMP comprises N 2 mmp_cores; the j-th mmp_core in the N 2 mmp_cores comprises M multiply-accumulate units (MACs) and a j-th weight storage part; wherein i is an integer greater than or equal to 0 and less than N; j is an integer greater than or equal to 0 and less than N 2 ; and M is an integer greater than 0.
20. The method of claim 19, wherein The i-th logical processing part performs non-convolution calculation according to data of the i-th channel; The i-th data storage part stores the data of the i-th channel.
21. The method of claim 19, wherein The M MACs perform convolution calculation; The j-th weight storage part stores the weight data of the convolution calculation.
22. The method of claim 21, wherein The M MACs are configured in the form of at least one column of mac trees; wherein each column of mac trees is composed of at least one MAC.
23. The method of claim 22, wherein The j-th mmp_core configures the input data form according to the structure parameters of the at least one column of mac trees. The jth mmp_core controls the at least one column mac tree to multiplex the weight data when performing convolution calculation.
24. The method of claim 19, wherein, The i-th vsp_core controls the i-th row mmp_core in the N 2 mmp_cores to perform convolution calculation according to the data of the i-th channel.
25. The method of claim 19, wherein, The PSUM comprises N psum_cores; The N 2 The first row and the ith column mmp_core in the mmp_core sends the convolution calculation result corresponding to the first row and the ith column mmp_core to the second row and the ith column mmp_core. The N 2 The mmp_core in the i-th column of the k-th row sends a sum value of k convolution calculation results corresponding to the mmp_core in the i-th column of the k-th row to the mmp_core in the i-th column of the (k+1)-th row; wherein k is an integer greater than 1 and less than N. The N 2 The Nth row i-th column mmp_core in the mmp_core sends the partial sum of the N convolution calculation results corresponding to the N rows i-th column mmp_core to the i-th psum_core in the N psum_core.
26. The method of claim 25, wherein, The i-th psum_core among the N psum_cores is related to the N 2 The partial sums of the convolution calculation results of each round of the i-th column of mmp_core are stored, and the total partial sums are accumulated to obtain the cumulative result.
27. The method of claim 26, wherein, An ith psum_core of the N psum_cores quantizes the accumulated result to obtain the convolution result, and sequentially passes the convolution result to the ith vsp_core through N different mmp_cores.
28. The method of claim 15, wherein, The SSP is composed of a configuration information storage module CMEM and a ring buffer module, The ring buffer module obtains the feature map and the neural network structure parameter from an external memory, and stores the feature map and the neural network structure parameter to the CMEM; The SSP reads the feature map and the neural network structure parameter from the CMEM.
29. A terminal, characterized by The terminal comprises a generating unit, a calculating unit, and a determining unit, The generating unit is configured to generate a control instruction according to the feature map and the neural network structure parameter, and generate a convolution calculation instruction and convolution calculation data based on the control instruction; The calculating unit is configured to perform convolution calculation according to the convolution calculation instruction, the convolution calculation data, and weight data, and obtain a first calculation result; The determining unit is configured to determine a convolution result according to the first calculation result; The terminal is further configured to transmit the determined convolution result.
30. A terminal, characterized by The terminal comprises a neural network inference chip, a processor, and a memory storing executable instructions, and when the instructions are executed, the method in any one of claims 15-28 is implemented.
Citation Information
Patent Citations
Method for extracting data features and related device
CN111914996A
Operation Accelerator, Processing Method, and Related Device
US20210224125A1