An in-memory accelerator for accelerating convolutional neural network inference

CN117852589BActive Publication Date: 2026-09-29INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311818871.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2026-09-29
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

[0004]然而,由于非易失性存储阵列是在模拟域进行运算的,而加速器的其他电路工作在数字域,因此在电路中还需要数字模拟转换器(DAC)、模拟数字转换器(ADC)、采样保持(S&H)、移位累加(S&A)等外围电路来完成运算,而深度神经网络加速器的片上内存容量有限,这些外围电路的存在可能会增加整个系统的复杂性和功耗,并且由于模型的大小和复杂度不断增加,可能会导致加速器上的内存不足以至于加速器的吞吐量降低,从而降低了深度神经网络的推理速度,因此,如何有效利用片上本地内存成为了一个关键问题

Benefits of technology

[0019]1)能够利用预设的遗传算法自动根据硬件资源量选取各处理层的复制倍数以及最佳的任务映射核心,充分利用硬件资源,确保各个核心之间任务均衡分配,不仅提高了计算的并行性,还能不影响计算任务的分配和结构冲突,使得加速器的吞吐量增加,加快了神经网络的推理速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117852589B_ABST
    Figure CN117852589B_ABST
Patent Text Reader

Abstract

The application provides an in-memory accelerator for accelerating convolutional neural network inference, the in-memory accelerator comprising a global memory, an on-chip router, and a plurality of cores connected to the global memory, each core comprising: a control unit configured to obtain an instruction stream, and control respective units to perform corresponding operations based on the instruction stream, wherein the instruction stream comprises: a computation operation and a memory access operation; a local memory unit configured to perform the memory access operation, access input data of the convolutional neural network in the global memory in order according to a size of a sliding window to obtain input data corresponding to the computation operation, and send the input data corresponding to the computation operation to an in-memory computation matrix unit; the in-memory computation matrix unit comprising a plurality of in-memory computation arrays, each in-memory computation array configured to perform the computation operation, and perform a matrix-vector multiplication calculation according to the input data corresponding to the computation operation; and a vector functional unit configured to perform the computation operation, and perform post-processing according to a result of the matrix-vector multiplication calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to neural network processor architecture and design methods, specifically to the field of hardware acceleration for neural network model computation, and more specifically, to an in-memory accelerator for accelerating convolutional neural network inference. Background Technology

[0002] In recent years, deep neural networks have achieved significant breakthroughs in various tasks. Due to the exponential growth in the parameter scale of neural network models, the industry expects substantial performance improvements in hardware to efficiently run deep neural network algorithms. Therefore, various deep neural network accelerators have been proposed. However, traditional accelerators based on Complementary Metal Oxide Semiconductor (CMOS) or following the von Neumann architecture may encounter challenges when processing large-scale neural network algorithms, such as the memory wall problem, leading to bottlenecks in improvements in storage, bandwidth, and energy efficiency.

[0003] Currently, in-memory computing is considered a crucial technology for solving the memory wall problem. It leverages the characteristics of non-volatile memory devices to combine computation and storage functions, effectively addressing the memory wall issue. Furthermore, in-memory computing has become a hot research area in deep neural network accelerator design. In-memory computing has various implementation methods, among which emerging non-volatile memory devices have great potential to challenge the dominance of CMOS. These emerging non-volatile memory devices integrate these devices into a cross-point array in a two-dimensional form, resulting in an in-memory computing array. This two-dimensional array structure composed of non-volatile memory is attracting increasing interest due to its high storage density and highly parallel in-situ computing characteristics. For ease of understanding, Figure 1 This diagram illustrates a common two-dimensional array composed of non-volatile memory devices. Due to their hardware characteristics, emerging non-volatile memory arrays can often efficiently perform matrix-vector multiplication operations using analog circuitry. Data is programmed as conductance into each node of the array, and input is applied as voltage to each row; according to Ohm's law, the current at each node is I. ij =G ij V j According to Kirchhoff's laws, the accumulated current value I that can be read in each column is... i =∑ j G ij V j Based on this, the array can perform parallel computation of the product operation between matrix G and vector V.

[0004] However, since non-volatile memory arrays operate in the analog domain while other circuits in the accelerator operate in the digital domain, peripheral circuits such as digital-to-analog converters (DACs), analog-to-digital converters (ADCs), sample-and-hold (S&H) circuits, and shift-and-accumulate (S&A) circuits are required to complete the computation. The on-chip memory capacity of deep neural network accelerators is limited, and the presence of these peripheral circuits may increase the complexity and power consumption of the entire system. Furthermore, as the size and complexity of the models continue to increase, there may be insufficient memory on the accelerator, leading to a decrease in the accelerator's throughput and thus reducing the inference speed of deep neural networks. Therefore, how to effectively utilize on-chip local memory has become a critical issue.

[0005] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solutions of the present invention, and does not imply that the relevant information is necessarily prior art. In the absence of evidence indicating that the relevant information was disclosed before the filing date of this invention, the relevant information should not be considered prior art. Summary of the Invention

[0006] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide an in-memory accelerator for accelerating inference in convolutional neural networks.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] According to a first aspect of the present invention, an in-memory accelerator for accelerating convolutional neural network inference is provided. The in-memory accelerator includes global memory, on-chip routing, and multiple cores connected to the global memory. The global memory is used to store data of the convolutional neural network. The on-chip routing is used to send data from one core to other cores or receive data sent by other cores. Each core includes: a control unit, a vector function unit, an in-memory computation matrix unit, and a local memory unit. The control unit is used to acquire an instruction stream and control each unit to execute corresponding operations based on the instruction stream. The instruction stream includes computation operations and memory access operations. The local memory unit is used to execute the memory access operation, accessing the input data of the convolutional neural network in the global memory sequentially according to the size of a sliding window. The system obtains the input data corresponding to the computational operation and sends the corresponding input data to the in-memory computation matrix unit. When accessing the input data of the convolutional neural network by sliding window, the system reassembles the newly added data of the sliding window to be accessed from global memory and the existing sliding window data in local memory, pixel by pixel, to obtain the sliding window data to be accessed. The in-memory computation matrix unit includes multiple in-memory computation arrays, each used to execute the computational operation and perform matrix-vector multiplication calculations based on the input data corresponding to the computational operation. The vector function unit executes the computational operation and performs post-processing based on the result of the matrix-vector multiplication calculation. Finally, through the cooperation of the in-memory computation matrix unit and the vector function unit, the inference result of the convolutional neural network is obtained.

[0009] In some embodiments of the present invention, when accessing the input data of a convolutional neural network by sliding window, the difference and intersection of the existing sliding window data and the data of the sliding window to be accessed are generated. The data of the difference set is accessed from global memory and the data of the intersection obtained from local memory unit are recombined to obtain the sliding window data to be accessed.

[0010] In some embodiments of the present invention, the data of the difference set accessed from global memory and the data of the intersection obtained from local memory units are recombined according to the order of sliding window data reading to obtain the sliding window data to be accessed.

[0011] In some embodiments of the present invention, the local memory unit is configured to divide the sliding window data accessed from the global memory according to the computation operation to obtain the input data corresponding to the computation operation.

[0012] In some embodiments of the present invention, the step of sequentially accessing the input data of the convolutional neural network in global memory according to the size of the sliding window to obtain the input data corresponding to the computation operation includes: sequentially acquiring the data contained in the first sliding window according to the size of the sliding window to obtain the first sliding window data, and loading it into the local memory unit; moving the first sliding window by one unit length and determining whether there is an intersection between the data in the current sliding window and the first sliding window data; if there is no intersection, sequentially acquiring the data contained in the current sliding window to obtain the second sliding window data, and loading it into the local memory unit; if there is an intersection, determining the difference between the data in the current sliding window and the first sliding window data, recombining the data of the difference obtained from global memory with the data of the intersection obtained from the local memory unit to obtain the second sliding window data, and loading it into the local memory unit; acquiring the input data of the convolutional neural network based on the acquisition method of the second sliding window data to obtain the input data corresponding to the computation operation.

[0013] In some embodiments of the present invention, the instruction stream is obtained as follows: a convolutional neural network is obtained, the convolutional neural network including multiple processing layers; the weight matrix of each processing layer is divided according to the size of the in-memory computing array to obtain multiple array groups corresponding to each processing layer; the weight replication factor of each processing layer is iteratively selected according to a preset genetic algorithm, and the multiple array groups corresponding to each processing layer are mapped to the corresponding cores; based on the array groups mapped to each core, it is determined whether each core has an array group with an incomplete computing task; if there is an array group with an incomplete computing task, a corresponding instruction stream is generated according to the incomplete computing task.

[0014] In some embodiments of the present invention, the in-memory computing unit is configured to: when there is a structural conflict between two matrix-vector multiplication operations, execute the first matrix-vector multiplication operation in sequence, and then execute the second matrix-vector multiplication operation; when there is a data dependency between two matrix-vector multiplication operations, execute the first matrix-vector multiplication operation first, and then execute the second matrix-vector multiplication operation based on the result of the first matrix-vector multiplication operation; when there is no structural conflict or data dependency between multiple matrix-vector multiplication operations, execute multiple matrix-vector multiplication operations in parallel.

[0015] In some embodiments of the present invention, the vector function unit is configured to perform post-processing based on the result of the matrix-vector multiplication calculation, the post-processing including activation processing, pooling processing and / or element-wise processing.

[0016] In some embodiments of the present invention, the element-by-element processing is to perform an accumulation calculation based on the result of the matrix-vector multiplication, and to perform an activation process on the accumulation result to obtain the inference result of the convolutional neural network.

[0017] According to a second aspect of the present invention, a method for accelerating convolutional neural network inference using the in-memory accelerator described in the first aspect is provided. The method includes: acquiring an instruction stream; loading data of the convolutional neural network in global memory according to the instruction stream to obtain corresponding input data; performing matrix-vector multiplication calculations based on the loaded input data to obtain the result of the matrix-vector multiplication calculation; and post-processing the result of the matrix-vector multiplication calculation to obtain the inference result of the convolutional neural network.

[0018] Compared with the prior art, the advantages of the present invention are as follows:

[0019] 1) It can automatically select the replication factor of each processing layer and the optimal task mapping core based on the amount of hardware resources using a preset genetic algorithm, making full use of hardware resources and ensuring a balanced distribution of tasks among the cores. This not only improves the parallelism of computation but also does not affect the allocation of computational tasks or structural conflicts, thereby increasing the throughput of the accelerator and speeding up the inference speed of the neural network.

[0020] 2) The data reading method using multiple sliding windows can avoid repeated data reading, thereby avoiding the waste of memory bandwidth. At the same time, it can reduce the number of instructions and the complexity of control signals, reduce the pressure on bandwidth and on-chip cache, and improve memory utilization.

[0021] 3) By utilizing the memory allocation strategy of array group reuse, the calculation results can be stored by reusing memory blocks, thereby classifying memory resources, optimizing memory usage, reducing chip area and system power consumption, and avoiding memory waste. Attached Figure Description

[0022] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0023] Figure 1 This is a schematic diagram of a common two-dimensional array composed of non-volatile storage devices;

[0024] Figure 2 This is a schematic diagram of the structure of an in-memory accelerator for accelerating convolutional neural network inference according to an embodiment of the present invention;

[0025] Figure 3 This is a flowchart illustrating a preset genetic algorithm according to an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of a data reading method according to an embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram illustrating a data reading method using multiple sliding windows according to an embodiment of the present invention;

[0028] Figure 6 This is a schematic diagram of a memory allocation strategy according to an embodiment of the present invention;

[0029] Figure 7 This is a schematic diagram illustrating the process of accelerating computation using an in-memory accelerator for accelerating convolutional neural network inference according to an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0031] As mentioned in the background section, existing in-memory computing accelerators require peripheral circuits to complete the computation. However, the on-chip memory capacity of deep neural network accelerators is limited. The presence of these peripheral circuits may increase the complexity and power consumption of the entire system. Furthermore, as the size and complexity of the models continue to increase, there may be insufficient memory on the accelerator, resulting in a reduction in the accelerator's throughput and thus reducing the inference speed of deep neural networks.

[0032] To address the aforementioned problems, this invention proposes an in-memory accelerator for accelerating inference in convolutional neural networks, such as... Figure 2As shown, the in-memory accelerator includes global memory, on-chip routing, and multiple cores connected to the global memory. Each core includes a control unit, a vector function unit, an in-memory computation matrix unit, and a local memory unit. The control unit is used to acquire instruction streams and control each unit to execute corresponding operations based on the instruction streams. The instruction streams include computation operations and memory access operations. The local memory unit is used to execute the memory access operations, accessing the input data of the convolutional neural network in the global memory sequentially according to the size of a sliding window to obtain the input data corresponding to the computation operation, and sending the corresponding input data to the in-memory computation matrix unit. When accessing the input data of the convolutional neural network by sliding window, the input data is retrieved from the global memory in pixels. The newly added data in the access window is recombined with the existing data in the local memory to obtain the data to be accessed. This data reading method avoids repeated data readings, thus preventing wasted memory bandwidth and reducing the amount of computational memory consumed during subsequent calculations, thereby improving memory utilization. The in-memory computation matrix unit includes multiple in-memory computation arrays, each used to execute the computation operation, performing matrix-vector multiplication based on the input data corresponding to the computation operation. The vector function unit executes the computation operation and performs post-processing based on the results of the matrix-vector multiplication. Through the cooperation of the in-memory computation matrix unit and the vector function unit, the inference result of the convolutional neural network is finally obtained.

[0033] To better understand the present invention, the technical solution of the present invention will be described in detail below with reference to specific embodiments.

[0034] This invention proposes an in-memory accelerator for accelerating convolutional neural network inference, comprising global memory, on-chip routing, and multiple cores connected to the global memory. The global memory stores the data of the convolutional neural network, and the on-chip routing sends data from one core to other cores or receives data from other cores. Multiple cores can be interconnected via on-chip routing or a bus for parallel and asynchronous computation. Each core includes a control unit, a vector function unit, an in-memory computation matrix unit, and a local memory unit. The in-memory computation matrix unit and the vector function unit can only access data in the local memory. The following details the global memory, on-chip routing, and the data interaction process between the multiple cores.

[0035] First, the hardware description file is read to initialize the in-memory accelerator. This hardware description file includes parameters such as: global memory capacity, bus bandwidth, on-chip routing bandwidth, number of cores, number of vector function units in the cores, number of in-memory computation matrix units in the cores, and local memory capacity. Next, a pre-trained convolutional neural network model is read and parsed to obtain the model description file. This model description file includes convolutional neural network node information and topology. The model description file uses nodes as basic elements, taking into account both topology and parameters. Table 1 shows an example of this model description file (taking the second node as an example). Here, bitwidth represents the precision of the node's weights, consumer represents the consumer name, provider represents the producer name, node_index represents the node index, node_operation represents the type of node operation, input_dim represents the input dimension of the node, output_dim represents the output dimension of the node, and param represents the parameter information of the node. More specifically, as shown in Table 1:

[0036] Table 1. Example of a model description file

[0037]

[0038]

[0039] According to one embodiment of the present invention, the pre-trained convolutional neural network model can be pre-trained or trained using an existing dataset (e.g., the ILSVRC2012 dataset). The training process based on the training data includes: obtaining a training set, which includes multiple image samples and labels indicating the categories to which the image samples belong; and training the convolutional neural network model once or multiple times using the training set to obtain the pre-trained convolutional neural network model. Indicatively, the convolutional neural network model can use existing models, such as AlexNet, VGG, ResNet, InceptionNet, etc., or a suitable model can be constructed based on actual research.

[0040] Secondly, in order to standardize the data interaction process between global memory, on-chip routing, and multiple cores, this invention provides a dedicated instruction set architecture. This allows the control unit of each core to control each unit to execute corresponding operations based on the instruction flow in the instruction set architecture. The instruction flow includes computation operations, memory access operations, and communication operations. The computation operations are executed by the in-memory computation matrix unit and vector function unit of each core. The memory access operations are executed by the local memory unit of each core. The communication operations are executed by the on-chip router (or bus). More specifically, as shown in Table 2:

[0041] Table 2 Dedicated instruction set architecture for in-memory accelerators

[0042]

[0043] The MVM instruction is a matrix-vector multiplication instruction that performs matrix-vector multiplication using in-memory computation matrix units. The instruction format is "mvm idx dst src len", which means: the in-memory computation array with index idx reads data at address src from local memory, with a data length of len, performs one matrix-vector multiplication operation, and then saves the result to local memory at address dst.

[0044] The VEC instruction is a post-processing instruction that utilizes the Vector Function Unit (VFU) for post-processing. The instruction format is "vec op dstsrc1 src2 len", which means: the VFU reads two vectors, src1 and src2, from local memory. The length of each vector is len. The VFU then performs post-processing on the two vectors and saves the result to local memory at address dst. Here, op represents different post-processing operations: op=0 indicates element-wise addition of the two vectors; op=1 indicates element-wise subtraction of the two zero vectors; op=2 indicates element-wise multiplication of the two vectors; op=3 indicates element-wise division of the two vectors; and op=4 indicates element-wise maximization of the two vectors. When src1 and src2 have the same address, the vector at address src1 can be activated. Correspondingly, op=5 indicates activation using the ReLU activation function, op=6 indicates activation using the tanh activation function, and op=7 indicates activation using the sigmoid activation function.

[0045] The LOAD instruction is a load instruction that reads data from global memory using a local memory location and writes the read data to local memory. The instruction format is "load dst src size", which means: load contiguous data at address src and length size from global memory into local memory at address dst and length size.

[0046] STORE is a save command that uses local memory locations to write data from local memory to global memory. The command format is "store dst src size", which means: write contiguous data at address src and length size from local memory to global memory at address dst and length size.

[0047] LMD is a move instruction that moves data from local memory. The instruction format is "lmd dst srcsize", which means: write contiguous data at address src with length size from local memory to local memory at address dst with length size.

[0048] SEND is a send command that uses the on-chip router (or bus) to send data to other cores. The command format is "send src size idx", which means: send a contiguous block of data with address src and length size from local memory to the core with sequence number idx.

[0049] RECV is a receive command that uses on-chip routing (or bus) to receive data from other cores. The command format is "recv dst size idx", which means: receive data sent from the core with sequence number idx and save it to local memory at address dst with length size.

[0050] Furthermore, research on in-memory accelerators revealed that existing technologies focus on designing specific hardware architectures without fully considering the details of deploying neural networks within in-memory computing accelerators, resulting in low throughput. Therefore, this invention provides a method for generating instruction streams, which are obtained as follows: A convolutional neural network (CNN) is acquired, comprising multiple processing layers (nodes); the weight matrix of each processing layer is divided according to the size of the in-memory computing array, resulting in multiple array groups corresponding to each processing layer; a preset genetic algorithm iteratively selects the weight replication factor for each processing layer and maps the multiple array groups corresponding to each processing layer to the corresponding core; based on the array groups mapped to each core, it is determined whether each core has array groups with incomplete computing tasks; if so, an instruction stream is generated based on the incomplete computing tasks.

[0051] Because existing technologies rely on manually mapping weight data onto the array, this ignores the impact of weight mapping on array parallelism. Furthermore, the size of the in-memory computation array within the in-memory computation matrix unit is limited, and a complete processing layer (convolutional layer or fully connected layer) in a convolutional neural network generally cannot be completely mapped to the same in-memory computation array. Therefore, it is necessary to divide the weights of these processing layers according to the size of the in-memory computation array. According to one embodiment of the present invention, the weights of each convolutional kernel in the processing layer of the convolutional neural network are reorganized into a column (the organization order needs to be consistent with the order in which the sliding window is reorganized into a column) to obtain the corresponding weight matrix. The weight matrices of each processing layer are then divided according to the size of the in-memory computation matrix to obtain multiple array groups corresponding to each processing layer. Illustratively, a fully connected layer in a convolutional neural network is considered as a convolutional layer with a kernel size of 1, a kernel channel number equal to the number of input elements of the fully connected layer, and a kernel count equal to the number of output elements of the fully connected layer. Reorganizing the weights of each kernel in the convolutional layer into a column yields a column of height k. w ×k h ×C in Width is C out The weight matrix, where k w Indicates the kernel length of the convolutional layer, k h Indicates the kernel width, C in Indicates the number of input channels, C out This indicates the number of output channels. The size of the internal computing array is: width W. xbar Height is H xbar By dividing the weight matrix of each processing layer according to the size of the in-memory computation matrix, we can obtain... The technical solution of this embodiment can achieve at least the following beneficial technical effects: by dividing the weight matrix by the size of the in-memory computing array, the resulting array group can adapt to the scale of in-memory computing hardware resources, reducing the pressure on bandwidth and on-chip cache.

[0052] Secondly, existing technical solutions often neglect to copy weights or adopt an intuitive approach by choosing the copying factor, such as copying the first few layers of the network several times to achieve inter-layer computational balance. However, this approach cannot effectively utilize resources. Furthermore, since the in-memory computation matrix unit in an in-memory accelerator is both a storage unit and a computation unit, and a key way to improve computational parallelism is to copy the weight data multiple times, core mapping can affect the allocation of computational tasks and structural conflicts. Therefore, to improve computational parallelism without affecting the allocation of computational tasks and structural conflicts, this invention presents a genetic algorithm to simultaneously solve these two problems. This genetic algorithm uses integer encoding, which balances flexibility and running efficiency, avoiding the increased running time caused by binary encoding. According to one embodiment of the invention, to ensure that the core mapping is not too dispersed, causing on-chip memory to become a limiting factor, this invention sets the number of processing layers that each core can accommodate. Therefore, the position of each gene in the chromosome determines the core index corresponding to these array groups. The execution steps of the genetic algorithm are as follows: Figure 3 As shown, step T1: Encode several array groups of each processing layer into integers; step T2: Randomly select a replication factor for the weights of each processing layer, and randomly select a core to be bound to the array group; step T3: Determine whether the selected replication factor and mapping core meet the preset requirements, the preset requirements being that the number of iterations of the genetic algorithm reaches the upper limit or the self-utilization rate meets the requirements. If yes, proceed to T4; otherwise, proceed to T5; step T4: If yes, replicate the weights according to the replication factor, and map multiple array groups of each processing layer to the corresponding cores according to the mapping cores; step T5: If no preset requirements are met, use the fitness function to evaluate the time required for the in-memory computing unit to perform one matrix-vector multiplication operation according to the initial replication factor and mapping cores; step T6: Select multiple processing layers with the shortest required time according to the evaluated time; step T7: Randomly select one processing layer from multiple processing layers, and use gene mutation in the genetic algorithm to select the weight replication factor and mapping core for the processing layer; return to step T3 to determine whether the selected weight replication factor and mapping core meet the preset requirements; if yes, execute step T4; otherwise, execute step T5. Schematic, its fitness function uses the total inference time as an indicator, where the total inference time is: T = max(T i ),T i =n i ×MVM time , where n i MVM represents the total number of matrix-vector multiplication operations in the i-th core. time The time required to perform one matrix-vector multiplication operation on an in-memory computing array.

[0053] According to one embodiment of the present invention, the gene mutation includes: 1) randomly selecting a treatment layer, increasing the replication fold, and mapping it to the corresponding core. For example, assuming the i-th treatment layer is randomly selected, its original replication fold is r. i After the mutation operation, the replication factor of the i-th processing layer is r. i +1.2) Randomly select a processing layer to reduce the replication factor, but increase its array resources. For example, suppose the i-th processing layer is randomly selected, and its original replication factor is r. i After the mutation operation, the replication factor of the i-th node is r. i -1. 3) Randomly select a processing layer and distribute its corresponding array group to other cores. For example, assuming the i-th processing layer is randomly selected and its array group is distributed across k cores, the core mapping process is to redistribute its array group to k+1 cores, thus increasing the dispersion of the array group during allocation. 4) Randomly select a processing layer and merge its corresponding array group into the same processing layer of other cores. For example, assuming the i-th processing layer is randomly selected and its array group is distributed across k cores, the core mapping process is to redistribute its corresponding array group to k-1 cores, thus reducing the dispersion of the array group during allocation. Since the genetic algorithm is an iterative optimization process, the weight replication factor may be selected in the k-th iteration, while the core mapping may only be selected in the k+1th operation. The technical solution of this embodiment can at least achieve the following beneficial technical effects: by selecting the weight replication factor and mapping core through the above genetic algorithm, both computational parallelism and resource utilization can be improved without affecting the allocation of computational tasks and structural conflicts.

[0054] According to one embodiment of the present invention, when binding the weight matrix and the in-memory computing array, arrays in the same array group are preferentially mapped to the same core. This is because arrays belonging to the same array group can be driven by the same instruction, thus reducing the number of instructions and the complexity of control signals. The technical solution of this embodiment can achieve at least the following beneficial effects: these arrays have identical inputs; if mapped to the same core, the input data can be broadcast to these arrays, thereby avoiding repeated reading of data corresponding to each array, and simultaneously reducing bandwidth and on-chip cache pressure.

[0055] Based on the weight replication factor and core mapping obtained by a preset genetic algorithm, it is determined whether each core has an array group with incomplete computation tasks according to the array groups of each core mapping. If there is an array group with incomplete computation tasks, a corresponding instruction stream is generated according to the incomplete computation tasks. According to one embodiment of the present invention, the total number of computation tasks in each array group is determined according to the weight replication factor, and the computation tasks of the array groups in the core are recorded by the data structure of each core. The array groups are mapped to the corresponding cores according to the core mapping. Each core will be mapped to multiple array groups, and each array group corresponds to multiple matrix-vector multiplication operation computation tasks. It is determined whether the core has an array group with incomplete computation tasks according to the data structure of each core. If there is an array group with incomplete computation tasks, a corresponding instruction stream is generated according to the incomplete computation tasks. The corresponding instruction stream consists of computation operations, memory access operations, and communication operations as shown in Table 2 above. Each time a matrix-vector multiplication operation computation task is executed, the computation task in its corresponding data structure is reduced by one.

[0056] The initial in-memory accelerator accelerates the convolutional neural network inference process according to the instruction stream generated above, including:

[0057] 1. By performing the memory access operation, the input data of the convolutional neural network in the global memory is accessed sequentially according to the size of the sliding window to obtain the input data corresponding to the calculation operation, and the corresponding input data is sent to the in-memory calculation matrix unit.

[0058] Existing methods for reading data typically load data from global memory into local memory in a left-to-right, top-to-bottom order based on the size of the sliding window. However, this method may repeatedly read some data, leading to excessively complex and large datasets during subsequent calculations, thus consuming more computational memory and wasting memory bandwidth. To improve memory utilization, this invention proposes a data reading method using multiple sliding windows. When accessing the input data of a convolutional neural network using a sliding window, the newly added data for the sliding window to be accessed is retrieved from global memory in pixels and recombined with the existing sliding window data in local memory to obtain the data for the sliding window to be accessed. Specifically, when accessing the input data of a convolutional neural network using a sliding window, the difference and intersection of the existing sliding window data and the sliding window data to be accessed are generated. The data from the difference set retrieved from global memory and the data from the intersection retrieved from local memory are recombined according to the order in which the sliding window data is read to obtain the data for the sliding window to be accessed. The order in which the sliding window data is read can be from left to right and top to bottom, or from top to bottom and left to right; this invention does not impose any restrictions on this.

[0059] According to an embodiment of the present invention, the step of sequentially accessing the input data of the convolutional neural network in global memory according to the size of the sliding window to obtain the input data corresponding to the computation operation includes: sequentially acquiring the data contained in the first sliding window according to the size of the sliding window to obtain the first sliding window data, and loading it into the local memory unit; moving the first sliding window by one unit length and determining whether there is an intersection (or duplicate data) between the data in the current sliding window and the data in the first sliding window; if there is no intersection, sequentially acquiring the data contained in the current sliding window to obtain the second sliding window data, and loading it into the local memory unit; if there is an intersection, determining the difference between the data in the current sliding window and the data in the first sliding window, recombining the data of the difference obtained from global memory with the data of the intersection obtained from the local memory unit to obtain the second sliding window data, and loading it into the local memory unit; acquiring the input data of the convolutional neural network based on the acquisition method of the second sliding window data to obtain the input data corresponding to the computation operation. It is worth noting that when the length of the sliding window is one unit, moving the sliding window one unit to the left or right does not result in duplicate data between the current sliding window and the first sliding window; the data of the current sliding window is still obtained in the same way as the data of the first sliding window. This embodiment achieves at least the following beneficial technical effects: by using multiple sliding windows for data reading, duplicate data reading can be avoided, thus preventing wasted memory bandwidth. Simultaneously, it reduces the number of instructions and the complexity of control signals, alleviating the pressure on bandwidth and on-chip cache, and improving memory utilization.

[0060] According to one example of the present invention, existing methods for reading data are as follows: Figure 4 As shown, the sliding window data corresponds to the data of input channels 1, 2, 3, 4, 5, 6, 7, 8, and 9 in the global memory. When reading the input data from the global memory to the local memory, the data of each input channel is read into the local memory in the order from left to right and from top to bottom, so that it becomes continuous input data (123456789). The sliding window data (123456789) accessed from the global memory is divided according to the calculation operation to obtain the input data corresponding to the calculation operation (e.g., the input data corresponding to array groups 0, 1, and 2).

[0061] According to one example of the present invention, the present invention provides a data reading method for multiple sliding windows, such as... Figure 5As shown, the first sliding window corresponds to the data of input channels 0, 1, 2, 5, 6, 7, 10, 11, and 12 in global memory. Based on the existing method of reading input data, the data corresponding to this sliding window is read into local memory to form continuous data, with the data arrangement order being 0, 1, 2, 5, 6, 7, 10, 11, and 12. The second sliding window corresponds to the data of input channels 1, 2, 3, 6, 7, 8, 11, 12, and 13 in global memory. The overlapping data between the second and first sliding windows can be obtained; this overlapping data is the data of input channels 1, 2, 6, 7, 11, and 12. In other words, the data of input channels 1, 2, 6, 7, 11, and 12 is the data of the first sliding window. The intersection of the moving window data and the second sliding window data, specifically the data from input channels 3, 8, and 13, represents the difference between the two sliding windows. When reading the data from the second sliding window, only the data from input channels (3, 8, 13) that have not been previously loaded is read from global memory. The data from previously loaded input channels (1, 2, 6, 7, 11, 12) is migrated to local memory using LMD instructions. The data from channels 3, 8, and 13 are then reassembled with the data from channels 1, 2, 6, 7, 11, and 12 according to the order in which the sliding windows are read, resulting in the data for the second sliding window (1, 2, 3, 6, 7, 8, 11, 12, 13). The data from global memory is then loaded into local memory according to the above data reading method. It should be understood that in this example, the sliding window size is 3×3. Those skilled in the art can adjust the size of the sliding window, for example, to 2×2, 4×4, or 3×1, to obtain other embodiments.

[0062] 2. Matrix-vector multiplication is performed using multiple in-memory computing arrays in the in-memory computing matrix unit based on the input data corresponding to the computing operation accessed by the local memory unit.

[0063] According to one embodiment of the present invention, for any two matrix-vector multiplication operations in the instruction stream of each core, the in-memory computation units in each core are configured as follows: If there is a structural conflict between the two matrix-vector multiplication operations, the preceding matrix-vector multiplication operation is executed sequentially, followed by the following matrix-vector multiplication operation. A structural conflict occurs when both matrix-vector multiplication operations are computations on the same array; therefore, the following matrix-vector multiplication operation must wait until the preceding matrix-vector multiplication operation has finished before starting. If there is a data dependency between the two matrix-vector multiplication operations, the preceding matrix-vector multiplication operation is executed first. Matrix-vector multiplication operations execute the next matrix-vector multiplication operation based on the result of the previous one. The data dependency means that the output of the previous matrix-vector multiplication operation is the input of the next; therefore, the next matrix-vector multiplication operation must wait until the previous one has finished before it begins. When multiple matrix-vector multiplication operations do not have structural conflicts or data dependencies, they can be executed in parallel. However, the order of computation must follow the order in the instruction stream, and the start time of two adjacent matrix-vector multiplication operations is constrained by the on-chip memory bandwidth.

[0064] According to one embodiment of the present invention, each sliding window is expanded to a size of k. w ×k h ×C in By using vectors, the convolution operation of a convolutional neural network can be transformed into a matrix-vector multiplication operation.

[0065] Third, the calculation operation is performed using the vector function unit, and the result of the matrix-vector multiplication calculation in the in-memory calculation matrix unit is post-processed.

[0066] According to one embodiment of the present invention, the vector functional unit is configured to perform post-processing based on the result of the matrix-vector multiplication calculation. The post-processing includes activation processing, pooling processing, and / or element-wise processing, wherein the post-processing is distributed among all cores for joint completion. More specifically, element-wise processing (e.g., addition, subtraction, multiplication, and / or division) is performed based on the result of the matrix-vector multiplication calculation, and activation processing is applied to the result after element-wise processing to obtain the inference result of the convolutional neural network. Illustratively, taking the addition of two vectors as an example, the result of the matrix-vector multiplication calculation is accumulated, and activation processing is applied to the accumulated result to obtain the inference result of the convolutional neural network.

[0067] According to one embodiment of the present invention, when multiple array groups of each processing layer are mapped to cores for storage, multiple array groups of processing layers may be mapped to one core, and array groups of one processing layer may be mapped to multiple cores. The computation results obtained from multiple array groups of the same processing layer need to be accumulated to obtain a complete convolution result. If array groups of the same layer are mapped to different cores, cross-core data accumulation is required. During cross-core data accumulation, on-chip routing is used to send the computation results of the same weight block in other cores to the core where the first array group of that weight block is located for accumulation. To reduce the synchronization overhead of inter-core transmission, each array group can perform several rounds of computation before inter-core transmission. Since on-chip local memory is insufficient to store the complete data of each layer, it is necessary to periodically move some input / output data between global memory and local memory.

[0068] Fourth, based on the cooperation between the in-memory computation matrix unit and the vector function unit, the inference result of the convolutional neural network is finally obtained.

[0069] According to one embodiment of the present invention, the inference results of the convolutional neural network are stored in the global memory through the local memory unit. In addition to performing memory access operations, the local memory unit can also perform operations such as supplementation, concatenation, and segmentation.

[0070] Since on-chip memory capacity is limited and accessing global memory is a very expensive operation, efficient utilization of on-chip local memory is crucial. Therefore, this invention provides a memory allocation strategy to allocate memory resources and achieve efficient memory utilization. According to one embodiment of the invention, as... Figure 6 As shown, it is necessary to save the calculation results and accumulated results of array groups 0, 1, and 2. This is achieved using both existing memory allocation strategies and the memory allocation strategy provided in this invention, as illustrated in the figure. Figure 6 (a) represents the existing memory allocation strategy, which allocates new memory for each operation to store the result of each operation. Schematably: First, a memory block is allocated for the result calculated by array group 0, and result A is stored in the first memory block; a memory block is allocated for the result calculated by array group 1, and result B is stored in the second memory block; the results of array group 0 and array group 1 are summed, and result C is stored in the third memory block; a memory block is allocated for the result calculated by array group 2, and result D is stored in the fourth memory block; finally, result D of array group 2 is summed with result C, and the final result E is stored in the fifth memory block. In the existing memory allocation strategy, much memory is only used once and then not accessed again, leading to resource waste. Based on this, this invention provides an accumulation and reuse memory allocation strategy, such as... Figure 6As shown in (b), its accumulation multiplexing saves memory by reusing the memory block of the accumulation operation to store the result of the new operation. Illustrated: First, a memory block is allocated to the result calculated by array group 0, and result A is stored in the first memory block; a memory block is allocated to the result calculated by array group 1, and result B is stored in the second memory block; the results of array group 0 and array group 1 are accumulated, and result C is stored in the third memory block; a memory block is allocated to the result calculated by array group 2, and result D is stored in the fourth memory block; finally, when result D of array group 2 is accumulated with result C, no new memory block is allocated, but the final accumulated result E is stored in the third memory block previously used to store the accumulated result C. The memory allocation strategy of accumulation multiplexing can reduce the memory blocks used to store intermediate results, but allocating memory for each array group still results in resource waste. Therefore, this invention provides a memory allocation strategy for array group multiplexing based on accumulation multiplexing, as follows: Figure 6 As shown in (c), array group reuse further reuses memory blocks for matrix-vector multiplication operations based on cumulative reuse to save memory. Schematic: First, a memory block is allocated to the result calculated by array group 0, and the result A is saved to the first memory block; a memory block is allocated to the result calculated by array group 1, and the result B is saved to the second memory block; then, the result C of the sum of the results of array group 0 and array group 1 is saved to the first memory block, because the contents saved in the first and second memory blocks (results A and B, respectively) will not be used subsequently; the result D of the calculation of array group 2 is saved to the second memory block; finally, the sum E of the result of array group 2 and result C in the first memory block is saved back to the first memory block. This embodiment's technical solution can achieve at least the following beneficial technical effects: by reusing memory blocks to save the results of new operations, memory usage is planned. The memory allocation strategy of array group reuse not only increases memory reuse and reduces on-chip memory usage, but also reduces chip area and system power consumption, avoiding memory waste.

[0071] Accordingly, the present invention also proposes a method for accelerating convolutional neural network inference using an in-memory accelerator according to the first aspect of the present invention. The method includes: acquiring an instruction stream; loading data of the convolutional neural network in global memory according to the instruction stream to obtain corresponding input data; performing matrix-vector multiplication calculations based on the loaded input data to obtain the result of the matrix-vector multiplication calculation; and post-processing the result of the matrix-vector multiplication calculation to obtain the inference result of the convolutional neural network.

[0072] To better understand the process of the technical solution of this invention, the relevant algorithms of this invention are given below:

[0073]

[0074] Based on the aforementioned algorithms, the following is combined with the appendix. Figure 7 The process of the technical solution of this invention will be described in detail. For example... Figure 7 As shown, the process of accelerating computation using an in-memory accelerator for accelerating convolutional neural network inference according to the present invention includes:

[0075] Step S1: Encode several array groups of each processing layer into integers;

[0076] Step S2: Randomly select a replication factor for the weight of each processing layer, and randomly select the cores to be bound to the array group;

[0077] Step S3: Determine whether the selected replication factor and mapping core meet the preset requirements. The preset requirements are that the number of iterations of the genetic algorithm reaches the upper limit or the self-use utilization rate reaches the requirement. If not, proceed to step T4; if they meet, proceed to step T7.

[0078] Step S4: Based on the selected replication factor and mapping core, use the fitness function to evaluate the time required for an in-memory computing unit to perform one matrix-vector multiplication operation;

[0079] Step S5: Select the processing layers with the shortest required time based on the evaluation time.

[0080] Step S6: Randomly select one processing layer from multiple processing layers, and use gene mutation in the genetic algorithm to select the weight replication factor and mapping core for the processing layer; return to step T3 to determine whether the selected weight replication factor and mapping core meet the preset requirements; if not, proceed to step T4; if they meet, proceed to step T7.

[0081] Step S7: Copy the weights according to the replication factor, and map the multiple array groups of each processing layer to the corresponding cores according to the mapping cores;

[0082] Step S8: Based on the multiple array groups mapped to each core, determine whether the core has any array groups with unfinished computing tasks. If yes, proceed to step S9; otherwise, proceed to step S12.

[0083] Step S9: Generate the corresponding instruction stream based on the unfinished computation tasks;

[0084] Step S10: Read the input data corresponding to the computing task from global memory according to the instruction stream;

[0085] Step S11: Perform a matrix-vector multiplication operation based on the input data corresponding to the computation task, and return to step S8. Determine whether there is an array group in the core that has not completed the computation task. If yes, proceed to step S9; otherwise, proceed to step S12.

[0086] Step S12: Post-process the calculation results of the matrix-vector multiplication operation;

[0087] Step S13: Save the post-processing results to global memory.

[0088] In summary, the in-memory accelerator for accelerating convolutional neural network inference proposed in this invention has the following advantages:

[0089] 1) It can automatically select the replication factor of each processing layer and the optimal task mapping core based on the amount of hardware resources using a preset genetic algorithm, making full use of hardware resources and ensuring a balanced distribution of tasks among the cores. This not only improves the parallelism of computation but also does not affect the allocation of computational tasks or structural conflicts, thereby increasing the throughput of the accelerator and speeding up the inference speed of the neural network.

[0090] 2) The data reading method using multiple sliding windows can avoid repeated data reading, thereby avoiding the waste of memory bandwidth. At the same time, it can reduce the number of instructions and the complexity of control signals, reduce the pressure on bandwidth and on-chip cache, and improve memory utilization.

[0091] 3) By utilizing the memory allocation strategy of array group reuse, the calculation results can be stored by reusing memory blocks, thereby classifying memory resources, optimizing memory usage, reducing chip area and system power consumption, and avoiding memory waste.

[0092] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0093] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0094] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0095] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An in-memory accelerator for accelerating inference in convolutional neural networks, characterized in that, The in-memory accelerator includes global memory, on-chip routing, and multiple cores connected to the global memory. The global memory is used to store the data of the convolutional neural network. The on-chip routing is used to send data from one core to other cores or receive data sent by other cores. Each core includes: a control unit, a vector function unit, an in-memory computation matrix unit, and a local memory unit, wherein: The control unit is used to acquire instruction streams and control each unit to perform corresponding operations based on the instruction streams. The instruction streams include: computation operations and memory access operations. The local memory unit is used to perform the memory access operation, access the input data of the convolutional neural network in the global memory in sequence according to the size of the sliding window, so as to obtain the input data corresponding to the calculation operation, and send the corresponding input data to the in-memory calculation matrix unit. When accessing the input data of the convolutional neural network by sliding window, the data to be accessed by the sliding window is obtained from the global memory in pixels and recombined with the data of the existing sliding window in the local memory to obtain the sliding window data to be accessed. The in-memory computation matrix unit includes multiple in-memory computation arrays, each in-memory computation array being used to perform the computation operation and to perform matrix-vector multiplication calculations based on the input data corresponding to the computation operation. The vector function unit is used to perform the calculation operation and perform post-processing based on the result of the matrix-vector multiplication calculation; The inference result of the convolutional neural network is finally obtained by cooperating with the in-memory computation matrix unit and the vector function unit.

2. The in-memory accelerator according to claim 1, characterized in that, When accessing the input data of a convolutional neural network by sliding window, the difference and intersection of the existing sliding window data and the data of the sliding window to be accessed are generated. The data of the difference set is accessed from global memory and the data of the intersection set is obtained from local memory unit and recombined to obtain the sliding window data to be accessed.

3. The in-memory accelerator according to claim 2, characterized in that, According to the order of reading data from the sliding window, the data of the difference set accessed from global memory and the data of the intersection obtained from local memory units are recombined to obtain the sliding window data to be accessed.

4. The in-memory accelerator according to claim 3, characterized in that, The local memory unit is configured to divide the sliding window data accessed from the global memory according to the calculation operation to obtain the input data corresponding to the calculation operation.

5. The in-memory accelerator according to claim 3, characterized in that, The step of accessing the input data of the convolutional neural network in global memory sequentially according to the size of the sliding window to obtain the input data corresponding to the computation operation includes: The input data of the convolutional neural network in global memory is sequentially obtained according to the size of the sliding window, and the data contained in the first sliding window is obtained to obtain the first sliding window data, which is then loaded into the local memory unit. Move the first sliding window by one unit length and determine whether there is an intersection between the data in the current sliding window and the data in the first sliding window; if there is no intersection, sequentially obtain the data contained in the current sliding window to obtain the data in the second sliding window, and load it into the local memory unit; If there is an intersection, determine the difference between the data in the current sliding window and the data in the first sliding window. Then, reassemble the data of the difference obtained from the global memory with the data of the intersection obtained from the local memory unit to obtain the second sliding window data, and load it into the local memory unit. The input data of the convolutional neural network is obtained based on the second sliding window data acquisition method, and the input data corresponding to the calculation operation is obtained.

6. The in-memory accelerator according to claim 1, characterized in that, The instruction stream is obtained in the following manner: Obtain a convolutional neural network, wherein the convolutional neural network includes multiple processing layers; The weight matrix of each processing layer is divided according to the size of the in-memory computing array to obtain multiple array groups corresponding to each processing layer; The weight replication factor of each processing layer is selected iteratively according to the preset genetic algorithm, and the multiple array groups corresponding to each processing layer are mapped to the corresponding core. Based on the array groups mapped to each core, determine whether each core has an array group with incomplete computing tasks. If there is an array group with incomplete computing tasks, generate the corresponding instruction stream based on the incomplete computing tasks.

7. The in-memory accelerator according to claim 1, characterized in that, The in-memory computing unit is configured as follows: In the event of a structural conflict between two matrix-vector multiplication operations, the first matrix-vector multiplication operation is executed in sequence, followed by the second matrix-vector multiplication operation. When there is a data dependency between two matrix-vector multiplication operations, the first matrix-vector multiplication operation is executed first, and the second matrix-vector multiplication operation is executed based on the result of the first matrix-vector multiplication operation. Multiple matrix-vector multiplication operations can be executed in parallel when there are no structural conflicts or data dependencies among them.

8. The in-memory accelerator according to claim 1, characterized in that, The vector function unit is configured to perform post-processing based on the result of the matrix-vector multiplication calculation, the post-processing including activation processing, pooling processing and / or element-wise processing.

9. The in-memory accelerator according to claim 8, characterized in that, The element-by-element processing involves accumulating the results of the matrix-vector multiplication calculation and then activating the accumulated results to obtain the inference result of the convolutional neural network.

10. A method for accelerating convolutional neural network inference using an in-memory accelerator as described in any one of claims 1-9, characterized in that, The method includes: Obtain the instruction stream; The corresponding input data is obtained by loading the data of the convolutional neural network in global memory according to the instruction stream; Perform matrix-vector multiplication based on the loaded input data to obtain the result of the matrix-vector multiplication calculation; The results of the matrix-vector multiplication calculation are post-processed to obtain the inference results of the convolutional neural network.