A resource allocation method and an accelerator

By extracting the parameters of the convolutional neural network and accelerator, determining the minimum demand for the convolutional layer and adjusting the resource allocation multiple, the problem of poor resource allocation of convolutional neural network accelerator in the existing technology is solved, and the accelerator performance and resource utilization efficiency are improved.

CN115712506BActive Publication Date: 2025-07-01INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211499120.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-07-01
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing memristor-based convolutional neural network accelerators have non-optimal solutions in computing resource allocation, and each convolutional neural network needs to design a resource allocation strategy separately, which is not very versatile.

Method used

By extracting the structural parameters of the convolutional neural network and the architectural parameters of the accelerator, the minimum demand for each convolutional layer is determined, and the resource allocation multiple is adjusted based on this to solve the resource allocation strategy with the optimal total processing delay.

Benefits of technology

Under predetermined constraints, it is realized that the resource allocation strategy with the optimal total processing delay of the convolutional neural network running on the accelerator is efficiently determined, which improves the performance and resource utilization efficiency of the accelerator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712506B_ABST
    Figure CN115712506B_ABST
Patent Text Reader

Abstract

The present invention provides a resource allocation method and an accelerator. The accelerator includes a plurality of memristor arrays. The method includes: obtaining the structural parameters of a convolutional neural network to be accelerated and the architecture parameters of the accelerator; determining the minimum requirement of each convolutional layer of the convolutional neural network model based on the structural parameters and the architecture parameters, where the minimum requirement is the minimum number of memristor arrays required to store all the weight parameters in the corresponding convolutional layer on the accelerator; according to a predetermined constraint condition, taking the allocation multiple of each convolutional layer based on its minimum requirement as an adjustment object, determining a resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on the accelerator, where the optimal resource allocation strategy indicates the final allocation multiple of each convolutional layer; and determining the number of memristor arrays allocated to each convolutional layer according to the final allocation multiple and the minimum requirement of each convolutional layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of in-memory computing, specifically to the technical field of hardware resource allocation based on in-memory computing, and more specifically to a resource allocation method and an accelerator. Background Art

[0002] With the development of artificial intelligence technology, convolutional neural networks have been widely applied in many fields, such as image recognition, object detection, face detection, etc. At the same time, due to convolutional neural networks containing a large amount of data and computations, and mostly being repetitive combinations of multiplication and summation operations, although the computing power of existing processors has been greatly improved compared to before, the parallel computing efficiency during the computing process of neural network models is still slightly insufficient. Therefore, some researchers have started to study hardware accelerators for neural network model computing, making neural network acceleration a current technological hotspot.

[0003] Among them, the convolutional neural network accelerator based on in-memory computing (or memristor) is a current popular technology. In-memory computing, as a new type of computing technology, its own structure is conducive to supporting the acceleration requirements of convolutional neural networks. As the name implies, in-memory computing refers to the integration of storage and computing, which can eliminate unnecessary data access between computing components and memories, thereby greatly reducing the energy and performance losses caused by data access. In the field of in-memory computing, the memristor is one of the common in-memory computing devices, and data can be stored in the form of resistance. For ease of understanding, Figure 1 Fig. shows a common crossbar array based on memristors. Among them, each bit line of the crossbar array is connected to a word line through a memristor unit. Assume that the conductance values of the memristors in the first column are G 1,1 , G 2,1 and G 3,1 , and voltages V1, V2, and V3 are applied to the three word lines respectively. Each unit will transfer V i ×G i,1 to the bit line of the first column; according to Kirchhoff's law, after the currents on the bit line of the first column are accumulated, the first column of memristor units in this crossbar array completes a set of dot product operations and obtains At the same time, parallel computing is performed between the columns of the crossbar array to obtain I2 corresponding to the second column and I3 corresponding to the third column. Therefore, the crossbar array composed of memristor units can complete a set of matrix-vector multiplications in O(1) time, which also provides good support for accelerating convolutional neural network operations.

[0004] In a memristor-based convolutional neural network accelerator, weights are stored in memristors in the form of resistance values. Through the structure of the crossbar array, high-parallel matrix-vector multiplication operations are completed with the input. Memristors serve as both computing resources and storage resources at the same time. This not only reduces the data access of weights but also significantly improves the computing density. Such characteristics support the multi-layer parallel computing of convolutional neural networks. Different from the process in traditional accelerators where convolutional layers are sequentially computed, the computing resources of memristor accelerators can support different convolutional layers to be computed simultaneously when the computing conditions are met. Therefore, the allocation of computing resources between different convolutional layers becomes a key factor restricting the performance of the accelerator. In current convolutional neural network accelerators composed of memristors, memristors are used as computing resources to be allocated and are allocated in a heuristic manner in each convolutional layer, obtaining a feasible but possibly non-optimal solution, which cannot fully guarantee the improvement of the accelerator performance. Moreover, for different convolutional neural networks, the parameter forms required for resource allocation extraction are not unified and there is no general criterion, resulting in the need to design a separate resource allocation strategy for each convolutional neural network, with poor generality. Therefore, it is necessary to improve the existing technology. Summary of the Invention

[0005] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a resource allocation method and an accelerator.

[0006] The object of the present invention is achieved by the following technical solutions:

[0007] According to a first aspect of the present invention, there is provided a method for allocating resources in an accelerator for a convolutional neural network based on memristors, the accelerator including a plurality of memristor arrays, the method comprising: obtaining the structural parameters of the convolutional neural network to be accelerated and the architecture parameters of the accelerator; determining the minimum demand of each convolutional layer of the convolutional neural network model based on the structural parameters and the architecture parameters, the minimum demand being the minimum number of memristor arrays required to store all the weight parameters in the corresponding convolutional layer on the accelerator; according to a predetermined constraint condition, taking the allocation multiple of each convolutional layer based on its minimum demand as an adjustment object, determining a resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on the accelerator, wherein the optimal resource allocation strategy indicates the final allocation multiple of each convolutional layer; and determining the number of memristor arrays allocated to each convolutional layer according to the final allocation multiple and the minimum demand of each convolutional layer.

[0008] In some embodiments of the present invention, the structural parameters include the following structural information related to each convolutional layer: the number of input channels of the convolutional layer, the number of output channels, the width and height of the output feature map, the kernel size of the convolutional layer, the kernel size of the pooling layer, the stride of the convolutional layer, the stride of the pooling layer, the padding size of the convolutional layer, and the padding size of the pooling layer. Among them, the structural information of each convolutional layer records the kernel size and the stride of the pooling layer of the subsequent pooling layer of this convolutional layer.

[0009] In some embodiments of the present invention, the architecture parameters include the total number of memristor arrays and the size of each memristor array, and this size indicates the number of bit lines and the number of word lines of the corresponding memristor array.

[0010] In some embodiments of the present invention, the minimum demand of each convolutional layer is determined in the following manner:

[0011]

[0012] Among them, K c ×K c represents the kernel size of the convolutional layer, C i represents the number of input channels of the convolutional layer, C o represents the number of output channels of the convolutional layer, M represents the number of word lines of the memristor array, and N represents the number of bit lines of the memristor array.

[0013] In some embodiments of the present invention, the method includes: calculating the corresponding total processing delay according to whether there is a subsequent pooling layer for the convolutional layer and the kernel size and / or the stride of the pooling layer.

[0014] In some embodiments of the present invention, the total processing delay of the convolutional neural network running on this accelerator is determined in the following manner: determining the corresponding resource allocation strategy based on the total number and size of the memristor arrays; obtaining the maximum cycle size among the cycle size for one calculation and the memory access cycle size of each convolutional layer under the corresponding resource allocation strategy, and determining the total processing delay of this resource allocation strategy according to the maximum cycle size and the number of cycles that the last convolutional layer runs.

[0015] In some embodiments of the present invention, the number of cycles that each convolutional layer runs is determined in the following manner:

[0016] Op i =max{NormalOp i +PreOp i ,Op i-1 +Tail i}

[0017] Among them, Op irepresents the number of cycles for the operation of convolutional layer i, max{·} represents taking the maximum value, and NormalOp i represents the number of cycles for the calculation of convolutional layer i, and PreOp i represents the number of cycles elapsed while convolutional layer i waits to be activated, and Op i-1 represents the number of cycles for the operation of the previous convolutional layer of convolutional layer i, and Tail i represents the number of cycles that convolutional layer i still needs to run after the previous convolutional layer stops calculating.

[0018] In some embodiments of the present invention, the predetermined constraint conditions include: the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator.

[0019] In some embodiments of the present invention, the predetermined constraint conditions include: the sum of the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator, and the allocation multiple of any one convolutional layer is less than or equal to the width multiplied by the height of the output feature map of that convolutional layer.

[0020] According to a second aspect of the present invention, there is provided an accelerator, which includes a plurality of chips, each chip includes one or more memristor arrays, and the accelerator is configured to: obtain the number of memristor arrays allocated to each convolutional layer of the convolutional neural network to be accelerated according to the method of the first aspect; allocate the required number of memristor arrays to the convolutional neural network according to the number of memristor arrays allocated to each convolutional layer among the plurality of chips to perform corresponding convolutional operations.

[0021] According to a third aspect of the present invention, there is provided an electronic device, including: one or more processors; and a memory, where the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the first aspect of the claims by executing the executable instructions.

[0022] Compared with the prior art, the advantages of the present invention are as follows:

[0023] The present invention extracts the structural parameters of the convolutional neural network to be accelerated and the architecture parameters of the accelerator through a predetermined parameter extraction format, and based on this, determines the minimum requirements of each convolutional layer of the convolutional neural network model. Taking the allocation multiple of each convolutional layer based on its minimum requirements as the adjustment object, it solves the resource allocation strategy with the optimal total processing delay under the predetermined constraint conditions. Thus, by extracting the general parameters required for resource allocation from the convolutional neural network and the accelerator respectively based on the predetermined parameter extraction format, and based on this determining the minimum requirements, and using the allocation multiple corresponding to the minimum requirements as the adjustment object, it efficiently determines the resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on this accelerator. Brief Description of the Drawings

[0024] The embodiments of the present invention will be further described below with reference to the accompanying drawings, where:

[0025] Figure 1 is a partial schematic diagram of an existing memristor array;

[0026] Figure 2 is a module schematic diagram of an existing accelerator;

[0027] Figure 3 is a flowchart of a resource allocation method according to an embodiment of the present invention;

[0028] Figure 4 is a schematic diagram of the number of cycles and related parameters of the operation of a convolutional layer according to an embodiment of the present invention;

[0029] Figure 5 is a flowchart of the solution process of a solver based on dynamic programming according to an embodiment of the present invention;

[0030] Figure 6 is a flowchart of a schematic calculation of the total processing delay according to an embodiment of the present invention. Detailed Embodiments

[0031] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] As mentioned in the background art section, the existing technology allocates computing resources in a convolutional neural network accelerator composed of memristors in a heuristic manner, which can obtain a feasible but possibly non-optimal solution, and cannot fully guarantee the improvement of the accelerator performance. In addition, the existing resource allocation strategies are suitable for limited and different application scenarios, and the generality needs to be improved. In this regard, the present invention extracts the structural parameters of the convolutional neural network to be accelerated and the architecture parameters of the accelerator, and based on this, determines the minimum requirements of each convolutional layer of the convolutional neural network model. Taking the allocation multiple of each convolutional layer based on its minimum requirements as the adjustment object, a resource allocation strategy with the optimal total processing delay under predetermined constraints is solved. Thus, by extracting the general parameters required for resource allocation from the convolutional neural network and the accelerator respectively based on a predetermined parameter extraction format, and based on this, determining the minimum requirements, and using the allocation multiple corresponding to the minimum requirements as the adjustment object, a resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on this accelerator is efficiently determined.

[0033] First, for the convenience of understanding the implementation background of the present invention, refer to Figure 2Introduce the structure of a conventional memristor-based convolutional neural network accelerator (hereinafter referred to as the convolutional neural network accelerator or accelerator). The accelerator includes multiple chips, and each chip includes a controller and one or more memristor arrays. The controller is used to control the overall control of the chip, including operations, storage, data transmission, etc. Data is transmitted between chips through a bus or on-chip network. The memristor array here is the component used to accelerate the convolutional operation of the convolutional layer and is also the resource that needs to be allocated and concerned in the present invention. Generally speaking, the accelerator is composed of multiple chips, and the memristor resources contained in each chip are the same. For any convolutional layer, the number of memristor resources it needs is different and may require one or more chips. Different convolutional layers are usually distributed on different chips, and data exchange between different convolutional layers needs to be completed through a bus. To complete other types of operations in the convolutional neural network besides the convolutional operation, each chip also includes components: digital-to-analog converter, analog-to-digital converter, sample-and-hold circuit, shift-and-accumulate circuit, pooling circuit, non-linear circuit, input buffer, and output buffer. The functions of each component are as follows:

[0034] Input buffer, used to buffer operation-related data, including input data.

[0035] Digital-to-analog converter, used to perform digital-to-analog conversion processing on the input data related to the convolutional operation (the sample data or the feature map obtained by processing the sample data in the present invention) to obtain an analog voltage signal; these data are converted into an analog voltage signal after passing through the digital-to-analog converter and are loaded onto each row (word line) of the memristor array.

[0036] Memristor array, used to store the weights related to the convolutional operation (i.e., the weight parameters in the convolutional layer) and complete the convolutional operation (or matrix-vector multiplication) of the input data and the corresponding weight data based on the corresponding analog voltage signal to obtain an intermediate operation result. In other words, the weights of the convolutional neural network need to be pre-mapped into the resistance values of the memristor array for storage (see again Figure 1 ), the input data completes the matrix-vector multiplication with the weights through the memristor array, and the resulting current indicating the intermediate operation result will be generated on the bit line.

[0037] Sample-and-hold circuit, used to convert the resulting current on the bit line into an analog voltage indicating the intermediate operation result. It is converted into an analog voltage through the sample-and-hold circuit.

[0038] Analog-to-digital converter, used to convert the analog voltage indicating the intermediate operation result into a digital signal indicating the intermediate operation result.

[0039] The shift-accumulation circuit performs shift-accumulation processing based on the intermediate operation result to obtain the operation result of the convolution operation. As is known to those skilled in the art, the number of bits calculated each time depends on the accuracy of the digital-to-analog converter. In order to obtain the operation result of the convolution operation, the result after the input is operated bit by bit (the intermediate operation result) needs to be added through the shift-accumulation circuit to obtain the operation result of the convolution operation.

[0040] The non-linear circuit (or activation function circuit) is used to perform activation processing on the operation result of the convolution operation according to the activation function set in the convolutional neural network to obtain the feature map after activation processing (the output of a convolutional layer). Multiple activation functions can be preset in the non-linear circuit, such as ReLU, Sigmoid, Leaky ReLU, etc.

[0041] The pooling circuit is used to perform pooling processing on the corresponding feature map (usually the feature map output by the previous convolutional layer) according to the pooling layer set in the convolutional neural network to obtain the pooled feature map. The result after pooling processing is stored back in the output buffer.

[0042] The output buffer is used to store the operation result. Among them, if a chip only performs convolution operations, the feature map after convolution operations is stored back in the output buffer; if the chip also performs pooling processing on the feature map after convolution operations, the pooled feature map is stored back in the output buffer. If the result after the output buffer is to be used as the input of other convolutional layers, it will be transmitted through the bus to the input buffer of the corresponding chip for subsequent calculations.

[0043] It should be understood that the above structure is only illustrative. Without conflict in principles, the present invention can be applied to other improved accelerators with different structural forms (such as adding some additional processing components to perform corresponding tasks) or any existing accelerators, and the present invention makes no limitation thereto.

[0044] Secondly, the process of the resource allocation method of the present invention is introduced. According to an embodiment of the present invention, with reference to Figure 3 , a method for allocating resources in an accelerator for a convolutional neural network based on memristors (or a resource allocation method) is provided, including steps S1, S2, S3, and S4. For a better understanding of the present invention, each step will be described in detail below in combination with specific embodiments.

[0045] Step S1: Obtain the structural parameters of the convolutional neural network to be accelerated and the architectural parameters of the accelerator.

[0046] According to an embodiment of the present invention, according to a predetermined parameter extraction format, the structural parameters of the convolutional neural network to be accelerated and the architecture parameters of the accelerator are obtained. Preferably, the predetermined parameter extraction format defines the structural parameters uniformly extracted for different convolutional neural networks and the architecture parameters of the accelerator to be obtained. Thus, the present invention can use the structural parameters to uniformly describe the structures of different convolutional neural networks, and in combination with the architecture parameters of the accelerator, to more quickly determine the resource allocation scheme, improve the resource allocation efficiency, save time costs, and reduce the implementation difficulty and the labor intensity of the implementer.

[0047] According to an embodiment of the present invention, in the predetermined parameter extraction format, the information contained in the structural parameters is specified. Preferably, the structural parameters include the following structural information related to each convolutional layer: the number of input channels, the number of output channels, the width and height of the output feature map, the kernel size of the convolutional layer, the kernel size of the pooling layer, the stride of the convolutional layer, the stride of the pooling layer, the padding size of the convolutional layer, and the padding size of the pooling layer of the convolutional layer. Among them, the structural information of each convolutional layer records the kernel size and the stride of the pooling layer of the subsequent pooling layer of this convolutional layer. Since in the existing convolutional neural networks, some pooling layers are set after some convolutional layers to reduce the size of the feature map and reduce the computational amount, while some convolutional layers are not followed by pooling layers. If the format forms are set in different cases, it will be inconvenient for unified description and calculation. Therefore, in order to unify the form of the structural parameters, when there is no subsequent pooling layer for a convolutional layer, the kernel size and / or the stride of the pooling layer in the structural information of the convolutional layer are set to 1. The technical solution of this embodiment can at least achieve the following beneficial technical effects: The present invention can describe the structure of the convolutional neural network in the form of predetermined structural parameters, which helps to efficiently and uniformly represent the network structures of each convolutional neural network, and is convenient for subsequent efficient calculation of the total processing delay corresponding to the resource allocation strategy based on the network structure and determination of the optimal resource allocation strategy.

[0048] According to an embodiment of the present invention, to more clearly illustrate the structural parameters of the present invention, the VGG-A convolutional neural network is taken as an example for illustration here. The numerical values of each parameter of each convolutional layer of the VGG-A convolutional neural network are as follows:

[0049]

[0050]

[0051] Among them, in the numerical values of each convolutional layer, each row has 8 numerical values, corresponding to 8 convolutional layers respectively. Taking the first convolutional layer of the VGG-A convolutional neural network as an example, its structural information can be represented by the first numerical value in the second column of rows 2-10 in the above table, that is, the number of input channels C of the convolutional layer iis 3, the number of output channels C of the convolutional layer O is 64, the length W of the output feature map O is 224, the height H of the output feature map of the convolutional layer O is 3, the kernel size K of the convolutional layer C is 3, the kernel size K of the pooling layer P is 2, the stride S of the convolutional layer C is 1, the stride S of the pooling layer P is 2, the padding size P of the convolutional layer C is 1, the padding size P of the pooling layer P is 0. Additionally, after the third convolutional layer comes the fourth convolutional layer, and there is no adjacent pooling layer. Therefore, for unified representation, it can be seen that in the information corresponding to the third convolutional layer, the kernel size K of the pooling layer P , the stride S of the pooling layer P are set to 1; in actual situations, the kernel size and stride of the pooling layer cannot be 1. When there is a pooling layer adjacent to a certain convolutional layer, the actual parameters of that pooling layer are used for filling, such as 2; thus, the format of the above structural parameters can distinguish the situations with and without pooling. The technical solution of this embodiment can at least achieve the following beneficial technical effects: For the parameters of the convolutional neural network, the present invention describes the convolutional layer and pooling layer in the convolutional neural network with a set of general parameters, thereby facilitating the description of convolutional neural network tasks, specifically including {C i , C O , W O , H O , K C , K P , S C , S P , P C , P P}, and its parameters respectively refer to the number of input channels C of the convolutional layer i , the number of output channels C O , the length W of the output feature map O and the height H O , the kernel size K of the convolutional layer C , the kernel size K of the pooling layer P , the stride S of the convolutional layer C , the stride S of the pooling layer P , the padding size P of the convolutional layer C , the padding size P of the pooling layer P ; thus, the structure of the convolutional neural network is described in the form of unified structural parameters.

[0052] According to an embodiment of the present invention, in the parameter extraction format, the information contained in the specified architecture parameters is described. The architecture parameters include the total number Total of memristor arrays and the size M×N of the memristor array, where this size indicates the number M of word lines and the number N of bit lines of a memristor array. For example, assume that an accelerator contains 20 chips, and each chip is provided with 4 memristor arrays, and the size of each memristor array is 128×128; correspondingly, the total number of memristor arrays is 20×4 = 80, the number M of word lines of a memristor array is 128, and the number N of bit lines is 128. It should be understood that this is only illustrative here, and accelerators in the art may be set in different architecture forms, resulting in different corresponding architecture parameters. For example: an accelerator contains 100 chips; each chip is provided with 1, 8, or 16 memristor arrays; the size of each memristor array is 64×64, etc. The present invention does not impose any restrictions on this.

[0053] Step S2: Determine the minimum requirement of each convolutional layer of the convolutional neural network model based on the structural parameters and the architecture parameters. The minimum requirement is the minimum number of memristor arrays required to store all the weight parameters in the corresponding convolutional layer on this accelerator;

[0054] According to an embodiment of the present invention, determine the minimum requirement set of each convolutional layer in the following manner:

[0055]

[0056] where K c ×K c represents the kernel size of the convolutional layer, C i represents the number of input channels of the convolutional layer, C o represents the number of output channels of the convolutional layer, M represents the number of word lines of the memristor array, and N represents the number of bit lines of the memristor array. The minimum requirement can also be called a derived parameter, which refers to the information derived from the above structural parameters and architecture parameters; it represents the number of arrays required to map all the weights of a convolutional layer into the memristor arrays. For a convolutional neural network accelerator, set memristor arrays can generate 1×1×C O outputs in one calculation, and such set memristor arrays can be called a group.

[0057] Step S3: According to the predetermined constraint conditions, take the allocation multiple of each convolutional layer based on its minimum requirement as the adjustment object, and determine the resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on this accelerator, where the optimal resource allocation strategy indicates the final allocation multiple of each convolutional layer.

[0058] In the design of memristor-based accelerators, existing designs use the weight replication technique, i.e., multiple groups of memristor arrays, to improve the intra-layer parallelism and computing throughput. Therefore, after determining the minimum requirement corresponding to each convolutional layer in the aforementioned manner, the allocation multiple of each convolutional layer based on its minimum requirement (let the allocation multiple of the i-th convolutional layer be R i ) can be used as the adjustment object; thus, under the conditions of the aforementioned set structural parameters and architectural parameters, the problem of memristor resource allocation can be reduced to solving the R of each convolutional layer under certain constraints i to optimize the performance of the convolutional neural network to be accelerated by the accelerator.

[0059] Without considering resource waste, according to an embodiment of the present invention, the predetermined constraints include: the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator. The corresponding formula is expressed as follows:

[0060]

[0061] where L represents the total number of convolutional layers, R i represents the allocation multiple of the i-th convolutional layer, set i represents the minimum requirement of the i-th convolutional layer, and Total represents the total number of memristor arrays;

[0062] Since the workload of each convolutional layer is fixed, if the resources allocated to a certain convolutional layer exceed its total workload requirement, it will lead to resource waste. Therefore, it can be further improved. According to an embodiment of the present invention, the predetermined constraints include: the sum of the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator, and the allocation multiple of any one convolutional layer is less than or equal to the width multiplied by the height of the output feature map of the convolutional layer. The corresponding formula is expressed as follows:

[0063]

[0064] where represents the width of the output feature map of the i-th convolutional layer, represents the height of the output feature map of the i-th convolutional layer. The technical solution of this embodiment can at least achieve the following beneficial technical effects: this embodiment can reduce resource waste, allocate resources to more needy convolutional layers, and improve computing efficiency; and it can also reduce the size of the search space corresponding to the resource allocation strategy (i.e., the resource allocation strategies that need to be judged), and improve the efficiency of resource allocation.

[0065] According to one embodiment of the present invention, to solve the optimal resource allocation strategy, the convolutional layers from 0 to L-1 can be traversed, and based on the constraints, the possible combinations of allocation multiples of each convolutional layer can be determined. Each combination corresponds to a resource allocation strategy, and the total processing delay of each resource allocation strategy is calculated, and the resource allocation strategy with the best total processing delay (hereinafter referred to as the optimal resource allocation strategy) is selected.

[0066] Furthermore, the above embodiment of solving the optimal resource allocation strategy solves each resource allocation strategy separately during the solution, and the solution efficiency is low, which can be further improved. According to an embodiment of the present invention, the method includes: solving the optimal resource allocation strategy based on the properties of the optimal substructure of dynamic programming according to predetermined constraints. Therefore, a solver based on dynamic programming can be used to search the search space corresponding to the resource allocation strategy and obtain the optimal solution. The workflow diagram of the solver based on dynamic programming is as follows: Figure 4 As shown in Figure 1. After inputting the structural parameters and architecture parameters obtained in the first step as task parameters, the solver will traverse each convolution layer. opt [i,G] represents the best sub-strategy obtained by the solver under the condition of G memristors for the first i convolutional layers. The result of the last layer is represented by R opt [L-1,Total] means. In order to solve R opt [L-1,Total], the solver needs to first traverse the convolutional layers from 0 to L-1, and the optimal solution corresponding to each layer when it has a different number of arrays. Following the properties of the optimal substructure of dynamic programming, this optimal solution can be obtained by the optimal solution R of the subproblem opt [i-1,GR i ]Traverse R i Find. R i is the number of memristor groups divided into the i-th layer (for the allocation multiple). The optimal solution is the resource allocation strategy that optimizes the accelerator performance, that is, minimizes the total processing delay of the convolutional neural network operation. Thus, a quantitative relationship between the performance of the memristor-based convolutional neural network accelerator and its memristor allocation strategy is established. Based on the above model, according to a solver based on a dynamic programming algorithm, the problem of how to allocate memristor resources to optimize the accelerator performance under limited resources is solved. This solver is different from a simple heuristic strategy. It can obtain an approximately optimal resource allocation strategy in a shorter time, ensuring the acceleration effect of the solution obtained by this framework on the convolutional neural network accelerator. In this solver, it satisfies the following optimal substructure properties:

[0067]

[0068] By traversing the number of memristors j allocated by convolutional layer i, R opt[i, G] is derived from R that minimizes the calculation time of the first i convolutional layers. Through this step, the solution of the resource allocation strategy that optimizes the accelerator performance can be solved in a shorter time. opt [i - 1, G - j].

[0069] Furthermore, the calculation method of the total processing delay also affects the solving efficiency. Its calculation process involves the number of cycles for each convolutional layer to run. It is necessary to analyze and summarize the architecture and pipeline design of existing accelerators, classify and analyze various operations and no-operations in the pipeline. According to an embodiment of the present invention, the inventor summarizes the following calculation formula to calculate the number of cycles for each convolutional layer to run:

[0070] Op i = NormalOp i + PreOp i + Nop i

[0071] Among them, Op represents the number of cycles for each convolutional layer to run. represents the number of cycles for this convolutional layer to calculate. PreOp i = PreOp i-1 + interval(i, i - 1) represents the number of cycles that the i-th convolutional layer waits to be activated. It is equal to the preparation cycle number of the (i - 1)-th convolutional layer plus the difference between the activation cycles of the i-th convolutional layer and the (i - 1)-th convolutional layer. The accurate preparation cycle number of the i-th convolutional layer can be obtained by traversing the first (i - 1) convolutional layers. Nop represents the no-operation caused by the fact that the input required by the convolutional layer is still not valid and cannot be accurately counted. If this formula is used, an estimated value can be set for each convolutional layer in advance.

[0072] In addition to using the above formula, the estimated value of Nop can also be considered in other ways. According to an embodiment of the present invention, the number of cycles for each convolutional layer to run is determined in the following way:

[0073] Op i = max{NormalOp i + PreOp i , Op i-1 + Tail i}

[0074] Among them, Op i represents the number of cycles for convolutional layer i to run. max{·} represents taking the maximum value. NormalOp i represents the number of cycles for convolutional layer i to calculate. PreOp i represents the number of cycles that convolutional layer i waits to be activated. PreOp i = PreOpi-1 + interval(i, i - 1), Op i-1 represents the number of cycles that the previous convolutional layer of convolutional layer i runs, Tail i represents the number of cycles that convolutional layer i still needs to run after the previous convolutional layer stops computing. Tail i Both interval(i, i - 1) and can be obtained by the above-extracted structural parameters and variable R i calculated. The technical solution of this embodiment can at least achieve the following beneficial technical effects: This embodiment can measure NormalOp i + PreOp i and Op i-1 + Tail i in the way of taking the maximum value to replace the estimation of Nop. Thus, the number of cycles that each convolutional layer runs can be estimated more accurately, improving the accuracy of the final predicted strategy and ensuring the performance of the accelerator. For ease of understanding, see Figure 4 schematically shows the relationship between each parameter, where it is assumed that there are 3 convolutional layers, and the 0th convolutional layer can be directly operated, and PreOp 0 and Tail 0 are both 0, and only NormalOp 0m is actually involved. The 1st convolutional layer involves PreOp 1 , NormalOp 1 , Op 0 , Tail 1 . The 2nd convolutional layer involves PreOp 2 , NormalOp 2 , Op 1 , Tail 2 , where Op 0 and Op 1 respectively represent the number of cycles that the 0th layer and the 1st layer run.

[0075] Based on the various structural parameters and architecture parameters set above, see Figure 5 , taking the solution of the optimal resource allocation strategy by the property of the optimal substructure of dynamic programming as an example to illustrate the schematic process of the solution, which includes the following steps:

[0076] K1. Start the solution according to the input task parameters (i.e., the structural parameters and architecture parameters of the present invention), and initialize the layer index i = 0;

[0077] K2. Determine whether the condition i < L holds? If so, go to step K3; if not, go to step K15;

[0078] K3. Initialize

[0079] K4. Determine whether the condition i > 0 holds. If yes, go to step K9; if no, go to step K5;

[0080] K5. Let Go to step K6;

[0081] K6. Update G = G + set i ;

[0082] K7. Determine whether the condition G < Total holds. If yes, go to step K4; if no, go to step K8

[0083] K8. i = i + 1, go to step K2;

[0084] K9. R i = 1;

[0085] K10. Construct a new solution R[i, G] = [R opt [i - 1, G - R i set i , R i , go to step K11;

[0086] K11. Calculate the current total processing delay Time i Ri = f(R[i, G]), go to step K12;

[0087] K12. R i = R i + 1, go to step K13;

[0088] K13. Determine whether R i satisfies the constraint condition. If yes, go to step K10; if no, go to step K14;

[0089] K14. Obtain the optimal solution based on the optimal substructure property Go to step K6;

[0090] K15. Output R opt [L - 1, Total]; where it contains the required R for each of the total L layers, that is, the allocation multiples for each of the 0 - (L - 1) layers.

[0091] According to an embodiment of the present invention, the total processing delay is also related to the cycle size of each layer. For the cycle size of each layer, it can be achieved without considering the memory access cycle size. In this regard, according to an embodiment of the present invention, the total processing delay is determined in the following manner: Obtain the maximum cycle size among the cycle sizes for each calculation of each convolutional layer, and determine the total processing delay based on the maximum cycle size and the number of cycles for the operation of the last convolutional layer.

[0092] It should be noted that not considering the size of the memory access cycle may have an impact when the input data volume of the convolutional layer is large, resulting in a significant difference between the actual performance and the expected performance of the accelerator. In this regard, the present invention also provides a total processing delay calculation scheme considering the size of the memory access cycle. According to an embodiment of the present invention, the total processing delay of the convolutional neural network running on the accelerator is determined in the following manner: determining the corresponding resource allocation strategy based on the total number and size of the memristor arrays; obtaining the maximum cycle size among the cycle size for one calculation and the memory access cycle size of each convolutional layer under the corresponding resource allocation strategy, and determining the total processing delay of the resource allocation strategy based on the maximum cycle size and the number of cycles for the last convolutional layer to run. For example, to solve the time required for a convolutional neural network containing L convolutional layers to complete one inference operation, each layer needs to be traversed. For each layer, the number of NormalOp, PreOp, and Tail of each layer needs to be calculated separately. Among them, for preparing the operands (corresponding to the input data), the convolutional layers before this layer need to be traversed to accurately solve the number of prepared operands. The number of operands required for each layer is jointly determined by the number of operands required by the previous layer and the situation of this layer. At the same time, assume T A 、T C and T cycle respectively represent the size of the memory access cycle, the size of the calculation cycle, and the size of the clock cycle required for one calculation by the memristor. Among them, the impact of bandwidth on performance is reflected in the impact of T A on T cycle . For this framework, T C is determined by the accelerator architecture, and for a task with a given accelerator architecture, this value is a constant. T A is closely related to the number of groups of memristor arrays allocated to each layer. Specifically, in this embodiment, different convolutional layers may occupy one to multiple chips. The current convolutional layer needs to collect input data from all the chips where the previous convolutional layer is located, and the current convolutional layer also needs to distribute data (such as intermediate operation results) to each array within its own chip. Therefore, T A is positively correlated with the number of memristor array resources allocated to the current layer and the previous layer at the same time. Schematically, considering the size of the memory access cycle, the total processing delay is calculated with reference to the process of Figure 6 , including the steps:

[0093] A1. According to the input task parameters (i.e., the structural parameters and architecture parameters of the present invention), start calculating the total processing delay, initialize the layer index i = 0, k = 0, and go to step A2;

[0094] A2. Calculate NormalOp i , and go to step A3;

[0095] A3. Determine whether the condition i == 0 holds. If so, go to step A4; if not, go to step A5;

[0096] A4. PreOp 0 = 0, Op -1 = 0, Tail 0 = 0, go to step A10;

[0097] A5. Calculate Temp k = Interval(i, k)+PreOp k , go to step A6;

[0098] A6. k = k + 1, go to step A7;

[0099] A7. Determine whether the condition k < i holds. If so, go to step A5; if not, go to step A8;

[0100] A8. Calculate PreOp i = max k Temp k , go to step A9;

[0101] A9. Calculate Tail i , go to step A10;

[0102] A10. Calculate Op i = max{NormalOp i +PreOp i , Op i-1 +Tail i}, go to step A11;

[0103] A11. Calculate the memory access cycle size of the current layer according to Op i Obtain the cycle size for one calculation of the current layer Go to step A12;

[0104] A12. i = i + 1, set k to 0, go to step A13;

[0105] A13. Determine whether the condition i < L holds. If so, go to step A2; if not, go to step A14;

[0106] A14. Calculate Go to step A14;

[0107] A15. Output Time is the total processing delay calculated considering the memory access cycle.

[0108] ​According to an embodiment of the present invention, the method includes: calculating the total processing delay in a corresponding manner according to whether the kernel size of the pooling layer and / or the stride of the pooling layer is 1. Wherein, if the kernel size of the pooling layer corresponding to a convolutional layer and / or the stride of the pooling layer is 1, the delay of the pooling of this convolutional layer is not calculated in the total processing delay; if the kernel size of the pooling layer corresponding to a convolutional layer and / or the stride of the pooling layer is a value greater than 1, the delay of the pooling of this convolutional layer is added to the total processing delay. The calculation method of the pooling delay is known to those skilled in the art and will not be elaborated here.

[0109] Step S4: Determine the number of memristor arrays allocated to this convolutional layer according to the final allocation multiple and the minimum requirement of each convolutional layer.

[0110] According to an embodiment of the present invention, the number of memristor arrays allocated to each convolutional layer is equal to the product of the final allocation multiple and the minimum requirement of this convolutional layer.

[0111] According to an embodiment of the present invention, there is provided an accelerator for a convolutional neural network based on memristors. The accelerator includes multiple chips, and each chip includes one or more memristor arrays. The accelerator is configured to: obtain the number of memristor arrays allocated to each convolutional layer of the convolutional neural network to be accelerated according to the resource allocation method of the foregoing embodiment; allocate the required number of memristor arrays to the convolutional neural network according to the number of memristor arrays allocated to each convolutional layer among the multiple chips to perform corresponding convolutional operations.

[0112] According to an embodiment of the present invention, there is provided an electronic device, including: one or more processors; and a memory, where the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method of the foregoing embodiment by executing the executable instructions. The process of calculating the number of memristor arrays allocated to each convolutional layer can be implemented on a general-purpose computer or an accelerator. If it is implemented on a general-purpose computer, it will send the number of memristor arrays allocated to each convolutional layer to the accelerator, and the accelerator will allocate corresponding resources according to this.

[0113] To verify the effect of the present invention, the inventor simulated mapping the VGG-A convolutional neural network to the Figure 2 accelerator shown in the hardware environment of an intel i5 processor and the software environment of Python 3.8. Experiments prove that the allocation strategy obtained by solving with the present invention can make the performance of the accelerator reach 94% of the quasi-optimal solution. Compared with the existing heuristic allocation method, the present invention can improve the performance of CNN inference by 14 times.

[0114] In summary, the resource allocation method for a memristor-based convolutional neural network accelerator proposed by the present invention is superior to existing heuristic allocation strategies in terms of generality, accuracy, solution time, and solution effect. Generally speaking, some embodiments of the present invention have the following advantages:

[0115] Some embodiments of the present invention unify different scenarios and tasks into a general expression form through a predetermined parameter extraction format, making the framework of the proposed solution allocation strategy have good universality;

[0116] Some embodiments of the present invention accurately model the performance of the convolutional neural network accelerator. Combining the extracted structural parameters and architecture parameters, corresponding solution formulas or methods are proposed for the number of cycles of each convolutional layer operation, completing the modeling of the objective function in the mathematical problem abstracted in the previous step, which is a prerequisite for ensuring the effect of the allocation strategy;

[0117] Some embodiments of the present invention propose a solver based on dynamic programming on the basis of analysis and observation, avoiding brute-force search of the problem and completing the task of obtaining an approximate optimal solution in a short time

[0118] The resource allocation strategy obtained by some embodiments of the present invention comprehensively considers the influence of computing and bandwidth and is comprehensive compared with existing heuristic methods.

[0119] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed as long as the required functions can be achieved.

[0120] The present invention can be a system, method, and / or computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0121] A computer-readable storage medium can be a tangible device that retains and stores instructions for use by an instruction execution device. A computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.

[0122] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for allocating resources in an accelerator for a memristor-based convolutional neural network, the accelerator comprising a plurality of memristor arrays, characterized in that, Including: Obtain the structural parameters of the convolutional neural network to be accelerated and the architectural parameters related to the memristor array in the accelerator, where the structural parameters include the structural information related to each convolutional layer; Determine the minimum requirement of each convolutional layer of the convolutional neural network model based on the structural information and architectural parameters, where the minimum requirement is the minimum number of memristor arrays required to store all the weight parameters in the corresponding convolutional layer on this accelerator; According to the predetermined constraint conditions, taking the allocation multiple of each convolutional layer based on its minimum requirement as the adjustment object, determine the resource allocation strategy with the optimal total processing delay for the convolutional neural network to run on this accelerator, where the optimal resource allocation strategy indicates the final allocation multiple of each convolutional layer; Determine the number of memristor arrays allocated to this convolutional layer according to the final allocation multiple and minimum requirement of each convolutional layer.

2. The method according to claim 1, wherein The structural parameters include the following structural information related to each convolutional layer: The number of input channels, the number of output channels, the width and height of the output feature map, the kernel size of the convolutional layer, the kernel size of the pooling layer, the stride of the convolutional layer, the stride of the pooling layer, the padding size of the convolutional layer, the padding size of the pooling layer of the convolutional layer, where the structural information of each convolutional layer records the kernel size and stride of the pooling layer after this convolutional layer.

3. The method according to claim 2, wherein The architectural parameters include the total number of memristor arrays and the size of each memristor array, and this size indicates the number of bit lines and word lines of the corresponding memristor array.

4. The method according to claim 3, wherein Determine the minimum requirement of each convolutional layer in the following manner: Among them, represents the kernel size of the convolutional layer, represents the number of input channels of the convolutional layer, represents the number of output channels of the convolutional layer, represents the number of word lines of the memristor array, represents the number of bit lines of the memristor array.

5. The method according to claim 1, characterized in that, The method includes: calculating the corresponding total processing delay according to whether there is a pooling layer after the convolutional layer and the kernel size and / or stride of the pooling layer.

6. The method according to claim 1, characterized in that, Determine the total processing delay of the convolutional neural network running on this accelerator in the following manner: Determine the corresponding resource allocation strategy based on the total number and size of the memristor arrays; Obtain the maximum cycle size among the cycle size for each convolutional layer to calculate once and the memory access cycle size under the corresponding resource allocation strategy, and determine the total processing delay of this resource allocation strategy according to the maximum cycle size and the number of cycles for the last convolutional layer to run.

7. The method according to claim 6, characterized in that, Determine the number of cycles for each convolutional layer to run in the following manner: Among them, represents the number of cycles for the convolutional layer to run, represents taking the maximum value, represents the convolutional layer the number of cycles for calculation, represents the convolutional layer the number of cycles passed while waiting to be activated, represents the convolutional layer the number of cycles for the previous convolutional layer of the convolutional layer to run, represents the convolutional layer the number of cycles that still need to run after the previous convolutional layer of the convolutional layer stops calculating.​ 8. The method according to any one of claims 1 to 7, characterized in that The predetermined constraint conditions include: the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator; Or The predetermined constraint conditions include: The sum of the number of memristor arrays allocated to all convolutional layers is less than or equal to the total number of memristor arrays contained in the accelerator, and the allocation multiple of any one convolutional layer is less than or equal to the width multiplied by the height of the output feature map of this convolutional layer.

9. An accelerator, the accelerator includes a plurality of chips, each chip includes one or more memristor arrays, characterized in that, This accelerator is configured to: Obtain the number of memristor arrays allocated to each convolutional layer of the convolutional neural network to be accelerated according to the method of any one of claims 1-8; Allocate the required number of memristor arrays for the convolutional neural network according to the number of memristor arrays allocated to each convolutional layer among the multiple chips to perform the corresponding convolutional operation.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method described in any one of claims 1 to 8.

11. An electronic device, characterized in that, Including: One or more processors; And A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1 to 8 by executing the executable instructions.

Citation Information

Patent Citations

  • Resource allocation method and device for DNN accelerator based on memristor

    CN112561049A

  • Optical convolutional neural network accelerator

    US20200019851A1