Model processing method, device, electronic device and readable medium
By fusing the computing-intensive layer with the non-computation-intensive layer into the fusion layer, and using the input and output of the fusion layer to perform subsequent processing in the on-chip cache, the problem of time-consuming data access in the non-computation-intensive layer in the neural network is solved, and processing performance and efficiency are improved.
Patent Information
- Application Number
- CN202111677862.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-12-31
AI Technical Summary
During the processing of existing neural networks, the processing performance is lower due to the long time to access data in non-computing-intensive layers and memory.
By fusing the computing-intensive layer with the non-computation-intensive layer into the fusion layer, the target output of the non-computation-intensive layer is calculated using the input and output of the fusion layer, and subsequent processing operations are performed in the on-chip cache, reducing the repeated access of data in memory.
It reduces the time-consuming process, improves processing performance, reduces the inventory of access, and improves the processing efficiency of neural networks.
Smart Images

Figure CN114418065B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of electronic devices, and in particular to a model processing method, device, electronic device, and readable medium. Background Art
[0002] At present, neural networks, as a very important method in the field of artificial intelligence, are widely used in applications such as image classification, pattern recognition, target detection, and speech recognition.
[0003] In related technologies, neural networks are often composed of multiple layers. Each layer, during processing, needs to read all of its inputs (data) from memory, then perform the layer's specified calculations based on the inputs (data), and then store the results back into memory. However, the data access process between some layers and memory is time-consuming, resulting in longer processing times and lower performance. Summary of the Invention
[0004] The embodiments of the present invention provide a model processing method, device, electronic device and readable medium to reduce the time consumption of the processing process and improve the processing performance.
[0005] In a first aspect, a model processing method is provided, which is applied to a target neural network, wherein the target neural network includes a fusion layer, wherein the fusion layer is used to implement the computational operations required to be performed by a computationally intensive layer and at least one non-computationally intensive layer in the original neural network, the method comprising:
[0006] For at least one fusion layer in the target neural network, the at least one fusion layer calculates a target output of a current layer of a non-computational intensive layer corresponding to the at least one fusion layer based on an input of the computational intensive layer corresponding to the at least one fusion layer and / or a computational output of the computational intensive layer;
[0007] Perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0008] In a second aspect, a model processing device is provided for use in a target neural network, the target neural network including a fusion layer, the fusion layer being configured to implement computational operations required to be performed by a computationally intensive layer and at least one non-computationally intensive layer in an original neural network, the device comprising:
[0009] a computing module configured to, for at least one fusion layer in the target neural network, calculate, by the at least one fusion layer, a target output of a current layer of a non-computational intensive layer corresponding to the at least one fusion layer based on an input of the computational intensive layer corresponding to the at least one fusion layer and / or a computational output of the computational intensive layer;
[0010] A processing module is used to perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0011] In a third aspect, a model processing method is provided, which is applied to a target neural network, wherein the target neural network includes a fusion layer, and the fusion layer is used to implement the computational operations required to be performed by the computation-intensive layer and at least one non-computation-intensive layer in the original neural network, the method comprising:
[0012] In response to the received preset instruction, determining, according to the preset instruction, a pre-processing operation and / or a post-processing operation to be performed corresponding to at least one fusion layer in the target neural network;
[0013] The at least one fusion layer performs the pre-processing operation and / or the post-processing operation based on the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer, so as to calculate the target output of the computation of the non-computationally intensive layer corresponding to the at least one fusion layer;
[0014] Perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0015] In a fourth aspect, a model processing device is provided for use in a target neural network, wherein the target neural network includes a fusion layer, wherein the fusion layer is configured to implement computational operations required to be performed by a computationally intensive layer and at least one non-computationally intensive layer in an original neural network, the device comprising:
[0016] A first determining module is configured to determine, in response to a received preset instruction, a pre-processing operation and / or a post-processing operation to be performed corresponding to at least one fusion layer in the target neural network according to the preset instruction;
[0017] an execution module, configured to execute, by the at least one fusion layer, the pre-processing operation and / or the post-processing operation based on the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer, so as to calculate a target output of the computation of the non-computationally intensive layer corresponding to the at least one fusion layer;
[0018] A processing module is used to perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0019] According to a fifth aspect, an electronic device is provided, including:
[0020] One or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform the model processing method.
[0021] In a sixth aspect, one or more machine-readable media are provided, on which instructions are stored, which, when executed by one or more processors, cause the processors to perform the model processing method.
[0022] In an embodiment of the present invention, based on the fusion layer in the target neural network for implementing the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, the target output of the computation of the non-computationally intensive layer corresponding to the fusion layer is calculated according to the input of the computationally intensive layer corresponding to the fusion layer and / or the computational output of the computationally intensive layer. Subsequent processing is performed based on the target output, and the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer. In this way, by fusing the non-computationally intensive layer that takes a long time to access data from the memory with the computationally intensive layer, the problem of repeatedly accessing data from the memory in the process of executing the non-computationally intensive layer layer by layer can be avoided to a certain extent, thereby reducing the amount of memory access, reducing the time consumption of the processing process, and improving processing performance.
[0023] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0025] Figure 1 This is a flowchart of the steps of a model processing method provided by an embodiment of the present invention;
[0026] Figure 2 is a structural diagram of an existing accelerator shown in an embodiment of the present invention;
[0027] Figure 3 is another structural diagram of an accelerator provided by an embodiment of the present invention;
[0028] Figure 4 is a flowchart of another model processing method provided by an embodiment of the present invention;
[0029] Figure 5 This is a structural block diagram of a model processing device provided by an embodiment of the present invention;
[0030] Figure 6 It is a structural block diagram of another model processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0032] Figure 1 This is a flow chart of the steps of a model processing method provided by an embodiment of the present invention. The method can be applied to a target neural network, wherein the target neural network includes a fusion layer, and the fusion layer is used to implement the operations of the computationally intensive layer and at least one non-computationally intensive layer in the original neural network. Figure 1 As shown, the method may include:
[0033] Step 101: For at least one fusion layer in the target neural network, the at least one fusion layer calculates a target output of a non-computational intensive layer corresponding to the at least one fusion layer based on an input of the computational intensive layer corresponding to the at least one fusion layer and / or a computational output of the computational intensive layer.
[0034] In an embodiment of the present invention, the target neural network can be obtained by fusing the computationally intensive layer and the non-computationally intensive layer in the original neural network. The computational operation required to be performed by each layer in the neural network is the computational operation specified by the layer. Fusion of the computationally intensive layer and the non-computationally intensive layer can refer to merging the computationally intensive layer and the non-computationally intensive layer into one layer. The fused layer obtained after merging can be used to implement the computational operation specified by the computationally intensive layer and part / all of the computational operation specified by the non-computationally intensive layer. In other words, the fused layer in the target neural network can be composed of part / all of the structure of the computationally intensive layer and the non-computationally intensive layer. For example, assuming that the fused layer A can be used to implement the computational operation specified by the computationally intensive layer a1 and the computational operation specified by the non-computationally intensive layer a2, then the computationally intensive layer a1 is the computationally intensive layer corresponding to the fused layer A, and the non-computationally intensive layer a2 is the non-computationally intensive layer corresponding to the fused layer A. Furthermore, assuming that fusion layer B can be used to implement the layer-specific computations specified by compute-intensive layer b1 and part of the layer-specific computations specified by non-compute-intensive layer b2, and fusion layer C can be used to implement the layer-specific computations specified by compute-intensive layer c1 and another part of the layer-specific computations specified by non-compute-intensive layer b2, then compute-intensive layer b1 is the compute-intensive layer corresponding to fusion layer B, and non-compute-intensive layer b2 is the non-compute-intensive layer corresponding to fusion layer A. Compute-intensive layer c1 is the compute-intensive layer corresponding to fusion layer C, and non-compute-intensive layer b2 is the non-compute-intensive layer corresponding to fusion layer C. In other words, the layer-specific computations specified by a non-compute-intensive layer can be broken down into multiple parts, with different fusion layers performing some of these parts, and multiple fusion layers working together to implement the layer-specific computations specified by the non-compute-intensive layer.
[0035] Furthermore, the computationally intensive layer corresponding to the fusion layer may be a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer. It should be noted that the relative positions before and after the processing are different. The processing may be a forward processing process or a backpropagation process. The forward processing may be a forward processing process in the prediction link or the training link. The backpropagation process may refer to the process of backpropagating the error in the training link to calculate the weight gradient corresponding to each sample and uniformly update the weights. The forward processing may be processed sequentially from the input layer to the output layer, and the reverse processing may be processed sequentially from the output layer to the input layer. In different processing processes, there are differences in the non-computationally intensive layers before and after the same non-computationally intensive layer.
[0036] Since the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer, when performing fusion calculations, the target output of the current layer calculation of the non-computationally intensive layer corresponding to the fusion layer can be further calculated based on the input of the computationally intensive layer corresponding to the fusion layer and / or the computational output of the computationally intensive layer. The computational output of the computationally intensive layer can be the output that can be obtained after executing the current layer calculation specified by the computationally intensive layer, or it can be the output that can be obtained after executing part of the current layer calculation specified by the computationally intensive layer. This is not limited in the embodiments of the present invention. The target output refers to the output that can be obtained after executing the current layer calculation specified by the non-computationally intensive layer.
[0037] Furthermore, the computationally intensive layers and the non-computationally intensive layers may be pre-divided, the computationally intensive layers may be layers whose computational time ratio is greater than a preset threshold, and the non-computationally intensive layers may be layers whose computational time ratio is not greater than a preset threshold. The computational time ratio can be understood as the ratio of the time required to perform the computation to the total operation time of the layer, and the total operation time of the layer refers to the sum of the data access time and the time required to perform the computation. For example, the computationally intensive layers may include convolutional layers, and the non-computationally intensive layers may include BN layers, fully connected layers, pooling layers, and non-linear layers, such as RELU layers. The same type of layers often contain multiple layers in a neural network, for example, multiple convolutional layers. Accordingly, different fusion layers may contain layers of the same type.
[0038] Step 102: performing subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0039] In the embodiment of the present invention, after obtaining the target output, subsequent processing can be performed based on the target output to ensure that the current processing process can be smoothly executed. The specific type of subsequent processing can be set according to actual needs and is not limited in the embodiment of the present invention.
[0040] In summary, the model processing method provided by the embodiment of the present invention is based on the fusion layer in the target neural network for realizing the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and calculates the target output of the current layer calculation of the non-computationally intensive layer corresponding to the at least one fusion layer according to the input of the computationally intensive layer and / or the computational output of the computationally intensive layer corresponding to the at least one fusion layer. Subsequent processing is performed based on the target output, and the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer. In this way, by fusing the non-computationally intensive layer that takes a long time to access data from the memory with the computationally intensive layer, the problem of repeatedly accessing data from the memory in the process of executing the non-computationally intensive layer layer by layer can be avoided to a certain extent, thereby reducing the amount of memory access, reducing the time consumption of the processing process, and improving processing performance.
[0041] Optionally, in one implementation, the operation of calculating the target output of the at least one corresponding non-computation-intensive layer based on the input of the computation-intensive layer corresponding to the at least one fusion layer and / or the computation output of the computation-intensive layer may specifically include:
[0042] Step S21: After executing the first data calculated by the current layer of the first computationally intensive layer corresponding to the first fusion layer, perform a first post-processing operation, wherein the first post-processing operation includes calculating the first intermediate data of the non-computationally intensive layer based on the first data; the first intermediate data is the intermediate data required to calculate the target output.
[0043] In this implementation, the computationally intensive layer corresponding to the first fusion layer may include a first computationally intensive layer, and the computationally intensive layer corresponding to the second fusion layer may include a second computationally intensive layer. The non-computationally intensive layers corresponding to the first fusion layer and the second fusion layer are the same. The aforementioned at least one fusion layer may include a first fusion layer and a second fusion layer. That is, at least one fusion layer may include a fusion layer for implementing the current layer calculations specified by the first computationally intensive layer and calculating the first intermediate data, and a fusion layer for implementing the current layer calculations specified by the second computationally intensive layer and calculating the target output based on the first intermediate data. Furthermore, the first intermediate data may be predetermined according to the calculation logic of the target output, and the first intermediate data may be the coefficients required to calculate the target output. The target neural network may include multiple fusion layers, and one fusion layer in the target neural network may be the first fusion layer, the second fusion layer, the second fusion layer, or the fourth fusion layer in the embodiment of the present invention.
[0044] In an embodiment of the present invention, calculating the first intermediate data of the non-computational intensive layer based on the first data can be regarded as a post-processing operation of the first computational intensive layer. After executing the current layer calculation of the first computational intensive layer to obtain the first data, the post-processing operation can be performed based on the first data still cached on-chip to obtain the first intermediate data.
[0045] Step S22: store the first intermediate data and the first data in a memory.
[0046] In an embodiment of the present invention, after obtaining the first intermediate data, the first fusion layer may store the first intermediate data and the first data in an off-chip memory. Accordingly, the layer following the first fusion layer may read the first intermediate data and the first data from the memory as inputs to the layer for subsequent processing. The layer following the first fusion layer is referred to as the second fusion layer.
[0047] Step S23: The second fusion layer calculates the target output based on the first intermediate data and the first data read from the memory.
[0048] Accordingly, the above-mentioned operation of performing subsequent processing based on the target output may specifically include: performing a first pre-processing operation, wherein the first pre-processing operation includes performing the current layer calculation of the second computationally intensive layer corresponding to the second fusion layer according to the target output.
[0049] Among them, the operation of calculating the target output based on the first intermediate data and the first data can be regarded as a pre-processing operation of the second computationally intensive layer. After the target output is calculated, the current layer calculation of the second computationally intensive layer corresponding to the second fusion layer can be performed based on the target output. Specifically, the current layer calculation of the second computationally intensive layer can be performed directly based on the target output, thereby completing the operations required to be performed by the second fusion layer. In the case where the second fusion layer also includes other non-computationally intensive layers, the current layer calculation of other non-computationally intensive layers can also be performed, and then the current layer operations of the second computationally intensive layer can be performed based on the results of the current layer calculation of other non-computationally intensive layers. Alternatively, after the current layer calculation of the second computationally intensive layer, based on the calculation results of the current layer calculation of the second computationally intensive layer, the operations of other non-computationally intensive layers can be further performed, thereby completing the operations required to be performed by the second fusion layer. For example, the other non-computationally intensive layers can be RELU layers.
[0050] Furthermore, after obtaining the calculation results of the current layer of the second computationally intensive layer, the calculation results may be stored in an off-chip memory for use by the next layer of the second fusion layer.
[0051] In an embodiment of the present invention, a non-computationally intensive layer is fused with the preceding and following computationally intensive layers, and the computation of the current layer specified by the non-computationally intensive layer is used as the post-processing operation of the first computationally intensive layer before the non-computationally intensive layer and the pre-processing operation of the second computationally intensive layer after the non-computationally intensive layer. The post-processing operation is a calculation performed on the output result of the computationally intensive layer, and can be executed immediately while the output result is still on-chip, without the need for additional memory access operations. The pre-processing operation is equivalent to adding a pre-processing link to the computationally intensive layer after executing the input, and there is no need to increase the memory access requirements of the layer. In this way, the problem of repeatedly accessing data from the off-chip memory in the process of executing the non-computationally intensive layer layer by layer can be eliminated through multiple fusion layers, thereby greatly reducing the amount of memory access.
[0052] Optionally, in another implementation, the fusion layer may also refer to a third fusion layer or a fourth fusion layer. It should be noted that there may be multiple third fusion layers in the embodiment of the present invention, and the computationally intensive layers and non-computationally intensive layers corresponding to different third fusion layers may be different. There may also be multiple fourth fusion layers, and the computationally intensive layers and non-computationally intensive layers corresponding to different fourth fusion layers may be different. The above-mentioned operation of calculating the target output of the current layer calculation of the at least one corresponding non-computationally intensive layer based on the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer may specifically include step S31 or step S41:
[0053] Step S31: After executing the current layer calculation of the third computationally intensive layer corresponding to the third fusion layer to obtain second data, perform a second post-processing operation, where the second post-processing operation includes calculating the target output based on the second data.
[0054] Among them, calculating the target output based on the second data can be regarded as a post-processing operation of the third computationally intensive layer. In this implementation, the calculation of this layer specified by the non-computationally intensive layer can be implemented based on the post-processing operation. For example, in this implementation, the non-computationally intensive layer may include a RELU layer, a fully connected layer, a pooling layer, and so on. Accordingly, the above-mentioned operation of subsequent processing based on the target output may specifically include: writing the target output to the memory for use by subsequent layers in the target neural network. By writing the target output to the memory, the next layer of the third fusion layer in this processing process can read the target output in the memory as the input of this layer, so as to perform subsequent processing.
[0055] Step S41: The fourth fusion layer reads the current layer input from the memory and performs a second pre-processing operation, where the second pre-processing operation includes calculating the target output based on the current layer input; the current layer input is the original input corresponding to the non-computation-intensive layer.
[0056] In this implementation, the operation of calculating the target output can serve as a preprocessing operation for the fourth compute-intensive layer corresponding to the fourth fusion layer. In this implementation, the preprocessing operation of the fourth compute-intensive layer can be used to perform the layer-specific computations specified by the non-computation-intensive layer. That is, the input of the fourth compute-intensive layer is no longer the target output of the non-computation-intensive layer, but rather the original input corresponding to the non-computation-intensive layer. The original input can be the information required by the non-computation-intensive layer to calculate the target output. Accordingly, the subsequent processing based on the target output can specifically include performing the layer-specific computations of the fourth compute-intensive layer corresponding to the fourth fusion layer based on the target output. Specifically, the layer-specific computations of the fourth compute-intensive layer can be performed directly based on the target output, thereby completing the operations required by the fourth fusion layer. For example, assume that the order of layers in this processing is: layer_A -> ReLU -> layer_B, where both layer_A and layer_B are compute-intensive layers. Then the ReLU layer can be fused with the computationally intensive layer (layer_B) behind it, and appear as the "pre-processing" in the execution of the fusion layer, that is, directly read the output value of layer_A, perform the required ReLU operation, obtain the result of ReLU, and then use the result of ReLU as the input value of layer_B to perform the required calculation. Furthermore, in the case where the fourth fusion layer also includes other non-computationally intensive layers, the current layer calculations of other non-computationally intensive layers can be performed, and then the current layer operations of the fourth computationally intensive layer can be performed based on the results of the current layer calculations of other non-computationally intensive layers. Alternatively, after the current layer calculations of the fourth computationally intensive layer, the operations of other non-computationally intensive layers can be further performed based on the calculation results of the current layer calculations of the fourth computationally intensive layer, thereby completing the operations required to be performed by the fourth fusion layer. For example, the other non-computationally intensive layers can be RELU layers.
[0057] In this embodiment of the present invention, a non-computationally intensive layer is integrated with the preceding or preceding computationally intensive layer, and the computations specified by the non-computationally intensive layer are used as post-processing or pre-processing operations. This eliminates the need to separate the computations of the non-computationally intensive layer and eliminates the problem of repeatedly accessing data from off-chip memory during the independent execution of the non-computationally intensive layer, thereby significantly reducing the amount of memory access.
[0058] Optionally, taking the first computationally intensive layer as a convolutional layer (CONV) and the non-computationally intensive layer as a BN layer as an example, when this processing process is a forward processing process of a training link, the operation of calculating the first intermediate data of the non-computationally intensive layer based on the first data may specifically include:
[0059] Step S51: Calculate the first parameter and the second parameter based on the output feature map of the same channel obtained by the convolutional layer calculation to serve as the first intermediate data; the output feature map is the first data;.
[0060] Among them, the output feature map of the same channel obtained by the convolution layer calculation is the first data calculated by the current layer of the convolution layer corresponding to the first fusion layer.
[0061] Correspondingly, the above-mentioned operation of calculating the target output based on the first intermediate data and the first data read from the memory may specifically include: performing a multiplication and addition operation based on the first parameter, the output feature map and the second parameter to obtain the target output.
[0062] In actual application scenarios, in the forward propagation algorithm of the BN layer, the calculation logic of the layer specified by the BN layer can be expressed as:
[0063] First, calculate the input value B of a certain part of the layer in the batch size (mini-batch) of the current iteration, which is {x1, x2, ..., x m The mean μ B and variance
[0064]
[0065]
[0066] Among them, X i represents the i-th value in the input value B, and m represents the total number of values included in the input value B.
[0067] Then, for the input value B={x1,x2,…,x m}Perform normalization operation to obtain the normalized result:
[0068]
[0069]
[0070] Finally, the normalized result is linearly transformed to obtain the output result y of the BN layer i :
[0071]
[0072] Among them, BN γ,β (x i ) is another representation of the output of the BN layer. ε represents a preset constant, γ and β represent predefined learnable parameters.
[0073] That is to say, for the connection mode of CONV layer->BN layer, the calculation of this layer specified by the BN layer is based on the output feature map X of the CONV layer. i As input data, based on the output feature map X i Calculate the target output y i Specifically, all input values on the same channel in the mini-batch samples can be selected as the input value set B for batch normalization operation, and the input values of the same channel use the same learnable parameters γ and β to calculate formula (4).
[0074] Furthermore, in order to split the operations specified by the BN layer into pre-processing operations and post-processing operations, the processing process of the BN layer can be optimized in the embodiment of the present invention. Specifically, formulas (3) and (4) can be combined to obtain:
[0075]
[0076] set up:
[0077]
[0078] Then, formula (11) can be expressed as y i =F1×x i +F2.
[0079] Wherein, F1 and F2 may represent first intermediate data, and F1 and F2 may represent first parameter and second parameter respectively.
[0080] In the embodiment of the present invention, the calculation operation of the BN layer in the forward processing process can be divided into two stages. The first calculation stage calculates the constant values F1 and F2; the second calculation stage uses the constant values calculated in the first calculation stage to perform a linear transformation operation, that is, a multiplication and addition operation, to obtain the target output y in this processing process. i It will be appreciated that the target neural network may perform multiple processing steps.
[0081] Since the first calculation stage calculates F1 and F2, it is necessary to calculate the mean μ B and variance According to formulas (1) and (2), as long as the input value (that is, the output feature map X i ) Cumulative sum and the sum of squares The mean and variance can be calculated relatively easily. Therefore, in the embodiment of the present invention, the cumulative sum and square sum of the output feature map values of the same channel can be calculated while performing the convolution calculation of the CONV layer. In this way, after the calculation of a certain channel of the output feature map of the mini-batch samples of the CONV layer is completed, the cumulative sum of this channel is and the sum of squares The final result can also be obtained. Accordingly, according to formula (1), the division operation can be performed to obtain μ B ; Perform division operation according to formula (2) to obtain Further multiplication and addition operations You can get Finally, F1 and F2 can be calculated based on the mean and variance. The other parameters in the calculation formula corresponding to the mean and variance are pre-set. It should be noted that during the forward processing of the inference phase, the first and second parameters can be directly calculated using the fixed mean and variance from the training phase.
[0082] Furthermore, layerL represents the second computationally intensive layer corresponding to the second fusion layer. When the inter-layer connection is CONV->BN->layerL, the linear operation of the BN layer on the input value (i.e., the second computation phase) can be fused with layerL into a new layer: layerL_New. The input of layerL_New is no longer the output result of the BN layer, but directly the output feature map of the CONV layer. When executing layerL_New, the output feature map of the CONV layer, F1 and F2 can be read in, and then the multiplication and addition operation F1×x of the second computation phase of the BN layer can be executed immediately. i +F2, continue to perform the calculation required by layerL using the result of the multiply-add operation (ie, the target output).
[0083] In an application scenario involved in an embodiment of the present invention, the present invention can be applied to the field of neural network acceleration. For example, convolutional neural networks (CNN), deep neural networks (DNN), etc. Among them, CNN may include ResNet, ResNext, Wide ResNet, VGG, etc. The task of a neural network includes two parts: training and inference. Training refers to the process of determining the weight values in the network model through a certain method. Inference refers to the process of using the trained network model to calculate the prediction results corresponding to new inputs. Currently, during the training process of DNN, the update of network parameters causes the distribution of activation values of each layer in the network to change, which in turn leads to the phenomenon of internal covariate shift (ICS). In order to overcome the problem of ICS and ensure the accuracy and convergence of training, it is necessary to use a lower learning rate in training and perform more accurate parameter initialization. However, this will greatly reduce the training speed. To this end, batch normalization (BN) is often used in neural networks. By normalizing the input values and fixing the mean and variance of the input, the internal covariate shift is reduced, allowing a higher learning rate to be used during training, accelerating network training. At the same time, batch normalization allows the network to no longer rely too much on parameter initialization, making it easier for the network to converge and achieve higher accuracy. Therefore, batch normalization is currently used in most neural network training processes to accelerate training. The BN layer has become an important component of neural network models.
[0084] In the embodiment of the present disclosure, the non-computational intensive layer is integrated with the computational intensive layers before and after it. Specifically, the calculation of the BN layer is divided into two stages. The first stage is integrated with the execution of the computational intensive layer before the BN layer, and appears as the "post-processing" of the calculation required to be performed by the computational intensive layer. The second stage is integrated with the first computational intensive layer after the BN layer, and appears as the "pre-processing" of the calculation required to be performed by the layer, thereby developing the parallelism of the BN layer and the computational intensive layer. In this way, the problem of repeatedly accessing data from the off-chip memory in the process of executing the non-computational intensive layer layer by layer can be avoided to a certain extent, which greatly reduces the amount of memory access, thereby improving the performance of training and reducing the energy consumption of training to a certain extent. At the same time, it can also save time in the reasoning process, improve the processing speed of the reasoning process, and accelerate the training and reasoning process of the neural network.
[0085] Optionally, in this embodiment of the present invention, the non-computationally intensive layer may further include a RELU layer. Accordingly, the aforementioned operation of performing the current layer calculation of the second computationally intensive layer corresponding to the second fusion layer based on the target output may specifically include: performing a RELU operation on the target output, and performing the current layer calculation of the second computationally intensive layer based on the result of the RELU operation. This can reduce the off-chip memory access overhead of multiple non-computationally intensive layers, thereby further saving time.
[0086] In other words, when the inter-layer connection is CONV->BN->ReLU->layerL, the linear operation on the input value of the BN layer can be combined with the ReLU layer and layerL to form a new layer, layerL_New. The input of layerL_New is no longer the output of the ReLU layer, but directly the output feature map of the CONV layer. When executing layerL_New, the output feature map of the CONV layer, F1 and F2, can be read, and then the multiplication and addition operation of the second calculation stage of the BN layer can be performed. The ReLU operation is then performed, and finally the calculation required by layerL is continued using the result of the ReLU operation.
[0087] As can be seen, in the connection layer_A->BN->(ReLU)->layer_B, both layer_A and layer_B are compute-intensive layers. In this embodiment of the present invention, the BN layer is divided into two stages. The first computational stage is integrated with the execution of the compute-intensive layer (layer_A) preceding the BN layer, appearing as a post-processing operation within the execution of this layer (layer_A). Post-processing operations represent required processing of the output of layer_A, such as calculating cumulative sums, sums of squares, divisions, or calculating F1 and F2. The linear mapping operation in the second computational stage can be integrated with the non-compute-intensive ReLU layer following the BN layer, and with the first compute-intensive layer (layer_B) following the BN layer, appearing as a pre-processing operation within the execution of this fused layer. Pre-processing operations directly read the output of layer_A and perform the required operations, such as the multiplication-addition operation in the second BN computational stage, or a multiplication-addition operation combined with a ReLU operation, to obtain the input value for layer_B.
[0088] Furthermore, in the forward processing, when the BN layer is executed separately, it is necessary to first read all the input values of the BN layer and calculate the mean μ B and variance Then re-read all the input values of the BN layer and calculate and yi; then y iStored back to off-chip memory. Assuming the total number of input values for the batch normalization layer from a mini-batch of samples is V, the memory access amount can be recorded as 3V. When executing the ReLU layer individually, all input values for the ReLU layer must first be read in, the ReLU result calculated, and then stored back off-chip. Assuming the total number of input values for the ReLU layer from a mini-batch of samples is V, the memory access amount can be recorded as 2V.
[0089] In the prior art, when the connection is layer_A->BN->layer_B, the amount of memory access performed by layer can be expressed as:
[0090] INPUT(layer_A)+OUTPUT(layer_A)+3*OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B).
[0091] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0092] As can be seen from the calculation of INPUT(layer_A)+OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B), compared to the layer-by-layer approach, this embodiment of the present invention can reduce the number of memory accesses by: 3*OUTPUT(layer_A). Each INPUT and OUTPUT operation can be considered as a single memory access.
[0093] When the connection is layer_A->BN->ReLU->layer_B, the amount of memory access performed by layer can be expressed as:
[0094] INPUT(layer_A)+OUTPUT(layer_A)+5*OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B).
[0095] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0096] It can be seen from INPUT(layer_A)+OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B) that, compared with the layer-by-layer execution method, the embodiment of the present invention can reduce the amount of memory access: 5*OUTPUT(layer_A).
[0097] When the connection is layer_A->ReLU->layer_B, the amount of memory access performed by layer can be expressed as:
[0098] INPUT(layer_A)+OUTPUT(layer_A)+2*OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B).
[0099] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0100] It can be seen from INPUT(layer_A)+OUTPUT(layer_A)+INPUT(layer_B)+OUTPUT(layer_B) that, compared with the layer-by-layer execution method, the embodiment of the present invention can reduce the amount of memory access: 2*OUTPUT(layer_A).
[0101] Optionally, taking the first computationally intensive layer as a convolutional layer (CONV) and the non-computationally intensive layer as a BN layer as an example, when the current processing process is a backpropagation process of a training link, the operation of calculating the first intermediate data of the non-computationally intensive layer based on the first data may specifically include:
[0102] Step S61: Calculate a third parameter and a fourth parameter based on the error value obtained by the convolution layer calculation to serve as the first intermediate data.
[0103] The error value obtained by the convolutional layer calculation is the first data obtained after executing the current layer calculation of the first computationally intensive layer corresponding to the first fusion layer.
[0104] Accordingly, the operation of calculating the target output based on the first intermediate data and the first data read from the memory may specifically include performing a multiplication-addition operation based on the third parameter, the fourth parameter, and the error value to obtain the target output. In this way, the fusion layer can reduce the amount of memory access during the backpropagation process, thereby reducing the time consumption to a certain extent.
[0105] In actual application scenarios, in the back propagation algorithm of the BN layer, the BN layer needs to calculate the error value. Accordingly, the calculation logic of this layer specified by the BN layer can be expressed as:
[0106]
[0107]
[0108]
[0109]
[0110] in, The symbols representing partial derivative operations, γ and β are learnable parameters, and their gradients can be calculated as follows:
[0111]
[0112]
[0113] When connected in CONV->BN mode, that is, during the back propagation process, the BN layer serves as the previous processing layer of the CONV, and the target output calculated by the BN layer is the output error value of the BN layer. When calculating the output error value of the BN layer, all data of the same channel in the mini-batch samples of the current iteration can be selected to participate in the corresponding accumulation operation.
[0114] Furthermore, in order to split the operations specified by the BN layer into pre-processing operations and post-processing operations, the processing process of the BN layer can be optimized in the embodiment of the present invention. Specifically, formulas (5)-(8) can be combined to obtain:
[0115]
[0116] set up:
[0117]
[0118] Then formula (12) can be expressed as:
[0119]
[0120] set up:
[0121] B=F1×F3×(μ B × S1-S2), C = -μ B ×B-F1×S1
[0122] Then formula (13) can be expressed as:
[0123] Wherein, B and C may represent the first intermediate data, and B and C may represent the third parameter and the fourth parameter respectively.
[0124] In the embodiment of the present invention, the calculation operation of the BN layer during the back propagation process can be divided into two stages. The first calculation stage calculates the constant values B and C; the second calculation stage uses the constant values calculated in the first calculation stage to perform multiple multiplication and addition operations to obtain the target output in this processing process.
[0125] Furthermore, the first calculation stage can only calculate B and C. Among them, the required F1, F3, mean μB These are constant values calculated in the previous processing (for CONV layers, each channel has its own constant value). After these constant values are calculated, they will be stored in the off-chip memory and can be read back on-chip to participate in the calculation directly during the back propagation process. Therefore, in order to calculate B and C, it is necessary to calculate S1 and S2, that is, to calculate the cumulative sum of the input error values of the BN layer. And the cumulative sum of the product of the input error value of the BN layer and the input feature map Among them, the input error value of the BN layer refers to the error value of the output of the layer above the BN layer during the back propagation process.
[0126] When the inter-layer connection is BN->CONV, CONV is before BN during the back propagation process, Refers to the error value output by the back propagation process of the CONV layer, x i Refers to the input feature map of the BN layer during the forward propagation process (i.e., the forward processing process). In other words, the inter-layer connection is specifically CONV->BN->CONV. While performing the error backpropagation calculation of the CONV layer, the cumulative sum of the output error values of the same channel and the cumulative sum of the products of the output error values and the input feature map can be calculated to calculate the third parameter and the fourth parameter. The input feature map can be obtained by performing the layer calculation specified by the CONV layer during the forward propagation process.
[0127] Furthermore, when the inter-layer connection is BN->ReLU->CONV, the order of execution in the back propagation process is CONV->ReLU->BN. It can refer to the output error value of the ReLU layer that should be output to the BN layer during the back propagation process. The output feature map of the BN layer during the forward propagation process can be recorded as y i ,y i =F1×x i +F2. The output value of the ReLU layer in the forward process can be recorded as Accordingly, based on the chain rule, we can derive: Accordingly, we can get in, is the error value output by the back propagation process of the CONV layer, ReLU′(y i ) represents the ReLU(y i ) to find the derivative,
[0128] It can be seen that in the back propagation process of the inter-layer connection BN->ReLU->CONV, the first fusion layer is used to implement the current layer calculation specified by the CONV layer, the current layer calculation specified by the ReLU layer, and the first calculation stage of the BN layer. Correspondingly, while performing the error back propagation calculation of the CONV layer, two cumulative sums can be calculated: and Thus, the third parameter and the fourth parameter are calculated.
[0129] Specifically, after the output error values of a certain channel of the mini-batch samples of the CONV layer are calculated, the cumulative calculation of this channel can also get the final result. Furthermore, the two cumulative sums are divided to obtain S1 and S2. Furthermore, the constant values F1, F2, F3 and μ for each channel calculated in the forward propagation process and stored in the off-chip memory are used. B , you can calculate the B and C corresponding to each channel.
[0130] Furthermore, in an embodiment of the present invention, the second computational stage of the BN layer in the back-propagation process can be considered to be integrated with the computationally intensive layer that is executed thereafter in the current processing process.
[0131] When the inter-layer connection is layerL->BN->CONV, the calculation formula of the second calculation stage can be expressed as: Accordingly, the target output can be calculated by two multiplication and addition operations. For example, the first calculation The second calculation of B×x i + tmp, thereby obtaining the target output in this processing process. Of course, a dedicated logic unit for calculating a*b+c*d+e may also be used to implement the multiplication and addition operation based on the third parameter, the fourth parameter, and the error value to obtain the target output, which is not limited in this embodiment of the present invention.
[0132] In this embodiment of the present invention, the multiple multiplication and addition operations of the second calculation stage are fused with layerL to form a new layer, layerL_New. The input of layerL_New is no longer the output error value of the BN layer, but directly the output error value of the CONV layer, the first intermediate data of the current calculation process, and the data obtained in the previous forward propagation process. When executing layerL_New, the multiple specified operations of the second calculation stage of the BN layer can be performed first, and the results of the second calculation stage (i.e., the target output of the current processing process) can be used to continue the calculations required by the backpropagation process of layerL.
[0133] When the inter-layer connection is layerL->BN->ReLU->CONV, the calculation formula of the second calculation stage can be expressed as: Accordingly, the implementation method can be: first perform a multiplication and addition operation F1×x i +F2 calculates y i , due to y i It is a floating point number, so the ReLU′(y i ) is 0 or 1, so we can directly know The result of the calculation is 0 or Then perform a multiplication and addition operation tmp2 = F1 × tmp1 + C, and then perform a multiplication and addition operation B × x i + tmp2, we can get the result of the second calculation stage. The multiple multiplication and addition operations of the second calculation stage are fused with layerL to form a new layer layerL_New. The input of layerL_New is no longer the output error value of the BN layer. Accordingly, when executing layerL_New, we can first perform the multiple specified operations of the second calculation stage of the BN layer, and use the results of the second stage to continue the calculation required by the backpropagation process of layerL.
[0134] In the case where the inter-layer connection is specifically layer_A->BN->(ReLU)->layer_B, where both layer_A and layer_B are computationally intensive layers, this embodiment of the present invention is equivalent to dividing the backpropagation process of the BN layer into two stages. The first computational stage is combined with the execution of the non-computationally intensive layer (ReLU layer) preceding the BN layer and the computationally intensive layer (layer_B) preceding the BN layer, appearing as a post-processing operation in the execution of this layer (layer_B). The post-processing operation here refers to the necessary processing of the output error value of the backpropagation process of layer_B, such as calculating the cumulative sum, calculating the cumulative sum of its product with the input feature map, calculating division, calculating B and C, etc. The specified operations of the second computational stage are combined with the first computationally intensive layer (layer_A) following the BN layer and appearing as a pre-processing operation in the execution of this fused layer. Among them, the pre-processing operation here represents directly reading the output error value obtained by executing the current layer calculation specified by layer_B in the back propagation process and the output feature map obtained by executing the current layer calculation specified by layer_A in the forward propagation process, and performing the operations required for the second calculation stage to obtain the input error value of layer_A (that is, the aforementioned target output).
[0135] It should be noted that when the inter-layer connection is layer_A->ReLU->layer_B, where both layer_A and layer_B are computationally intensive layers. In the embodiment of the present invention, the ReLU layer can be fused with the computationally intensive layer (layer_B) preceding it, and appear as a post-processing operation in the execution of the layer (layer_B), that is, the product of the output error value of layer_B and ReLU'(*) (* represents the output feature map of layer_A) is calculated. Similarly, due to the functional characteristics of ReLU', the sign of the output feature map of layer_A can be used to directly determine whether the result of the product is 0 or the output error value of layer_B, without the need for actual multiplication operations, thereby improving processing efficiency.
[0136] In an embodiment of the present invention, during the backpropagation process, common non-computationally intensive layers in the entire neural network model, such as BN layers and ReLU layers, can be fused with the preceding and following computationally intensive layers according to the above scheme, and appear as pre-processing operations or post-processing operations. The post-processing operation is a calculation of the output error value of the computationally intensive layer, which can be executed immediately while the output error value is still on-chip, without the need for additional memory access operations. The pre-processing operation is equivalent to adding a pre-processing link after executing the input value of the computationally intensive layer, because there is no need to increase the memory access requirements of the layer. Therefore, the problem of repeatedly accessing data from the off-chip memory in the process of executing the non-computationally intensive layer layer by layer can be eliminated, thereby reducing the amount of memory access.
[0137] Furthermore, during the back propagation process, when the BN layer is executed individually, it is necessary to first read in and x i , calculate and And store them off-chip respectively; then read them back in and x i , calculate Then Store back to the off-chip memory. Assuming that the total input value of the mini-batch samples for the BN layer is V, the memory access amount can be recorded as 5V.
[0138] When executing a ReLU layer individually, it first reads in all the ReLU layer's input error values and the input feature map of the ReLU layer's forward pass. It then calculates the ReLU output error values and stores them off-chip. Assuming the total number of ReLU layer input values for a mini-batch of samples is V, the memory access amount can be recorded as 3V.
[0139] In the prior art, when the inter-layer connection is layer_A->BN->layer_B, the memory access amount performed by layer can be expressed as:
[0140] INPUT(layer_B)+OUTPUT(layer_B)+OUTPUT(layer_B)+5*OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_A).
[0141] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0142] INPUT(layer_B)+OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_A). It can be seen that compared with the layer-by-layer execution method, the embodiment of the present invention can reduce the amount of memory access: 4*OUTPUT(layer_B).
[0143] When the connection is layer_A->BN->ReLU->layer_B, the amount of memory access performed by layer can be expressed as:
[0144] INPUT(layer_B)+OUTPUT(layer_B)+OUTPUT(layer_B)+3*OUTPUT(layer_B)+5*OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_A).
[0145] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0146] INPUT(layer_B)+OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_A). It can be seen that compared with the layer-by-layer execution method, the embodiment of the present invention can reduce the memory access amount by 7*OUTPUT(layer_B).
[0147] When the connection is layer_A->ReLU->layer_B, the amount of memory access performed by layer can be expressed as:
[0148] INPUT(layer_B)+OUTPUT(layer_B)+OUTPUT(layer_B)+3*OUTPUT(layer_B)+INPUT(layer_A)+OUTPUT(layer_A).
[0149] The memory access amount after fusion in the embodiment of the present invention can be expressed as:
[0150] INPUT(layer_B)+OUTPUT(layer_B)+INPUT(layer_A)+INPUT(layer_A)+OUTPUT(layer_A). It can be seen that compared with the layer-by-layer execution method, the embodiment of the present invention can reduce the amount of memory access: 3*OUTPUT(layer_B).
[0151] Optionally, in an embodiment of the present invention, the target neural network can be deployed in an accelerator. Specifically, one layer in the target neural network can correspond to one accelerator, and the accelerator corresponding to the fusion layer includes a pre-processing unit and / or a post-processing unit; the pre-processing operation is performed based on the pre-processing unit, and the post-processing operation is performed based on the post-processing unit, that is, the pre-processing unit is used to perform the pre-processing operation, and the post-processing unit is used to perform the post-processing operation. The pre-processing operation can be the first pre-processing operation or the second pre-processing operation mentioned above. The post-processing operation can be the first post-processing operation or the second post-processing operation mentioned above.
[0152] For example, Figure 2 FIG. 1 is a structural diagram of an existing accelerator according to an embodiment of the present invention. Figure 2 As shown, the main components of the accelerator include a computing unit (i.e., the operation unit in the figure), an input cache for storing inputs, a register stack for storing intermediate calculation values (i.e., the operation intermediate value register in the figure), a cache area for storing outputs (i.e., the output cache), and a control unit (i.e., the memory controller and control logic module in the figure).
[0153] Figure 3 is another accelerator structure diagram provided by an embodiment of the present invention, Figure 3 The accelerator shown can be a device that implements the above-mentioned optimization scheme of fusing non-computation-intensive layers with the previous and subsequent computation-intensive layers as "pre-processing" or "post-processing". Figure 2 The same as in the above, is used to perform calculations of the computationally intensive layer, and the further added pre-processing unit can be used to implement the pre-processing operation. The added post-processing unit can be used to continue to perform post-processing operations with the result as the source operand after the operation unit calculates the result of the computationally intensive layer and stores it in the operation intermediate value register. The local components included in the pre-processing unit and the post-processing unit can be set according to the pre-processing operation and post-processing operation to be performed, and the embodiment of the present invention is not limited to this. For example, the pre-processing unit can be composed of a multiplication and addition unit. The post-processing unit can be composed of a multiplication and addition unit and a division and square root unit.
[0154] In the embodiment of the present invention, by setting a pre-processing unit and a post-processing unit, a basis for implementing the pre-processing operations and post-processing operations required to be performed after fusion can be provided, thereby ensuring that the operations of the fusion layer can be performed smoothly.
[0155] Optionally, the following steps may also be performed in an embodiment of the present invention: based on the received preset instructions, determine the current processing process, the pre-processing operation and the post-processing operation; the preset instructions include at least a forward and reverse process indication field, a pre-processing category field and a post-processing category field.
[0156] In an embodiment of the present invention, the preset instructions may be pre-designed based on the above-mentioned model processing method. Specifically, the accelerator may be connected to the central processing unit (CPU), and the CPU may send preset instructions to the accelerator. In actual application scenarios, the preset instructions may include five parts: a layer category field, a forward and reverse process indication field, a pre-processing category field, a post-processing category field, and other information fields. Among them, the layer category field is used to indicate the layers corresponding to the computational operations required to be performed by the fusion layer in the original neural network, that is, to indicate which computationally intensive layers and non-computationally intensive layers correspond to the fusion layer. For example, it may indicate convolutional layers, fully connected layers, pooling layers, BN layers, RELU layers, and so on. The forward and reverse process indication field is used to characterize the current processing process, that is, to indicate whether the current processing process is a forward processing process or a back-propagation process. The pre-processing category field is used to indicate the pre-processing operations required to be performed in the fusion layer. Specifically, the pre-processing category field can be used to indicate whether the currently executed layer has a pre-processing operation, and if so, which specific type it is. The Post-Processing Category field is used to indicate the post-processing operations required for the fusion layer. Specifically, the Post-Processing Category field can be used to indicate whether the currently executed layer has any post-processing operations, and if so, what type of post-processing operations they are. The Other Information field is used to indicate other information required for the fusion layer calculation. Specifically, the Other Information field can contain other information required for executing this layer, such as the address of the input data, the address of the output data, the length of the input data, the length of the output data, etc.
[0157] In the embodiment of the present invention, preset instructions are pre-designed to characterize the current processing process, pre-processing operations and post-processing operations based on the preset instructions, thereby providing a basis for the calculation process of the fusion layer and ensuring that the fusion layer can be executed normally.
[0158] Optionally, in an embodiment of the present invention, the following operations may also be performed: writing the calculated operation results of the partial first post-processing operation into the first register division, and performing the current layer calculation of the computationally intensive layer based on the second register division to obtain new first data, and calculating the remaining first intermediate data based on the new first data. Specifically, for a fusion layer with a post-processing operation, the computational intermediate value register may be set to a ping-pong structure, with one part being used to store the partial results being calculated to support the normal computational process of the computationally intensive layer. The other part stores the final result that has been calculated, thereby supporting the post-processing operation. In this way, the execution time of the post-processing operation can be hidden in the computational process of the computationally intensive layer through ping-pong storage, thereby eliminating the time overhead of the post-processing operation.
[0159] Optionally, the following operations may also be performed in an embodiment of the present invention: partially storing the results of some pre-processing operations based on the third register, and calculating the remaining target outputs based on the data subsequently read in by the second fusion layer based on the fourth register distribution. Specifically, the data read out of the input cache may be stored in an input data register stack using a ping-pong structure, thereby using one part of the input data registers to store the results of the pre-processing operations for subsequent calculations, while reading the input values to be used later into another part of the input data registers for pre-processing operations. In this way, the execution time of the pre-processing operations can be hidden in the calculation process of the computationally intensive layer through ping-pong storage, thereby eliminating the time overhead of the pre-processing operations.
[0160] It should be noted that, in the embodiment of the present invention, the processing results of the pre-processing operation can be stored in the on-chip cache, for example, the first intermediate data can be stored in the on-chip cache. In the subsequent processing process, the first intermediate data in the on-chip cache can be reused, thereby improving the processing efficiency. For example, for the fusion layer with "pre-processing", the result after "pre-processing" of the input data can be stored in the on-chip cache area, for example, stored back in the input cache. In the subsequent execution process of this fusion layer, the "pre-processed" value stored in the on-chip cache area can be directly used as input for subsequent calculations. In this way, the computational time overhead for "pre-processing" only occurs when the input value is processed for the first time, which can further reduce time consumption and improve processing efficiency.
[0161] Since the BN layer needs to calculate the mean and variance of the entire mini-batch during training, the method of starting the processing of the next layer immediately after calculating a value cannot be adapted to the training process. In an embodiment of the present invention, the processing is continued only after the calculation of the mini-batch samples on the same channel in the CONV layer is completed, which can ensure that it can adapt to the forward processing process in the training process and the reasoning process. At the same time, in an embodiment of the present invention, during the training process, the result of the CONV layer will be used in the back propagation stage. Therefore, the first intermediate data and the first data must be stored off-chip, which can ensure that the required data can continue to be used in the back propagation stage, for example, using the F1 obtained in the forward processing process, thereby ensuring that the training process can proceed smoothly.
[0162] Furthermore, the use of dedicated on-chip cache and dedicated arithmetic units for the BN layer to perform calculations, thereby accelerating processing, is limited by the capacity of the on-chip cache and cannot accommodate all input data. Therefore, it cannot truly accelerate the BN layer. Increasing the size of the on-chip cache to accommodate all the data required to pass between layers during training, thereby eliminating memory access operations for non-computationally intensive layers, is also not feasible due to hardware and capacity limitations.
[0163] In an embodiment of the present invention, the non-computational intensive layer and the computational intensive layer are integrated, and the parallelism of the BN layer and the computational intensive layer is developed, so that their parallelism can be utilized to adopt a corresponding structure to improve performance, eliminate a large number of off-chip memory access operations of the non-computational intensive layer, reduce training time, reduce pressure on memory bandwidth, improve execution performance, and reduce energy consumption during operation.
[0164] Furthermore, embodiments of the present invention can reduce the amount of data required to be stored in off-chip memory during training, so that within the same memory capacity constraints, larger models can be trained or larger mini-batches can be used for training, thereby increasing the flexibility and scalability of training.
[0165] Figure 4 is a flowchart of the steps of another model processing method provided by an embodiment of the present invention. The method can be applied to a target neural network, wherein the target neural network includes a fusion layer, and the fusion layer is used to implement the computational operations required to be performed by the computation-intensive layer and at least one non-computation-intensive layer in the original neural network. The method includes:
[0166] Step 201: In response to a received preset instruction, determine the pre-processing operation and / or post-processing operation to be performed corresponding to at least one fusion layer in the target neural network according to the preset instruction.
[0167] In the embodiment of the present invention, the pre-processing operation may include the first pre-processing operation and the second pre-processing operation mentioned above. The post-processing operation may include the first post-processing operation and the second post-processing operation mentioned above.
[0168] Step 202: The at least one fusion layer performs the pre-processing operation and / or post-processing operation based on the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer to calculate the target output of the current layer calculation of the non-computationally intensive layer corresponding to the at least one fusion layer.
[0169] In an embodiment of the present invention, the fusion layer can perform pre-processing operations and / or post-processing operations based on the input of the corresponding computationally intensive layer and / or the computational output of the computationally intensive layer, thereby obtaining the target output of the corresponding non-computationally intensive layer's computation.
[0170] Step 203: Perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0171] It should be noted that the implementation method of each step can refer to the above-mentioned related description and will not be repeated here.
[0172] In summary, the model processing method provided by the embodiment of the present invention responds to the received preset instructions and determines the pre-processing operation and / or post-processing operation to be performed on at least one fusion layer in the target neural network according to the preset instructions. At least one fusion layer performs pre-processing operations and / or post-processing operations based on the input of the computationally intensive layer and / or the computational output of the computationally intensive layer corresponding to the at least one fusion layer to calculate the target output calculated by the non-computationally intensive layer corresponding to the at least one fusion layer. Subsequent processing is performed based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer. In this way, by fusing the non-computationally intensive layer that takes a long time to access data from the memory with the computationally intensive layer, the problem of repeatedly accessing data from the memory during the execution of the non-computationally intensive layer layer by layer can be avoided to a certain extent, thereby reducing the amount of memory access, reducing the time consumption of the processing process, and improving processing performance.
[0173] Optionally, the following operations may also be performed in an embodiment of the present invention: Step S71, determining the current processing process according to the preset instruction; the current processing process includes a forward processing process and a backpropagation process. Accordingly, the above-mentioned execution of the pre-processing operation and / or post-processing operation based on the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer may specifically include: executing the pre-processing operation and / or post-processing operation based on the current processing process, the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer.
[0174] Specifically, this step may include, after executing the first data obtained by the current layer calculation of the first computationally intensive layer corresponding to the first fusion layer, performing a first post-processing operation according to this processing process, wherein the first post-processing operation includes calculating the first intermediate data of the non-computationally intensive layer based on the first data; the first intermediate data is the intermediate data required for calculating the target output; the first intermediate data and the first data are stored in the memory; and the second fusion layer performs a first pre-processing operation, wherein the first pre-processing operation includes calculating the target output based on the first intermediate data and the first data read from the memory. Accordingly, the subsequent processing based on the target output may include: performing the current layer calculation of the second computationally intensive layer corresponding to the second fusion layer according to the target output.
[0175] Furthermore, when the first computationally intensive layer is a convolutional layer and the non-computationally intensive layer is a batch normalization layer, the above-mentioned processing process performs a first post-processing operation, which may specifically be: if the current processing process is a forward processing process, based on the output feature map of the same channel obtained by the convolutional layer, calculate the first parameter and the second parameter as the first intermediate data; the output feature map is the first data. Accordingly, the calculation of the target output based on the first intermediate data read from the memory and the first data includes: performing a multiplication and addition operation based on the first parameter, the output feature map, and the second parameter to obtain the target output.
[0176] Alternatively, the above-mentioned processing procedure performs the operation of the first post-processing operation, which may specifically be: if the current processing procedure is a back-propagation process, then based on the error value obtained by the convolution layer calculation, the third parameter and the fourth parameter are calculated as the first intermediate data. Accordingly, the calculation of the target output based on the first intermediate data and the first data read from the memory may include: performing a multiplication and addition operation based on the third parameter, the fourth parameter and the error value to obtain the target output. It should be noted that the implementation method of each step and the technical effect that can be achieved can refer to the above-mentioned relevant description and will not be repeated here.
[0177] Optionally, the above-mentioned determination of the pre-processing operation and / or post-processing operation to be performed corresponding to at least one fusion layer in the target neural network according to the preset instruction may specifically include: determining the pre-processing operation and / or post-processing operation to be performed corresponding to the fusion layer according to the pre-processing category domain and the post-processing category domain in the preset instruction; the pre-processing category domain and the post-processing category domain are respectively used to indicate the pre-processing operation and post-processing operation to be performed corresponding to the fusion layer. Further, the above-mentioned determination of the operation of the current processing process according to the preset instruction may specifically include: determining the current processing process according to the forward and reverse process indication domain in the preset instruction; the forward and reverse process indication domain is used to characterize the current processing process. In this way, the current processing process, pre-processing operation and / or post-processing operation can be conveniently determined through the different information domains contained in the preset instruction, thereby ensuring processing efficiency to a certain extent.
[0178] Specifically, the preset instructions can be parsed to determine the specific content of each information field in the preset instructions. According to the specific content of the pre-processing category field and the post-processing category field, the pre-processing operation and / or post-processing operation are determined. According to the specific content of the forward and reverse process indication field, the current processing process is determined. For example, different identifiers can be set for different pre-processing operations, different identifiers can be set for different post-processing operations, and different identifiers can be set for different processing processes. According to the identifier in the pre-processing category field, the pre-processing operation corresponding to the identifier is determined, according to the identifier in the post-processing category field, the post-processing operation corresponding to the identifier is determined, and according to the identifier in the forward and reverse process indication, the current processing process corresponding to the identifier is determined.
[0179] Figure 5 This is a structural block diagram of a model processing device provided by an embodiment of the present invention. The device is applied to a target neural network, wherein the target neural network includes a fusion layer, and the fusion layer is used to implement the computational operations required to be performed by the computation-intensive layer and at least one non-computation-intensive layer in the original neural network. The device includes:
[0180] A calculation module 301 is configured to calculate, for at least one fusion layer in the target neural network, a target output of a non-computational intensive layer corresponding to the at least one fusion layer based on an input of the computational intensive layer corresponding to the at least one fusion layer and / or a computational output of the computational intensive layer;
[0181] The processing module 302 is used to perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0182] In summary, the model processing device provided by the embodiment of the present invention is based on the fusion layer in the target neural network for realizing the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and calculates the target output of the current layer calculation of the non-computationally intensive layer corresponding to the fusion layer according to the input of the computationally intensive layer corresponding to the fusion layer and / or the computational output of the computationally intensive layer. Subsequent processing is performed based on the target output, and the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer. In this way, by fusing the non-computationally intensive layer that takes a long time to access data from the memory with the computationally intensive layer, the problem of repeatedly accessing data from the memory in the process of executing the non-computationally intensive layer layer by layer can be avoided to a certain extent, thereby reducing the amount of memory access, reducing the time consumption of the processing process, and improving processing performance.
[0183] Optionally, the calculation module 301 is specifically configured to:
[0184] After executing the first computation-intensive layer corresponding to the first fusion layer to obtain first data, performing a first post-processing operation, wherein the first post-processing operation includes calculating first intermediate data of the non-computation-intensive layer based on the first data; the first intermediate data is intermediate data required for calculating the target output;
[0185] storing the first intermediate data and the first data in a memory;
[0186] performing, by the second fusion layer, a first pre-processing operation, wherein the first pre-processing operation includes calculating the target output based on the first intermediate data read from the memory and the first data;
[0187] The performing subsequent processing based on the target output includes: performing current layer calculation of the second computationally intensive layer corresponding to the second fusion layer according to the target output.
[0188] Optionally, the calculation module 301 is specifically configured to:
[0189] After executing the current layer calculation of the third computationally intensive layer corresponding to the third fusion layer to obtain second data, performing a second post-processing operation, wherein the second post-processing operation includes calculating the target output based on the second data; and performing subsequent processing based on the target output includes writing the target output to a memory for use by a subsequent layer in the target neural network;
[0190] Alternatively, the fourth fusion layer reads the current layer input from the memory and performs a second pre-processing operation, wherein the second pre-processing operation includes calculating the target output based on the current layer input; the current layer input is the original input corresponding to the non-computational intensive layer; and the subsequent processing based on the target output includes: executing the current layer calculation of the fourth computational intensive layer corresponding to the fourth fusion layer according to the target output.
[0191] Optionally, the first computationally intensive layer is a convolutional layer, and the non-computationally intensive layer is a BN layer; when this processing process is a forward processing process, the computing module 201 is further specifically configured to:
[0192] Calculating a first parameter and a second parameter based on an output feature map of the same channel obtained by the convolution layer to serve as the first intermediate data; the output feature map is the first data;
[0193] The calculating the target output based on the first intermediate data and the first data read from the memory includes: performing a multiplication and addition operation based on the first parameter, the output feature map and the second parameter to obtain the target output.
[0194] Optionally, the non-computational intensive layer also includes a RELU layer; the computing module 201 is further specifically used to: perform a RELU operation on the target output, and perform the current layer calculation of the second computational intensive layer based on the result of the RELU operation.
[0195] Optionally, the first computationally intensive layer is a convolutional layer, and the non-computationally intensive layer is a BN layer; when the current processing process is a back propagation process, the computing module 201 is further specifically configured to:
[0196] Calculating a third parameter and a fourth parameter based on the error value obtained by the convolutional layer calculation to serve as the first intermediate data;
[0197] The calculating the target output based on the first intermediate data and the first data read from the memory includes: performing a multiplication and addition operation based on the third parameter, the fourth parameter and the error value to obtain the target output.
[0198] Optionally, the accelerator corresponding to the fusion layer includes a pre-processing unit and / or a post-processing unit; the pre-processing operation is performed based on the pre-processing unit, and the post-processing operation is performed based on the post-processing unit.
[0199] Optionally, the device further includes:
[0200] The instruction fetch and decode module is configured to determine the current processing procedure, the pre-processing operation, and the post-processing operation based on a received preset instruction; the preset instruction includes at least a forward and reverse process indication field, a pre-processing category field, and a post-processing category field; the forward and reverse process indication field is used to characterize the current processing procedure, and the pre-processing category field and the post-processing category field are used to indicate the pre-processing operation and post-processing operation to be performed in the fusion layer, respectively. Instruction fetch and decode can refer to obtaining the preset instruction and parsing it to determine the current processing procedure, pre-processing operation, and post-processing operation.
[0201] Figure 6 is a structural block diagram of another model processing device provided by an embodiment of the present invention. The device is applied to a target neural network, wherein the target neural network includes a fusion layer, and the fusion layer is used to implement the computational operations required to be performed by the computation-intensive layer and at least one non-computation-intensive layer in the original neural network. The device includes:
[0202] A first determining module 401 is configured to determine, in response to a received preset instruction, a pre-processing operation and / or a post-processing operation to be performed corresponding to at least one fusion layer in the target neural network according to the preset instruction;
[0203] An execution module 402 is configured to execute, by the at least one fusion layer, the pre-processing operation and / or the post-processing operation based on the input of the computation-intensive layer corresponding to the at least one fusion layer and / or the computation output of the computation-intensive layer, so as to calculate a target output of the current layer of the non-computation-intensive layer corresponding to the at least one fusion layer;
[0204] The processing module 403 is used to perform subsequent processing based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
[0205] The model processing device provided by an embodiment of the present invention responds to a received preset instruction and determines the pre-processing operation and / or post-processing operation to be performed on at least one fusion layer in the target neural network according to the preset instruction. At least one fusion layer performs pre-processing operation and / or post-processing operation based on the input of the computationally intensive layer and / or the computational output of the computationally intensive layer corresponding to the at least one fusion layer to calculate the target output calculated by the non-computationally intensive layer corresponding to the at least one fusion layer. Subsequent processing is performed based on the target output; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer. In this way, by fusing the non-computationally intensive layer that takes a long time to access data from the memory with the computationally intensive layer, the problem of repeatedly accessing data from the memory during the execution of the non-computationally intensive layer layer by layer can be avoided to a certain extent, thereby reducing the amount of memory access, reducing the time consumption of the processing process, and improving processing performance.
[0206] Optionally, the device further includes: a second determining module, configured to determine a current processing process according to the preset instruction; the current processing process includes a forward processing process and a back-propagation process;
[0207] The execution module 402 is specifically configured to:
[0208] The pre-processing operation and / or post-processing operation is performed according to the current processing process, the input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computational output of the computationally intensive layer.
[0209] Optionally, the first determining module 401 is specifically configured to:
[0210] According to the pre-processing category field and the post-processing category field in the preset instruction, the pre-processing operation and / or post-processing operation required to be performed by the fusion layer is determined; wherein the pre-processing category field and the post-processing category field are respectively used to indicate the pre-processing operation and post-processing operation required to be performed by the fusion layer.
[0211] Optionally, the second determining module is specifically configured to
[0212] The current processing process is determined according to the forward and reverse process indication fields in the preset instruction; wherein the forward and reverse process indication fields are used to represent the current processing process.
[0213] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0214] Preferably, an embodiment of the present invention also provides an electronic device, comprising one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enables the electronic device to execute the model processing method provided in the above embodiment.
[0215] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various processes of the model processing method provided in the above embodiment and can achieve the same technical effect. To avoid repetition, the details are not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0216] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0217] It is easy for those skilled in the art to think that any combination of the above embodiments is feasible, so any combination of the above embodiments is an implementation scheme of the present invention. However, due to space limitations, this specification will not describe them in detail here.
[0218] The method provided herein is not inherently relevant to any particular computer, virtual system or other equipment. Various general-purpose systems may also be used together with the teachings based thereon. According to the above description, it is apparent that the structure required for the system having the scheme of the present invention is constructed. In addition, the present invention is not directed to any specific programming language either. It should be understood that various programming languages can be utilized to realize the content of the present invention described herein, and the above description of specific languages is for the purpose of disclosing the best mode of the present invention.
[0219] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0220] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the claims, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Accordingly, the claims that follow the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0221] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition may be divided into multiple submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed herein may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0222] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of the present invention and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0223] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the model processing method according to an embodiment of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing a part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0224] It should be noted that the above embodiments illustrate rather than limit the invention, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
Claims
1. A model processing method, characterized in that: Applied to a target neural network, specifically in the field of neural network acceleration, the target neural network is deployed in an accelerator, the target neural network includes a fusion layer, and the accelerator corresponding to the fusion layer includes an arithmetic unit, an input buffer, an intermediate value register, an output buffer, a control unit, a pre-processing unit, and a post-processing unit; The fusion layer is used to implement the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and the method includes: For at least one fusion layer in the target neural network, the at least one fusion layer executes the computation of the computation-intensive layer through the operation unit, performs a pre-processing operation on the data in the input buffer through the pre-processing unit and / or performs a post-processing operation on the data in the intermediate value register through the post-processing unit, calculates a target output result of the current layer calculation of the non-computation-intensive layer corresponding to the at least one fusion layer according to the memory access data input of the computation-intensive layer corresponding to the at least one fusion layer and / or the calculation output result of the computation-intensive layer, and outputs the target output result through the output buffer; Perform subsequent processing based on the target output result; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
2. The method according to claim 1, characterized in that The calculating, based on the memory access data input of the computation-intensive layer corresponding to the at least one fused layer and / or the computation output result of the computation-intensive layer, a target output result of the computation of the non-computation-intensive layer corresponding to the at least one fused layer, includes: After executing the first computation-intensive layer corresponding to the first fusion layer to obtain first data, performing a first post-processing operation, wherein the first post-processing operation includes calculating first intermediate data of the non-computation-intensive layer based on the first data; the first intermediate data is intermediate data required for calculating the target output result; storing the first intermediate data and the first data in a memory; performing, by the second fusion layer, a first pre-processing operation, wherein the first pre-processing operation includes calculating the target output result based on the first intermediate data and the first data read from the memory; The performing subsequent processing based on the target output result includes: performing current layer calculation of the second computationally intensive layer corresponding to the second fusion layer according to the target output result.
3. The method according to claim 1, characterized in that The calculating, based on the memory access data input of the computation-intensive layer corresponding to the at least one fusion layer and / or the computation output result of the computation-intensive layer, a target output result of the computation of the at least one corresponding non-computation-intensive layer comprises: After executing the current layer calculation of the third computationally intensive layer corresponding to the third fusion layer to obtain second data, performing a second post-processing operation, wherein the second post-processing operation includes calculating the target output result based on the second data; performing subsequent processing based on the target output result includes: writing the target output result into a memory for use by a subsequent layer in the target neural network; Alternatively, the fourth fusion layer reads the current layer input from the memory and performs a second pre-processing operation, wherein the second pre-processing operation includes calculating the target output result based on the current layer input; the current layer input is the original input corresponding to the non-computational intensive layer; and the subsequent processing based on the target output result includes: executing the current layer calculation of the fourth computational intensive layer corresponding to the fourth fusion layer according to the target output result.
4. The method according to claim 2, characterized in that The first computationally intensive layer is a convolutional layer, and the non-computationally intensive layer is a batch normalization layer. When the current processing is a forward processing process, calculating the first intermediate data of the non-computationally intensive layer based on the first data includes: Calculating a first parameter and a second parameter based on an output feature map of the same channel obtained by the convolution layer to serve as the first intermediate data; the output feature map is the first data; The calculating the target output result based on the first intermediate data and the first data read from the memory includes: performing a multiplication and addition operation based on the first parameter, the output feature map and the second parameter to obtain the target output result.
5. The method according to claim 4, characterized in that The non-computational intensive layer also includes a RELU layer; and performing the current-layer calculation of the second computational intensive layer corresponding to the second fusion layer according to the target output result includes: performing a RELU operation on the target output result, and performing the current-layer calculation of the second computational intensive layer based on the result of the RELU operation.
6. The method according to claim 2, characterized in that The first computationally intensive layer is a convolutional layer, and the non-computationally intensive layer is a batch normalization layer. When the current processing is a back propagation process, the first intermediate data of the non-computationally intensive layer is calculated based on the first data, including: Calculating a third parameter and a fourth parameter based on the error value obtained by the convolutional layer calculation to serve as the first intermediate data; The calculating the target output result based on the first intermediate data and the first data read from the memory includes: performing a multiplication and addition operation based on the third parameter, the fourth parameter and the error value to obtain the target output result.
7. The method according to claim 1, characterized in that The method further comprises: Based on the received preset instructions, the current processing process, the pre-processing operation and the post-processing operation are determined; the preset instructions include at least a forward and reverse process indication field, a pre-processing category field and a post-processing category field; the forward and reverse process indication field is used to characterize the current processing process, and the pre-processing category field and the post-processing category field are respectively used to indicate the pre-processing operation and the post-processing operation required to be performed in the fusion layer.
8. A model processing method, characterized in that: Applied to a target neural network, specifically in the field of neural network acceleration, the target neural network is deployed in an accelerator, the target neural network includes a fusion layer, and the accelerator corresponding to the fusion layer includes an arithmetic unit, an input buffer, an intermediate value register, an output buffer, a control unit, a pre-processing unit, and a post-processing unit; The fusion layer is used to implement the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and the method includes: In response to the received preset instruction, determining, according to the preset instruction, a pre-processing operation and / or a post-processing operation to be performed corresponding to at least one fusion layer in the target neural network; The at least one fusion layer performs the pre-processing operation and / or the post-processing operation based on the memory access data input of the computation-intensive layer corresponding to the at least one fusion layer and / or the computation output result of the computation-intensive layer, so as to calculate the target output result of the current layer computation of the non-computation-intensive layer corresponding to the at least one fusion layer; Perform subsequent processing based on the target output result; the computationally intensive layer corresponding to the fusion layer includes the computationally intensive layer before and / or after the non-computationally intensive layer corresponding to the fusion layer.
9. The method according to claim 8, characterized in that The method further includes: determining a current processing process according to the preset instruction; the current processing process includes a forward processing process and a back propagation process; The performing of the pre-processing operation and / or the post-processing operation according to the memory access data input of the computation-intensive layer corresponding to the at least one fusion layer and / or the calculation output result of the computation-intensive layer includes: The pre-processing operation and / or post-processing operation is performed according to the current processing process, the memory access data input of the computationally intensive layer corresponding to the at least one fusion layer and / or the calculation output result of the computationally intensive layer.
10. The method according to claim 8 or 9, characterized in that The determining, according to the preset instruction, the pre-processing operation and / or post-processing operation to be performed corresponding to at least one fusion layer in the target neural network includes: According to the pre-processing category field and the post-processing category field in the preset instruction, the pre-processing operation and / or post-processing operation required to be performed by the fusion layer is determined; wherein the pre-processing category field and the post-processing category field are respectively used to indicate the pre-processing operation and post-processing operation required to be performed by the fusion layer.
11. The method according to claim 9, characterized in that The step of determining the current processing according to the preset instruction includes: The current processing process is determined according to the forward and reverse process indication fields in the preset instruction; wherein the forward and reverse process indication fields are used to represent the current processing process.
12. A model processing device, characterized in that: Applied to a target neural network, specifically in the field of neural network acceleration, the target neural network is deployed in an accelerator, the target neural network includes a fusion layer, and the accelerator corresponding to the fusion layer includes an arithmetic unit, an input buffer, an intermediate value register, an output buffer, a control unit, a pre-processing unit, and a post-processing unit; The fusion layer is used to implement the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and the device includes: a computing module configured to, for at least one fusion layer in the target neural network, execute the computation-intensive layer computation by the at least one fusion layer through the computing unit, perform pre-processing operations on the data in the input buffer through the pre-processing unit and / or perform post-processing operations on the data in the intermediate value register through the post-processing unit, calculate a target output result of the current layer computation of the non-computation-intensive layer corresponding to the at least one fusion layer based on the memory access data input of the computation-intensive layer corresponding to the at least one fusion layer and / or the computation output result of the computation-intensive layer, and output the target output result through the output buffer; A processing module is used to perform subsequent processing based on the target output result; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
13. A model processing device, characterized in that: Applied to a target neural network, specifically in the field of neural network acceleration, the target neural network is deployed in an accelerator, the target neural network includes a fusion layer, and the accelerator corresponding to the fusion layer includes an arithmetic unit, an input buffer, an intermediate value register, an output buffer, a control unit, a pre-processing unit, and a post-processing unit; The fusion layer is used to implement the computational operations required to be performed by the computationally intensive layer and at least one non-computationally intensive layer in the original neural network, and the device includes: A first determining module is configured to determine, in response to a received preset instruction, a pre-processing operation and / or a post-processing operation to be performed corresponding to at least one fusion layer in the target neural network according to the preset instruction; an execution module, configured to execute, by the at least one fusion layer, the pre-processing operation and / or the post-processing operation based on the memory access data input of the computationally intensive layer corresponding to the at least one fusion layer and / or the computation output result of the computationally intensive layer, so as to calculate a target output result calculated by the non-computationally intensive layer corresponding to the at least one fusion layer; A processing module is used to perform subsequent processing based on the target output result; the computationally intensive layer corresponding to the fusion layer includes a computationally intensive layer located before and / or after the non-computationally intensive layer corresponding to the fusion layer.
14. An electronic device, characterized in that: include: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform the method according to any one of claims 1 to 11.
15. One or more machine-readable media, characterized in that Instructions are stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method for realizing single-broadcast multi-operation based on deep learning accelerator
CN108960414A
Neural network calculation graph optimization method
CN110321999A