Calculation acceleration method and device of feedforward neural network and main control board card

By blocking the input data and weight data of the feedforward neural network and reasonably allocating it on a multi-core processor, the problem of long execution time of the feedforward neural network is solved, and the effect of improving the computing speed is achieved.

CN120031085APending Publication Date: 2025-05-23太初(无锡)电子科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411851318.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The execution of feedforward neural networks consumes a lot of network execution time, resulting in slower computing speed.

Method used

By blocking the network input data and model weight data, and allocating the data to multiple computing cores in turn, the low-speed memory area is only needed when reading the data chunking, the starting input and the final output data, and the intermediate execution data is always stored in the high-speed memory area of ​​multi-core.

Benefits of technology

It effectively reduces the execution time of the feedforward neural network, improves the computing speed, and avoids the problem of insufficient high-speed storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031085A_ABST
    Figure CN120031085A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing, and discloses a computing acceleration method and device for a feedforward neural network and a main control board card, and the method comprises the steps: carrying out the blocking processing of network input data and model weight data of a first linear transformation layer, obtaining data blocks, and sequentially distributing the data blocks to a plurality of computing cores; reading the data blocks from the low-speed storage area by utilizing the computing core, performing linear transformation on the data blocks to obtain output data of a first linear transformation layer, and storing the output data of the first linear transformation layer to a high-speed storage area of the computing core; performing nonlinear transformation on the output data of the first linear transformation layer to obtain output data of a nonlinear transformation layer; and performing block processing on the output data of the nonlinear transformation layer, performing linear transformation on the output data after block processing to obtain a feedforward neural network operation result, and storing the feedforward neural network operation result in a low-speed storage area. According to the invention, the execution time of the feedforward neural network is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular to a method, device and main control board for accelerating the calculation of a feedforward neural network. Background Art

[0002] Feedforward Neural Network (FNN) is a basic neural network architecture that is widely used in various machine learning tasks, such as classification, regression, and feature learning. In a feedforward neural network, input data is passed from the input layer to the first hidden layer. The neurons in each hidden layer perform linear transformation (weight matrix multiplication) and nonlinear transformation (activation function) on the input data. The transformed result is passed to the next hidden layer until it reaches the output layer. Through forward propagation and back propagation of multiple layers of neurons, complex nonlinear relationships can be learned.

[0003] The execution of feedforward neural networks requires a large amount of storage resources and computing resources, resulting in a longer network execution time. Summary of the invention

[0004] In view of this, the present invention provides a method, device and main control board for accelerating the calculation of a feedforward neural network to solve the problem that the execution of the feedforward neural network consumes a lot of network execution time.

[0005] In a first aspect, the present invention provides a method for accelerating the calculation of a feedforward neural network, wherein the feedforward neural network includes a first linear transformation layer, a nonlinear transformation layer, and a second linear transformation layer; the method includes:

[0006] Obtain network input data and model weight data corresponding to the first linear transformation layer, perform block processing on the network input data and the model weight data respectively to obtain data blocks, and sequentially distribute the data blocks to multiple computing cores;

[0007] Using the computing core to read the data blocks from the low-speed storage area, performing linear transformation on the data blocks to obtain output data of the first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core;

[0008] Performing nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer;

[0009] The output data of the nonlinear transformation layer is processed in blocks, and the output data after block processing is linearly transformed to obtain the feedforward neural network operation result, and the feedforward neural network operation result is stored in the low-speed storage area.

[0010] The present embodiment provides a method for accelerating the calculation of a feedforward neural network, which performs block processing on network input data and model weight data respectively to obtain data blocks, uses a computing core to read the data blocks from a low-speed storage area, performs linear transformation on the data blocks to obtain output data of a first linear transformation layer, and stores the output data of the first linear transformation layer in a high-speed storage area of ​​the computing core, performs nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer; performs block processing on the output data of the nonlinear transformation layer, performs linear transformation on the block-processed output data to obtain a feedforward neural network operation result, and stores the feedforward neural network operation result in a low-speed storage area; for two linear transformations and one nonlinear transformation in the feedforward neural network, by reasonably allocating the network input data and the model weight data on a multi-core processor, the low-speed storage area only needs to be accessed when reading the data blocks, the initial input and the final output data of the feedforward neural network, and the intermediate execution data of the feedforward neural network is always stored in the high-speed storage area of ​​the multi-core, which effectively reduces the execution time of the feedforward neural network and improves the calculation speed of the feedforward neural network.

[0011] In an optional implementation, the network input data and the model weight data are processed in blocks to obtain data blocks, and the data blocks are sequentially allocated to multiple computing cores, including:

[0012] Divide the network input data into blocks according to the matrix dimension to obtain input data blocks;

[0013] Divide the model weight data into blocks according to the weight dimension to obtain weight data blocks;

[0014] The input data blocks and weight data blocks are used as data blocks, and the data blocks are allocated to multiple computing cores in sequence.

[0015] This embodiment provides a method for accelerating the calculation of a feedforward neural network. The method divides the network input data and model weight data into blocks, and distributes the data blocks to multiple computing cores in turn. Each computing core only reads the data blocks allocated to itself, thereby avoiding insufficient high-speed storage space and improving the memory access speed of the computing core.

[0016] In an optional implementation, the model weight data is processed in blocks according to the weight dimension to obtain weight data blocks, including:

[0017] Processing the model weight data in blocks according to the first weight dimension to obtain a plurality of initial data blocks;

[0018] The multiple initial data blocks are divided into blocks according to the second weight dimension to obtain weighted data blocks.

[0019] A method for accelerating the calculation of a feedforward neural network provided in this embodiment divides the model weight data into blocks according to the first weight dimension and the second weight dimension respectively due to the large amount of data of the model weight data, ensuring that each computing core can read the weight data block into the high-speed storage area.

[0020] In an alternative embodiment, using a computing core to read a data block from a low-speed storage area, performing a linear transformation on the data block to obtain the output data of the first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of the computing core includes:

[0021] Using the computing core to iteratively read and inter-core broadcast the input data block to obtain an input data matrix; wherein, the input data matrix is the matrix corresponding to the network input data;

[0022] Performing a blocked matrix multiplication operation and an accumulation reduction process on the input data matrix and the weight data block to obtain the output data of the first linear transformation layer;

[0023] Storing the output data of the first linear transformation layer in the high-speed storage area of the computing core.

[0024] A method for accelerating the calculation of a feedforward neural network provided in this embodiment uses a computing core to iteratively read and inter-core broadcast the input data block, performs a blocked matrix multiplication operation and an accumulation reduction process on the input data matrix and the weight data block to obtain the output data of the first linear transformation layer, realizes the linear transformation of the input data block, and stores the output data of the first linear transformation layer in the high-speed storage area of the computing core, ensuring that the execution of the feedforward neural network is all in the high-speed storage space and improving the calculation speed of the feedforward neural network.

[0025] In an alternative embodiment, the using the computing core to iteratively read and inter-core broadcast the input data block to obtain a data matrix includes:

[0026] Using the computing core to iteratively read the input data block from the low-speed storage area and sequentially broadcast the input data block to other computing cores until the iteration loop count is reached and the broadcast stops; wherein, the iteration loop count is the total number of sub-blocks of the input data block in the computing core;

[0027] Receiving the input data blocks broadcast by the other computing cores and storing the input data blocks broadcast by the other computing cores in the broadcast data storage area;

[0028] Determining the data matrix based on the input data blocks in the broadcast data storage area.

[0029] The present embodiment provides a method for accelerating the calculation of a feedforward neural network. Since the input data blocks are discretely stored in each computing core, the computing core is used to iteratively read the input data blocks from a low-speed storage area, and the input data blocks are broadcasted to other computing cores in sequence, and the input data blocks broadcasted by the other computing cores are stored in a broadcast data storage area; the data matrix is ​​determined based on the input data blocks in the broadcast data storage area, thereby realizing iterative reading and inter-core broadcasting of the input data blocks, realizing complete acquisition of the input data blocks, and laying a foundation for subsequent block matrix multiplication operations.

[0030] In an optional implementation, performing a nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain the output data of the nonlinear transformation layer includes:

[0031] An activation function operation is performed on the output data of the first linear transformation layer to obtain output data of the nonlinear transformation layer.

[0032] In an optional implementation, the output data of the nonlinear transformation layer is processed in blocks, the output data after the block processing is linearly transformed to obtain the feedforward neural network operation result, and the feedforward neural network operation result is stored in a low-speed storage area, including:

[0033] Using the output data of the nonlinear transformation layer as the input data of the second linear transformation layer, decomposing the input data of the second linear transformation layer to obtain an input data matrix and a weight data matrix;

[0034] The input data matrix and the weight data matrix are processed in blocks respectively to obtain a first matrix block and a second matrix block;

[0035] Performing linear transformation on the first matrix block and the second matrix block to obtain output data of the second linear transformation layer;

[0036] The output data of the second linear transformation layer is used as the feedforward neural network operation result, and the feedforward neural network operation result is written out to the low-speed storage area using the computing core.

[0037] The present embodiment provides a method for accelerating the calculation of a feedforward neural network. By reasonably dividing the input data of the second linear transformation layer into blocks and effectively utilizing the high-speed storage area on the multi-core processor, when the feedforward neural network is executed, there is no need to temporarily store and transfer the result data of the intermediate layer in the low-speed storage, thereby accelerating the execution of the feedforward neural network.

[0038] In a second aspect, the present invention provides a computing acceleration device for a feedforward neural network, the device comprising:

[0039] A block module is used to obtain the network input data and model weight data corresponding to the first linear transformation layer, respectively block the network input data and the model weight data to obtain data blocks, and sequentially distribute the data blocks to multiple computing cores;

[0040] A first linear transformation module, used for reading data blocks from the low-speed storage area using the computing core, performing linear transformation on the data blocks to obtain output data of the first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core;

[0041] A nonlinear transformation module, used for performing nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer;

[0042] The second linear transformation module is used to perform block processing on the output data of the nonlinear transformation layer, perform linear transformation on the block-processed output data to obtain the feedforward neural network operation result, and store the feedforward neural network operation result in the low-speed storage area.

[0043] In a third aspect, the present invention provides a main control board, comprising: a multi-core processor, a high-speed storage device and a low-speed storage device, wherein the multi-core processor is used to execute the computational acceleration method of the feedforward neural network of the above-mentioned first aspect or any corresponding embodiment thereof.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the computational acceleration method for a feedforward neural network according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0046] Figure 1 is a schematic diagram of a multi-layer memory access structure of a modern processor according to an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of a feedforward neural network performing low-speed storage access according to an embodiment of the present invention;

[0048] Figure 3 is a schematic flow chart of a method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention;

[0049] Figure 4 is a schematic diagram of storage access performed by a feedforward neural network according to an embodiment of the present invention;

[0050] Figure 5 is a flow chart of another method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention;

[0051] Figure 6 is a schematic diagram of input data blocks after the first linear transformation layer performs block processing on network input data according to an embodiment of the present invention;

[0052] Figure 7 is a schematic diagram of weight data blocks after the first linear transformation layer performs block processing on model weight data according to an embodiment of the present invention;

[0053] Figure 8 is a schematic diagram of broadcasting an input data block across computing cores according to an embodiment of the present invention;

[0054] Fig. 9 is a flow chart of another method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention;

[0055] Fig.10 is a schematic diagram of first matrix blocks after the second linear transformation layer processes the input data matrix blocks according to an embodiment of the present invention;

[0056] Fig.11 is a schematic diagram of second matrix blocks after the second linear transformation layer processes the weight data matrix blocks according to an embodiment of the present invention;

[0057] Fig.12 It is a flow chart of a method for accelerating the calculation of a feedforward neural network in a model reasoning and decoding stage according to an embodiment of the present invention;

[0058] Fig.13 is a structural block diagram of a computing acceleration device for a feedforward neural network according to an embodiment of the present invention;

[0059] Fig.14 It is a schematic diagram of the hardware structure of the main control board according to an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0061] The relevant language models often have a huge number of parameters, which in turn brings a large amount of network execution calculations. Therefore, the network is usually executed in parallel through processors with multiple computing cores such as multi-core processors, GPUs (Graphics Processing Units), and AI (Artificial Intelligence) processors.

[0062] The Transformer model (a deep learning model architecture) structure is widely used in the field of natural language processing. For large Transformer-based language models, the time consumed by the feedforward neural network accounts for a relatively high proportion of the total time consumed in the end-to-end reasoning process. Therefore, accelerating the computation of the feedforward neural network part is of great significance for improving the reasoning speed of the model.

[0063] Among them, in the feedforward neural network module of the Transformer model, the input data needs to undergo two linear transformations and one nonlinear transformation. The nonlinear transformation is between the two linear transformations, and the linear transformation is implemented using matrix multiplication.

[0064] like Figure 1 As shown, in modern processors, due to the "memory wall" problem, a multi-layer memory access structure is implemented. The closer the storage is to the computing core, the faster the memory access speed, but the smaller the capacity and the higher the price; the farther the storage is from the computing core, the slower the memory access speed, but the larger the capacity and the lower the price. Registers are located inside the processor and have the fastest access speed, but the smallest number / capacity.

[0065] When performing calculations on a feedforward neural network, the processor's high-speed storage capacity is small and therefore usually cannot accommodate the intermediate result data of linear transformation / nonlinear transformation. Each time a linear and nonlinear transformation in the network is executed, data must be obtained from a low-speed storage far away from the computing core. After the calculation is completed, the result is written back to the low-speed storage.

[0066] like Figure 2 As shown, the execution of a feedforward neural network involves multiple transfers and movements of data in low-speed storage, which consumes more network execution time.

[0067] An embodiment of the present invention provides a method for accelerating the calculation of a feedforward neural network, which is applied to a multi-core processor. The multi-core processor includes multiple computing cores. For two linear transformations and one nonlinear transformation in the feedforward neural network, the input data of the network is reasonably distributed on the multi-core processor. Only when reading the neural network weight data, the initial input and the final output data, it is necessary to access the low-speed storage. The intermediate execution data of the network is always stored in the high-speed storage of the multi-core, thereby effectively reducing the execution time of the network.

[0068] According to an embodiment of the present invention, an embodiment of a method for accelerating the calculation of a feedforward neural network is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0069] In this embodiment, a method for accelerating the calculation of a feedforward neural network is provided, which can be used in the above-mentioned multi-core processor. Figure 3 is a flow chart of a method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0070] Step S301, obtain the network input data and model weight data corresponding to the first linear transformation layer, respectively block the network input data and the model weight data to obtain data blocks, and distribute the data blocks to multiple computing cores in turn.

[0071] Specifically, Figure 4 As shown, the feedforward neural network includes a first linear transformation layer, a nonlinear transformation layer and a second linear transformation layer; the operation result of the first linear transformation layer will be used as the input data of the nonlinear transformation layer, and the output data of the nonlinear transformation layer will be used as the input data of the second linear transformation layer, and then the second linear transformation layer outputs the operation result of the feedforward neural network.

[0072] Step S302, using the computing core to read the data blocks from the low-speed storage area, performing linear transformation on the data blocks to obtain output data of the first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core.

[0073] Specifically, the low-speed storage area has a large storage capacity and a slow access speed, while the high-speed area has a small storage capacity and a fast access speed. Since the amount of model weight data is large, it can only be read in blocks from the low-speed storage area.

[0074] Furthermore, in the decoding stage of model reasoning, the amount of network input data is positively correlated with the number of batches involved in model reasoning. When a multi-core processor is used to perform matrix multiplication, the high-speed storage of a single core is usually unable to accommodate the left matrix data and matrix operation results, but the high-speed storage areas of all cores in the multi-core processor are sufficient to accommodate the data; or in the case of a small batch number, the high-speed storage of a single core is sufficient to accommodate it, and the high-speed storage of multiple cores can support the execution of a larger number of batches when the high-speed storage of multiple cores is superimposed.

[0075] Step S303: Perform nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer.

[0076] Specifically, an activation function operation is performed on the output data of the first linear transformation layer to obtain output data of the nonlinear transformation layer.

[0077] Step S304, block processing is performed on the output data of the nonlinear transformation layer, and linear transformation is performed on the output data after block processing to obtain a feedforward neural network operation result, and the feedforward neural network operation result is stored in a low-speed storage area.

[0078] The present embodiment provides a method for accelerating the calculation of a feedforward neural network, which performs block processing on network input data and model weight data respectively to obtain data blocks, uses a computing core to read the data blocks from a low-speed storage area, performs linear transformation on the data blocks to obtain output data of a first linear transformation layer, and stores the output data of the first linear transformation layer in a high-speed storage area of ​​the computing core, performs nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer; performs block processing on the output data of the nonlinear transformation layer, performs linear transformation on the block-processed output data to obtain a feedforward neural network operation result, and stores the feedforward neural network operation result in a low-speed storage area; for two linear transformations and one nonlinear transformation in the feedforward neural network, by reasonably allocating the network input data and the model weight data on a multi-core processor, the low-speed storage area only needs to be accessed when reading the data blocks, the initial input and the final output data of the feedforward neural network, and the intermediate execution data of the feedforward neural network is always stored in the high-speed storage area of ​​the multi-core, which effectively reduces the execution time of the feedforward neural network and improves the calculation speed of the feedforward neural network.

[0079] In this embodiment, a method for accelerating the calculation of a feedforward neural network is provided, which can be used in the above-mentioned multi-core processor. Figure 5 is a flow chart of a method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention, wherein the feedforward neural network includes a first linear transformation layer, a nonlinear transformation layer, and a second linear transformation layer. Figure 5 As shown, the process includes the following steps:

[0080] Step S501, obtain the network input data and model weight data corresponding to the first linear transformation layer, respectively block the network input data and the model weight data to obtain data blocks, and distribute the data blocks to multiple computing cores in turn.

[0081] Specifically, the above step S501 includes:

[0082] Step S5011, dividing the network input data into blocks according to the matrix dimension to obtain input data blocks.

[0083] Specifically, assuming that in a feedforward neural network, the matrix dimension of the network input data of the matrix multiplication corresponding to the first linear transformation layer is [M, K1], the weight dimension of the model weight data is [K1, N1], the number of computing cores in the multi-core processor is N, the network input data is blocked in the K1 dimension, the block size is K1b, and then each input data block is allocated to the N computing cores in turn, one input data block at a time; for example, when the total number of blocks of the input data block is 2N, each core is allocated 2 blocks.

[0084] Furthermore, assuming that the number of computing cores is 4, namely core 0 to core 3, the input data block is as follows Figure 6 shown.

[0085] Step S5012, dividing the model weight data into blocks according to the weight dimension to obtain weight data blocks.

[0086] In some optional implementations, the above step S5012 includes:

[0087] Step a1, dividing the model weight data into blocks according to the first weight dimension to obtain multiple initial data blocks.

[0088] Specifically, the model weight data of the first linear transformation layer is divided into blocks in the N1 dimension, with a block size of N1b, and then each block is sequentially allocated to the N computing cores, one block at a time.

[0089] Step a2, dividing the multiple initial data blocks into blocks according to the second weight dimension to obtain weighted data blocks.

[0090] Specifically, due to the large amount of model weight data, it needs to be further divided into blocks in the K1 dimension before it can be read into high-speed storage. The size of the block is K1b, that is, for the model weight data, each computing core only reads a block of size [K1b, N1b] at a time. The weight data block is as follows: Figure 7 shown.

[0091] Step S5013, taking the input data block and the weight data block as data blocks, and allocating the data blocks to multiple computing cores in sequence.

[0092] Step S502, using the computing core to read the data blocks from the low-speed storage area, performing linear transformation on the data blocks to obtain output data of the first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core.

[0093] Specifically, after the allocation is completed, each computing core reads the data blocks from the low-speed storage area to the high-speed storage area, and only reads the data blocks allocated to itself to avoid insufficient high-speed storage space.

[0094] Furthermore, the above step S502 includes:

[0095] Step S5021, using the computing core to iteratively read the input data block and broadcast it between cores to obtain an input data matrix; wherein the input data matrix is ​​a matrix corresponding to the network input data.

[0096] Specifically, according to the characteristics of matrix multiplication, in order to perform operations correctly, each computing core needs to obtain complete network input data. The acquisition of complete network input data is achieved through iterative reading of data blocks and on-chip broadcasting between computing cores.

[0097] In some optional implementations, the above step S5021 includes:

[0098] Step b1, use the computing core to iteratively read the input data block from the low-speed storage area, broadcast the input data block to other computing cores in turn until the number of iterative cycles is reached, and stop broadcasting; wherein the number of iterative cycles is the total number of blocks of the input data block in the computing core.

[0099] Specifically, assuming K1=K1n*K1b, the total number of blocks is K1n, and a total of K1n cycles are performed. Since the network input data is discretely stored on each computing core, the computing core that stores the input data block to be circulated is used to broadcast the input data block to other computing cores in each iterative cycle. The broadcasted input data block is stored in a fixed broadcast data storage space in the high-speed storage, and the size is K1b*N1b; assuming that K1n (that is, the total number of blocks of the input data block) is 8, and the number of computing cores is 4, the first two data broadcasts are shown, and the input data block is broadcast across computing cores as follows: Figure 8 shown.

[0100] Step b2: receiving the input data blocks broadcasted by the other computing cores, and storing the input data blocks broadcasted by the other computing cores into the broadcast data storage area.

[0101] Step b3, determining the data matrix based on the input data blocks in the broadcast data storage area.

[0102] Step S5022, performing block matrix multiplication and accumulation reduction processing on the input data matrix and the weight data block to obtain output data of the first linear transformation layer.

[0103] Specifically, in the K1n loop, after each computing core obtains the input data matrix and the weight data block, it performs the current block matrix multiplication operation and accumulates the operation results to the result data area of ​​the current computing core.

[0104] Step S5023, storing the output data of the first linear transformation layer into the high-speed storage area of ​​the computing core.

[0105] Specifically, after the operation of the first linear transformation layer is completed, the result of the operation is stored in the high-speed storage area of ​​each computing core in the form of blocks, and is not written back to the low-speed storage area.

[0106] Furthermore, the input and output dimensions of the nonlinear transformation layer are both [M, N1], where the structure of the feedforward neural network essentially limits N1 to K2, that is, the nonlinear transformation layer does not affect the dimensions of the input and output data, and the output dimension of the first linear transformation layer serves as the input dimension of the second linear transformation layer.

[0107] Step S503: Perform nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer.

[0108] Specifically, each computing core directly performs nonlinear transformation, namely activation function operation, on the block data assigned to it, without reading input data from the low-speed storage area and writing back the result data.

[0109] Step S504, block processing is performed on the output data of the nonlinear transformation layer, and the block-processed output data is linearly transformed to obtain the feedforward neural network operation result, and the feedforward neural network operation result is stored in the low-speed storage area. For details, please refer to Figure 3 Step S304 of the illustrated embodiment will not be described in detail here.

[0110] The present embodiment provides a method for accelerating the calculation of a feedforward neural network, which processes network input data and model weight data in blocks, and distributes the data blocks to multiple computing cores in turn, so that each computing core only reads the data blocks allocated to itself, thereby avoiding insufficient high-speed storage space and improving the memory access speed of the computing core; secondly, the computing core is used to iteratively read and broadcast the input data blocks between cores, and the input data matrix and the weight data block are subjected to block matrix multiplication and accumulation reduction processing to obtain the output data of the first linear transformation layer, thereby realizing the linear transformation of the input data block, and the output data of the first linear transformation layer is stored in the high-speed storage area of ​​the computing core, thereby ensuring that the execution of the feedforward neural network is in the high-speed storage space, thereby improving the calculation speed of the feedforward neural network.

[0111] In this embodiment, a method for accelerating the calculation of a feedforward neural network is provided, which can be used in the above-mentioned multi-core processor. Fig. 9 is a flow chart of a method for accelerating calculation of a feedforward neural network according to an embodiment of the present invention, wherein the feedforward neural network includes a first linear transformation layer, a nonlinear transformation layer, and a second linear transformation layer. Fig. 9 As shown, the process includes the following steps:

[0112] Step S901, obtain the network input data and model weight data corresponding to the first linear transformation layer, respectively block the network input data and model weight data to obtain data blocks, and sequentially distribute the data blocks to multiple computing cores. For details, please refer to Figure 5 Step S501 of the illustrated embodiment will not be described in detail here.

[0113] Step S902: Use the computing core to read the data blocks from the low-speed storage area, perform linear transformation on the data blocks, obtain the output data of the first linear transformation layer, and store the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core. Figure 5 Step S502 of the illustrated embodiment will not be described in detail here.

[0114] Step S903, perform nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer. Figure 5 Step S503 of the illustrated embodiment will not be described in detail here.

[0115] Step S904, the output data of the nonlinear transformation layer is processed in blocks, the output data after the block processing is linearly transformed to obtain the feedforward neural network operation result, and the feedforward neural network operation result is stored in the low-speed storage area.

[0116] Specifically, the above step S904 includes:

[0117] Step S9041: Use the output data of the nonlinear transformation layer as the input data of the second linear transformation layer, decompose the input data of the second linear transformation layer, and obtain an input data matrix and a weight data matrix.

[0118] Specifically, it is assumed that the matrix dimension of the input data matrix of the matrix multiplication corresponding to the second linear transformation layer is [M, K2], and the weight dimension of the weight data matrix is ​​[K2, N2].

[0119] Step S9042, performing block processing on the input data matrix and the weight data matrix respectively to obtain a first matrix block and a second matrix block.

[0120] Specifically, Figure 10-11 As shown, for the second linear transformation layer, block processing is performed in a similar manner to the first linear transformation layer (ie, steps S5011 to S5013), as shown in FIG. Fig.11As shown, when the weight data matrix of the second linear transformation layer is processed in blocks, K2b=N1b is taken. Therefore, the calculation result of the first linear transformation layer can be directly used as the left matrix input of the second linear transformation layer after the nonlinear transformation is completed, without the need for data transfer through low-speed storage.

[0121] Step S9043: linearly transform the first matrix block and the second matrix block to obtain output data of the second linear transformation layer.

[0122] Specifically, assuming that K2=K2n*K2b, where K2b is the block size in the K2 direction, and K2n is the total number of blocks, the second linear transformation layer is executed for a total of K2n loops, and the calculation process is similar to that of the first linear transformation layer (i.e., step S5021-step S5023), including the broadcasting of the first matrix blocks, the matrix multiplication operation of the blocks and the accumulation reduction of the results. After the loop is completed, the operation of the second linear transformation layer is completed.

[0123] Step S9044, using the output data of the second linear transformation layer as the feedforward neural network operation result, and using the computing core to write the feedforward neural network operation result to the low-speed storage area.

[0124] Specifically, the calculation result of the second linear transformation layer is the calculation result of the entire feedforward neural network. At this time, the blocks are stored in the high-speed storage area of ​​each computing core, and each computing core writes its own block data to the corresponding position in the low-speed storage, thereby completing the execution of the feedforward neural network.

[0125] The present embodiment provides a method for accelerating the calculation of a feedforward neural network. By reasonably dividing the input data of the second linear transformation layer into blocks and effectively utilizing the high-speed storage area on the multi-core processor, when the feedforward neural network is executed, there is no need to temporarily store and transfer the result data of the intermediate layer in the low-speed storage, thereby accelerating the execution of the feedforward neural network.

[0126] The following is a specific example to illustrate the specific steps of a method for accelerating the calculation of a feedforward neural network.

[0127] Embodiment 1:

[0128] like Fig.12 As shown, the calculation acceleration method of the feedforward neural network in the model reasoning decoding stage includes:

[0129] S1. Divide the input data of the first linear transformation layer into blocks according to the specified mode, and each computing core only processes the allocated blocks:

[0130] The left matrix of the first linear transformation layer (i.e., network input data) is divided into blocks in the K1 dimension, with a block size of K1b, and then each block is allocated to N computing cores in turn, one block at a time. Each computing core reads the blocks from the low-speed storage area to the high-speed storage area, and only reads the left matrix blocks allocated to itself to avoid insufficient high-speed storage space;

[0131] The right matrix of the first linear transformation layer (i.e., the model weight data) is divided into blocks in the N1 dimension with a block size of N1b, and then each block is allocated to the N computing cores in turn, one block at a time; it needs to be further divided into blocks in the K1 dimension before it can be read into high-speed storage, with a block size of K1b, that is, for the right matrix, each core only reads a block of [K1b, N1b] size at a time.

[0132] S2. Broadcast the left matrix data of the first linear transformation layer between cores and circulate the data into blocks:

[0133] To perform operations correctly, each computing core needs to obtain the complete left matrix data, which is obtained through iterative reading of blocks and on-chip broadcasting between cores;

[0134] In the K1n loop, each computing core, after obtaining the corresponding left matrix and right matrix blocks, performs the current block matrix multiplication operation and accumulates the results to the result data area of ​​the current computing core; after the loop is completed, the operation of the first linear transformation layer is completed.

[0135] S3. Each computing core performs nonlinear transformation on the data block on the current core:

[0136] Each computing core directly performs nonlinear transformations, namely activation function operations, on the block data assigned to it, without having to read input data from low-speed storage and write back result data.

[0137] S4. The second linear transformation layer receives the data blocks in the high-speed storage area as input data and divides them into blocks according to the execution mode:

[0138] When dividing the right matrix of the second linear transformation layer into blocks, K2b=N1b is taken. After the nonlinear transformation, the calculation result of the first linear transformation layer can be directly used as the left matrix input of the second linear transformation layer without the need for data transfer through low-speed storage.

[0139] S5. Broadcast the left matrix data of the second linear transformation layer between cores, and loop the data blocks to complete the network operation.

[0140] The calculation result of the second linear transformation layer is the calculation result of the entire feedforward neural network. Each computing core writes its own block data to the corresponding position in the low-speed storage, thus completing the execution of the feedforward neural network.

[0141] In this embodiment, a computing acceleration device for a feedforward neural network is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0142] This embodiment provides a computing acceleration device for a feedforward neural network, such as Fig.13 As shown, including:

[0143] The block module 1301 is used to obtain the network input data and the model weight data corresponding to the first linear transformation layer, respectively block the network input data and the model weight data to obtain data blocks, and sequentially distribute the data blocks to multiple computing cores;

[0144] The first linear transformation module 1302 is used to read the data blocks from the low-speed storage area using the computing core, perform linear transformation on the data blocks to obtain output data of the first linear transformation layer, and store the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core;

[0145] The non-linear transformation module 1303 is used to perform non-linear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the non-linear transformation layer;

[0146] The second linear transformation module 1304 is used to perform block processing on the output data of the nonlinear transformation layer, perform linear transformation on the block-processed output data to obtain a feedforward neural network operation result, and store the feedforward neural network operation result in a low-speed storage area.

[0147] In some optional implementations, the block module 1301 includes:

[0148] A first block division unit is used to divide the network input data into blocks according to the matrix dimension to obtain input data blocks;

[0149] The second block division unit is used to divide the model weight data into blocks according to the weight dimension to obtain weight data blocks;

[0150] The allocation unit is used to use the input data block and the weight data block as data blocks, and allocate the data blocks to multiple computing cores in sequence.

[0151] In some optional implementations, the second block unit includes:

[0152] A first block subunit is used to perform block processing on the model weight data according to a first weight dimension to obtain a plurality of initial data blocks;

[0153] The second blocking subunit is used to block the multiple initial data blocks according to the second weight dimension to obtain weighted data blocks.

[0154] In some optional implementations, the first linear transformation module 1302 includes:

[0155] A broadcast unit is used to iteratively read the input data block and broadcast it between cores using the computing core to obtain an input data matrix; wherein the input data matrix is ​​a matrix corresponding to the network input data;

[0156] A computing unit, used for performing block matrix multiplication and accumulation reduction processing on the input data matrix and the weight data block to obtain output data of the first linear transformation layer;

[0157] The first storage unit is used to store the output data of the first linear transformation layer into the high-speed storage area of ​​the computing core.

[0158] In some optional implementations, the first broadcast unit includes:

[0159] The broadcast subunit is used to iteratively read the data blocks from the low-speed storage area using the computing core, and broadcast the data blocks to other computing cores in sequence until the number of iteration cycles is reached, and then the broadcast is stopped; wherein the number of iteration cycles is the total number of data blocks in the computing core;

[0160] A storage subunit, used for receiving data blocks broadcasted by other computing cores, and storing the data blocks broadcasted by other computing cores in a broadcast data storage area;

[0161] The determination subunit is used to determine the data matrix based on the data blocks in the broadcast data storage area.

[0162] In some optional implementations, the nonlinear transformation module 1303 is specifically configured to perform an activation function operation on the output data of the first linear transformation layer to obtain output data of the nonlinear transformation layer.

[0163] In some optional implementations, the second linear transformation module 1304 includes:

[0164] A decomposition unit, used to use the output data of the nonlinear transformation layer as the input data of the second linear transformation layer, decompose the input data of the second linear transformation layer, and obtain an input data matrix and a weight data matrix;

[0165] A third block division unit is used to perform block processing on the input data matrix and the weight data matrix respectively to obtain a first matrix block and a second matrix block;

[0166] A linear transformation unit, used for performing a linear transformation on the first matrix block and the second matrix block to obtain output data of a second linear transformation layer;

[0167] The second storage unit is used to use the output data of the second linear transformation layer as the feedforward neural network operation result, and use the computing core to write the feedforward neural network operation result to the low-speed storage area.

[0168] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0169] At the hardware level, the multi-core processor can be integrated into a main control board.

[0170] like Fig.14 As shown, a main control board includes: a multi-core processor, a high-speed storage device and a low-speed storage device, and the multi-core processor is used to execute the calculation acceleration method of the feedforward neural network.

[0171] Among them, the multi-core processor is the execution device of the computational acceleration method of the feedforward neural network, and the computing cores can communicate on-chip; the high-speed storage device is used to store the intermediate execution data of the feedforward neural network; the low-speed storage device is used to store the weight data, initial input and final output data of the feedforward neural network.

[0172] In addition, an embodiment of the present invention further provides a computer-readable storage medium, and the above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or is implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0173] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for accelerating the calculation of a feedforward neural network, characterized in that: The feedforward neural network includes a first linear transformation layer, a nonlinear transformation layer, and a second linear transformation layer; the method includes: Obtaining network input data and model weight data corresponding to the first linear transformation layer, performing block processing on the network input data and the model weight data respectively to obtain data blocks, and sequentially allocating the data blocks to multiple computing cores; Using the computing core to read the data blocks from the low-speed storage area, performing linear transformation on the data blocks to obtain output data of a first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core; Performing nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer; The output data of the nonlinear transformation layer is processed in blocks, the output data after the block processing is linearly transformed to obtain a feedforward neural network operation result, and the feedforward neural network operation result is stored in the low-speed storage area.

2. The method according to claim 1, characterized in that The processing of the network input data and the model weight data in blocks respectively to obtain data blocks, and sequentially allocating the data blocks to a plurality of computing cores comprises: Processing the network input data in blocks according to the matrix dimension to obtain input data blocks; Processing the model weight data in blocks according to the weight dimension to obtain weight data blocks; The input data block and the weight data block are used as the data blocks, and the data blocks are sequentially allocated to the multiple computing cores.

3. The method according to claim 2, characterized in that The block processing of the model weight data according to the weight dimension to obtain the weight data block includes: Processing the model weight data in blocks according to the first weight dimension to obtain a plurality of initial data blocks; The multiple initial data blocks are divided into blocks according to the second weight dimension to obtain the weighted data blocks.

4. The method according to claim 2, characterized in that: The method of using the computing core to read the data blocks from the low-speed storage area, performing linear transformation on the data blocks to obtain output data of a first linear transformation layer, and storing the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core includes: Using the computing core to iteratively read and broadcast the input data block between cores to obtain an input data matrix; wherein the input data matrix is ​​a matrix corresponding to the network input data; Performing block matrix multiplication and accumulation reduction processing on the input data matrix and the weight data block to obtain output data of the first linear transformation layer; The output data of the first linear transformation layer is stored in the high-speed storage area of ​​the computing core.

5. The method according to claim 4, characterized in that The step of using the computing core to iteratively read and broadcast the input data block between cores to obtain a data matrix includes: Iteratively read the input data block from the low-speed storage area using the computing core, and broadcast the input data block to other computing cores in sequence until the iterative cycle number is reached, and then stop broadcasting; wherein the iterative cycle number is the total number of blocks of the input data block in the computing core; receiving input data blocks broadcasted by the other computing cores, and storing the input data blocks broadcasted by the other computing cores in a broadcast data storage area; The data matrix is ​​determined based on input data blocks in the broadcast data storage area.

6. The method according to claim 1, characterized in that The performing nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer includes: An activation function operation is performed on the output data of the first linear transformation layer to obtain output data of the nonlinear transformation layer.

7. The method according to claim 1, characterized in that The step of processing the output data of the nonlinear transformation layer in blocks, performing linear transformation on the output data after the block processing to obtain a feedforward neural network operation result, and storing the feedforward neural network operation result in the low-speed storage area includes: Using the output data of the nonlinear transformation layer as the input data of the second linear transformation layer, decomposing the input data of the second linear transformation layer to obtain an input data matrix and a weight data matrix; Performing block processing on the input data matrix and the weight data matrix respectively to obtain a first matrix block and a second matrix block; Performing a linear transformation on the first matrix block and the second matrix block to obtain output data of a second linear transformation layer; The output data of the second linear transformation layer is used as the feedforward neural network operation result, and the feedforward neural network operation result is written out to the low-speed storage area using the computing core.

8. A computing acceleration device for a feedforward neural network, characterized in that: The device comprises: A block module is used to obtain network input data and model weight data corresponding to the first linear transformation layer, respectively perform block processing on the network input data and the model weight data to obtain data blocks, and sequentially distribute the data blocks to multiple computing cores; A first linear transformation module, configured to use the computing core to read the data blocks from the low-speed storage area, perform linear transformation on the data blocks to obtain output data of a first linear transformation layer, and store the output data of the first linear transformation layer in the high-speed storage area of ​​the computing core; A nonlinear transformation module, used for performing nonlinear transformation on the output data of the first linear transformation layer in the high-speed storage area to obtain output data of the nonlinear transformation layer; The second linear transformation module is used to perform block processing on the output data of the nonlinear transformation layer, perform linear transformation on the output data after block processing to obtain a feedforward neural network operation result, and store the feedforward neural network operation result in the low-speed storage area.

9. A main control board, characterized in that: The method comprises: a multi-core processor, a high-speed storage device and a low-speed storage device, wherein the multi-core processor is used to execute the calculation acceleration method of the feedforward neural network according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the calculation acceleration method of the feedforward neural network according to any one of claims 1 to 7.