Method for Processing Samples in a Neural Network, Computing Device, and Storage Medium
By adjusting the sample processing method according to the feature map size of each layer and the hardware channel width in the neural network, the problem of unbalanced resource consumption is solved and more efficient computing and resource utilization is achieved.
Patent Information
- Application Number
- CN202210087807.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-01-25
AI Technical Summary
In neural networks, due to the different sizes of feature maps of each layer, the resource consumption of cache or access input and output data is unbalanced, which affects the computing efficiency.
By processing different numbers of samples at each stage of the neural network and splicing the samples appropriately to maximize the utilization of on-chip cache access excitation values while making the sample processing suitable for the hardware channel width of the calculation core.
Effectively utilize on-chip cache, reduce dependence on external storage, improve computing efficiency, and optimize resource use to avoid resource waste.
Smart Images

Figure CN114418079B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of neural network computing, and more particularly, to a method for processing samples in a neural network, a computing device, and a computer-readable storage medium. Background Art
[0002] Currently, neural networks, especially neural networks for computer vision, have been widely applied in fields such as image classification and object recognition. In these fields, a neural network can be trained using pre-acquired image samples, etc., to obtain a corresponding trained neural network model. This trained neural network model can be used to identify or classify new image data.
[0003] The training of a neural network is a complex process. A small change in the front layer of the network will accumulate and amplify to the subsequent layers, so the update of the training parameters of the front layer will cause a change in the input data distribution of the subsequent layers. For this reason, the concept of batch normalization (BN) is introduced in the neural network. Before processing the input data of each layer, these input data are first batch-normalized to forcefully transform the features of the input data into a mathematical model with a mean of 0 and a variance of 1.
[0004] The sizes of the various layers of a neural network may be different, so the output data sizes for each layer running once are different, making the input data sizes of the next layer also different accordingly, thus resulting in unbalanced resource consumption for caching or accessing input and output data between different layers. Summary of the Invention
[0005] In view of the above problems, the present disclosure provides a method for processing samples in a neural network, which maximizes the use of on-chip cache to access the excitation values generated in each stage and makes the sample processing suitable for the hardware channel width of the computing core by making the number of samples processed in each intermediate stage of the neural network different when running once and appropriately splicing the samples.
[0006] According to one aspect of the present disclosure, there is provided a method for processing samples in a neural network. The method includes: determining the number of samples included in a batch of input data based on the minimum feature map size among multiple stages of the neural network and the size of the on-chip cache of the computing core; receiving a batch of input data containing the number of samples at the first stage of the neural network; in each stage of the neural network, dividing the batch of input data into one or more groups of sub-batch data based on the feature map size of the stage and the hardware channel width of the computing core; and sequentially performing batch normalization processing and activation processing on each group of sub-batch data.
[0007] According to another aspect of the present disclosure, a computing device is provided. The computing device includes: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, which when executed by the at least one processor, cause the computing device to perform the steps according to the above method.
[0008] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program code is stored, and the computer program code performs the method as described above when run.
[0009] In some implementations, the neural network includes a computer vision network, and the feature map sizes of the multiple stages of the computer vision network gradually decrease or first decrease and then increase.
[0010] In some implementations, determining the number of samples included in a batch of input data includes: determining the number of samples included in the batch of input data such that the generated excitation values can be fully cached in the on-chip cache at the stage corresponding to the minimum feature map size.
[0011] In some implementations, dividing the batch of input data into one or more groups of sub-batch data includes: at each stage of the neural network, determining all candidate division methods based on the number of samples included in the batch of input data, where the number of groups of sub-batch data is different in each candidate division method; and selecting a division method from all the candidate division methods based on the feature map size of the stage and the hardware channel width, where the selected division method makes the size of each group of sub-batch data an integer multiple of the hardware channel width.
[0012] In some implementations, the method further includes: determining whether the stage is the stage with the smallest feature map size of the neural network; if it is determined that the stage is not the stage with the smallest feature map size of the neural network, dividing each group of sub-batch data into at least two groups of micro-batch data, and the samples in each group of micro-batch data have a dependency relationship in two adjacent stages; and if it is determined that the stage is the stage with the smallest feature map size of the neural network, taking each group of sub-batch data as a group of micro-batch data.
[0013] In some implementations, the method further includes: for each group of micro-batch data at each stage, determining whether the feature map size of the stage is a multiple of the hardware channel width of the computing core; if it is determined that the feature map size of the stage is not a multiple of the hardware channel width of the computing core, splicing the samples in the group of micro-batch data so that the size of the spliced micro-batch data is a multiple of the hardware channel width.
[0014] In some implementations, performing batch normalization processing and activation processing on each group of sub-batch data in sequence includes: performing batch normalization processing and activation processing on each concatenated micro-batch data in each group of sub-batch data in sequence.
[0015] In some implementations, the method further includes: at each stage of the neural network, caching the excitation values generated each time activation processing is performed into the on-chip cache.
[0016] In some implementations, the method further includes: if the number of groups of sub-batch data in a stage is greater than two, then the operation can directly proceed to the next stage after running twice in this stage. Description of the Drawings
[0017] By referring to the description of the specific embodiments of the present disclosure given in the following drawings, the present disclosure will be better understood, and other objects, details, features, and advantages of the present disclosure will become more obvious.
[0018] Figure 1 Shows a schematic structural diagram of a neural network according to the present disclosure.
[0019] Figure 2 Shows a schematic diagram of a computing device for implementing a method for processing samples in a neural network according to an embodiment of the present disclosure.
[0020] Figure 3 Shows a flowchart of a method for processing samples in a neural network according to an embodiment of the present disclosure.
[0021] Figure 4 Shows a flowchart of a process for dividing batch input data according to some embodiments of the present disclosure.
[0022] Figure 5 Shows a timing diagram of processing samples in a neural network according to some embodiments of the present disclosure. Detailed Description of the Embodiments
[0023] The preferred embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0024] As used herein, the term "comprising" and its variations denote open-ended inclusion, i.e., "including but not limited to". Unless otherwise specified, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one embodiment" and "some embodiments" mean "at least one exemplary embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0025] Figure 1 FIG. shows a schematic structural diagram of a neural network 100 according to the present disclosure. As Figure 1 shown, the neural network 100 may include an input layer 110, a plurality of intermediate stages 120 ( Figure 1 four stages 120-1, 120-2, 120-3, and 120-4 are schematically shown therein) and a fully connected layer 130 as an output layer. Among them, each intermediate stage 120 includes a convolutional layer 122, a pooling layer 124, a batch normalization layer 126, and an activation layer 128. The convolutional layer 122 is used to perform convolutional processing on the input data, the pooling layer 124 is used to downsample the convolutional result to ensure feature invariance, the batch normalization layer 126 is used to perform batch normalization on the input data, and the activation layer 128 is used to apply an activation function (such as the Sigmoid or Relu function) to the normalized input data to generate activation values.
[0026] The neural network 100 can be implemented on a computing core, and the computing core may include an on-chip cache 140 for caching the activation values (activations) generated by the activation layer 128 of each stage 120 of the neural network 100. The capacity of the on-chip cache 140 is limited, usually only a few megabytes.
[0027] In the existing sample processing method, a specific number of batches of input data are input into the input layer 110 of the neural network 100 each time for training or prediction. However, due to the different sizes of the convolutional kernels of the convolutional layers 122 of each stage 120 of the neural network 100, the sizes of the feature maps generated by each convolutional layer 122 are different, and thus the sizes of the activation values of the activation layers 126 of each stage 120 are also different. Therefore, the activation values generated by some stages 120 running once may not all be cached in the on-chip cache 140, or may not occupy the entire on-chip cache 140. For the former, the activation values generated by this stage 120 may need to be stored in an off-chip memory ( Figure 1 exemplarily shown as L2 in FIG.), so that the next stage 120 needs to access these activation values from the off-chip memory L2 to input them into its convolutional layer 122, which significantly increases the time consumed for accessing the activation values. For the latter, it will cause the storage space of the on-chip cache 140 not to be fully utilized, resulting in waste of storage resources.
[0028] In particular, for neural networks in the field of computer vision, as the number of stages increases, the size of the feature map becomes smaller, the size of the channel becomes larger, and the size of the excitation value also becomes smaller. When the number of samples in the first stage 120 can be processed in one go and the generated excitation values fit into the on-chip cache 140, the fewer the excitation values generated in the subsequent stages 120, the more free space there will be in the on-chip cache 140.
[0029] To address this issue, the present disclosure proposes a method for processing samples in a neural network, where the number of samples processed in one run (i.e., the batch size) is different at each stage of the neural network. The maximum batch size is determined by the stage 120 with the smallest feature map size, so that all samples can be processed in just one run at this stage 120. For other stages 120, the data volume of the maximum batch size can be divided into multiple groups, and only one group is processed in each run. In addition, for each group of data, the samples in the group can be axially concatenated according to the hardware channel width of the computing core, so that the concatenated data fits the hardware channel width.
[0030] Figure 2 A schematic diagram of a computing device 20 for implementing the method of processing samples in a neural network 100 according to an embodiment of the present disclosure is shown. As Figure 2 shown, the computing device 20 may include at least one processor 210 and at least one memory 220 coupled to the at least one processor 210. Instructions 230 executable by the at least one processor 210 are stored in the memory 220. When the instructions 230 are executed by the at least one processor 210, at least a part of the method 300 described below is performed. Specifically, the instructions 230 stored in the memory 220 may include instruction codes for constructing the neural network 100 as shown above in connection with Figure 1 shown, to perform the method 300 of processing samples as described below in connection with Figure 3 shown at runtime.
[0031] Figure 3 A flowchart of the method 300 for processing samples in a neural network 100 according to an embodiment of the present disclosure is shown.
[0032] As Figure 3 shown, at block 310, the computing device 20 determines the number of samples included in a batch of input data based on the smallest feature map size among the multiple stages 120 of the neural network 100 and the size of the on-chip cache 140 of the computing core. Batch input data refers to the data input to the input layer 110 of the neural network 100 as a batch, which may include one or more samples.
[0033] For example, for a neural network in the field of computer vision, the sizes of the feature maps at each stage are shown in Table 1 below:
[0034] Table 1
[0035] Stage 120-1 120-2 120-3 120-4 Feature map size 56x56 28x28 14x14 7x7
[0036] At block 310, based on the minimum feature map size and the size of the on-chip cache 140, the number of samples included in the batch input data can be determined such that the generated excitation values at the stage corresponding to the minimum feature map size (e.g., stage 120-4 shown in Table 1) can be fully cached in the on-chip cache 140. That is, the excitation values generated at this stage fill or almost fill the on-chip cache 140. For example, in the case of the feature map sizes shown in Table 1, when the maximum feature map size at stage 120-1 is 56*56 and the number of channels is 256, the data volume of the excitation values generated for processing one sample is 56*56*256*2 = 1568 KiB. When the minimum feature map size at stage 120-4 is 7*7 and the number of channels is 2048, the data volume of the excitation values generated for processing one sample is 7*7*2048*2 = 196 KiB. Therefore, when the excitation values generated by processing 1 sample in the activation layer at stage 120-1 basically fill the on-chip cache 140, the excitation values generated by processing 8 samples in the activation layer 128 at stage 120-4 basically fill the on-chip cache 140. In this case, the number of samples included in a batch input data of the neural network 100 can be determined to be 8.
[0037] At block 320, the computing device 20 can receive a batch input data containing that number of samples at the first stage 120-1 of the neural network 100. Here, it is assumed that the batch input data received at the first stage 120-1 of the neural network 100 contains 8 samples.
[0038] At block 330, the computing device 20 can divide the batch input data into one or more groups of sub-batch data at each stage 120 of the neural network 100 based on the feature map size at that stage and the hardware channel width of the computing core.
[0039] As described above, since the feature map sizes are different for each stage, when all samples of the batch input data can be processed at once in the stage with the smallest feature map size (such as stage 120-4) and all the generated excitation values are cached in the on-chip cache 140, stages with larger feature map sizes than this stage (such as stages 120-3, 120-2, and 120-1) will not be able to process all samples at once and cache all the generated excitation values in the on-chip cache 140. For these stages, the samples included in the batch input data can be divided into multiple groups, with each group being a sub-batch of data. In these stages, only one group of sub-batch data can be run each time.
[0040] Figure 4 FIG. shows a flowchart of a process 330 for dividing batch input data according to some embodiments of the present disclosure.
[0041] As Figure 4 shown, at block 331, at each stage 120 of the neural network 100, the computing device 20 can determine all candidate partitioning schemes based on the number of samples included in the batch input data determined at block 310. In each candidate partitioning scheme, the number of groups of sub-batch data is different.
[0042] For example, in the case where each batch input data contains 8 samples, there can be the following partitioning schemes: divided into 2 groups of sub-batch data, with each group containing 4 samples, that is, processing 4 samples each time; or divided into 4 groups of sub-batch data, with each group containing 2 samples, that is, processing 2 samples each time.
[0043] At block 332, the computing device 20 can also select a partitioning scheme from all candidate partitioning schemes based on the feature map size and the hardware channel width of each stage 120, and the selected partitioning scheme makes the size of each group of sub-batch data an integer multiple of the hardware channel width. Here, the hardware channel width refers to the number of hardware threads that the computing core can support, and this number can be 4, 8, 16, etc. In this article, the case where the hardware channel width is 8 is used as an example for description.
[0044] For stage 120-1, its feature map size is 56x56. Therefore, both candidate schemes can make the size of the sub-batch data an integer multiple of the hardware channel width. However, if 4 samples are processed each time, the generated excitation values cannot be completely cached in the on-chip cache 140. Therefore, at block 332, the partitioning scheme of processing 2 samples each time and running 4 times to process 8 samples should be selected, that is, divided into 4 groups of sub-batch data, with each group containing 2 samples.
[0045] For stage 120-2, the size of the feature map is 28x28. Whether processing 2 samples or 4 samples each time, the generated excitation values can be fully cached in the on-chip cache 140. However, considering that 28 is not an integer multiple of 8. To make the amount of data processed each time an integer multiple of 8, two samples need to be concatenated together. Therefore, the division method of processing 4 samples each time and running 2 times to process 8 samples should be selected, that is, divided into 2 groups of sub-batch data, with each group containing 4 samples.
[0046] Similarly, for stage 120-3, the division method of processing 8 samples each time and running 1 time to process 8 samples can be selected, that is, divided into 1 group of sub-batch data, with each group containing 8 samples.
[0047] In this way, for Figure 1 the example of Table 1, the way to divide the batch input data containing 8 samples into sub-batch data can be as shown in Table 2 below:
[0048] Table 2
[0049] Stage 120-1 120-2 120-3 120-4 Feature map size 56x56 28x28 14x14 7x7 Sub-batch data (number of samples) 2 4 8 8 Number of sub-batch data groups (number of runs) 4 2 1 1
[0050] In this way, the samples of the batch input data can be divided into appropriate groups, and at each stage 120 of the neural network 100, only one group of sub-batch data is processed each time, and the generated excitation values can be fully cached in the on-chip cache 140.
[0051] Furthermore, in some embodiments of the present disclosure, the sample dependence characteristics of micro-batch processing can be further used to divide each group of sub-batch data into micro-batch data.
[0052] Continuing to refer to Figure 4 , at block 333, the computing device 20 determines whether the current stage 120 is the stage with the smallest feature map size of the neural network 100. For example, as Figure 1 shown in Table 1, it can be determined whether the current stage is the stage 120-4 with the smallest feature map size.
[0053] If it is determined that this stage is not the stage with the smallest feature map size of the neural network 100 (the judgment at block 333 is "no"), for example, the current stage is stage 120-1, 120-2 or 120-3 of the neural network 100, then at block 334, the computing device 20 can divide each group of sub-batch data into at least two groups of micro-batch data, and the samples in each group of micro-batch data have a dependence relationship in two adjacent stages.
[0054] For example, for stage 120-1, each group of sub-batch data (containing 2 samples) can be divided into 2 groups of micro-batch data, with each group of micro-batch data containing 1 sample.
[0055] For stage 120-2, each group of sub-batch data (containing 4 samples) can be divided into 2 groups of mini-batch data, and each group of mini-batch data contains 2 samples.
[0056] For stage 120-3, each group of sub-batch data (containing 8 samples) can be divided into 2 groups of mini-batch data, and each group of mini-batch data contains 4 samples.
[0057] On the other hand, if it is determined that this stage is the stage with the smallest feature map size of neural network 100 (the judgment in block 333 is "yes"), for example, the current stage is stage 120-4 of neural network 100, then in block 335, computing device 20 can use each group of sub-batch data as a group of mini-batch data.
[0058] In this way, for Figure 1 and the example in Table 1, the way to divide the batch input data containing 8 samples into mini-batch data can be as shown in Table 3 below:
[0059] Table 3
[0060] Stage 120-1 120-2 120-3 120-4 Feature map size 56x56 28x28 14x14 7x7 Number of mini-batch data groups 2 2 2 1 Size of mini-batch data (number of samples) 1 2 4 8 Number of sub-batch data groups (number of runs) 4 2 1 1
[0061] In this way, the dependency relationship between samples in different stages can be utilized as much as possible to achieve stream processing.
[0062] Furthermore, when processing each group of mini-batch data, the size of each group of mini-batch data and the hardware channel width can also be considered to splice the mini-batch data to utilize the hardware channel width as much as possible.
[0063] Continuing to refer to Figure 4 , in block 336, for each group of mini-batch data in each stage, computing device 20 can determine whether the feature map size of this stage is a multiple of the hardware channel width of the computing core.
[0064] If it is determined that the feature map size of this stage is not a multiple of the hardware channel width of the computing core, then computing device 20 can splice the samples in this group of mini-batch data so that the size of the spliced mini-batch data is a multiple of the hardware channel width. The specific splicing method is as described above.
[0065] Note that in the above combination Figure 4In the description of [the operation of block 330], although it is described in terms of first dividing the mini-batch data (blocks 332 to 335) and then concatenating the input data to a multiple of the hardware channel width (blocks 336 to 337), in some embodiments, these two aspects are optional and, depending on the actual situation, only one of these two aspects may be achievable. In such a case, block 330 may further include an evaluation process. Specifically, a cost model or actual measurement can be used to evaluate the performance data of only implementing the aspect of dividing the mini-batch data (blocks 332 to 335) and the aspect of concatenating the input data (the divided mini-batch data or the undivided sub-batch data) to a multiple of the hardware channel width (blocks 336 to 337), and only the aspect with the best performance is selected for implementation.
[0066] Continue Figure 3 , at block 340, computing device 20 may sequentially perform batch normalization processing and activation processing on each group of sub-batch data (e.g., via corresponding batch normalization layer 126 and activation layer 128).
[0067] In the case where the sub-batch data is further divided into mini-batch data, computing device 20 may perform batch normalization processing and activation processing on each concatenated mini-batch data at each stage.
[0068] Figure 5 FIG. 500 shows a timing diagram of processing samples in neural network 100 according to some embodiments of the present disclosure. Here, it is assumed that each batch of input data includes 8 samples, and the feature map sizes, the number of mini-batch data groups, sizes, and the number of sub-batch data groups at each stage 120 are as shown in Table 3.
[0069] As Figure 5 shown, stage 120-1 runs 4 times (i.e., blocks 501, 502, 504, and 505), processing two samples each time; stage 120-2 runs 2 times (i.e., blocks 506 and 507), processing four samples each time; stage 120-3 runs 1 time (i.e., block 507), processing eight samples each time; stage 120-4 runs 1 time (i.e., block 508), processing eight samples each time.
[0070] Furthermore, after each run of each stage 120, the excitation values generated by that run need to be cached in on-chip cache 140 for use by the next stage. The excitation values cached in on-chip cache 140 that have not been read by the next stage will be stored in off-chip memory L2. To enable the fastest possible data reading when the next stage makes a call, the processing order of the mini-batch data at each stage can be adjusted to maximize the number of times of reading excitation values from on-chip cache 140.
[0071] Therefore, if the number of groups of the sub-batch data in stage 120 (i.e., the number of runs in this stage) is greater than 2, the operation can directly enter the next stage after every 2 runs in stage 120. For example, as Figure 5 shown in, the number of groups of the sub-batch data in stage 120-1 is 4 (i.e., this stage needs to run 4 times), then the operation can enter the next stage 120-2 after the operation of block 502. Specifically, after the operation of block 501 is completed, the generated excitation value is written to the off-chip memory L2, and after the operation of block 502 is completed, the generated excitation value is cached in the on-chip cache 140. Thus, when the next stage 120-2 is running (block 503), the operation result of block 501 can be read from the off-chip memory L2, and the operation result of block 502 can be read from the on-chip cache 140. Similarly, after the operations of blocks 504 and 505, the excitation value generated by the operation of block 504 is stored in the off-chip memory L2, and the excitation value generated by the operation of block 505 is cached in the on-chip cache 140. Thus, when the next stage 120-2 is running (block 506), the operation result of block 504 can be read from the off-chip memory L2, and the operation result of block 505 can be read from the on-chip cache 140.
[0072] In this way, when stage 120-2 is running, data needs to be read twice from the off-chip memory L2 and twice from the on-chip cache 140. Compared with entering the operation of stage 120-2 after stage 120-1 runs completely 4 times, the number of times of reading data from the off-chip memory L2 is reduced by one, thereby further shortening the processing time.
[0073] The method 300 described above can be executed, for example, by the processor 210 of the computing device 20. For example, in some embodiments, the method 300 can be implemented as a computer software program, which is tangibly included in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computing device 20 via the memory 220 to perform one or more operations of the method 300 described above.
[0074] Those skilled in the art can understand that Figure 2 the computing device 20 shown is only illustrative. In some embodiments, the computing device 20 may include more or fewer components.
[0075] The method 300 for processing samples in the neural network 100 according to the present disclosure and the computing device 20 capable of executing this method have been described above in conjunction with the accompanying drawings. However, those skilled in the art can understand that the execution of the steps of the method 300 is not limited to the order shown in the figures and described above, but can be executed in any other reasonable order. In addition, the computing device 20 does not necessarily include Figure 2All components shown may include only some or more of the components necessary to perform the functions described in this disclosure, and the connection manners of these components are not limited to the forms shown in the figures.
[0076] This disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of this disclosure.
[0077] In one or more exemplary designs, the functions described in this disclosure may be implemented using hardware, software, firmware, or any combination thereof. For example, if implemented using software, the functions may be stored on a computer-readable medium as one or more instructions or codes, or transmitted as one or more instructions or codes on a computer-readable medium.
[0078] Each unit of the apparatus disclosed herein may be implemented using discrete hardware components or may be integrally implemented on a hardware component, such as a processor. For example, a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination for performing the functions described herein may be used to implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure.
[0079] Those of ordinary skill in the art should also understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments of this disclosure may be implemented as electronic hardware, computer software, or a combination of the two.
[0080] The foregoing description of this disclosure enables any ordinary skill in the art to implement or use this disclosure. For those of ordinary skill in the art, various modifications to this disclosure are obvious, and the general principles defined herein may be applied to other variations without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.
Claims
1. A method for processing samples in a neural network, comprising: Determine the number of samples included in a batch of input data based on the minimum feature map size in multiple stages of the neural network and the size of the on-chip cache of the computing core for implementing the neural network, so that the generated excitation values can be fully cached into the on-chip cache at the stage corresponding to the minimum feature map size; Receive a batch of input data containing the number of samples at the first stage of the neural network; At each stage of the neural network, divide the batch of input data into one or more groups of sub-batch data based on the feature map size of the stage and the hardware channel width of the computing core; and Perform batch normalization processing and activation processing on each group of sub-batch data in sequence.
2. The method according to claim 1, wherein the neural network includes a computer vision network, and the feature map sizes of the multiple stages of the computer vision network gradually decrease or first decrease and then increase.
3. The method according to claim 1, wherein dividing the batch input data into one or more groups of sub-batch data includes: At each stage of the neural network, determine all candidate partitioning methods based on the number of samples included in the batch of input data, where the number of groups of sub-batch data is different in each candidate partitioning method; and Based on the feature map size of the stage and the hardware channel width, select a partitioning method from all the candidate partitioning methods, where the selected partitioning method makes the size of each group of sub-batch data an integer multiple of the hardware channel width.
4. The method according to claim 3, further comprising: Determine whether the stage is the stage with the smallest feature map size of the neural network; If it is determined that the stage is not the stage with the smallest feature map size of the neural network, divide each group of sub-batch data into at least two groups of micro-batch data, and the samples in each group of micro-batch data have a dependency relationship in two adjacent stages; and If it is determined that the stage is the stage with the smallest feature map size of the neural network, regard each group of sub-batch data as a group of micro-batch data.
5. The method according to claim 4, further comprising: For each group of micro-batch data at each stage, determine whether the feature map size of the stage is a multiple of the hardware channel width of the computing core; If it is determined that the feature map size of the stage is not a multiple of the hardware channel width of the computing core, splice the samples in this group of micro-batch data so that the size of the spliced micro-batch data is a multiple of the hardware channel width.
6. The method according to claim 4, wherein sequentially performing batch normalization processing and activation processing on each group of sub-batch data includes: Perform batch normalization processing and activation processing on each spliced micro-batch data in each group of sub-batch data in sequence.
7. The method according to claim 1, further comprising: At each stage of the neural network, cache the excitation values generated each time activation processing is performed into the on-chip cache.
8. The method according to claim 7, further comprising: If the number of groups of sub-batch data in a stage is greater than two, directly enter the operation of the next stage after running twice in this stage.
9. A computing device, comprising: At least one processor; And At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, and when the instructions are executed by the at least one processor, the computing device is caused to execute the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, on which computer program code is stored, and the computer program code, when run, executes the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Model operation method and device, electronic equipment and storage medium
CN112862074A