Data processing method and device, neural network processing device
By receiving and splicing the combined output feature data through the processing unit array of the neural network processor, the inefficiency problem of traditional processors when processing highly parallel neural networks is solved, and more efficient data flow and bus bandwidth utilization are achieved.
Patent Information
- Application Number
- CN202111590435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-12-23
AI Technical Summary
When traditional CPUs and GPUs process neural networks, they find it difficult to effectively handle the computational characteristics of large parameters, large computational complexity, and high parallelism, resulting in low system efficiency.
A neural network processor (NPU) is used to receive and combine the output feature data subsets through the processing unit array, and then reorganize them and write them into the storage device to optimize data flow and bus bandwidth utilization.
It improves the system efficiency of the neural network processor, enhances the bus bandwidth utilization, takes into account the read and write efficiency between the front and back network layers, and optimizes the mapping of data format and shape.
Smart Images

Figure CN114330687B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to a data processing method and apparatus, and a neural network processing apparatus. Background Art
[0002] Artificial intelligence (AI) is a cutting-edge, interdisciplinary discipline that integrates computer science, statistics, neuroscience, and social science. AI applications include robotics, speech recognition, image recognition, natural language processing, and expert systems. Currently, deep learning technology has achieved remarkable results in applications such as image recognition, speech recognition, and autonomous driving. Deep learning techniques, such as convolutional neural networks (CNNs), deep neural networks (DNNs), and recurrent neural networks (RNNs), all of which are characterized by a large number of parameters, high computational complexity, and a high degree of parallelism.
[0003] However, the traditional method of using CPU, GPU, etc. to process neural networks is not suitable for the above-mentioned computing characteristics of large number of parameters, large amount of calculation, and high degree of parallelism. Therefore, it becomes very necessary to design a dedicated processor for the field of deep learning. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a data processing method for a neural network, wherein the neural network includes multiple network layers, the multiple network layers including a first network layer, and the method includes: receiving at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer from the processing unit array; splicing and combining the at least two output feature data subsets according to positions of the at least two different parts in the first network layer to obtain a recombined data subset; and writing the recombined data subset to a destination storage device.
[0005] At least one embodiment of the present disclosure provides a data processing device for a neural network, wherein the neural network includes multiple network layers, including a first network layer. The device includes a receiving unit, a reassembly unit, and a writing unit. The receiving unit is configured to receive, from a processing unit array, at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer; the reassembly unit is configured to concatenate and combine the at least two output feature data subsets based on the positions of the at least two different parts in the first network layer to obtain a reassembled data subset; and the writing unit is configured to write the reassembled data subset to a destination storage device.
[0006] At least one embodiment of the present disclosure provides a data processing device for a neural network, comprising a processing unit and a memory, wherein the memory stores one or more computer program modules; the one or more computer program modules are configured to execute the above data processing method when executed by the processing unit.
[0007] At least one embodiment of the present disclosure provides a non-transitory readable storage medium, wherein the non-transitory readable storage medium stores computer instructions, wherein when the computer instructions are executed by a processor, the above data processing method is performed.
[0008] At least one embodiment of the present disclosure provides a neural network processing device, comprising the above-mentioned data processing device, a processing unit array, and a destination storage device, wherein the data processing device is coupled to the processing unit array and the destination storage device. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0010] Figure 1A Abstractly showing the input and output of a neuron in a convolutional neural network;
[0011] Figure 1B A schematic diagram showing a convolutional layer performing multi-channel convolution operations;
[0012] Figure 1C It shows the characteristics of data reuse in convolution operation;
[0013] Figure 2A A schematic diagram of the architecture of a neural network processor is shown;
[0014] Figure 2B Shown for Figure 2A Three different data multiplexing methods employed by the processing unit array shown;
[0015] Figure 2C Shown for Figure 2A The exemplary mapping method used by the processing unit array shown is
[0016] Figure 3 A schematic diagram of a neural network processing device provided according to at least one embodiment of the present disclosure is shown;
[0017] Figure 4 A schematic diagram illustrating a data processing method for a neural network provided according to at least one embodiment of the present disclosure is shown;
[0018] Figure 5 A schematic diagram of a neural network processing device provided according to another embodiment of the present disclosure is shown;
[0019] Figure 6 A schematic diagram showing a data processing device for a neural network according to another embodiment of the present disclosure is shown;
[0020] Figure 7 A schematic diagram showing coupling of a data processing device and a data bus according to at least one embodiment of the present disclosure is shown;
[0021] Figure 8 A schematic diagram of a data processing device provided according to at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0022] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0023] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning understood by persons of ordinary skill in the field to which the present disclosure pertains. The terms "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are merely used to distinguish between different components. Similarly, terms such as "include" or "comprise" and the like mean that the elements or objects preceding the term encompass the elements or objects listed following the term and their equivalents, without excluding other elements or objects. Terms such as "connect" or "connected" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0024] In order to make the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and components.
[0025] Neural networks are mathematical computational models inspired by the structure of neurons in the brain and the principles of neural transmission. The implementation of intelligent computing based on these models is called brain-inspired computing. For example, neural networks include various network structures, such as back propagation (BP) neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs). Convolutional neural networks can also be further divided into fully convolutional networks, deep convolutional networks, and U-nets.
[0026] For example, a common convolutional neural network usually includes an input end, an output end, and multiple processing layers. For example, the input end is used to receive data to be processed, such as an image to be processed, and the output end is used to output processing results, such as a processed image. These processing layers may include convolution layers, pooling layers, batch normalization layers (Batch Normalization, abbreviated as BN), fully connected layers, etc. Depending on the structure of the convolutional neural network, the processing layers may include different contents and combinations. After the input data is input into the convolutional neural network, it passes through several processing layers to obtain the corresponding output. For example, the input data can pass through several processing layers to complete operations such as convolution, upsampling, downsampling, standardization, full connection, and flattening.
[0027] The convolutional layer is the core layer of a convolutional neural network. It applies several filters to the input data (input image or input feature map), which is then used to extract various types of features. The result of applying a filter to the input data is called a feature map, and the number of feature maps is equal to the number of filters. The feature map output by a convolutional layer can be fed into the next convolutional layer for further processing to produce a new feature map. The pooling layer is an intermediate layer sandwiched between consecutive convolutional layers, used to reduce the size of the input data and, to a certain extent, mitigate overfitting. There are many ways to implement pooling, including but not limited to max-pooling, avg-pooling, random pooling, decimation (e.g., selecting fixed pixels), and demux (splitting the input image into multiple smaller images). Typically, the last subsampling layer or convolutional layer is connected to one or more fully connected layers, the output of which is the final output, resulting in a one-dimensional matrix, or vector.
[0028] Figure 1AThe input and output of a neuron in a convolutional neural network are abstractly shown. As shown in FIG1A , C1, C2 to Cn refer to different signal channels. For a certain local receptive field (the local receptive field contains multiple channels), different filters are used to convolve the data on the C1 to Cn signal channels of the local receptive field. The convolution result is input into the stimulation node, and the stimulation node performs calculations according to the corresponding function to obtain feature information. For example, a convolutional neural network is typically a deep convolutional neural network and may include at least five convolutional layers. For example, the VGG-16 neural network has 16 layers, and the GoogLeNet neural network has 22 layers. Of course, other neural network structures may have more processing layers. The above content is only an exemplary introduction to the neural network, and the present disclosure does not limit the structure of the neural network.
[0029] Figure 1B Figure 2 shows a schematic diagram of a convolutional layer performing multi-channel convolution operations. Figure 1B As shown in the figure, M groups of R×S convolution kernels with C channels are used to perform convolution operations on N groups of H×W input images (or input feature maps) with C channels, and N groups of E×F output feature maps with M channels are obtained respectively. Therefore, the output feature maps generally include multiple dimensions of F / E / M.
[0030] Convolution operations are characterized by high parallelism and high data reuse. High parallelism is reflected in the fact that multiple convolution kernels can operate on multiple input feature maps simultaneously. Figure 1C It shows the characteristics of data reuse in convolution operation. Figure 1C As shown in Figure 2, high data reuse in convolution operations is reflected in the following aspects:
[0031] (a) Reuse of convolution operations: a convolution kernel operates on multiple pixels in an input feature map (e.g., a multi-channel input feature map);
[0032] (b) Input feature map reuse: one input feature map is operated with multiple convolution kernels (e.g., multi-channel convolution kernel);
[0033] (c) Convolution kernel reuse: a convolution kernel (e.g., a multi-channel convolution kernel) operates on multiple input feature maps (e.g., multi-channel input feature maps).
[0034] Moreover, due to the large amount of computation required by neural networks, especially for convolutional layers with large-scale input feature maps, it is often necessary to decompose the computational operations of a convolutional layer in the neural network. For example, the convolution operations of different parts of the same convolutional layer can be performed independently of each other. These decomposed tasks are assigned to multiple processing units to perform calculations in parallel. The calculation results of these processing units are then merged to obtain the calculation results of the entire convolutional layer. The calculation results of this network layer can then be used as the input of the next convolutional layer.
[0035] A neural-network processing unit (NPU) is a type of microprocessor or computing system dedicated to hardware acceleration of artificial intelligence (especially artificial neural networks, machine vision, machine learning, etc.), sometimes also called an artificial intelligence accelerator (AI Accelerator).
[0036] Figure 2A A schematic diagram of the architecture of a neural network processor is shown. Figure 2A As shown, the neural network processor includes a processing element (PE) array 110, a global cache 120, and a memory 130. The processing element array 110 includes multiple rows and columns (e.g., 12 rows × 12 columns) of processing elements, which are coupled to each other through on-chip interconnection and share a global cache 120. The on-chip interconnection is, for example, a network on chip (NoC). Each processing element has a computing function and, for example, may also have its own local cache, such as a cache or register array including a multiply-accumulator (MAC) and a vector (or matrix) for caching inputs. Each PE can access other PEs around it, the PE's own local cache, and the global cache. The global cache 120 is further coupled to the memory 130, for example, via a bus.
[0037] During operation, for example, the convolution kernel (Flt) and input feature map (Ifm) data required for calculations of a network layer (e.g., a convolutional layer) are read from memory 130 into global cache 120. The convolution kernel (Flt) and input image (Img) are then input from global cache 120 into processing unit array 110 for calculation. Computational tasks for different image pixels are assigned to different processing units (i.e., mapping is performed). The partial cumulative sum (Psum1) generated during the calculation process is temporarily stored in the global cache. If a subsequent calculation requires further accumulation of the previously generated partial cumulative sum (Psum1), the required partial cumulative sum (Psum2) can be read from global cache 120 and returned to processing unit array 110 for calculation. The output feature map (Ofm) obtained after completing the calculations of a convolutional layer can be output from global cache 120 to memory 130 for storage, for example, to be used for calculations of the next network layer (e.g., a convolutional layer).
[0038] For example, data generated by processing unit array 110, particularly when involving sparse matrices, can be compressed and stored. One compression method for sparse matrices is RLC coding, which encodes consecutive zeros as the number of zeros, thereby saving storage space. When data is stored from processing unit array 110 into memory 130, an encoder (not shown) can be used to compress the data. Correspondingly, when data is read from memory 130 into processing unit array 110, a decoder (not shown) can be used to decompress the data.
[0039] Figure 2B Shown for Figure 2A The processing unit array shown in FIG3 employs three different data reuse methods. The neural network processor may, for example, employ a row stationary (RS) data flow. The row stationary data flow RS can reduce the movement of all data types (e.g., Ifm, Flt, and Psum / Ofm) and has a higher data reuse rate. In particular, for the different properties of the convolution kernel, the input image (or input feature map), and the accumulated sum, such as Figure 2B As shown in the figure, a 5×5 input feature map, a 3×3 convolution kernel, and a 3×3 output feature map are used as an example. Figure 2A The processing unit array shown uses three different data multiplexing methods:
[0040] (a) The weight data of the convolution kernel is horizontally reused between PEs;
[0041] (b) The data of the input feature map is diagonally multiplexed among PEs;
[0042] (c) Output row data (accumulated sum) is multiplexed vertically between PEs (vertical accumulation).
[0043] More specifically, Figure 2BIn the 3×3 PE array, each PE can, for example, perform a one-dimensional convolution operation. As shown in the figure, the weight data of the convolution kernel is multiplexed horizontally between PEs. The three processing units (PE1.1, PE 1.2, and PE 1.3) in the first row of the PE array are respectively input to the first row of the convolution kernel; the three processing units in the second row of the PE array are respectively input to the second row of the convolution kernel; and the three processing units in the third row of the PE array are respectively input to the second row of the convolution kernel. The data of the input feature map is multiplexed diagonally between PEs. The first row of the input feature map is input to PE 1.1; the second row of the input feature map is input to PE 2.1 and PE 1.2; ..., and so on. Correspondingly, the output row data (accumulated sum) is multiplexed vertically between PEs. The first row of the accumulated sum (output feature map) is output in the first column of the PE array; the second row of the accumulated sum is output in the second column of the PE array; and the third row of the accumulated sum is output in the third column of the PE array. In the above data reuse method, the mapping relationship between the input feature map, convolution kernel and output feature map and the PE array is different.
[0044] In addition, although the processing element (PE) array 110 has multiple rows and columns of processing units, such as 12 rows × 12 columns of processing units, relative to some larger network layers or smaller network layers, or network layers with multiple channel outputs, in order to perform calculations, the input feature maps and / or convolution kernels of the network layers need to be split and mapped separately.
[0045] Figure 2C Shown for Figure 2A The exemplary mapping method used by the processing unit array shown in FIG. Figure 2C As shown on the left side of the figure, for a single-channel 5-row × 10-column input feature map (Ifm), the image pixels of the input feature map can be directly mapped one by one to the processing elements of the processing element array. At the same time, to parallelize the operation, the input feature map can be replicated (repeated) once vertically, or input feature maps of other channels can be input, thereby improving the utilization of the PE. However, there are still idle processing elements in the row and column directions.
[0046] like Figure 2C As shown on the right side of the figure, for a single-channel 5-row × 20-column input feature map (Ifm), it needs to be split horizontally into two sub-input feature maps Ifm-1 and Ifm-2 with sizes of 5 rows × 12 columns and 5 rows × 8 columns respectively, so as to be input into the processing unit array 110 and mapped to the processing units of the processing unit array accordingly. At this time, there are still idle processing units in the column direction.
[0047] There are significant differences in the various parameters after mapping between different neural network algorithms or different network layers of the same neural network algorithm. For example, refer again to Figure 1B , the output feature map always includes multiple dimensions of F / E / M, so in order to accurately locate a pixel point, it is necessary to know the coordinate value of this pixel point in the F / E / M dimension. At the same time, refer again Figure 1A The shapes (length, width and channels) of data feature maps are different between different network layers. The output feature map of a network layer is written to the memory after calculation. However, in the same network layer, even different parts of the output feature map corresponding to adjacent and continuous parts may be written to non-continuous storage locations in the memory due to different times generated during the calculation process or corresponding to spaced PE units. When the next network layer is to be calculated and the output feature map is read from the memory as input image data, it is necessary to continuously jump addresses to read these non-continuous storage locations. On the other hand, the bandwidth requirements of neural network processors are usually very high, so a large bit width (such as 512 bits) bus is usually used. If the address jump is too frequent, it means that a large amount of unnecessary data may appear in the large bit width bus, which in turn greatly reduces the effective bandwidth, reduces system efficiency and increases power consumption.
[0048] At least one embodiment of the present disclosure provides a data processing method for a neural network, wherein the neural network includes multiple network layers, the multiple network layers including a first network layer, and the method includes: receiving at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer from the processing unit array; splicing and combining the at least two output feature data subsets according to positions of the at least two different parts in the first network layer to obtain a recombined data subset; and writing the recombined data subset to a destination storage device.
[0049] At least one embodiment of the present disclosure provides a data processing device for a neural network, wherein the neural network includes multiple network layers, including a first network layer. The device includes a receiving unit, a reassembly unit, and a writing unit. The receiving unit is configured to receive, from a processing unit array, at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer; the reassembly unit is configured to concatenate and combine the at least two output feature data subsets based on the positions of the at least two different parts in the first network layer to obtain a reassembled data subset; and the writing unit is configured to write the reassembled data subset to a destination storage device.
[0050] At least one embodiment of the present disclosure provides a data processing device for a neural network, comprising a processing unit and a memory, wherein the memory stores one or more computer program modules; the one or more computer program modules are configured to execute the above data processing method when executed by the processing unit.
[0051] At least one embodiment of the present disclosure provides a non-transitory readable storage medium, wherein the non-transitory readable storage medium stores computer instructions, wherein when the computer instructions are executed by a processor, the above data processing method is performed.
[0052] At least one embodiment of the present disclosure provides a neural network processing device, comprising the above-mentioned data processing device, a processing unit array, and a destination storage device, wherein the data processing device is coupled to the processing unit array and the destination storage device.
[0053] The processing method and processing device for a neural network provided by at least one embodiment of the present disclosure can improve system efficiency and increase bus bandwidth utilization; and, in at least one embodiment, can simultaneously take into account the read and write efficiency between the front and rear network layers in the neural network, thereby allowing each network layer to be mapped according to the optimal processing unit utilization when the data formats or shapes of the upper network layer and the lower network layer of the neural network are independent of each other.
[0054] Several embodiments of the processing method and processing device for a neural network disclosed in the present invention will be described below.
[0055] Figure 3 A schematic diagram of a neural network processing device provided according to at least one embodiment of the present disclosure is shown.
[0056] like Figure 3 As shown, the neural network processing device includes one or more processing element (PE) arrays 210 , a global cache 220 , a storage device 330 and a reorganization processing device 240 .
[0057] For example, the neural network processing device may include multiple processing unit arrays 210. These processing unit arrays 210 may share a global cache 220, a storage device 230, and a reassembly processing device 240, or each processing unit array 210 may be provided with a global cache 220, a storage device 230, and a reassembly processing device 240. For example, each processing unit array 210 includes multiple rows and columns (e.g., 12 rows x 12 columns or other dimensions) of processing units. These processing units are coupled to each other via an on-chip interconnect, such as a network-on-chip (NoC), and share a global cache 220. For example, each processing unit has computational functionality, such as an arithmetic logic unit (ALU), and may also have its own local cache. Each processing unit (PE) can access other surrounding PEs, its own local cache, and the global cache. The global cache 220 is further coupled to a storage device 230, such as a memory or other storage device, via a data bus (arrows in the figure). The memory may be a dynamic random access memory (DRAM). The embodiments of the present disclosure have no limitation on the type of the data bus, for example, it can be a PCIE bus.
[0058] During operation, for example, the convolution kernel (Flt) and input feature map (Ifm) data required for calculations in a network layer (e.g., a convolutional layer) are read from storage device 230 into global cache 220. The convolution kernel (Flt) and input image (Img) are then transferred from global cache 220 to processing unit array 210 for calculations. Computational tasks for different image pixels are assigned to different processing units. The partial cumulative sum (Psum1) generated during the calculation process is temporarily stored in global cache 220. If a subsequent calculation requires further accumulation of the previously generated partial cumulative sum (Psum1), the required partial cumulative sum (Psum2) can be read from global cache 220 into processing unit array 210. The output feature map (Ofm) obtained after completing the calculation of a convolution layer can be output from the global cache 220 to the reorganization processing device 240, in which different parts of the received output feature map are reorganized, and then the reorganized output feature map is written to the storage device 230 for storage through, for example, a data bus, and the stored reorganized output feature map is used, for example, for the calculation of the next network layer (for example, a convolution layer).
[0059] For example, the neural network processing device may adopt a row stationary (RS) data stream. The row stationary data stream RS can reduce the movement of all data types (such as Ifm, Flt and psum / Ofm) and has a higher data reuse rate. For this, please refer to Figure 2B-2CEtc., where the mapping relationship between the input feature map, convolution kernel and output feature map and the PE array is different and will not be repeated here.
[0060] Similarly, for example, data generated by the processing unit array 210, particularly when involving sparse matrices, can be compressed and stored. One compression method for sparse matrices is RLC coding, which can encode consecutive zeros as the number of zeros, thereby saving storage space. When data is stored from the processing unit array 210 into the storage device 230, an encoder (not shown) can be used to compress the data. Correspondingly, when data is read from the storage device 230 into the processing unit array 210, a decoder (not shown) can be used to decompress the data.
[0061] The neural network that the neural network processing device is configured to process may include multiple network layers, and the multiple network layers include a first network layer. For example, the neural network may be a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), etc., and the embodiments of the present disclosure are not limited to this. For example, the first network layer may be an input layer, a convolution layer, a pooling layer, etc. The input data of the network layer (such as an input image or an input feature map) may be single-channel or multi-channel, and the output data of the network layer may be single-channel or multi-channel. For example, the size of the input data of the network layer (such as rows × columns) is larger than the size of the processing unit array (such as rows × columns), so in order to perform calculations, the input feature map and / or convolution kernel of the network layer needs to be split and mapped accordingly.
[0062] Due to the aforementioned splitting and mapping, different portions of the output feature map received by the reassembly processing device 240 from the global cache 220 may no longer be continuous due to arriving at the global cache 220 at different times or using different locations in the processing unit array. The reassembly processing device 240 can reassemble these portions to make them continuous again, allowing them to be continuously written to the storage device 230 via the bus. This can improve the system efficiency of the neural network processor and increase bus bandwidth utilization. Furthermore, through this reassembly operation, in at least one embodiment, even if the input feature maps of the two network layers have different shapes, the read and write efficiency between the two network layers in the neural network can be maintained.
[0063] Figure 4 A schematic diagram of a data processing method for a neural network provided according to at least one embodiment of the present disclosure is shown.
[0064] For example, the data processing method can be executed by the reorganization processing device 240 of the above-mentioned neural network processing device. Figure 4 As shown, and including steps S101 to S102:
[0065] Step S101: receiving from a processing unit array at least two subsets of output feature data obtained by the processing unit array for at least two different parts of a first network layer;
[0066] Step S102: combining at least two output feature data subsets according to positions of at least two different parts in the first network layer to obtain a recombined data subset;
[0067] Step S103: writing the reorganized data subset into the destination storage device.
[0068] For example, the processing unit array may be as follows: Figure 3 The processing unit array 210 is shown, and the destination storage device can be Figure 3 As described above, the neural network involved in the method may include multiple network layers, and the multiple network layers include a first network layer. The first network layer is mapped to the processing unit array during the calculation process, and at least two output feature data subsets are obtained for at least two different parts of the first network layer.
[0069] For example, the at least two different parts of the first network layer may be at least two different sub-matrices, or at least two different data elements (e.g., image pixels), thereby obtaining at least two corresponding output feature data subsets after calculation. The output feature data subsets may be, for example, sub-matrices in an output feature map, or data elements (e.g., feature map pixels). The embodiments of the present disclosure are not limited in this regard.
[0070] When splicing and combining at least two output feature data subsets, the at least two output feature data subsets are spliced and combined according to the positions of the at least two different parts in the first network layer to obtain a recombined data subset. Figure 2B , the mapping relationship between the input feature map, the convolution kernel and the output feature map and the PE array is different. When the output feature map calculated by the current network layer is used as the input for calculation by the next network layer, the mapping relationship between the output feature map and the PE array will change; for another example, as shown on the right side of reference Figure 2C, the elements in the 12th and 13th columns that originally belonged to the same row in the input feature map are separated and no longer continuous. After the processing unit array performs calculations, the output feature data obtained by them also become no longer continuous. Therefore, through the method of at least one embodiment of the present disclosure, the output feature data obtained by them are spliced and combined (reorganized), thereby, for example, becoming continuous.
[0071] For example, in at least one example of the embodiments of the present disclosure, step S101, i.e., receiving from the processing unit array at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer, a specific example may include: obtaining at least two output feature data subsets from different processing units of the processing unit array in the same operation cycle or different operation cycles, or obtaining at least two output feature data subsets from the same processing unit in different operation cycles; performing address mapping according to a mapping pattern of different processing units of the processing unit array or the same processing unit and the input feature map of the first network layer to the processing unit array, to obtain the positions of the at least two different parts in the first network layer.
[0072] For example, the mapping mode of the input feature map of the first network layer to the processing unit array can be seen in Figure 2C (But not limited to this case), according to the size of the processing unit array and the input feature map, the input feature map can be split and mapped separately. For example, different parts of the input feature map can also be mapped to multiple different processing unit arrays of the neural network processor.
[0073] For example, in at least one example of the embodiments of the present disclosure, the above-mentioned step S101, i.e., receiving from the processing unit array at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer, a specific example thereof may also include: in response to obtaining at least two output feature data subsets in different operation cycles, caching at least two output feature data subsets for splicing and combining.
[0074] When at least two output feature data subsets are obtained in different operation cycles, in order to be able to reassemble them, it is usually necessary to first cache the output feature data subset that arrived first, and then splice and combine them. For example, by analyzing the position (coordinates) of the output feature data subset that arrived first in the output feature map, it can be determined that the output feature data subset is discontinuous with the currently cached output feature data subset and therefore needs to be reassembled with other output feature data subsets. In this case, the output feature data subset can be cached.
[0075] For example, in at least one example of the embodiments of the present disclosure, step S102, that is, splicing and combining at least two output feature data subsets according to the positions of at least two different parts in the first network layer to obtain a recombined data subset, a specific example may include: performing coordinate scanning according to the positions of at least two different parts in the first network layer to obtain the coordinates of the at least two output feature data subsets in the output feature map of the first network layer; splicing and combining the at least two output feature data subsets according to the result of the coordinate scanning to obtain the recombined data subset.
[0076] As described above, in order to analyze the position (coordinates) of the output feature data subset in the output feature map, a coordinate scan can be performed based on the position of the output feature data subset in the first network layer. Then, based on the results of the coordinate scan, at least two output feature data subsets can be spliced and combined to obtain a recombined data subset.
[0077] For example, in at least one example of the embodiments of the present disclosure, the multiple network layers also include a second network layer adjacent to the first network layer, and the output feature map of the first network layer serves as the input feature map of the second network layer. In this case, splicing and combining at least two output feature data subsets according to the positions of the at least two different parts in the first network layer to obtain a recombined data subset may include: determining the second positions of the at least two output feature data subsets in the output feature map of the first network layer according to the first positions of the at least two different parts in the first network layer, and splicing and combining the at least two output feature data subsets according to the second positions to obtain the recombined data subset.
[0078] In this example, the output feature map of the first network layer serves as the input feature map of the second network layer. Therefore, when splicing and combining at least two output feature data subsets to obtain a recombined data subset, it is necessary to refer to the shape of the input feature map, so as to make the input operation smoother. Therefore, even if the data formats or shapes of the previous network layer (first network layer) and the next network layer (second network layer) of the neural network are independent of each other, each network layer can be mapped according to the optimal processing unit utilization, so that the mapping of the previous and next network layers is decoupled.
[0079] In this example, for example, splicing and combining at least two output feature data subsets according to the second position to obtain a recombined data subset may include: in response to the at least two output feature data subsets being correspondingly arranged continuously in the input feature map of the second network layer, continuously arranging the at least two output feature data subsets to obtain the recombined data subset.
[0080] In this example, for example, splicing and combining at least two output feature data subsets according to the second position to obtain a recombined data subset may include: splicing and combining at least two output feature data subsets according to the pattern in which the input feature map of the second network layer is input to the processing unit array and the second position to obtain a recombined data subset.
[0081] In this example, for example, according to the mapping pattern and the second position of the input feature map of the second network layer being input into the processing unit array, at least two output feature data subsets are spliced and combined to obtain a recombined data subset, which may include: in response to at least two output feature data subsets being continuously input into the processing unit array according to the above pattern and the second position, at least two output feature data subsets are continuously arranged to obtain a recombined data subset.
[0082] Similarly, for example, the mapping mode of the input feature map of the second network layer to the processing unit array can be seen in Figure 2C (but not limited to this case).
[0083] For example, in at least one example of the embodiments of the present disclosure, step S103, i.e., writing the reorganized data subset into the destination storage device, may include, in a specific example: writing the reorganized data subset obtained by consecutively arranging at least two output feature data subsets into the destination storage device through a single storage operation on the data bus.
[0084] If the data bus width permits, a recombined data subset obtained by consecutively arranging at least two output feature data subsets can be written to the destination storage device in a single storage operation via the data bus. For example, if the data bus width is 512 bits (or alternatively, 256 bits, 128 bits, etc.), a single access operation can transmit 512 bits of data. If the length of the recombined data subset obtained by consecutively arranging at least two output feature data subsets is less than 512 bits, the recombined data subset can be written to the destination storage device in a single storage operation. In this case, the difference between the length of the recombined data subset and the data bus width is filled with other data.
[0085] In the above example, step S103, i.e., writing the reorganized data subset to the destination storage device, may further include obtaining the storage address and valid data length of the reorganized data subset. Correspondingly, performing a single storage operation to write to the destination storage device via the data bus may include performing a single storage operation to write to the destination storage device using the storage address and the valid data length. For example, the valid data length may be the data bus width, or 1 / 2 or 1 / 4 of the data bus width.
[0086] Figure 5 A schematic diagram of a neural network processing device provided according to another embodiment of the present disclosure is shown.
[0087] like Figure 5 As shown, the neural network processing device includes one or more processing unit arrays 310, a storage device 330, and a reorganization processing device 340, without the need to include a global cache.
[0088] For example, the neural network processing device may include multiple processing unit arrays 310. For example, these processing unit arrays 310 may share a storage device 330 and a reassembly processing device 340, or each processing unit array 210 may be provided with a storage device 330 and a reassembly processing device 340. For example, each processing unit array 310 includes multiple rows and columns (e.g., 12 rows × 12 columns or other dimensions) of processing units. These processing units are coupled to each other via an on-chip interconnect, such as a network on chip (NoC), and share a global cache 220. For example, each processing unit has a computational function, such as an arithmetic logic unit (ALU), and also has a cache BF. Each PE can access its own cache as well as other surrounding PEs and their cache BFs. In at least one example, in the processing unit array 310, multiple processing units may also be clustered, and in each cluster, multiple PEs share a cache BF. The processing unit array 310 is coupled to a storage device 330 via a data bus, for example. The storage device 330 is a memory, such as a dynamic random access memory (DRAM). This embodiment has no limitation on the type of the data bus, which may be a PCIE bus, for example.
[0089] Similarly, due to the aforementioned splitting, mapping, etc., different parts of the output feature map received by the reassembly processing device 340 from the processing unit array 310 may no longer be continuous due to different times of arrival at the cache BF or different locations of the processing unit array used. The reassembly processing device 340 may reassemble these parts so that they become continuous again, thereby enabling them to be continuously written to the storage device 330 via the bus, thereby improving system efficiency and increasing bus bandwidth utilization. Furthermore, through this reassembly operation, in at least one embodiment, even if the input feature maps of the previous and next network layers have different shapes, the read and write efficiency between the previous and next network layers in the neural network can be taken into account simultaneously.
[0090] Figure 6 A schematic diagram of a data processing device for a neural network according to another embodiment of the present disclosure is shown.
[0091] like Figure 6As shown, the data processing device 600 includes a receiving unit 610, a reassembling unit 620, and a writing unit 630. The receiving unit 610 is configured to receive, from the processing unit array, at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer; the reassembling unit 620 is configured to splice and combine the at least two output feature data subsets according to the positions of the at least two different parts in the first network layer to obtain a reassembled data subset; and the writing unit 630 is configured to write the reassembled data subset into a destination storage device.
[0092] The data processing device 600 of this embodiment is used, for example, to implement the reorganization processing device in the neural network processing device of the above-mentioned embodiment, such as the above-mentioned reorganization processing device 240 or the reorganization processing device 340.
[0093] For example, in at least one example, writing the reorganized data subset into the destination storage device includes: writing the reorganized data subset obtained by continuously arranging at least two output feature data subsets into the destination storage device through a single storage operation on a data bus.
[0094] For example, in at least one example, the write unit 630 is further configured to obtain the storage address and valid data length of the reorganized data subset, and use the storage address and valid data length to perform a single storage operation to write to the destination storage device.
[0095] For example, in at least one example, the receiving unit 610 is further configured to obtain at least two subsets of output feature data from different processing units of the processing unit array in the same operating cycle or different operating cycles, or to obtain at least two subsets of output feature data from the same processing unit in different operating cycles, and to perform address mapping according to a mapping pattern of different processing units of the processing unit array or the same processing unit and the input feature map of the first network layer input to the processing unit array to obtain the positions of at least two different parts in the first network layer.
[0096] For example, in at least one example, Figure 6 As shown, the data processing device 600 may further include a cache unit 640 , which is configured to cache at least two output feature data subsets for splicing and combining.
[0097] For example, in at least one example, Figure 6As shown, the reassembly unit 620 includes a coordinate scanning subunit 621 and a coordinate combining subunit 622. The coordinate scanning subunit 621 is configured to perform coordinate scanning according to the positions of at least two different parts in the first network layer to obtain coordinates of at least two output feature data subsets in the output feature map of the first network layer; the coordinate combining subunit 622 is configured to splice and combine the at least two output feature data subsets according to the results of the coordinate scanning to obtain a reassembled data subset.
[0098] For example, in at least one example, the cache unit 640 includes multiple sub-cache units, and correspondingly, the reassembly unit 620 is also configured to perform splicing and combination based on the bit width of each cache sub-unit to obtain a reassembled data subset; the bus width of the write unit 630 to write to the destination storage device is a multiple of the bit width of each cache sub-unit.
[0099] Figure 7 A schematic diagram illustrating coupling of a data processing device to a data bus according to at least one embodiment of the present disclosure is shown.
[0100] like Figure 7 As shown, the data processing device receives the initial output feature map calculated by the network layer through a data bus with a width of 512 bits, and then sends the reassembled initial output feature map to the destination storage device (such as memory) through a data bus with a width of 512 bits. For example, the cache unit 640 of the data processing device includes 16 sub-cache units 641, each sub-cache unit 641 corresponds to the 32-bit width of the data bus. Therefore, for example, the data processing device can splice and combine each sub-cache unit 641 as a unit to achieve reassembly. Therefore, during the reassembly process, the address can be easily jumped without affecting bandwidth utilization.
[0101] Figure 8 A schematic diagram of a data processing device provided according to at least one embodiment of the present disclosure is shown.
[0102] like Figure 8 As shown, the processing device 800 of the instruction pipeline includes a processing unit 810 and a memory 820, and the memory 820 stores one or more computer program modules 821; when the computer program module 821 is executed by the processing unit 810, it is used to execute the data processing method of any of the above embodiments.
[0103] At least one embodiment of the present disclosure provides a non-transitory readable storage medium having computer instructions stored thereon. When the computer instructions are executed by a processor, the data processing method of any of the above embodiments is executed.
[0104] For example, the non-transitory readable storage medium is implemented as a memory, such as a volatile memory and / or a non-volatile memory.
[0105] In the above-mentioned embodiment, the memory may be a volatile memory, for example, a random access memory (RAM) and / or a cache memory. Non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. The memory may also store various applications and various data, as well as various data used and / or generated by the applications.
[0106] Regarding this disclosure, the following points need to be explained:
[0107] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0108] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0109] The foregoing description is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of protection of the present disclosure. The scope of protection of the present disclosure is determined by the appended claims.
Claims
1. A data processing method for a neural network, wherein the neural network includes a plurality of network layers, the plurality of network layers including a first network layer and a second network layer adjacent to the first network layer, wherein an output feature map of the first network layer serves as an input feature map of the second network layer, the method comprising: receiving, from a processing unit array, at least two subsets of output feature data obtained by the processing unit array for at least two different parts of the first network layer; Determining second positions of the at least two output feature data subsets in the output feature graph of the first network layer according to the first positions of the at least two different parts, and splicing and combining the at least two output feature data subsets according to the second positions to obtain a recombined data subset; Writing the reorganized data subset into a destination storage device, wherein writing the reorganized data subset into the destination storage device comprises: The recombined data subset obtained by continuously arranging the at least two output feature data subsets is written into the destination storage device through a single storage operation via a data bus.
2. The data processing method according to claim 1, wherein: The step of splicing and combining the at least two output feature data subsets according to the second position to obtain the recombined data subset includes: In response to the at least two output feature data subsets being correspondingly arranged continuously in the input feature map of the second network layer, the at least two output feature data subsets are continuously arranged to obtain the recombined data subset.
3. The data processing method according to claim 1, wherein: The step of splicing and combining the at least two output feature data subsets according to the second position to obtain the recombined data subset includes: According to the mode in which the input feature map of the second network layer is input to the processing unit array and the second position, the at least two output feature data subsets are spliced and combined to obtain the recombined data subset.
4. The data processing method according to claim 3, wherein: The step of concatenating the at least two output feature data subsets to obtain the recombined data subset according to the mapping mode in which the input feature map of the second network layer is input to the processing unit array and the second position comprises: In response to the at least two output feature data subsets being successively input into the processing unit array according to the pattern and the second position, the at least two output feature data subsets are successively arranged to obtain the recombined data subset.
5. The data processing method according to claim 1, wherein: Writing the reorganized data subset into a destination storage device further includes: Obtaining the storage address and valid data length of the reorganized data subset; The step of performing a single storage operation via a data bus to write data into the destination storage device includes: Using the storage address and the valid data length, a single storage operation is performed to write into the destination storage device.
6. The data processing method according to claim 1, wherein: The receiving, from the processing unit array, at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer, comprises: Obtaining the at least two subsets of output feature data from different processing units of the processing unit array in the same operation cycle or different operation cycles, or obtaining the at least two subsets of output feature data from the same processing unit in different operation cycles; Address mapping is performed according to a mapping pattern of different processing units or the same processing unit in the processing unit array and the input feature map of the first network layer input to the processing unit array to obtain the positions of the at least two different parts in the first network layer.
7. The data processing method according to claim 6, wherein: The receiving, from the processing unit array, at least two output feature data subsets obtained by the processing unit array for at least two different parts of the first network layer further includes: In response to obtaining the at least two output feature data subsets in the different operation cycles, the at least two output feature data subsets are buffered for performing the splicing combination.
8. The data processing method according to claim 6, wherein: The step of splicing and combining the at least two output feature data subsets according to positions of the at least two different parts in the first network layer to obtain a recombined data subset includes: Performing coordinate scanning according to positions of the at least two different parts in the first network layer to obtain coordinates of the at least two output feature data subsets in an output feature map of the first network layer; The at least two output feature data subsets are spliced and combined according to the result of the coordinate scanning to obtain a recombined data subset.
9. A data processing device for a neural network, the neural network comprising a plurality of network layers, the plurality of network layers comprising a first network layer and a second network layer adjacent to the first network layer, the output feature map of the first network layer serving as the input feature map of the second network layer, the device comprising: a receiving unit configured to receive, from a processing unit array, at least two subsets of output feature data obtained by the processing unit array for at least two different parts of the first network layer; a recombining unit configured to determine, based on the first positions of the at least two different parts in the first network layer, second positions of the at least two output feature data subsets in the output feature graph of the first network layer, and concatenate and combine the at least two output feature data subsets according to the second positions to obtain a recombined data subset; The writing unit is configured to write the reorganized data subset into a destination storage device, wherein writing the reorganized data subset into the destination storage device comprises: The recombined data subset obtained by continuously arranging the at least two output feature data subsets is written into the destination storage device through a single storage operation via a data bus.
10. The data processing apparatus according to claim 9, wherein: The writing unit is further configured to obtain a storage address and a valid data length of the reorganized data subset, and use the storage address and the valid data length to perform a single storage operation to write the data into the destination storage device.
11. The data processing apparatus according to claim 9, wherein: The receiving unit is further configured to obtain the at least two output feature data subsets from different processing units of the processing unit array in the same operation cycle or different operation cycles, or to obtain the at least two output feature data subsets from the same processing unit in different operation cycles, and to perform address mapping according to a mapping pattern of different processing units of the processing unit array or the same processing unit and the input feature map of the first network layer input to the processing unit array to obtain the positions of the at least two different parts in the first network layer.
12. The data processing apparatus according to claim 9, further comprising: A cache unit is configured to cache the at least two output feature data subsets for use in the splicing combination.
13. The data processing apparatus according to claim 12, wherein: The cache unit includes multiple sub-cache units. The reassembly unit is further configured to perform the splicing and combination based on the bit width of each cache sub-unit to obtain a reassembled data subset. The bus width used by the write unit to write into the destination storage device is a multiple of the bit width of each of the cache sub-units.
14. The data processing apparatus according to claim 9, wherein: The recombination unit comprises: a coordinate scanning subunit, configured to perform coordinate scanning according to positions of the at least two different parts in the first network layer to obtain coordinates of the at least two output feature data subsets in the output feature map of the first network layer; and The coordinate combining subunit is configured to, according to the result of the coordinate scanning, combine the at least two output feature data subsets to obtain a recombined data subset.
15. A data processing device for a neural network, comprising processing unit, a memory having one or more computer program modules stored thereon; in, The one or more computer program modules are configured to, when executed by the processing unit, perform the data processing method according to any one of claims 1 to 8.
16. A non-transitory readable storage medium, wherein: The non-transitory readable storage medium stores computer instructions, wherein the computer instructions, when executed by a processor, perform the data processing method according to any one of claims 1 to 8.
17. A neural network processing device, comprising: The data processing device according to any one of claims 9 to 14; the processing unit array; the destination storage device, The data processing device is coupled to the processing unit array and the destination storage device.
Citation Information
Patent Citations
Method for accelerating convolution neutral network hardware and AXI bus IP core thereof
CN104915322A
Processing system and processing method for neural network
CN107818367A