Convolution operation method, convolution operation device, electronic device, and storage medium
By adjusting the array format of input data to match the number of channels of the convolution kernel, the method optimizes the utilization of matrix operation units in hardware accelerators, reducing data transmission and computation time in convolutional neural networks.
Patent Information
- Application Number
- JP2024570935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-31
- Filing Date
- 2023-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-05-30
AI Technical Summary
The utilization rate of matrix operation units in hardware accelerators for convolution operations in convolutional neural networks is low due to mismatched channel numbers, leading to increased storage space, data transmission time, and prolonged computation times, especially in the first layer of convolution operations.
A method and device that adjusts the array format of input data based on the number of channels of a convolution kernel, converting it to target data with equal channels for the convolution operation, using a convolution kernel represented by [1, 1, (C×R×S), K], to maximize the computing power of the matrix operation unit.
This approach enhances the utilization rate of the matrix operation unit, reduces data transmission time, and significantly shortens the convolution operation time, improving efficiency and resource utilization.
Smart Images

Figure 2025520154000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims the priority of a Chinese patent application with application number 202210610935.6 filed on May 31, 2022, and the entire content disclosed in that application is incorporated herein by reference in its entirety.
[0002] [Technical Field] Embodiments of the present disclosure relate to a method for convolution operation, an apparatus for convolution operation, an electronic device, and a storage medium.
Background Art
[0003] With the development of technology, artificial intelligence (AI) technology has been widely used in many fields. Deep learning is one of the important technologies of artificial intelligence technology. Deep learning technology based on artificial neural networks has already made great progress in fields such as object classification, text processing, image search, and human - machine interaction. Convolutional Neural Network (CNN) is a widely used deep learning technology that can directly input image data without requiring complex pre - processing and has great advantages in image processing and the like.
Summary of the Invention
[0004] At least one embodiment of the present disclosure is a method for a convolution operation, the method comprising: determining a convolution kernel for the operation, the convolution kernel for the operation being obtained based on an initial convolution kernel, the initial convolution kernel being represented by [R, S, C, K], the convolution kernel for the operation being represented by [1, 1, (C×R×S), K], where R, S, C, and K are all integers greater than 0; adjusting an array format of input data based on the number of channels of the convolution kernel for the operation to obtain target data, the size and the number of channels of the target data being different from those of the input data, and the number of channels of the target data being equal to the number of channels of the convolution kernel for the operation; performing a convolution operation based on the target data and the convolution kernel for the operation to obtain a result of the convolution operation, the result of the convolution operation between the target data and the convolution kernel for the operation being equal to the result of the convolution operation between the input data and the initial convolution kernel. A method for a convolution operation is provided that includes the above steps.
[0005] At least one embodiment of the present disclosure is an apparatus for convolution operation. The apparatus includes a determination unit for determining a convolution kernel for operation, where the convolution kernel for operation is obtained based on an initial convolution kernel. The initial convolution kernel is represented by [R, S, C, K], and the convolution kernel for operation is represented by [1, 1, (C×R×S), K], where R, S, C, and K are all integers greater than 0. The apparatus further includes an adjustment unit for adjusting the array format of input data based on the number of channels of the convolution kernel for operation to obtain target data. The size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolution kernel for operation. The apparatus also includes a calculation unit for performing a convolution operation based on the target data and the convolution kernel for operation to obtain a result of the convolution operation, where the result of the convolution operation between the target data and the convolution kernel for operation is equal to the result of the convolution operation between the input data and the initial convolution kernel. An apparatus for convolution operation is further provided.
[0006] At least one embodiment of the present disclosure further provides an electronic device including the apparatus for convolution operation provided by any embodiment of the present disclosure.
[0007] At least one embodiment of the present disclosure further provides an electronic device including a processor and a memory including at least one computer program module. The at least one computer program module is stored in the memory and configured to be executed by the processor. The at least one computer program module is used to implement the method for convolution operation provided by any embodiment of the present disclosure.
[0008] At least one embodiment of the present disclosure further provides a storage medium storing non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are executed by a computer, a method for convolution operation provided by any embodiment of the present disclosure is executed.
Brief Description of the Drawings
[0009] To more clearly explain the technical solutions according to the embodiments of the present disclosure, the drawings of the embodiments are briefly introduced below. Obviously, the drawings in the following description are only related to some embodiments of the present disclosure and do not limit the present disclosure.
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Although specific embodiments of the present disclosure are shown in the drawings, the present disclosure may be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and do not limit the protection scope of the present disclosure.
[0012] It should be understood that each step described in the implementation form of the method of the present disclosure can be executed in a different order and / or in parallel. Also, the method embodiment may include additional steps and / or omit the execution of the illustrated steps. The scope of the present disclosure is not limited in this regard.
[0013] As used herein, the term "comprising" and variations thereof are non-limiting, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment", the term "another embodiment" means "at least one additional embodiment", and the term "some embodiments" means "at least some embodiments". Related definitions of other terms are given in the following description.
[0014] Note that concepts such as "first" and "second" referred to in the present disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0015] Note that the modifications of "one" and "a plurality" referred to in the present disclosure are illustrative and not limiting, and those skilled in the art should understand them as "one or more" unless the context clearly indicates otherwise.
[0016] The names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0017] The input data of a convolutional neural network is usually an image with 3 channels. For example, the input image for the convolution of the first layer of the Residual Network ResNet50 is [1, 224, 224, 3], that is, the input image has 3 channels, the image size of each channel is 224×224, and the shape of the convolution kernel used in the convolution of the first layer of the Residual Network ResNet50 is [7, 7, 3, 64]. Generally used neural network accelerators are provided with a matrix operation unit (Matrix), and the matrix operation unit is mainly responsible for accelerating the matrix operations and convolution operations in the neural network. In order to accelerate the matrix operations, it is common for the matrix operation unit to increase the computational parallelism by increasing the computational scale. For example, the computational scale is set to 64×64, 128×128, etc. However, since the number of channels of the input data for the convolution of the first layer of the convolutional neural network is small (e.g., 3 channels), the computing power of the matrix operation unit on the hardware accelerator is low, and the computing time for the convolution of the first layer is relatively long, and the acceleration effect is not obvious. Furthermore, strictly following the data array method with aligned channels (Channel Align Tensor Layout) will significantly increase the storage space of the data and increase the data transmission time.
[0018] As shown in FIG. 1, a hardware accelerator is usually mounted on the host's PCIe (Peripheral Component Interconnect Express) node as a slave device of the host. PCIe is a high-speed serial computer expansion bus standard capable of realizing high-speed data transmission. To the host side, the hardware accelerator is the device side. When performing a convolution operation, it is necessary to transmit the data input in the first-layer convolution from the host to the hardware accelerator via PCIe, and this process is called Host2Device. For example, a central processing unit (CPU) reads data from memory, transmits it to the hardware accelerator on the device side via PCIe, and stores it in the memory (e.g., DDR) on the hardware accelerator. Then, the hardware accelerator can perform a convolution operation using these data.
[0019] Taking the convolution of the first layer of the residual network ResNet50 as an example, in the convolution of the first layer, it is necessary to realize the operation of the input data and the convolution kernel. This is represented by [1, 224, 224, 3] × [7, 7, 3, 64] = [1, 112, 112, 64]. Here, the input data is represented by [1, 224, 224, 3], that is, it is 3-channel data and the size is 224×224. The convolution kernel is represented by [7, 7, 3, 64], that is, there are 64 groups, each group has 3 convolution kernels, the size of the convolution kernel is 7×7, and the obtained result is 64-channel data with a size of 112×112.
[0020] Assuming that the matrix operation unit on the hardware accelerator is 64×64, due to the limitation of the data array method with aligned channel numbers, the input data of the first-layer convolution needs to be expanded from [1, 224, 224, 3] to [1, 224, 224, 64] on the host, and all redundant channel data is filled with 0. Regarding the storage space, it is necessary to increase by 21.33 times. Similarly, the time for transmitting data from the host side to the hardware accelerator also increases by 21.33 times. In this case, the utilization rate of the computing power of the matrix operation unit is only 4.68%. Regarding the convolution operation time, it takes 614656 cycle times for the matrix operation unit to complete the first-layer convolution operation.
[0021] For the convolution operation of the first layer of the convolutional neural network, since the number of channels of the input data is small and the matrix operation unit of the hardware accelerator is larger, the operation requirements do not match the characteristics of the hardware, which leads to the following problems in the convolution operation of the first layer of the convolutional neural network. First, regarding the input data, it is necessary to readjust the array method using the host's CPU, so the occupied storage space increases and CPU Time is required. Second, the amount of the rearranged input data increases, and the PCIe transmission time of Host2Device increases. Third, the utilization rate of the matrix operation unit of the hardware accelerator is low, and its computing power cannot be fully utilized, resulting in waste of hardware resources. Fourth, the matrix operation unit of the hardware accelerator takes a long time for the first-layer convolution operation, and the purpose of the hardware accelerator cannot be achieved.
[0022] At least one embodiment of the present disclosure provides a convolution operation method, a convolution operation device, an electronic device, and a storage medium. The convolution operation method can improve the utilization rate of the matrix operation unit, effectively utilize the computing power of the matrix operation unit, shorten the time of the convolution operation, improve the operation efficiency, and save the data transmission time.
[0023] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals in different drawings are used to refer to the same elements described.
[0024] At least one embodiment of the present disclosure provides a convolution operation method. The convolution operation method includes a step of determining a convolution kernel for operation, where the convolution kernel for operation is obtained based on an initial convolution kernel. The initial convolution kernel is represented by [R, S, C, K], and the convolution kernel for operation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0; a step of adjusting the array format of input data based on the number of channels of the convolution kernel for operation to obtain target data, where the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolution kernel for operation; and a step of performing a convolution operation based on the target data and the convolution kernel for operation to obtain the result of the convolution operation. The result of the convolution operation between the target data and the convolution kernel for operation is equal to the result of the convolution operation between the input data and the initial convolution kernel.
[0025] FIG. 2 is a schematic flowchart of a convolution operation method provided by some embodiments of the present disclosure. As shown in FIG. 2, in some embodiments, the convolution operation method includes steps S10 to S30. Step S10: A step of determining a convolution kernel for operation, where the convolution kernel for operation is obtained based on an initial convolution kernel. The initial convolution kernel is represented by [R, S, C, K], and the convolution kernel for operation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0. Step S20: Based on the number of channels of the convolutional kernel for calculation, adjust the array format of the input data to obtain target data. The size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolutional kernel for calculation. Step S30: Based on the target data and the convolutional kernel for calculation, perform a convolution operation to obtain the result of the convolution operation. The result of the convolution operation between the target data and the convolutional kernel for calculation is equal to the result of the convolution operation between the input data and the initial convolutional kernel.
[0026] For example, the convolution operation method may be used for the convolution operation of the first layer of a convolutional neural network. Of course, the embodiments of the present disclosure are not limited thereto. The convolution operation method may be used not only for convolutional neural networks but also for convolution operations of other types of networks. It may be used not only for the convolution operation of the first layer (the convolution operation of the first layer) but also for convolution operations of other layers, and may be determined according to actual needs. The embodiments of the present disclosure are not limited thereto.
[0027] For example, in step S10, the initial convolutional kernel is the convolutional kernel required for the convolution operation of the first layer, and is represented by [R, S, C, K]. Taking the convolution of the first layer of the Residual Network ResNet50 as an example, the operation required for the convolution of the first layer is [1, 224, 224, 3] × [7, 7, 3, 64] = [1, 112, 112, 64]. Then the initial convolutional kernel [R, S, C, K] is [7, 7, 3, 64]. In this example, R = 7, S = 7, C = 3, and K = 64. Convert the parameters of the initial convolutional kernel to obtain the convolutional kernel for calculation [1, 1, (C × R × S), K]. In the above example, the convolutional kernel for calculation can be obtained based on the initial convolutional kernel, and the convolutional kernel for calculation is [1, 1, 147, 64]. Hereinafter, the conversion principle of the convolutional kernel will be briefly described with reference to FIG. 3.
[0028] Figure 3 is a schematic diagram showing the principle of the convolution operation. As shown in Figure 3, the size of the input data is [1, 3, 3, 5], the size of the convolution kernel is [2, 2, 5, 4], and the size of the output data is [1, 2, 2, 4]. For example, for point M, its calculation method is shown in Figure 3. Since the size of the convolution kernel is 2×2 and the number of channels is 5, point M is the result of multiplying and accumulating 20 points each of the input data and the convolution kernel. By utilizing the characteristics of the convolution operation, the convolution kernel can be converted from R×S×C×K to 1×1×(C×R×S)×K, and the input data can be adjusted accordingly, so the overall calculation result of the convolution remains unchanged. Through such a conversion operation, the number of channels can be increased. For the first layer network of the convolutional neural network, the convolution kernel is adjusted from [7, 7, 3, 64] to [1, 1, 147, 64], and the number of channels is expanded from 3 to 3×7×7 = 147, thereby maximizing the utilization of the computing power of the matrix operation unit (Matrix). Therefore, in step S10, the initial convolution kernel [R, S, C, K] can be converted to obtain the convolution kernel for operation [1, 1, (C×R×S), K], thereby realizing the change in the array method of the convolution kernel.
[0029] For example, the array of convolution kernels may be changed offline. This is because the convolution kernel used by the neural network model is fixed at the deployment stage and will not be changed even if the input changes. Therefore, the convolution kernel can be preprocessed into the required array format. In an embodiment of the present disclosure, at the deployment stage of the neural network model, the convolution kernel [R, S, C, K] used as the convolution kernel for later use is set to [1, 1, (C×R×S), K]. For example, at the model compilation stage, the code corresponding to the initial convolution kernel [R, S, C, K] can be changed and adjusted using a high-level language (e.g., Python) to obtain the convolution kernel [1, 1, (C×R×S), K] for computation. Of course, the embodiments of the present disclosure are not limited to this. Before each convolution operation is performed, the initial convolution kernel [R, S, C, K] used can also be adjusted to obtain the convolution kernel [1, 1, (C×R×S), K] actually used in this operation.
[0030] Returning to FIG. 2, in step S20, based on the number of channels of the convolutional kernel for calculation, the array format of the input data is adjusted to obtain target data. For example, the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolutional kernel for calculation. Taking the convolution of the first layer of the residual network ResNet50 as an example, the calculation performed is [1, 224, 224, 3]×[7, 7, 3, 64]=[1, 112, 112, 64], and since the convolutional kernel is adjusted to [1, 1, 147, 64], it is necessary to adjust the array format of the input data so that the calculation result remains unchanged, and change the calculation to [1, 112, 112, 147]×[1, 1, 147, 64]=[1, 112, 112, 64]. From this, it can be seen that the data obtained after adjusting the input data is [1, 112, 112, 147]. For example, the data obtained after adjusting the array format of the input data is called target data, and the target data is ultimately the data on which the convolutional operation is performed with the convolutional kernel for calculation. For example, due to the adjustment of the array format, the size and number of channels of the target data will be different from those of the input data. From the above calculation formula, it can be seen that the number of channels of the target data is equal to the number of channels of the convolutional kernel for calculation (for example, 147 in both of the above examples), which facilitates both convolutional operations.
[0031] For example, since the number of channels of the target data is larger than the number of channels of the input data and the number of channels of the convolutional kernel for calculation is larger than the number of channels of the initial convolutional kernel, the number of channels increases, maximizing the computing power of the matrix operation unit. For example, in the above example, both the number of channels of the input data and the number of channels of the initial convolutional kernel are 3, and both the number of channels of the target data and the number of channels of the convolutional kernel for calculation become 147, increasing the number of channels. For example, the conversion of the array format of the input data needs to be completed online, that is, in the inference stage of the neural network, it is necessary to adjust the array format of the data input each time.
[0032] Figure 4 is a schematic flowchart of step S20 in FIG. 2. For example, in some embodiments, as shown in FIG. 4, step S20 may further include steps S21 to S23. Step S21: Store the input data in a static memory (Static Memory) in row units, and each row of the input data is stored in the corresponding N storage rows (Row-Store) in the static memory, where N is an integer greater than 0. Step S22: Perform a padding operation on the input data stored in the static memory to obtain extended data. Step S23: Adjust the array format of the extended data, change the size and number of channels of the extended data to obtain target data.
[0033] For example, in step S21, first, store the input data in the static memory provided in the hardware accelerator, and the static memory is, for example, a static random access memory (Static Random Access Memory, SRAM). The input data is stored in the static memory in row units, that is, each row of the input data is stored in the corresponding N storage rows in the static memory, where N is an integer greater than 0. For example, using the data flow shown in FIG. 1, the input data may be transmitted to the static memory.
[0034] As shown in FIG. 5, step S21 may further include steps S211 and S212. Step S211: Pack the input data and store it in the memory. The input data includes multiple channels, and packing means that multiple channels of the same data point are stored adjacent to each other in order in the memory. Step S212: Using direct memory access (DMA), transmit the input data in the memory to the static memory of the hardware accelerator, store the first data point of each row of the input data in the first column of different rows of the static memory, and store each row of the input data in the corresponding N storage rows in the static memory.
[0035] For example, in step S211, arranging densely is, for example, a data array method (Channel Align Tensor Layout) with the number of channels aligned. The input data includes a plurality of channels, and arranging densely means that a plurality of channels of the same data point are stored adjacent to each other in order in the memory. As shown in FIG. 6, in some examples, the input data of the first-layer convolution is [1, 224, 224, 3]. The number 1 represents Batch Size = 1, one of the numbers 224 represents Height = 224, the other of the numbers 224 represents Width = 224, and the number 3 represents Channel = 3. That is, the input data is a 3-channel image with a size of 224×224. In order to reduce the storage space of the data on the host, a storage method of arranging the data densely is adopted. For example, for the pixel point at the first row and the first column of the input data, the values of its three channels are stored continuously in order in the memory space. Then, continue to store the pixel point at the first row and the second column, and the values of its three channels are stored continuously in order in the memory space, and so on.
[0036] For example, the value of each channel of each pixel is represented in the FP16 data format and occupies a 2-byte address space. The input image of the first layer is 224×224×3×2 Byte = 301056 Byte in total, that is, it occupies 294 KB. Here, both the adopted data format and the occupied address space are exemplary and do not limit the embodiments of the present disclosure.
[0037] For example, in step S212, using the Direct Memory Access (DMA) method, the input data in the memory is transmitted to the static memory of the hardware accelerator. The data transfer method may use the method shown in FIG. 1. The input data of the first-layer convolution is transferred from the host side to the DDR on the device side via PCIe and is stored in the DDR in the same way as the method of storing the input data in the host memory, thereby realizing the Host2Device process of the 1D Tensor. For example, the input data is continuously stored in the DDR and occupies, for example, 294 KB. Next, it is necessary to transfer the input data to the static memory, that is, it is necessary to transfer this 294 KB of data from the DDR to the SRAM in the processing engine (also called the PE, hardware accelerator).
[0038] For example, when storing, the first data point of each row of the input data is stored in the first column of different rows of the static memory, and each row of the input data is stored in the corresponding N storage rows in the static memory.
[0039] As shown in FIG. 7, in some examples, the configuration format of the SRAM may be abstractly considered as a table of M rows and N columns, and one piece of data is stored in each table. Since the size of the input data is 224×224, the input data is logically divided into 224 rows, and the starting position of each row is at the first column of a certain row of the SRAM. Since the number of SRAM columns is limited, it is difficult for all of one row of the input data to be stored in one row of the SRAM. Therefore, one row of the input data will be distributed among multiple rows of the SRAM, that is, different SRAM addresses. For the input data of [1,224,224,3], considering the data filled by subsequent filling operations, there are 229 points in each row, and there are 3 channels for each point. For the SRAM with a storage space of 1024 bits per row, the number of SRAM rows required to store one row of data is ceil(229*3 / 64)=11, where ceil represents rounding up. That is, one row of data points of the input data is stored in 11 storage rows of the SRAM. In this example, N = 11, and the entire input data occupies 224×11 = 2464 rows of SRAM.
[0040] As shown in FIG. 7, the left side shows that the input data is continuously stored in the DDR without the concepts of H, W, and C. The right side shows that after being transferred to the SRAM by the DMA, the data is divided according to the rows of the input data in the SRAM, and each row of the data occupies a certain amount of SRAM space (for example, 11 storage rows). Thereby, the data transfer from the DDR to the SRAM in the PE is realized, and the conversion from the one-dimensional tensor (1D Tensor) to the two-dimensional tensor (2D Tensor) is completed.
[0041] The DMA transfer process is briefly described as follows.
[0042] Assuming that the input data is stored in a continuous DDR space starting from the source address (source_address), first, it is necessary to transfer the 224×3×2Byte = 1344Byte data in the first row to a continuous SRAM space starting from the destination address (destiny_address). Since one row of SRAM is 128Byte, it is necessary to store this 1344Byte in ceil(1344 / 128) = 11 rows of SRAM. That is, the DMA needs to continuously transfer 11×128Byte of data. After the data transfer of the first row is completed, the DMA jumps the read address from source_address to the position of source_address + 1344Byte, that is, the DDR address at the beginning of the second row of the actual input data, and then needs to continuously transfer 11×128Byte into the SRAM space starting from destination_address + 11. By analogy, after 224 transfers are completed, all the input data has been transferred from DDR to SRAM in the processing engine, that is, the conversion from a one-dimensional tensor to a two-dimensional tensor is completed.
[0043] Note that the amount of data of 11×128Byte transmitted each time is more than the amount of data in each actual row, that is, the data in the next row is also included. However, since the first address transmitted each time is accurate, repeatedly transmitting the data will not affect the data itself, and these redundant data will not affect the subsequent processing.
[0044] Returning to FIG. 4, for example, in step S22, a padding operation is performed on the input data stored in the static memory to obtain extended data. Here, the extended data refers to the data obtained after the padding operation. For example, in some cases, assuming that the actual convolution operation to be completed is [1,224,224,3]×[7,7,3,64]=[1,112,112,64], next, it is necessary to perform a padding operation on the input data in all directions of up, down, left, and right. When padding, three points need to be padded on the left and top of the input data (padding 3 columns on the left and 3 rows on the top), and two points need to be padded on the right and bottom of the input data (padding 2 columns on the right and 2 rows on the bottom), and the size of the extended data obtained after the padding operation is [1,229,229,3].
[0045] FIG. 8 is a schematic flowchart of step S22 in FIG. 4. As shown in FIG. 8, step S22 may further include steps S221 to S223. Step S221: In the static memory, fill the memory rows corresponding to the front and back of the storage position of the input data with a first preset value to obtain first intermediate data. Here, the first intermediate data includes the input data and the filled first preset value. Step S222: Transmit the first intermediate data to the vector calculation unit, and use the shift instruction and padding instruction of the vector calculation unit to fill both ends of each row corresponding to the first intermediate data with a second preset value to obtain second intermediate data. Here, the second intermediate data includes the first intermediate data and the filled second preset value. Step S223: Transmit the second intermediate data to the corresponding storage position in the static memory to obtain extended data. The extended data has the same content as the second intermediate data.
[0046] For example, in step S221, in the static memory, the first preset value is filled in the memory rows before and after the storage position of the input data, and the first intermediate data is obtained. The first intermediate data includes the input data and the filled first preset value. In this step, for example, filling operations are performed above and below the input data. For example, in some cases, it is necessary to prepare in advance the SRAM space required for filling above near the destination address of the SRAM, that is, it is necessary to insert several lines of data before actually inputting the first line of data. The padding operations above and below are performed using the Vector calculation unit of the hardware accelerator. For example, since the filled first preset value is usually 0, to obtain the first intermediate data, it is necessary to write the value of all 0s to several addresses before and after the storage space of the input data in the SRAM, thereby obtaining the first intermediate data. The first intermediate data is the data with padding operations performed above and below it, and no padding operations have been performed on the left and right yet.
[0047] For example, in step S222, the first intermediate data is transmitted to the vector calculation unit, and the second preset value is filled at both ends of each row corresponding to the first intermediate data using the shift instruction (for example, vshiftri instruction) and the padding instruction (for example, SI2V instruction) of the vector calculation unit, and the second intermediate data is obtained. The second intermediate data includes the first intermediate data and the filled second preset value. In this step, for example, padding operations are performed on the left and right of the first intermediate data.
[0048] For example, in some cases, the data in the 2464 - address space in the SRAM is grouped every 11 rows, transmitted to the vector calculation unit in sequence, and stored in the storage space vmem in the vector calculation unit. Next, the vector calculation unit uses the vshiftri instruction to shift the entire data to the right, leaving space for left - hand padding. Then, using the SI2V instruction, these positions are written to the corresponding second preset value (for example, usually set to 0). In the case of right - hand padding, after the entire data is shifted to the right, the corresponding second preset value is written after the last column of the first row of the input data. If the amount of data to be filled is too large, it is necessary to add vmem space as needed. For example, using a pipeline method, left - and right - hand filling operations can be performed on multiple groups of 11 - row data to improve processing efficiency.
[0049] For example, in step S223, the second intermediate data is transmitted to the corresponding storage location in the static memory to obtain extended data. The extended data has the same content as the second intermediate data. That is, the second intermediate data for which the filling operation in vmem is completed is written back into the corresponding address space of the SRAM, and the data stored in the SRAM for which the filling operation is completed is called extended data.
[0050] Note that if there is no need to perform a filling operation on the input data, step S22 may be omitted. Furthermore, in the embodiments of the present disclosure, when a filling operation is required, the up - and - down filling may be performed first and then the left - and - right filling, or the left - and - right filling may be performed first and then the up - and - down filling. The specific filling order is not limited. Note that the instructions used for the filling operation are not limited to the vshiftri instruction and the SI2V instruction. Any other appropriate instructions may be used as long as they can implement the filling operation, and the embodiments of the present disclosure do not limit this.
[0051] Returning to FIG. 4, for example, in step S23, the arrangement method of the extended data is adjusted, the size and the number of channels of the extended data are changed, and target data is obtained. That is, in order to ensure that the calculation result remains unchanged according to the convolutional kernel, it is necessary to adjust the arrangement method of the extended data and change its size and the number of channels. For example, the number of channels of the target data obtained after adjustment is equal to the number of channels of the convolutional kernel, and the target data is represented by [1, ht, wt, (C×R×S)], where both ht and wt are integers greater than 0.
[0052] FIG. 9 is a schematic flowchart of step S23 in FIG. 4. As shown in FIG. 9, step S23 may further include steps S231 and S232. Step S231: Read the data of the R×N storage rows in the static memory one by one and transmit it to the vector calculation unit. Step S232: The vector calculation unit converts the data of the R*N storage rows received each time into the data of wt*ceil((C×R×S) / L) storage rows and obtains the target data.
[0053] For example, in step S231, the data of the R*N storage rows in the static memory is read each time and transmitted to the vector calculation unit, and the vector calculation unit converts the data in the R×N storage rows received each time. For example, the start address read each time moves str*N storage rows according to the preset hop count str. The preset hop count is the hop count in the row direction and the column direction of the sliding window required for the convolution operation of the input data and the initial convolutional kernel, and the total number of times of reading data from the static memory is equal to ht.
[0054] For example, in some cases, the convolution of the first layer of the Residual Network ResNet50 is taken as an example. Since the initial convolution kernel is [R, S, C, K] = [7, 7, 3, 64], R = 7. The input data is [1, 224, 224, 3], and one row of the input data is stored in N memory rows of the SRAM. When the space of one row of the SRAM is 128 bytes, N = 11. Therefore, the data in R * N = 77 memory rows in the static memory is read each time and transmitted to the vector calculation unit. The data stored in these 77 memory rows corresponds to one row of the 224×224 input data. For example, for the convolution operation of the input data [1, 224, 224, 3] and the initial convolution kernel [7, 7, 3, 64], the number of hops in the row direction and column direction of the sliding window required is set to 2. Therefore, the preset number of hops str is 2. For example, the start address read each time is shifted by str * N (i.e., 2×11 = 22) memory rows according to the preset number of hops str, so that the read data matches the data included in the sliding window when performing the convolution operation. According to the calculation formula [1, 224, 224, 3]×[7, 7, 3, 64]=[1, 112, 112, 147]×[1, 1, 147, 64]=[1, 112, 112, 64], the convolution kernel is converted to [1, 1, 147, 64], and the obtained target data [1, ht, wt, (C×R×S)] needs to be converted to [1, 112, 112, 147]. Therefore, it can be seen that ht = 112 and wt = 112. For example, the total number of times of reading data from the static memory is ht (for example, 112), and after the data read each time is converted, it corresponds to one row of the target data.
[0055] For example, in step S232, the vector calculation unit converts the data of R*N memory rows received each time into the data of wt*ceil((C×R×S) / L) memory rows, and the converted data is the target data. That is, the array format of the data is adjusted, and the size and number of channels of the data are changed. For example, in the calculation formula, L represents the number of data points that can be stored in each memory row of the static memory, and ceil((C×R×S) / L) represents rounding up to the nearest integer for (C×R×S) / L. For example, in some examples, the vector calculation unit receives the data of 7×11 memory rows each time. When the space of one row of the SRAM is 128 Byte, the number of data points L that can be stored in each memory row is 64. For the initial convolution kernel [7, 7, 3, 64], R = 7, S = 7, C = 3. When the target data [1, ht, wt, (C×R×S)] = [1, 112, 112, 147], wt = 112. Therefore, wt*ceil((C×R×S) / L) = 112×3, that is, the vector calculation unit converts the data of 7×11 memory rows received each time into the data of 112×3 memory rows.
[0056] FIG. 10 is a schematic flowchart of step S232 in FIG. 9. For example, in some examples, the above step S232 further includes steps S2321 to S2323. Step S2321: Divide the data in R*N memory rows into a plurality of groups according to a preset number of hops. Step S2322: For the data of each group, determine the initial position information parameter and the target position information parameter of each row of data of the sliding window corresponding to the data of the group. Step S2323: The vector calculation unit stores the data of each group in the corresponding position of the target memory in the converted array format according to the initial position information parameter and the target position information parameter, and obtains the target data.
[0057] For example, in step S2321, the data in R*N memory rows is divided into multiple groups according to a preset number of hops. The data of each group corresponds to a sliding window in the row direction, and the number of data in the multiple groups is equal to wt. For example, in some examples, the data in 7×11 memory rows is divided into 112 groups according to a preset number of hops str = 2, and wt = 112. The data in 7×11 memory rows corresponds to one row of data in the input data 224×224, and the data in the 112 divided groups corresponds to 112 sliding windows created by taking 224 data in one row with 2 hops.
[0058] For example, in step S2322, for the data of each group, the initial position information parameter and the target position information parameter of the data in each row of the sliding window corresponding to the data of the group are determined. The initial position information parameter is used to determine the source address where the data row is located within the sliding window, and the target position information parameter is used to determine the destination address to which these data are transferred.
[0059] The operation modes of steps S2321 to S2322 are described below with examples.
[0060] For example, after performing padding operations on the input data [1, 224, 224, 3] in the up, down, left, and right directions, the size of the entire input data becomes [1, 229, 229, 3], occupying 229×11 address spaces in the SRAM. Next, it is necessary to convert the shape of the input data to [1, 112, 112, 147]. Basically, for the convolution operation, each sliding window (Feature Window) where the convolution kernel slides needs to complete the conversion from [7×7×3] to [1, 1, 147], as shown in Figure 11.
[0061] Since the sliding window over which each convolutional kernel slides corresponds to 7 rows and 7 columns of the original data (input data or input image), the sliding windows swept by the convolutional kernel during the process of sliding from the upper left to the lower right overlap. Regarding the overlap in the column direction, in order to avoid repeatedly reading data from the SRAM, each time 7×11 = 77 data in the address space are read from the SRAM and transmitted to the vector calculation unit for processing. Regarding the overlap in the row direction, when the sliding window slides from left to right, the data overlapping in the row direction are repeatedly read. Overall, the data in the SRAM are divided into 112 groups, corresponding to 112 rows after conversion. The data of each group occupy 7×11 address spaces before conversion. The vector calculation unit reads and processes the data of one group and outputs the data of 112×3 address spaces. Here, 112 corresponds to the data width after conversion, and 3 corresponds to the space occupied by 147 channels (147 channels need to occupy 3 SRAM memory rows, that is, 3 SRAM address spaces).
[0062] After the vector calculation unit acquires the data, the data on these 7×11 SRAM memory rows (entries) are temporarily stored in the vmem in the vector calculation unit, and the data array remains unchanged. Then, using the instruction operation of the vector calculation unit, it is converted into the data of 112×3 vmem memory rows. Then, the result is written back to the SRAM.
[0063] The converted data width is 112 points in the row direction, with 147 channels for each point, and it is distributed among 3 vmem memory rows. For each point in the row direction, it is necessary to convert the original 7×7×3 sliding window to 1×1×147. Therefore, it is necessary to search for 7 rows of data corresponding to each sliding window and reconstruct them into a new data array. As shown in Figure 12, for the first sliding window, the data width of the sliding window is 7×3 = 21 channels, and there are 7 rows of data (stored in 7×11 memory rows). It is necessary to determine the memory addresses of these 7 rows of data and convert their array format. Referring to the right side of Figure 12, these 7 rows of data are rearranged into 3 rows. The original rows 0, 1, and 2 form the new row 1, the original rows 3 and 4 form the new row 2, and the original rows 5 and 6 form the new row 3. A total of 147 data points included in the sliding window corresponding to the original 7×7×3 are stored in these 3 new rows. By analogy, for the next sliding window in the row direction, the array format of the data is converted in a similar manner until all the sliding windows corresponding to 1 row of data for 7 rows of data are converted, and then the data of the next group of 7×11 memory rows is read.
[0064] To determine the initial position and target position of the data for each row within the sliding window, it is necessary to define initial position information parameters and target position information parameters for the data of each row within the sliding window.
[0065] The initial position information parameters include the first start boundary coordinate, the first end boundary coordinate, the first start address, the first end address, the first start number, and the first end number.
[0066] The first start boundary coordinate represents the relative coordinate in the row direction of the extended data of the start boundary of the corresponding sliding window, and the first end boundary coordinate represents the relative coordinate in the row direction of the extended data of the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. Since the data on which the filling operation is performed is the extended data, these coordinates and parameters are defined for the extended data. In other cases where the filling operation is not required, these coordinates and parameters may be directly defined for the input data. As shown in FIG. 12, the start boundary is, for example, the left boundary of the sliding window, and the first start boundary coordinate is the relative coordinate in the row direction of the 229×3 extended data of the left boundary of the sliding window, and the end boundary is, for example, the right boundary of the sliding window, and the first end boundary coordinate is the relative coordinate in the row direction of the 229×3 extended data of the right boundary of the sliding window.
[0067] The calculation formula for the first start boundary coordinate is src_row_start_index = i*str*ch. src_row_start_index represents the first start boundary coordinate, i represents the number of the corresponding data point in the size wt of the target data of the corresponding sliding window (for example, which one of the total 112 sliding windows in one row, that is, which one of the output data width wt = 112), str represents the number of hops in the row direction of the sliding window (for example, 2), and ch represents the number of channels of the input data (for example, 3).
[0068] The calculation formula for the first end boundary coordinate is src_row_end_index = src_row_start_index+(kernel_w*ch - 1). src_row_end_index represents the first end boundary coordinate, kernel_w represents the width of the sliding window (for example, 7), and the size of the sliding window is equal to the size of the initial convolution kernel (for example, both are 7x7).
[0069] The first start address represents the address in the memory of the vector calculation unit (for example, vmem) of the first start boundary coordinate, and the first end address represents the address in the memory of the vector calculation unit (for example, vmem) of the first end boundary coordinate. The first start number represents the number of the data point corresponding to the first start address of the first start boundary coordinate, and the first end number represents the number of the data point corresponding to the first end address of the first end boundary coordinate. Since vmem is stored in row units, a certain memory row in vmem can be found according to the first start address or the first end address, and the first start number or the first end number indicates which data in the corresponding memory row the corresponding data is.
[0070] The calculation formula for the first start address is src_row_start_address = src_row_start_index / vmem_lane + j*N. src_row_start_address represents the first start address, vmem_lane represents the number of data points that can be stored in each memory row in each memory of the vector calculation unit, and j represents the row number (for example, values from 1 to 7) of the corresponding data in the sliding window.
[0071] The calculation formula for the first end address is src_row_end_address = src_row_end_index / vmem_lane + j*N. src_row_end_address represents the first end address.
[0072] The calculation formula for the first starting number is src_row_start_lane = src_row_start_index % vmem_lane. src_row_start_lane represents the first starting number. For example, % represents the modulo operation.
[0073] The calculation formula for the first ending number is src_row_end_lane = src_row_end_index % vmem_lane. src_row_end_lane represents the first ending number.
[0074] Once the above parameters are determined, the positions of the source data in vmem required to convert 7×7×3 can be determined. To transfer these source data to the destination address in vmem, it is necessary to determine the corresponding destination address and related parameters, that is, it is necessary to further determine the target position information parameters.
[0075] The target position information parameters include the second starting boundary coordinates, the second ending boundary coordinates, the second starting address, the second ending address, the second starting number, and the second ending number.
[0076] The second start boundary coordinate represents the relative coordinate within the data size of [1, 1, (C × R × S)] of the start boundary of the corresponding sliding window, and the second end boundary coordinate represents the relative coordinate within the data size of [1, 1, (C × R × S)] of the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. For example, the target data represents [1, ht, wt, (C × R × S)]. In some examples, the target data is [1, 112, 112, 147], and it is necessary to convert the data size corresponding to each sliding window from [7, 7, 3] to [1, 1, 147]. As shown in FIG. 12, the start boundary is, for example, the left boundary of the sliding window, the second start boundary coordinate is the relative coordinate within the data size of [1, 1, 147] of the left boundary of the sliding window, the end boundary is, for example, the right boundary of the sliding window, and the second end boundary coordinate is the relative coordinate within the data size of [1, 1, 147] of the right boundary of the sliding window.
[0077] The calculation formula for the second start boundary coordinate is dst_row_start_index = j * kernel_w * ch. dst_row_start_index represents the second end boundary coordinate, j represents the row number (for example, values from 1 to 7) within the sliding window of the corresponding data, kernel_w represents the width of the sliding window (for example, 7), the size of the sliding window is equal to the size of the initial convolution kernel (for example, both are 7x7), and ch represents the number of channels of the input data (for example, 3).
[0078] The calculation formula for the second end boundary coordinate is dst_row_end_index = dst_row_start_index + (kernel_w * ch - 1). dst_row_end_index represents the second end boundary coordinate.
[0079] The second start address represents the address in the memory of the vector calculation unit (e.g., vmem) of the second start boundary coordinates, and the second end address represents the address in the memory of the vector calculation unit (e.g., vmem) of the second end boundary coordinates. The second start number represents the number of the data point corresponding to the second start address of the second start boundary coordinates, and the second end number represents the number of the data point corresponding to the second end address of the second end boundary coordinates. Since vmem is stored in row units, a certain memory row in vmem can be found according to the second start address or the second end address, and the second start number or the second end number indicates which data in the corresponding memory row the corresponding data is.
[0080] The calculation formula for the second start address is dst_row_start_address = dst_row_start_index / vmem_lane. dst_row_start_address represents the second start address, and vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit.
[0081] The calculation formula for the second end address is dst_row_end_address = dst_row_end_index / vmem_lane. dst_row_end_address represents the second end address.
[0082] The calculation formula for the second start number is dst_row_start_lane = dst_row_start_index % vmem_lane. dst_row_start_lane represents the second start number.
[0083] The calculation formula for the second end number is dst_row_end_lane = dst_row_end_index % vmem_lane. dst_row_end_lane represents the second end number.
[0084] After the initial position information parameter and the target position information parameter are determined, the source address and the destination address required for data transfer can be determined, and based on these parameters, the source data is transferred to the destination address.
[0085] For example, in step S2323, after the initial position information parameter and the target position information parameter are determined, the vector calculation unit stores the data of each group in the corresponding position of the target memory in the converted array format. The destination address indicated by the target position information parameter is the address in the target memory, and thus the target data is obtained. For example, the target memory stores data in row units, is transferred to the target memory, and the data stored on the target memory is the target data. For example, the target memory may be the above-mentioned static memory (in this case, the data before conversion and the data after conversion are stored at different addresses in the static memory), or may be another storage device different from the above-mentioned static memory. The embodiments of the present disclosure are not limited thereto.
[0086] For example, step S2323 may further include a step of concatenating data of each group in a converted array format according to an initial position information parameter and a target position information parameter using a circular shift instruction by a vector calculation unit according to a preset enable signal in a predicate register, storing the data in a corresponding position of a target memory, and obtaining target data. For example, in some examples, the vshiftri instruction of the vector instruction set architecture (Vector ISA) of the vector calculation unit is used to circularly shift the data of the source address to the right by several positions, and then write the data to the destination address according to a write enable signal in a vector predicate register (also called a VPR or VP register). The aforementioned preset enable signal may be, for example, a write enable signal. In the data concatenation process of converting data corresponding to a 7×7×3 sliding window into 1×1×147 data, it is necessary to determine the VP register to be used based on the second start number dst_row_start_lane and the second end number dst_row_end_lane. For the usage of the vshiftri instruction and the VP register, a conventional design may be referred to, but it will not be described in detail here.
[0087] By the above method, the conversion from a 2D tensor (2DTensor) to a 3D tensor (3DTensor) is completed using a vector calculation unit.
[0088] After the processing of each step, the input data [1, 224, 224, 3] is converted into target data [1, 112, 112, 147], and the convolution kernel for calculation determined based on the initial convolution kernel [7, 7, 3, 64] is [1, 1, 147, 64], whereby the number of channels increases from 3 to 147.
[0089] Returning to FIG. 2, in step S30, a convolution operation is performed based on the target data and the arithmetic convolution kernel to obtain a convolution result. The result of the convolution operation between the target data and the arithmetic convolution kernel is equal to the result of the convolution operation between the input data and the initial convolution kernel. For example, step S30 may further include a step of performing a convolution operation on the target data and the arithmetic convolution kernel using a matrix operation unit (Matrix).
[0090] For example, in some embodiments, taking the convolution of the first layer of the Residual Network ResNet50 as an example, the operation of the input data and the convolution kernel that needs to be realized is [1, 224, 224, 3] × [7, 7, 3, 64] = [1, 112, 112, 64]. Since there are only three channels, the computing power of the matrix operation unit is not fully utilized. By using the convolution operation method provided by the embodiments of the present disclosure, the arithmetic convolution kernel obtained based on the initial convolution kernel [7, 7, 3, 64] is [1, 1, 147, 64]. Since the target data obtained after adjusting the array format of the input data [1, 224, 224, 3] is [1, 112, 112, 147], the actual operation of the target data and the convolution kernel is [1, 112, 112, 147] × [1, 1, 147, 64] = [1, 112, 112, 64]. The result of this convolution operation is consistent with the result of the originally required convolution operation. The number of channels increases to 147, the computing power of the matrix operation unit can be maximally utilized, the utilization rate of the matrix operation unit is improved, the time of the convolution operation is shortened, and the operation efficiency is improved. Also, since there is no need to re-adjust the array format of the input data on the host CPU and there is no need to expand the number of channels on the host, the occupied data space does not increase significantly, and the data transmission volume from the host to the device does not increase significantly. Therefore, the PCIe transmission time of Host2Device does not increase, and the data transmission time can be saved.
[0091] The convolution operation method provided by the embodiments of the present disclosure helps to achieve the purpose of hardware acceleration, can realize the acceleration of the convolution operation of the first layer of the convolutional neural network (CNN), and has the characteristics of small storage space, short transmission time, high utilization rate of hardware modules, and short calculation time. For example, the time required to perform the convolution of the first layer of the residual network ResNet50 using the conventional convolution operation method is 614,656 cycle times, while the theoretical time required to perform the convolution of the first layer of the residual network ResNet50 using the convolution operation method provided by the embodiments of the present disclosure is 37,632 cycle times, which is reduced to 6.1% of the previous level, significantly shortening the convolution operation time of the first layer of the convolutional neural network (CNN).
[0092] It should be noted that in the embodiments of the present disclosure, the convolution operation method provided by each of the above embodiments of the present disclosure may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be clearly understood that the flow of the convolution operation method described above includes a plurality of operations performed in a specific order, but the order of the plurality of operations is not limited. The convolution operation method described above may be performed once or multiple times according to predetermined conditions.
[0093] It should be noted that the above description takes the convolution of the first layer of the residual network ResNet50 as an example, but this does not constitute a limitation on the embodiments of the present disclosure. The convolution operation method provided by the embodiments of the present disclosure may be applied to any applicable convolution operation, and the sizes and channel numbers of various data, as well as the sizes and channel numbers of various convolution kernels, may be determined according to actual needs and are not limited to the above specific values.
[0094] At least one embodiment of the present disclosure further provides a convolutional arithmetic unit. The convolutional arithmetic unit can improve the utilization rate of the matrix arithmetic unit, effectively utilize the computing power of the matrix arithmetic unit, shorten the time of convolutional arithmetic, improve the arithmetic efficiency, and save the data transmission time.
[0095] FIG. 13 is a schematic block diagram of a convolutional arithmetic unit provided by some embodiments of the present disclosure. As shown in FIG. 13, in some embodiments, the convolutional arithmetic unit 100 includes a determination unit 110, an adjustment unit 120, and a calculation unit 130.
[0096] The determination unit 110 determines a convolutional kernel for arithmetic. For example, the convolutional kernel for arithmetic is obtained based on an initial convolutional kernel. The initial convolutional kernel is represented by [R, S, C, K], and the convolutional kernel for arithmetic is represented by [1, 1, (C×R×S), K], where R, S, C, and K are all integers greater than 0. For example, the determination unit 110 may perform step S10 of the convolutional arithmetic method shown in FIG. 2.
[0097] The adjustment unit 120 adjusts the array format of the input data based on the number of channels of the convolutional kernel for arithmetic to obtain target data. For example, the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolutional kernel for arithmetic. For example, the adjustment unit 120 may perform step S20 of the convolutional arithmetic method shown in FIG. 2.
[0098] The calculation unit 130 performs a convolutional arithmetic based on the target data and the convolutional arithmetic kernel to obtain a convolutional arithmetic result. For example, the result of the convolutional arithmetic between the target data and the convolutional kernel for arithmetic is equal to the result of the convolutional arithmetic between the input data and the initial convolutional kernel. For example, the calculation unit 130 may perform step S30 of the convolutional arithmetic method shown in FIG. 2.
[0099] For example, the determination unit 110, the adjustment unit 120, and the calculation unit 130 may be hardware, software, firmware, or any possible combination thereof. For example, the determination unit 110, the adjustment unit 120, and the calculation unit 130 may be a dedicated or general-purpose circuit, chip, or device, etc., or may be a combination of a processor and a memory. The specific implementation forms of the determination unit 110, the adjustment unit 120, and the calculation unit 130 are not limited in the embodiments of the present disclosure.
[0100] In addition, in the embodiments of the present disclosure, each part of the convolution operation device 100 corresponds to each step of the above-mentioned convolution operation method. For the specific functions of the convolution operation device 100, reference may be made to the relevant descriptions of the above-mentioned convolution operation method, but they will not be repeated here. It should be noted that the components and structures of the convolution operation device 100 shown in FIG. 13 are merely examples and are not limiting. If necessary, the convolution operation device 100 may be provided with other components and structures.
[0101] At least one embodiment of the present disclosure further provides an electronic device. The electronic device can improve the utilization rate of the matrix operation unit, effectively utilize the computing power of the matrix operation unit, shorten the time of convolution operation, improve the operation efficiency, and save the data transmission time.
[0102] FIG. 14 is a schematic block diagram of an electronic device provided by some embodiments of the present disclosure. As shown in FIG. 14, the electronic device 200 includes a convolution operation device 210. The convolution operation device 210 may be a convolution operation device provided by any embodiment of the present disclosure, for example, the above-mentioned convolution operation device 100. The electronic device 200 may be any device with computing functions such as a server, a terminal device, a personal computer, etc., but the embodiments of the present disclosure are not limited thereto.
[0103] FIG. 15 is a schematic block diagram of another electronic device provided by an embodiment of the present disclosure. As shown in FIG. 15, the electronic device 300 includes a processor 310 and a memory 320 for implementing a client or a server. The memory 320 is used to non-temporarily store computer-executable instructions (e.g., at least one (one or more) computer program modules). The processor 310 is configured to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor 310, one or more steps in the above convolution operation method can be executed, whereby the above convolution operation method is executed. The memory 320 and the processor 310 may be interconnected by a bus system and / or other forms of connection mechanisms (not shown).
[0104] For example, the processor 310 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) may be X86 or ARM architecture, etc. The processor 310 may be a general-purpose processor or a dedicated processor that can control other components in the electronic device 300 to execute desired functions.
[0105] For example, the memory 320 may include any combination of at least one (e.g., one or more) computer program products, and the computer program products may include various forms of computer-readable storage media such as, for example, volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, and the like. The computer-readable storage media may store at least one (e.g., one or more) computer program modules, and the processor 310 may execute at least one (e.g., one or more) computer program modules to implement various functions of the electronic device 300. The computer-readable storage media may further store various application programs, various data, and various data used and / or generated by the application programs.
[0106] Note that in the embodiments of the present disclosure, for the specific functions and technical effects of the electronic device 300, reference may be made to the above description regarding the convolution operation method, but it will not be repeated here.
[0107] FIG. 16 is a schematic block diagram of another electronic device provided according to an embodiment of the present disclosure. The electronic device 400 is suitable for implementing, for example, the convolution operation method provided according to an embodiment of the present disclosure. The electronic device 400 may be a terminal device for implementing a client or a server, etc. The electronic device 400 may include mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (tablets), PMPs (Portable Multimedia Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc., but is not limited thereto. Note that the electronic device 400 shown in FIG. 16 is merely an example and does not limit the functions and usage ranges of the embodiments of the present disclosure.
[0108] As shown in FIG. 16, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 410, and may execute various appropriate operations and processes according to a program stored in a read-only memory (ROM) 420 or a program loaded from a storage device 480 into a random access memory (RAM) 430. Various programs and data necessary for the operation of the electronic device 400 are further stored in the RAM 430. The processing device 410, the ROM 420, and the RAM 430 are interconnected via a bus 440. An editing / output (I / O) interface 450 is also connected to the bus 440.
[0109] Generally, the following devices can be connected to the I / O interface 450. For example, input devices 460 such as touchscreens, touch pads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc., output devices 470 such as liquid crystal displays (LCDs), speakers, vibrators, etc., storage devices 480 such as magnetic tapes, hard disks, etc., and communication devices 490. The communication device 490 may allow the electronic device 400 to communicate with other electronic devices wirelessly or wiredly to exchange data. FIG. 16 shows the electronic device 400 equipped with various devices, but it should be understood that it is not necessary to implement or include all the illustrated devices. Instead, the electronic device 400 can implement or include more or fewer devices.
[0110] For example, according to an embodiment of the present disclosure, the above convolution operation method may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program incorporated in a non-transitory computer-readable medium, and the computer program includes program code for executing the above convolution operation method. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 490, or installed from the storage device 480, or from the ROM 420. When the computer program is executed by the processing device 410, the functions defined by the convolution operation method provided by the embodiment of the present disclosure may be executed.
[0111] At least one embodiment of the present disclosure further provides a storage medium. Using the storage medium, the utilization rate of the matrix operation unit can be improved, the computing power of the matrix operation unit can be effectively utilized, the time of the convolution operation can be shortened, the operation efficiency can be improved, and the data transmission time can be saved.
[0112] FIG. 17 is a schematic block diagram of a storage medium provided by some embodiments of the present disclosure. For example, as shown in FIG. 17, the storage medium 500 may be a non-transitory computer-readable storage medium storing non-transitory computer-readable instructions 510. When the non-transitory computer-readable instructions 510 are executed by a processor, the convolution operation method described in the embodiments of the present disclosure can be realized. For example, when the non-transitory computer-readable instructions 510 are executed by a processor, one or more steps of the above convolution operation method can be executed.
[0113] For example, the storage medium 500 may be applied to the above-described electronic device. For example, the storage medium 500 may include the memory 320 in the electronic device 300.
[0114] For example, the storage medium may include a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, or other suitable storage media.
[0115] For example, the description of the storage medium 500 may refer to the description of the memory in the embodiments of the electronic device, and the same will not be repeatedly described. For the specific functions and technical effects of the storage medium 500, reference may be made to the above description regarding the convolution operation method, but will not be repeatedly described here.
[0116] As described above, the convolution operation method, convolution operation device, electronic device, and storage medium provided by the embodiments of the present disclosure have been described in conjunction with FIGS. 1 to 17. The convolution operation method provided by the embodiments of the present disclosure may be used for the convolution operation of the first layer of the convolutional neural network. By adjusting the data arrangement method, the number of channels can be increased, and the convolution operation between the target data having more channels and the convolutional kernel having more channels can be performed, thereby improving the utilization rate of the matrix operation unit, effectively utilizing the computing power of the matrix operation unit, shortening the time of the convolution operation, improving the operation efficiency, and saving the data transmission time.
[0117] In the context of the present disclosure, a computer-readable medium may be a tangible medium that includes a program used by or in combination with an instruction execution system, apparatus, or device, or is capable of storing the program. The computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable medium may be any tangible medium that includes or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated within a baseband or as part of a carrier wave and having computer-readable program code incorporated therein. The data signal thus propagated can take many forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may be any computer-readable medium other than the computer-readable storage medium and can transmit, propagate, or transfer a program used by or in combination with an instruction execution system, apparatus, or device.Program code incorporated on a computer-readable medium may be transmitted using any suitable medium including, but not limited to, wires, optical cables, RF (radio frequency), or any suitable combination thereof.
[0118] In some embodiments, the client and server can communicate using any network protocol known currently, such as HTTP (Hyper Text Transfer Protocol), or developed in the future, and can perform digital data communication (e.g., communication network) interconnection in any form or medium. Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), and networks known currently, or developed in the future.
[0119] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0120] The computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user's computer via any type of network (local area network (LAN) or wide area network (WAN)), or connected to an external computer (e.g., an Internet connection via an Internet service provider).
[0121] Flowcharts and block diagrams in the drawings illustrate the possible architectures, functionality, and operation of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram can represent a module, program segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions shown within a block may be executed in a different order than that shown in the figures. For example, two blocks shown sequentially may actually be executed substantially in parallel or in the reverse order depending on the related functions. Each block of the block diagram and / or flowchart diagram, and combinations of blocks in the block diagram and / or flowchart diagram, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of special hardware and computer instructions.
[0122] The units according to the embodiments of the present disclosure can be implemented in software or hardware. The name of a unit does not necessarily limit the module itself in some cases.
[0123] The functions described herein can be executed, at least in part, by one or more hardware logic components. For example, but not limited to, typical types of available hardware logic components include field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on chip (SOC), complex programmable logic devices (CPLD), and the like.
[0124] According to one or more embodiments of the present disclosure, the convolution operation method includes: determining a convolution kernel for operation, where the convolution kernel for operation is obtained based on an initial convolution kernel, the initial convolution kernel is represented by [R, S, C, K], the convolution kernel for operation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0; adjusting the array format of input data based on the number of channels of the convolution kernel for operation to obtain target data, where the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolution kernel for operation; and performing a convolution operation based on the target data and the convolution kernel for operation to obtain a result of the convolution operation, where the result of the convolution operation between the target data and the convolution kernel for operation is equal to the result of the convolution operation between the input data and the initial convolution kernel.
[0125] According to one or more embodiments of the present disclosure, the number of channels of the target data is greater than the number of channels of the input data, and the number of channels of the convolution kernel for operation is greater than the number of channels of the initial convolution kernel.
[0126] According to one or more embodiments of the present disclosure, the step of adjusting the array format of the input data based on the number of channels of the convolution kernel for operation to obtain target data includes: storing the input data row by row in static memory, where each row of the input data is stored in corresponding N storage rows in the static memory, and N is an integer greater than 0; performing a padding operation on the input data stored in the static memory to obtain extended data; and adjusting the array format of the extended data, and changing the size and number of channels of the extended data to obtain the target data.
[0127] According to one or more embodiments of the present disclosure, the step of storing the input data in the static memory in row units is a step of storing the input data in a dense arrangement in the memory, where the input data includes a plurality of channels, and the dense arrangement means that a plurality of channels of the same data point are stored adjacent to each other in order in the memory, and a step of transmitting the input data in the memory to the static memory of the hardware accelerator using direct memory access, and storing the first data point of each row of the input data in the first column of different rows of the static memory, and storing each row of the input data in the corresponding N storage rows in the static memory.
[0128] According to one or more embodiments of the present disclosure, the step of performing the padding operation on the input data stored in the static memory and obtaining the extended data includes a step of padding the storage rows before and after the storage position of the input data in the static memory with a first preset value, and obtaining first intermediate data including the input data and the padded first preset value, and transmitting the first intermediate data to a vector calculation unit, and padding both ends of each row corresponding to the first intermediate data with a second preset value using the shift instruction and the padding instruction of the vector calculation unit, and obtaining second intermediate data including the first intermediate data and the padded second preset value, and transmitting the second intermediate data to the corresponding storage position in the static memory to obtain the extended data having the same content as the second intermediate data.
[0129] According to one or more embodiments of the present disclosure, the target data is represented by [1, ht, wt, (C×R×S)], where both ht and wt are integers greater than 0. The step of adjusting the array format of the extended data and changing the size and number of channels of the extended data to obtain the target data includes: successively reading the data of R×N storage rows in the static memory and transmitting it to the vector calculation unit. The starting address read each time is shifted by str*N storage rows according to a preset hop count str. The preset hop count is the hop count in the row direction and column direction of the sliding window required for the convolution operation between the input data and the initial convolution kernel. The total number of times of reading data from the static memory is equal to ht. The vector calculation unit converts the data in the R*N storage rows received each time into the data of wt*ceil((C×R×S) / L) storage rows to obtain the target data, where L represents the number of data points that can be stored in each storage row in the static memory, and ceil((C×R×S) / L) represents rounding up to the nearest integer for (C×R×S) / L. The converted data is the target data.
[0130] According to one or more embodiments of the present disclosure, the vector calculation unit converts the data in each received R*N storage rows into the data in wt*ceil((C×R×S) / L) storage rows, and the step of obtaining the target data is a step of dividing the data in R*N storage rows into a plurality of groups according to the preset number of hops, where the data in each group corresponds to a sliding window in the row direction, and the number of data in the plurality of groups is equal to wt; for the data in each group, determining the initial position information parameter and the target position information parameter of the data in each row within the sliding window corresponding to the data in the group; the vector calculation unit stores the data in each group at the corresponding position in the target memory in the converted array format according to the initial position information parameter and the target position information parameter, and the step of obtaining the target data, where the target memory is stored in row units, transmitted to the target memory, and the data stored on the target memory is target data.
[0131] According to one or more embodiments of the present disclosure, the initial position information parameter includes a first start boundary coordinate, a first end boundary coordinate, a first start address, a first end address, a first start number, and a first end number. The first start boundary coordinate represents the relative coordinate in the row direction of the extended data of the start boundary of the corresponding sliding window, the first end boundary coordinate represents the relative coordinate in the row direction of the extended data of the end boundary of the corresponding sliding window, the start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data, the first start address represents the address in the memory (e.g., vmem) of the vector calculation unit of the first start boundary coordinate, the first end address represents the address in the memory of the vector calculation unit of the first end boundary coordinate, the first start number represents the number of the data point corresponding to the first start address of the first start boundary coordinate, and the first end number represents the number of the data point corresponding to the first end address of the first end boundary coordinate.
[0132] According to one or more embodiments of the present disclosure, the calculation formula for the first start boundary coordinate is src_row_start_index = i * str * ch, where src_row_start_index represents the first start boundary coordinate, i represents the number of the corresponding data point in the size wt of the target data of the corresponding sliding window, str represents the number of hops in the row direction of the sliding window, ch represents the number of channels of the input data. The calculation formula for the first end boundary coordinate is src_row_end_index = src_row_start_index + (kernel_w * ch - 1), where src_row_end_index represents the first end boundary coordinate, and kernel_w represents the width of the sliding window. The size of the sliding window is equal to the size of the initial convolution kernel. The calculation formula for the first start address is src_row_start_address = src_row_start_index / vmem_lane + j * N, where src_row_start_address represents the first start address, vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit, j represents the row number of the corresponding data in the sliding window. The calculation formula for the first end address is src_row_end_address = src_row_end_index / vmem_lane + j * N, where src_row_end_address represents the first end address. The calculation formula for the first start number is src_row_start_lane = src_row_start_index % vmem_lane, where src_row_start_lane represents the first start number, and % represents the modulo operation. The calculation formula for the first end number is src_row_end_lane = src_row_end_index % vmem_lane, where src_row_end_lane represents the first end number.
[0133] According to one or more embodiments of the present disclosure, the target position information parameter includes a second start boundary coordinate, a second end boundary coordinate, a second start address, a second end address, a second start number, and a second end number. The second start boundary coordinate represents the relative coordinate within the data size of [1, 1, (C × R × S)] of the start boundary of the corresponding sliding window, and the second end boundary coordinate represents the relative coordinate within the data size of [1, 1, (C × R × S)] of the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. The second start address represents the address in the memory of the vector calculation unit of the second start boundary coordinate, the second end address represents the address in the memory of the vector calculation unit of the second end boundary coordinate, the second start number represents the number of the data point corresponding to the second start address of the second start boundary coordinate, and the second end number represents the number of the data point corresponding to the second end address of the second end boundary coordinate.
[0134] According to one or more embodiments of the present disclosure, the calculation formula for the second start boundary coordinate is dst_row_start_index = j * kernel_w * ch, where dst_row_start_index represents the second start boundary coordinate, j represents the row number within the sliding window of the corresponding data, kernel_w represents the width of the sliding window, the size of the sliding window is equal to the size of the initial convolution kernel, ch represents the number of channels of the input data, the calculation formula for the second end boundary coordinate is dst_row_end_index = dst_row_start_index + (kernel_w * ch - 1), where dst_row_end_index represents the second end boundary coordinate, the calculation formula for the second start address is dst_row_start_address = dst_row_start_index / vmem_lane, where dst_row_start_address represents the second start address, vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit, the calculation formula for the second end address is dst_row_end_address = dst_row_end_index / vmem_lane, where dst_row_end_address represents the second end address, the calculation formula for the second start number is dst_row_start_lane = dst_row_start_index % vmem_lane, where dst_row_start_lane represents the second start number, % represents the modulo operation, and the calculation formula for the second end number is dst_row_end_lane = dst_row_end_index % vmem_lane, where dst_row_end_lane represents the second end number.
[0135] According to one or more embodiments of the present disclosure, the vector calculation unit stores data of each group in a corresponding position of the target memory in a converted array format according to the initial position information parameter and the target position information parameter, and the step of obtaining the target data is that according to the initial position information parameter and the target position information parameter, the vector calculation unit uses a circular shift instruction and concatenates data of each group in a converted array format according to a preset enable signal in a predicate register, stores the data in a corresponding position of the target memory, and includes the step of obtaining the target data.
[0136] According to one or more embodiments of the present disclosure, the step of performing a convolution operation based on the target data and the operation convolution kernel is including the step of performing a convolution operation on the target data and the convolution kernel for operation by a matrix operation unit.
[0137] According to one or more embodiments of the present disclosure, the convolution operation method is used for the convolution operation of the first layer of a convolutional neural network.
[0138] According to one or more embodiments of the present disclosure, the convolutional operation device includes a determination unit that determines a convolutional kernel for operation, where the convolutional kernel for operation is obtained based on an initial convolutional kernel, the initial convolutional kernel is represented by [R, S, C, K], the convolutional kernel for operation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0; an adjustment unit that adjusts the array format of the input data based on the number of channels of the convolutional kernel for operation to obtain target data, where the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolutional kernel for operation; and a calculation unit that performs a convolutional operation based on the target data and the convolutional kernel for operation to obtain the result of the convolutional operation, where the result of the convolutional operation between the target data and the convolutional kernel for operation is equal to the result of the convolutional operation between the input data and the initial convolutional kernel.
[0139] According to one or more embodiments of the present disclosure, the number of channels of the target data is greater than the number of channels of the input data, and the number of channels of the convolutional kernel for operation is greater than the number of channels of the initial convolutional kernel.
[0140] According to one or more embodiments of the present disclosure, the adjustment unit includes a first sub-adjustment unit, a second sub-adjustment unit, and a third sub-adjustment unit. The first sub-adjustment unit is configured to store the input data in a static memory row by row, and each row of the input data is stored in corresponding N storage rows in the static memory, where N is an integer greater than 0. The second sub-adjustment unit performs a padding operation on the input data stored in the static memory to obtain extended data. The third sub-adjustment unit adjusts the array format of the extended data, changes the size and number of channels of the extended data, and obtains the target data.
[0141] According to one or more embodiments of the present disclosure, the first sub-adjustment unit includes a first storage unit and a second storage unit. The first storage unit is configured to densely arrange the input data and store it in a memory. The input data includes a plurality of channels, and the dense arrangement means that a plurality of channels of the same data point are stored adjacent to each other in order in the memory. The second storage unit uses direct memory access to transmit the input data in the memory to the static memory of the hardware accelerator, and stores the first data point of each row of the input data in the first column of different rows of the static memory, and stores each row of the input data in the corresponding N storage rows in the static memory.
[0142] According to one or more embodiments of the present disclosure, the second sub-adjustment unit includes a first filling unit, a second filling unit, and a third filling unit. The first filling unit fills the storage rows before and after the storage position of the input data in the static memory with a first preset value to obtain first intermediate data. Here, the first intermediate data includes the input data and the filled first preset value. The second filling unit transmits the first intermediate data to a vector calculation unit, and is configured to fill both ends of each row corresponding to the first intermediate data with a second preset value using the shift instruction and the filling instruction of the vector calculation unit to obtain second intermediate data. Here, the second intermediate data includes the first intermediate data and the filled second preset value. The third filling unit transmits the second intermediate data to the corresponding storage position in the static memory to obtain the extended data. The extended data has the same content as the second intermediate data.
[0143] According to one or more embodiments of the present disclosure, the target data is represented by [1, ht, wt, (C×R×S)], where both ht and wt are integers greater than 0. The third sub-adjustment unit includes a first modification unit and a second modification unit. The first modification unit is configured to sequentially read the data of R×N storage rows in the static memory and transmit it to the vector calculation unit. The start address read each time is moved by str*N storage rows according to a preset hop count str. The preset hop count is the hop count in the row direction and column direction of the sliding window required for the convolution operation of the input data and the initial convolution kernel. The total number of times of reading data from the static memory is equal to ht. The second modification unit is configured to convert the data of R*N storage rows received each time into the data of wt*ceil((C×R×S) / L) storage rows to obtain the target data. L represents the number of data points that can be stored in each storage row in the static memory, and ceil((C×R×S) / L) represents rounding up to the nearest integer for (C×R×S) / L. The converted data is the target data.
[0144] According to one or more embodiments of the present disclosure, the second modification part includes a grouping part, a parameter determination part, and a vector calculation part. The grouping part is configured to divide the data in R*N storage rows into a plurality of groups according to the preset number of hops. The data of each group corresponds to a sliding window in the row direction, and the number of data in the plurality of groups is equal to wt. The parameter determination part is configured to determine, for the data of each group, the initial position information parameter and the target position information parameter of the data in each row within the sliding window corresponding to the data of the group. The vector calculation part is configured to store the data of each group at the corresponding position in the target memory in the converted array format according to the initial position information parameter and the target position information parameter, and obtain the target data. The target memory is stored in row units, transmitted to the target memory, and the data stored on the target memory is the target data.
[0145] According to one or more embodiments of the present disclosure, the initial position information parameter includes a first start boundary coordinate, a first end boundary coordinate, a first start address, a first end address, a first start number, and a first end number. The first start boundary coordinate represents the relative coordinate in the row direction of the extended data of the start boundary of the corresponding sliding window. The first end boundary coordinate represents the relative coordinate in the row direction of the extended data of the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. The first start address represents the address in the memory of the vector calculation part of the first start boundary coordinate. The first end address represents the address in the memory of the vector calculation part of the first end boundary coordinate. The first start number represents the number of the data point corresponding to the first start address of the first start boundary coordinate. The first end number represents the number of the data point corresponding to the first end address of the first end boundary coordinate.
[0146] According to one or more embodiments of the present disclosure, the calculation formula for the first start boundary coordinate is src_row_start_index = i * str * ch, where src_row_start_index represents the first start boundary coordinate, i represents the number of the corresponding data point in the size wt of the target data of the corresponding sliding window, str represents the number of hops in the row direction of the sliding window, ch represents the number of channels of the input data, the calculation formula for the first end boundary coordinate is src_row_end_index = src_row_start_index + (kernel_w * ch - 1), where src_row_end_index represents the first end boundary coordinate, kernel_w represents the width of the sliding window, the size of the sliding window is equal to the size of the initial convolution kernel, the calculation formula for the first start address is src_row_start_address = src_row_start_index / vmem_lane + j * N, where src_row_start_address represents the first start address, vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit, j represents the row number of the corresponding data in the sliding window, the calculation formula for the first end address is src_row_end_address = src_row_end_index / vmem_lane + j * N, where src_row_end_address represents the first end address, the calculation formula for the first start number is src_row_start_lane = src_row_start_index % vmem_lane, where src_row_start_lane represents the first start number, % represents the modulo operation, and the calculation formula for the first end number is src_row_end_lane = src_row_end_index % vmem_lane, where src_row_end_lane represents the first end number.
[0147] According to one or more embodiments of the present disclosure, the target position information parameter includes a second start boundary coordinate, a second end boundary coordinate, a second start address, a second end address, a second start number, and a second end number. The second start boundary coordinate represents the relative coordinate within the data size of [1, 1, (C×R×S)] of the start boundary of the corresponding sliding window, the second end boundary coordinate represents the relative coordinate within the data size of [1, 1, (C×R×S)] of the end boundary of the corresponding sliding window, the start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data, the second start address represents the address in the memory of the vector calculation unit of the second start boundary coordinate, the second end address represents the address in the memory of the vector calculation unit of the second end boundary coordinate, the second start number represents the number of the data point corresponding to the second start address of the second start boundary coordinate, and the second end number represents the number of the data point corresponding to the second end address of the second end boundary coordinate.
[0148] According to one or more embodiments of the present disclosure, the calculation formula for the second start boundary coordinate is dst_row_start_index = j * kernel_w * ch, where dst_row_start_index represents the second end boundary coordinate, j represents the row number within the sliding window of the corresponding data, kernel_w represents the width of the sliding window, the size of the sliding window is equal to the size of the initial convolution kernel, ch represents the number of channels of the input data, the calculation formula for the second end boundary coordinate is dst_row_end_index = dst_row_start_index + (kernel_w * ch - 1), where dst_row_end_index represents the second end boundary coordinate, the calculation formula for the second start address is dst_row_start_address = dst_row_start_index / vmem_lane, where dst_row_start_address represents the second start address, vmem_lane represents the number of data points that can be stored in each memory row within the memory of the vector calculation unit, the calculation formula for the second end address is dst_row_end_address = dst_row_end_index / vmem_lane, where dst_row_end_address represents the second end address, the calculation formula for the second start number is dst_row_start_lane = dst_row_start_index % vmem_lane, where dst_row_start_lane represents the second start number, % represents the modulo operation, and the calculation formula for the second end number is dst_row_end_lane = dst_row_end_index % vmem_lane, where dst_row_end_lane represents the second end number.
[0149] According to one or more embodiments of the present disclosure, the vector calculation unit is further configured to concatenate data of each group in a converted array format according to the initial position information parameter and the target position information parameter, using a circular shift instruction and according to a preset enable signal in a predicate register, store the data in a corresponding position of the target memory, and obtain the target data.
[0150] According to one or more embodiments of the present disclosure, the calculation unit includes the sub-calculation unit, and the sub-calculation unit is configured to perform a convolution operation on the target data and the convolution kernel for calculation using a matrix operation unit.
[0151] According to one or more embodiments of the present disclosure, the convolution operation device is used for the convolution operation of the first layer of a convolutional neural network.
[0152] According to one or more embodiments of the present disclosure, an electronic device includes a convolution operation device provided by any embodiment of the present disclosure.
[0153] According to one or more embodiments of the present disclosure, an electronic device includes a processor and a memory including at least one computer program module, the at least one computer program module is stored in the memory and configured to be executed by the processor, and the at least one computer program module is used to implement the convolution operation method provided by any embodiment of the present disclosure.
[0154] According to one or more embodiments of the present disclosure, there is a storage medium storing non-temporary computer-readable instructions, and when the non-temporary computer-readable instructions are executed by a computer, the convolution operation method provided by any embodiment of the present disclosure is realized.
[0155] The above description is only an explanation of more preferred embodiments of the present disclosure and the applied technical principles. Those skilled in the art should understand that the disclosure scope of the present disclosure is not limited to the technical means consisting of specific combinations of the above technical features, and other technical means formed by any combination of the above technical features or their equivalent features should also be covered without departing from the concepts disclosed above. For example, technical means formed by replacing the above features with technical features having similar functions (but not limited to this) disclosed in the present disclosure are also covered.
[0156] Also, although various operations are shown in a specific order, this should not be understood as requiring that these operations be performed in the specific order shown or in a sequential order. In certain situations, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above description, these should not be construed as limiting the scope of the present disclosure. Specific features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0157] Although this subject matter has been described in language specific to structural features and / or logical operations of methods, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the above specific features and operations are merely exemplary forms for implementing the claims.
[0158] There are still some points that need to be explained regarding this disclosure. (1) The drawings of the embodiments of the present disclosure only include the structures related to the embodiments of the present disclosure, and common designs can be referred to for other structures. (2) The embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain a new embodiment without contradiction.
[0159] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and the protection scope of the present disclosure should follow the protection scope of the claims.
Claims
1. A method for convolution operation, the method comprising: Determining a convolution kernel for operation, wherein the convolution kernel for operation is obtained based on an initial convolution kernel, the initial convolution kernel is represented by [R, S, C, K], the convolution kernel for operation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0; Adjusting an array format of input data based on the number of channels of the convolution kernel for operation to obtain target data, wherein a size and the number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolution kernel for operation; Performing a convolution operation based on the target data and the convolution kernel for operation to obtain a result of the convolution operation, wherein the result of the convolution operation between the target data and the convolution kernel for operation is equal to the result of the convolution operation between the input data and the initial convolution kernel. The method.
2. The method according to claim 1, wherein the number of channels of the target data is greater than the number of channels of the input data, and the number of channels of the convolution kernel for operation is greater than the number of channels of the initial convolution kernel.
3. The step of adjusting an array of the input data based on the number of channels of the convolution kernel for operation to obtain target data comprises: Storing the input data in a static memory row by row, wherein each row of the input data is stored in corresponding N storage rows in the static memory, and N is an integer greater than 0; Performing a padding operation on the input data stored in the static memory to obtain extended data; Adjusting an array format of the extended data, and changing a size and the number of channels of the extended data to obtain the target data. The method according to claim 1 or 2.
4. The step of storing the input data in the static memory row by row comprises: A step of densely arranging the input data and storing it in a memory, wherein the input data includes a plurality of channels, and the dense arrangement means that a plurality of channels of the same data point are stored adjacent to each other in order in the memory, the step; Using direct memory access to transmit the input data in the memory to the static memory of the hardware accelerator, storing the first data point of each row of the input data in the first column of different rows of the static memory, and storing each row of the input data in the corresponding N storage rows in the static memory, the method according to claim 3, comprising:
5. Performing the padding operation on the input data stored in the static memory, the step of obtaining the extended data is: In the static memory, filling the storage rows before and after the storage position of the input data with a first preset value, and obtaining first intermediate data including the input data and the filled first preset value; Transmitting the first intermediate data to a vector calculation unit, and using the shift instruction and the padding instruction of the vector calculation unit to fill both ends of each row corresponding to the first intermediate data with a second preset value, and obtaining second intermediate data including the first intermediate data and the filled second preset value; Transmitting the second intermediate data to the corresponding storage position in the static memory to obtain extended data having the same content as the second intermediate data, the method according to claim 3, comprising:
6. The target data is represented by [1, ht, wt, (C×R×S)], where both ht and wt are integers greater than 0. Adjusting the array method of the extended data, changing the size and the number of channels of the extended data, the step of obtaining the target data is: A step of successively reading data of R*N memory rows in the static memory and transmitting the data to the vector calculation unit, wherein the start address read each time is shifted by str*N memory rows according to a preset hop count str, and the preset hop count str is the hop count in the row direction and the column direction of a sliding window required for the convolution operation of the input data and the initial convolution kernel, and the total number of times of reading data from the static memory is equal to ht, and the step; A step of converting, by the vector calculation unit, the data of R*N memory rows received each time into data of wt*ceil((C×R×S) / L) memory rows to obtain the target data, where L represents the number of data points storable in each memory row in the static memory, ceil((C×R×S) / L) represents rounding up to the nearest integer for (C×R×S) / L, and the converted data is the target data, and the step; The method according to claim 5, comprising the step.
7. The step in which the vector calculation unit converts the data of R*N memory rows received each time into data of wt*ceil((C×R×S) / L) memory rows to obtain target data is as follows: A step of dividing the data in the R*N memory rows into a plurality of groups according to a preset hop count, wherein the data of each group corresponds to a sliding window in the row direction, and the number of the plurality of groups of data is equal to wt, and the step; A step of determining, for the data of each group, an initial position information parameter and a target position information parameter of the data of each row in the sliding window corresponding to the data of the group; A step of storing, by the vector calculation unit, the data of each group in a corresponding position of the target memory in a converted array format according to the initial position information parameter and the target position information parameter to obtain the target data, wherein the target memory is stored in row units, transmitted to the target memory, and the data stored in the target memory is target data, and the step; The method according to claim 6, comprising the step.
8. The initial position information parameter includes a first start boundary coordinate, a first end boundary coordinate, a first start address, a first end address, a first start number, and a first end number. The first start boundary coordinate represents the relative coordinate in the row direction of the extended data at the start boundary of the corresponding sliding window. The first end boundary coordinate represents the relative coordinate in the row direction of the extended data at the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. The first start address represents the address in the memory of the vector calculation unit corresponding to the first start boundary coordinate. The first end address represents the address in the memory of the vector calculation unit corresponding to the first end boundary coordinate. The first start number represents the number of the data point corresponding to the first start address corresponding to the first start boundary coordinate. The first end number represents the number of the data point corresponding to the first end address corresponding to the first end boundary coordinate. The method according to claim 7. **Claim 9** The calculation formula for the first start boundary coordinate is src_row_start_index = i * str * ch. src_row_start_index represents the first start boundary coordinate. i represents the number of the corresponding data point in the size wt of the target data of the corresponding sliding window. str represents the number of hops in the row direction of the sliding window. ch represents the number of channels of the input data. The calculation formula for the first end boundary coordinate is src_row_end_index = src_row_start_index + (kernel_w * ch - 1). src_row_end_index represents the first end boundary coordinate. kernel_w represents the width of the sliding window. The size of the sliding window is equal to the size of the initial convolution kernel. The calculation formula for the first start address is src_row_start_address = src_row_start_index / vmem_lane + j * N, where src_row_start_address represents the first start address, vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit, j represents the row number of the corresponding data in the sliding window, The calculation formula for the first end address is src_row_end_address = src_row_end_index / vmem_lane + j * N, where src_row_end_address represents the first end address, The calculation formula for the first start number is src_row_start_lane = src_row_start_index % vmem_lane, where src_row_start_lane represents the first start number, and % represents the modulo operation, The calculation formula for the first end number is src_row_end_lane = src_row_end_index % vmem_lane, where src_row_end_lane represents the first end number, according to the method described in claim 8.
10. The target position information parameters include a second start boundary coordinate, a second end boundary coordinate, a second start address, a second end address, a second start number, and a second end number. The second start boundary coordinate represents the relative coordinate in the data size of [1, 1, (C × R × S)] of the start boundary of the corresponding sliding window, and the second end boundary coordinate represents the relative coordinate in the data size of [1, 1, (C × R × S)] of the end boundary of the corresponding sliding window. The start boundary of the corresponding sliding window and the end boundary of the corresponding sliding window are located at different positions in the row direction of the extended data. The second start address represents the address in the memory of the vector calculation unit of the second start boundary coordinate, and the second end address represents the address in the memory of the vector calculation unit of the second end boundary coordinate. The second start number represents the number of the data point corresponding to the second start address of the second start boundary coordinate, and the second end number represents the number of the data point corresponding to the second end address of the second end boundary coordinate, according to the method of claim 7.
11. The calculation formula for the second start boundary coordinate is dst_row_start_index = j * kernel_w * ch, where dst_row_start_index represents the second start boundary coordinate, j represents the row number in the sliding window of the corresponding data, kernel_w represents the width of the sliding window, the size of the sliding window is equal to the size of the initial convolution kernel, and ch represents the number of channels of the input data. The calculation formula for the second end boundary coordinate is dst_row_end_index = dst_row_start_index + (kernel_w * ch - 1), where dst_row_end_index represents the second end boundary coordinate. The calculation formula for the second start address is dst_row_start_address = dst_row_start_index / vmem_lane, where dst_row_start_address represents the second start address, and vmem_lane represents the number of data points that can be stored in each memory row in the memory of the vector calculation unit. The calculation formula for the second end address is dst_row_end_address = dst_row_end_index / vmem_lane, where dst_row_end_address represents the second end address. The calculation formula for the second start number is dst_row_start_lane = dst_row_start_index % vmem_lane, where dst_row_start_lane represents the second start number, and % represents the modulo operation. The calculation formula for the second end number is dst_row_end_lane = dst_row_end_index % vmem_lane, where dst_row_end_lane represents the second end number, according to the method of claim 10.
12. The vector calculation unit stores data of each group in corresponding positions of the target memory in a converted array format according to the initial position information parameter and the target position information parameter, and the step of obtaining the target data is The method according to claim 7, wherein the vector calculation unit concatenates data of each group in a converted array format according to the initial position information parameter and the target position information parameter, in accordance with a preset enable signal in a predicate register, using a circular shift instruction, stores the data in corresponding positions of the target memory, and includes the step of obtaining the target data.
13. The step of performing a convolution operation based on the target data and the convolution kernel for calculation is The method according to any one of claims 1 to 12, further including the step of performing a convolution operation on the target data and the convolution kernel for calculation by a matrix operation unit.
14. The convolution operation method is used for the convolution operation of the first layer of a convolutional neural network. The method according to any one of claims 1 to 13.
15. An apparatus for convolution operation, the apparatus comprising A determination unit for determining a convolution kernel for calculation, wherein the convolution kernel for calculation is obtained based on an initial convolution kernel, the initial convolution kernel is represented by [R, S, C, K], the convolution kernel for calculation is represented by [1, 1, (C×R×S), K], and R, S, C, and K are all integers greater than 0; a determination unit An adjustment unit for adjusting the array format of input data based on the number of channels of the convolution kernel for calculation to obtain target data, wherein the size and number of channels of the target data are different from those of the input data, and the number of channels of the target data is equal to the number of channels of the convolution kernel for calculation; an adjustment unit A calculation unit for performing a convolution operation based on the target data and the convolution kernel for calculation and obtaining a result of the convolution operation, wherein the result of the convolution operation between the target data and the convolution kernel for calculation is equal to the result of the convolution operation between the input data and the initial convolution kernel; a calculation unit, comprising Apparatus.
16. An electronic device including the apparatus according to claim 15.
17. An electronic device, the electronic device comprising: a processor; and a memory including at least one computer program module, wherein the at least one computer program module is stored in the memory and configured to be executed by the processor, and the at least one computer program module is used to implement the method according to any one of claims 1 to 14. An electronic device.
18. A storage medium storing non-transitory computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 14 is executed.
Citation Information
Patent Citations
Method and electronic device for performing convolution calculations in neutral network
JP2019109896A
Method and System for Multi-Scale Vision Transformer Architecture
US20230401825A1