A design method for an independent convolution of 3×3 based on 4-bit to 6-bit
By adopting an independent convolution 3×3 design method based on 4bit to 6bit on the Beijing Junzheng chip, the convolution kernel data is cross-stored and the feature map data is adjusted in sequence, and the convolution calculation is optimized using the simd instruction, the problems of slow image processing speed and low computing efficiency in the existing technology are solved, and significant speed improvement and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202110672459.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-06-17
AI Technical Summary
In the prior art, when the chip produced by Beijing Junzheng performs image processing, the use of C programs is slow, limited instructions leads to a long run time, and the storage and calculation efficiency of convolutional kernel data is low.
Using the 3×3 design method of independent convolution based on 4bit to 6bit, the convolution kernel data is cross-stored and feature map data is adjusted in sequence, and the simd instruction is used to multiply and add operations to optimize the convolution calculation process.
This significantly improves the speed of image processing, which is nearly 40 times faster than using C programs, reduces calculation time and improves efficiency.
Smart Images

Figure CN115495156B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a design method based on 4-bit to 6-bit independent convolution 3×3. Background Art
[0002] Integrated circuit technology is now becoming more and more the focus of technology, and many chip manufacturers are also developing their own chips. In chip applications, each chip design will also have its own problems. For example, the chips produced by Beijing Ingenic Integrated Circuit Co., Ltd., on Beijing Ingenic's chips, directly using C programs is slow, and the unreasonable use of limited instructions leads to slow running time. In addition, for example, the registers of Beijing Ingenic's chips T30 and T31 are 128-bit registers, and the number of registers is limited. This problem of the number of registers must be considered in the optimization design; the SIMD instruction set is limited, and some operations require the use of several instructions to implement their operations. Moreover, the limited data of the convolution kernel data is 4bit, 5bit or 6bit, and it is stored as 8bit. The depth of the output data, that is, the feature map, is a multiple of 16. The step size used in the convolution calculation is 1, 2 or 3. In actual use, the step size does not exceed the length and width of the convolution kernel or loses its meaning. The width of the convolution kernel is 3 and the length is 3. Therefore, on Beijing Ingenic chips, directly using C programs is very slow, and the unreasonable use of limited instructions also leads to slow running time.
[0003] In addition, the commonly used terms in the prior art are as follows:
[0004] 1. SIMD instruction: Single instruction stream multiple data stream, that is, one operation instruction can execute multiple data streams, which can improve the operation speed of the program. In a more popular understanding, it is a kind of vector (vector) calculation. Different chips have different specific instruction sets.
[0005] 2. Convolution kernel: The convolution kernel is a matrix used for image processing and a parameter for operating with the original image. The convolution kernel is usually composed of a column matrix (for example, a 3*3 matrix), and each square in the area has a weight value. The matrix shape is generally 1×1, 3×3, 5×5, 7×7, 1×3, 3×1, 2×2, 1×5, 5×1, ...
[0006] 3. Convolution: Place the center of the convolution kernel on the pixel to be calculated, calculate the product of each element in the kernel and the image pixel value it covers and sum them up. The resulting structure is the new pixel value at that position. This process is called convolution.
[0007] 4. Independent convolution: One convolution kernel is responsible for one channel, and one channel is convolved by only one convolution kernel.
[0008] 5. Feature Map: The result obtained after the input data is calculated by convolution is called the feature map (or output data), and the result generated after the data passes through the fully connected layer is also called the feature map (or output data). The size of the feature map is generally expressed as length × width × depth, or 1 × depth. Summary of the Invention
[0009] In order to solve the above problems in the prior art, the purpose of the present application is to: achieve a multiple increase in speed, especially compared to using a C program to increase speed.
[0010] This optimization method is an optimization method designed based on the simd instruction set in T series chips such as T30 and T31 produced by chip manufacturers, especially Beijing Junzheng. This method is suitable for the operation of vector (vector) instructions.
[0011] Specifically, the present invention provides a design method for an independent convolution 3×3 based on 4-bit to 6-bit. The method is to pre-transform the order of the data that needs to be optimized and transformed, that is, during the storage process of the convolution kernel data, cross-store the data of two adjacent depths, and cross-store the data of the last depth with 0; and during the subsequent convolution calculation process, adjust the order of the original feature map data, and then use the simd instruction of multiplying and then adding adjacent ones, so that the data after multiplying and adding 8-bit data becomes 16-bit; adjust the order of the convolution kernel data and store it in the order required during use. The convolution kernel data has been cross-processed using relevant methods before use or before calculation. When using specific simd instructions, the operation cannot be performed according to the original feature number data order. The order must be adjusted to the data order required by the instructions to be used first in order to obtain the correct result. Here, the loaded data is adjusted in order, and then the simd instruction operation is performed. There is another instruction, which is to perform the simd instruction operation while retaining the order of the data in the original feature map. However, a lot of processing and transformation are required later, and the efficiency of this instruction is very low. Therefore, this application does not use it.
[0012] The method further includes the following steps:
[0013] S1, input data and store it. Among them, the processing of optimizing the storage method of convolution kernel data:
[0014] Since the subsequent convolution calculation is an independent convolution, after 9 pairs of data are multiplied and then accumulated together, a single data is generated; and in the instruction set, only one instruction can satisfy that the number of bits of the result of multiplying and then adding adjacent ones is 16 bits; multiplying two 4-bit numbers and then adding them, and then accumulating 9 times later (the convolution kernel is 3x3), it will not exceed 16 bits, but may exceed 8 bits. Therefore, it is sufficient to require the result to be 16 bits. At the same time, the operation efficiency can be improved. The amount of loaded data will increase.
[0015] According to the use of this instruction, the data needs to be cross-stored, and a 0 needs to be added;
[0016] Cross-store the data of two adjacent depths, and cross-store the data of the last depth with 0; S2, use the simd instruction to design the simd instruction optimization algorithm to implement convolution calculation:
[0017] For the 3×3 depth data in the feature map for convolution, every two depths are grouped for cross-processing, which can form 4 groups. At the same time, there is one remaining depth data, and the remaining one depth data is combined with 0. At this time, 5 groups of depth data are formed;
[0018] S3, perform corresponding calculations on the 5 groups of depth data formed in step S2 and the convolution kernel data that has been processed in step S1. All data are multiplied corresponding to each other and then accumulated.
[0019] In step S1, the input data is stored in the order of depth first, width second, and height last; during the calculation, the spatial structure of the data is considered, and in the storage, it is a vector storage method.
[0020] Initial storage of the convolution kernel in step S1: Set a convolution kernel data-related information, the output depth out_dep is 4, the convolution kernel width is 3, and the convolution kernel height is 3; the data continuity method is continuous in the output depth direction, and then there are 3 groups in the output depth direction, with 4 data in each direction, and 12 data form a group of data in the width direction. Finally, 3 groups of data in the width direction form the data in the height direction.
[0021] In actual use, the output depth used by the method is a multiple of 16.
[0022] Step S1 further includes:
[0023] Set the input data indata as a group of data with an input depth in_depth of 32, a width in_width of 256, and a height in_height of 256;
[0024] The convolution kernel data filter_data is a group of data with an output depth out_depth of 32, a convolution kernel width ft_w of 3, and a convolution kernel height ft_h of 3;
[0025] Set the structure of the output data, that is, the feature map outdata:
[0026] The depth is out_depth, which is the same as the input depth of the input data here, which is 32, the width is out_width, and the height is out_height;
[0027] In convolution calculation, there is a stride, let the stride be stride:
[0028] The 3×3 input data in the input depth direction are:
[0029] [ax1, ax2, ax3, ax4, ax5, ax6, ax7, ax8, ax9, ax10, ax11, ax12, ax13, ax14, ax15, ax16, … ax32]
[0030] [bx1, bx2, bx3, bx4, bx5, bx6, bx7, bx8, bx9, bx10, bx11, bx12, bx13, bx14, bx15, bx16, … bx32]
[0031] [cx1, cx2, cx3, cx4, cx5, cx6, cx7, cx8, cx9, cx10, cx11, cx12, cx13, cx14, cx15, cx16, … cx32]
[0032] [dx1, dx2, dx3, dx4, dx5, dx6, dx7, dx8, dx9, dx10, dx11, dx12, dx13, dx14, dx15, dx16, … dx32]
[0033] [ex1, ex2, ex3, ex4, ex5, ex6, ex7, ex8, ex9, ex10, ex11, ex12, ex13, ex14, ex15, ex16, … ex32]
[0034] [fx1, fx2, fx3, fx4, fx5, fx6, fx7, fx8, fx9, fx10, fx11, fx12, fx13, fx14, fx15, fx16, … fx32]
[0035] [gx1, gx2, gx3, gx4, gx5, gx6, gx7, gx8, gx9, gx10, gx11, gx12, gx13, gx14, gx15, gx16, … gx32]
[0036] [hx1, hx2, hx3, hx4, hx5, hx6, hx7, hx8, hx9, hx10, hx11, hx12, hx13, hx14, hx15, hx16, … hx32]
[0037] [jx1,jx2,jx3,jx4,jx5,jx6,jx7,jx8,jx9,jx10,jx11,jx12,jx13,jx14,jx15,jx16,…jx32]-----(7)
[0038] The original data of the convolution kernel is:
[0039] [a1,a2,…,a32;
[0040] b1,b2…,b32;
[0041] c1,c2,…,c32;
[0042] d1,d2,…,d32;
[0043] e1,e2,…,e32;
[0044] f1,f2,…,f32;
[0045] g1,g2,…,g32;
[0046] h1,h2,…,h32;
[0047] j1,j2,…,j32]; -----(8)
[0048] The optimized data of the convolution kernel is:
[0049] [a1,b1,a2,b2,…,a16,b16;
[0050] c1,d1,c2,d2,…,c16,d16;
[0051] e1,f1,e2,f2,…,e16,f16;
[0052] g1,h1,g2,h2,…,g16,h16;
[0053] j1,0,j2,0,…,j16,0];
[0054] [a17,b17,a18,b18,…,a32,b32;
[0055] c17,d17,c18,d18,…,c32,d32;
[0056] e17,f17,e18,f18,…,e32,f32;
[0057] g17,h17,g18,h18,…,g32,h32;
[0058] [j17,0,j18,0,…,j32,0]; -----(9).
[0059] In step S2, for the 3×3 depth data in the feature map for convolution, such as data (7);
[0060] The composition of 5 groups of depth data, such as data (10): The data structure is:
[0061] [ax1,bx1,ax2,bx2,…,ax16,bx16,ax17,bx17,…,ax32,bx32]
[0062] [cx1,dx1,cx2,dx2,…,cx16,dx16,cx17,dx17,…,cx32,dx32]
[0063] [ex1,fx1,ex2,fx2,…,ex16,fx16,ex17,fx17,…,ex32,fx32]
[0064] [gx1,hx1,gx2,hx2,…,gx16,hx16,gx17,hx17,…,gx32,hx32]
[0065] [jx1,0,jx2,0,…,jx16,0,jx17,0,…,jx32,0]-----(10).
[0066] For step S3, bring it forward to the convolution kernel storage, adjust the data order, and the pseudo-code for the entire implementation stored in the required order is as follows:
[0067] SIMD type variable registers:
[0068] sum_0,sum_1;
[0069] in_value0, in_value1, in_value2, in_value3, in_value4, in_value5, in_value6, in_value7, in_value8, in_value9;
[0070] in_0, in_1, in_2, in_3, in_4,in_5, in_6, in_7, in_8, in_9;
[0071] in_value, cvft_0, cvft_1, cvft_2, cvft_3, cvft_4, cvft_5, cvft_6, cvft_7, cvft_8, cvft_9;
[0072] stride is the step size;
[0073] in_data is the input feature map data;
[0074] input_height is the height of the input feature map data;
[0075] input_width is the width of the input feature map data;
[0076] input_depth is the depth of the input feature map data;
[0077] output_data is the output feature map data;
[0078] output_height is the height of the output feature map data;
[0079] output_width is the width of the output feature map data;
[0080] output_depth is the depth of the output feature map data, output_depth = input_depth;
[0081] ft_data is the convolution kernel data stored in the required order;
[0082] output_depth is the depth of the output feature map data;
[0083] Other parameters are pointers or regular data;
[0084] The first layer of loop: Determine whether h < output_height holds. If not, jump out of the loop. If it holds, perform the following calculations within the first layer of the loop, and then increment out_h and re-determine whether h < output_height holds. Continue:
[0085] int out_w = 0;
[0086] j = out_h * output_width;
[0087] int in_h = out_h * stride;
[0088] int8_t *in_ptr0 = in_data + in_h * input_width * input_depth; The second loop: The second loop is included within the first loop. It checks if out_w < output_width holds. If not, the loop is exited. If it holds, the following calculations within the second loop are performed, and out_w++, j++ are executed, then it checks again if out_w < output_width holds and continues:
[0089] int in_w = out_w * stride;
[0090] int8_t *in_ptr1 = in_ptr0 + in_w * input_depth;
[0091] int8_t *out_ptr = output_data + j * output_depth;
[0092] ft_ptr = ft_data;
[0093] The third loop: The third loop is included within the second loop. dj = 0. It checks if dj <= output_depth – 16 holds. If not, the loop is exited. If it holds, the following calculations within the third loop are performed, and dj += 16 is executed, then it checks again if dj <= output_depth – 16 holds and continues:
[0094] sum_0 = [0, 0, …, 0]; Initialized to 0;
[0095] sum_1 = [0, 0, …, 0]; Initialized to 0;
[0096] in_locr = in_ptr;
[0097] in_locr2 = in_ptr + 1 * (input_width * input_depth);
[0098] in_locr3 = in_ptr + 2 * (input_width * input_depth);
[0099] Load 16 data from in_locr into register in_value0;
[0100] Load 16 data from (in_locr + input_depth) into register in_value1; Load 16 data from (in_locr + input_depth * 2) into register in_value2;
[0101] Load 16 data from in_locr2 into register in_value3;
[0102] Load 16 data from (in_locr2 + input_depth) into register in_value4; Load 16 data from (in_locr2 + input_depth * 2) into register in_value5;
[0103] Load 16 data from in_locr3 into register in_value6;
[0104] Load 16 data from (in_locr3 + input_depth) into register in_value7; Load 16 data from (in_locr3 + input_depth * 2) into register in_value8;
[0105] Select the first 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_0;
[0106] Select the last 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_1;
[0107] Select the first 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_2;
[0108] Select the last 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_3;
[0109] Select the first 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_4;
[0110] Select the last 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_5;
[0111] Select the first 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_6;
[0112] Select the last 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_7;
[0113] Select the first 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_8. The data stored in zero is 16 8-bit 0s;
[0114] Select the last 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_9. The data stored in zero is 16 8-bit 0s;
[0115] Next, copy the data. Each time, copy 16 8-bit data and load them in sequence according to the order in ft_ptr. Specifically as follows:
[0116] The data in cvft_0 is: a1, b1, a2, b2, …, a8, b8;
[0117] The data in cvft_1 is: a9, b9, a10, b10, …, a16, b16;
[0118] The data in cvft_2 is: c1, d1, c2, d2, …, c8, d8;
[0119] The data in cvft_3 is: c9, d9, c10, d10, …, c16, d16;
[0120] ……
[0121] The data in cvft_8 is: j1, 0, j2, 0, …, j8, 0
[0122] The data in cvft_9 is: j9, 0, j10, 0, …, j16, 0
[0123] The process of sum_0 and sum_1:
[0124] Corresponding to in_0 to in_9, cvft_0 to cvft_9, execute the simd instruction in sequence. The result of multiplying and adding the two registers is stored in the 128-bit registers sum_0 and sum_1. Each of sum_0 and sum_1 stores 8 16-bit data. The finally obtained sum_0 and sun_1 are the results generated by the independent convolution calculation. The simd instruction operation can be regarded as a function. The input inside is an array, and the output is also data. This array is a bit special. The total number of bits of the array members is a fixed value of 128 bits. The array members can be 8 bits, 16 bits, or 32 bits. See the following text for the specific calculation process.
[0125] In the said pseudocode, the adjacent sums are calculated after multiplying the main instructions. The calculation process is as follows:
[0126] First iteration
[0127] sum_0 = [0, 0, …, 0]; Initialized to 0;
[0128] sum_1 = [0, 0, …, 0]; Initialized to 0;
[0129] In_0 = [ax1, bx1, ax2, bx2, …, ax8, bx8];
[0130] In_1 = [ax9, bx9, ax10, bx10, …, ax16, bx16];
[0131] cvft_0 = [a1, b1, a2, b2, …, a8, b8];
[0132] cvft_1 = [a9, b9, a10, b10, …, a16, b16];
[0133] Results after using the adjacent sum after multiplication instruction:
[0134] sum_0 = [a1 * ax1 + b1 * bx1, a2 * ax2 + b2 * bx2, …, a8 * ax8 + b8 * bx8]
[0135] sun_1 = [a9 * ax9 + b9 * bx9, a10 * ax10 + b10 * bx10, …, a16 * ax16 +
[0136] b16 * bx16];
[0137] In_2 = [cx1, dx1, cx2, dx2, …, cx8, dx8];
[0138] In_3 = [cx9, dx9, cx10, dx10, …, cx16, dx16];
[0139] cvft_2 = [c1, d1, c2, d2, …, c8, d8];
[0140] cvft_3 = [c9, d9, c10, d10, …, c16, d16];
[0141] Results after using the adjacent sum after multiplication instruction:
[0142] sum_0 = [a1 * ax1 + b1 * bx1 + c1 * cx1 + d1 * dx1, a2 * ax2 + b2 * bx2 +
[0143] c2*cx2 + d2*dx2, …, a8*ax8 + b8*bx8 + c8*cx8 + d8*dx8
[0144] sun_1 = [a9*ax9 + b9*bx9 + c9*cx9 + d9*dx9, a10*ax10 + b10*bx10 +
[0145] c10*cx10 + d10*dx10, …, a16*ax16 + b16*bx16 + c16*cx16 + d16*
[0146] dx16];
[0147] Final iteration result:
[0148] sum_0 = [(a1*ax1 + b1*bx1 + c1*cx1 + d1*dx1 + … + j1*jx1), (a2*
[0149] ax2 + b2*bx2 + c2*cx2 + d2*dx2 + j2*jx2), …, a8*ax8 + b8*bx8 +
[0150] c8*cx8 + d8*dx8 + j8*jx8]
[0151] sun_1 = [(a9*ax9 + b9*bx9 + c9*cx9 + d9*dx9 + j9*jx9), (a10*ax10 +
[0152] b10*bx10 + c10*cx10 + d10*dx10 + j10*jx10), …, (a16*ax16 + b16*
[0153] bx16 + c16*cx16 + d16*dx16 + j16*jx16)],
[0154] The result obtained is the convolution result of ordinary calculation.
[0155] The method described is a method that can use any step size. If the step size is fixed, a method with a fixed window can also be used for optimization.
[0156] Therefore, the advantages of this application are as follows: It provides a new way of storing convolution kernel data, a design method for independent convolution 3×3 based on 4-bit to 6-bit. This optimization method is suitable for the operation of vector instructions. Since simd calculates 16 data at a time, the data reusability is very high, reducing the time for loading data, thus greatly reducing the time. The implementation speed is increased by several times, and it can be increased by about 40 times compared with the C program. Brief Description of the Drawings
[0157] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0158] Figure 1 It is a flowchart of the method of the present invention.
[0159] Figure 2 It is a schematic diagram of the storage structure of the planar graph of the input data features.
[0160] Figure 3 It is a schematic diagram of the three-dimensional structure of the data space of the input data. Detailed implementation manners
[0161] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0162] As Figure 1 shown, the present invention relates to a design method based on an independent convolution of 3×3 from 4bit to 6bit, and the method includes the following steps:
[0163] S1. Input data and store it, wherein, the processing of optimizing the storage mode of convolution kernel data:
[0164] Since the subsequent convolution calculation is an independent convolution, after 9 pairs of data are multiplied and then added together, a single data is generated; and in the instruction set, only one instruction can satisfy that the number of bits of the result of multiplying and then adding adjacent data is 16bit;
[0165] According to the use of this instruction, it is necessary to store the data crosswise and add a 0;
[0166] Cross-store the data of two adjacent depths, and cross-store the data of the last depth with 0; S2. Use simd instructions to design an optimized simd instruction algorithm to implement convolution calculation:
[0167] For the 3×3 depth data in the feature map for convolution, every two depths are grouped for cross-processing, which can form 4 groups, and there is still one depth data left. The remaining one depth data is combined with 0. At this time, 5 groups of depth data are formed;
[0168] S3. Perform corresponding calculations on the 5 groups of depth data formed in step S2 and the convolution kernel data that has been processed in step S1, multiply all the data correspondingly, and then add them up.
[0169] In the step S1, a storage order of the input data and the convolution kernel data:
[0170] The input data is stored in the order of depth first, width second, and height last. During calculation, the spatial structure of the data is considered. In storage, it is a vector storage method. Understanding diagram of the storage structure of the feature map plane graph (as Figure 2 ). In the figure, the input depth refers to the specific number of data depths, the input data width refers to how many input depths there are, and the input data height refers to how many input data widths there are. Three-dimensional space such as Figure 3 shown.
[0171] 2) Storage method of the input convolution kernel data.
[0172] a) Initial storage method of the convolution kernel:
[0173] Let's assume some information related to a convolution kernel data. The output depth out_dep is 4, the convolution kernel width is 3, and the convolution kernel height is 3. The data is continuous in the output depth direction. Then, a group of data in the width direction is composed of 3 groups of data in the output depth direction (4 data in each direction), and finally, a group of data in the height direction is composed of 3 groups of data in the width direction. Specifically as follows:
[0174] The output depth out_dep is 4, which can be expressed as:
[0175] [a1,a2,a3,a4]; -----(1)
[0176] The data on the convolution kernel width 3 is expressed as:
[0177] [a1,a2,a3,a4;
[0178] b1,b2,b3,b4;
[0179] c1,c2,c3,c4]; -----(2)
[0180] The data on the convolution kernel height 3, table
[0181] shown as:
[0182] [a1,a2,a3,a4;
[0183] b1,b2,b3,b4;
[0184] c1,c2,c3,c4;
[0185] d1,d2,d3,d4;
[0186] e1,e2,e3,e4;
[0187] f1,f2,f3,f4;
[0188] g1,g2,g3,g4;
[0189] h1, h2, h3, h4;
[0190] j1, j2, j3, j4]; -----(3)
[0191] Stored in the computer as:
[0192] [a1, a2, a3, a4, b1, b2, b3, b4, c1, c2, c3, c4, d1, d2, d3, d4, e1, e2, e3, e4, f1, f2, f3, f4, g1, g2, g3, g4, h1, h2, h3, h4, j1, j2, j3, j4]; -----(3)
[0193] b) Optimization of the storage method of convolution kernel data:
[0194] If conventional data storage is used, the convolution kernel data needs to be transformed according to the optimization method, and the data order needs to be transformed every time it is used, which greatly consumes time, reduces the optimization efficiency, and increases the computational amount. However, if the order that needs to be optimized and transformed is pre-transformed and directly used during use, both the optimization purpose can be achieved and the computational amount can be not increased.
[0195] Since the subsequent convolution calculation is an independent convolution, after 9 pairs of data are multiplied and then added together, a single data is generated. And in the instruction set, only one instruction can satisfy that the number of bits of the result of multiplying and then adding adjacent numbers is 16 bits.
[0196] According to the use of this instruction, we need to perform cross-storage on the data. The subsequent algorithm implementation will specifically implement it, and a 0 needs to be added. Cross-store the data of two adjacent depths, and cross-store the data of the last depth with 0. The specific storage is as follows.
[0197] [a1, b1, a2, b2, a3, b3, a4, b4;
[0198] c1, d1, c2, d2, c3, d3, c4, d4;
[0199] e1, f1, e2, f2, e3, f3, e4, f4;
[0200] g1, h1, g2, h2, g3, h3, g4, h4;
[0201] j1, 0, j2, 0, j3, 0, j4, 0]; -----(4)
[0202] Stored in the computer as:
[0203] [a1, b1, a2, b2,
[0204] a3, b3, a4, b4, c1, d1, c2, d2, c3, d3, c4, d4, e1, f1, e2, f2, e3, f3, e4, f4, g1, h1, g2, h2, g3, h3, g4, h4, j1, 0, j2, 0, j3, 0, j4, 0]; -----(5)
[0205] In actual use, the output depths used are all multiples of 16. The above assumes that the output depth out_dep is 4 for illustrative purposes of understanding.
[0206] Since in the design of the simd optimization algorithm, 16 output results are continuously calculated each time in terms of depth, and in the use of convolutional kernels, the number used each time is 16×(3×3 + 1), therefore, we also store the data in the order of using the convolutional kernels. When the output depth exceeds 16 (out_dep is a multiple of 16), the storage method is as follows:
[0207] [a1, b1, a2, b2, …, a16, b16;
[0208] c1, d1, c2, d2, …, c16, d16;
[0209] e1, f1, e2, f2, …, e16, f16;
[0210] g1, h1, g2, h2, …, g16, h16;
[0211] j1, 0, j2, 0, …, j16, 0];
[0212] [a17, b17, a18, b18, …, a32, b32;
[0213] c17, d17, c18, d18, …, c32, d32;
[0214] e17, f17, e18, f18, …, e32, f32;
[0215] g17, h17, g18, h18, …, g32, h32;
[0216] j17, 0, j18, 0, …, j32, 0];
[0217] ……
[0218] [a(out_dep - 15), b(out_dep - 15), a(out_dep - 14), b(out_dep - 14), …, a(out_dep), b(out_dep);
[0219] c(out_dep - 15), d(out_dep - 15), c(out_dep - 14), d(out_dep - 14), …, c(out_dep), d(out_dep);
[0220] e(out_dep - 15), f(out_dep - 15), e(out_dep - 14), f(out_dep - 14), …, e(out_dep), f(out_dep);
[0221] g(out_dep - 15), h(out_dep - 15), g(out_dep - 14), h(out_dep - 14), …, g(out_dep), h(out_dep); j(out_dep - 15), 0, j(out_dep - 14), 0, …, j(out_dep), 0]; -----(6)
[0222] 2. SIMD instruction algorithm.
[0223] 1) Introduction to SIMD instructions
[0224] The SIMD instructions involved are as follows.
[0225] a) SIMD instruction for multiplying and then adding adjacent values:
[0226] vrd = ingenic_muladd_h(vrd, vrs, vrt);
[0227] The input variables are vrd, vrs, and vrt, and the output variable is vrd. vrd, vrs, and vrt are SIMD - type variables, and these variables are 128 - bit registers. vrd stores 8 int16_t data, and vrs and vrt store 16 int8_t data. Since there are multiplications and additions in the operation, and 4 - bit data requires 16 - bit storage after the operation, we store the 4 - bit input data as 8 - bit.
[0228] vrd = [vrd0, vrd1, vrd2, vrd3, vrd4, vrd5, vrd6, vrd7];
[0229] vrs = [vrs0, vrs1, vrs2, vrs3, vrs4, vrs5, vrs6, vrs7, vrs8, vrs9, vrs10, vrs11, vrs12, vrs13, vrs14, vrs15];
[0230] vrt = [vrt0, vrt1, vrt2, vrt3, vrt4, vrt5,
[0231] vrt6, vrt7, vrt8, vrt9, vrt10, vrt11, vrt12, vrt13, vrt14, vrt15];
[0232] Equivalent operations:
[0233] vrd0 := vrd0 + vrs0 * vrt0 + vrs1 * vrt1;
[0234] vrd1 := vrd1 + vrs2 * vrt2 + vrs3 * vrt3; ......
[0236] vrd7 := vrd7 + vrs14 * vrt14 + vrs15 * vrt15;
[0237] b) Load data simd instruction: For the data to be loaded as input, currently it is a pointer to the data. Start loading 128-bit data from the position pointed to by this data in memory. If it is 8-bit data, load 16 of them; if it is 16-bit data, load 8 of them. The data is loaded into the variable vrd.
[0238] vrd = ingenic_load(indata)
[0239] c) Specify selection simd instruction: Select 4 or 8 or 16 data from variables vrs and vrt according to the number set in vri. When using this instruction, it is necessary to occupy a permanent register vri for the instruction to select data at specific positions.
[0240] vrd = ingenic_choise_h(vrs, vrt, vri);
[0241] 3) The simd instruction implements convolution calculation.
[0242] There are many methods to implement the calculation of convolution. When using simd instructions, the data can be first converted to 16-bit and then calculated using multiplication and addition. It is also possible to directly perform multiplication on the data, then convert the result after multiplication to 16-bit, and then perform cumulative addition to implement convolution calculation. However, different algorithms have different execution efficiencies and consume different amounts of time. The following algorithm design will minimize redundant calculations and improve efficiency.
[0243] Let the input data indata be a set of data with an input depth in_depth of 32, an input width in_width of 256, and an input height in_height of 256; the convolutional kernel data filter_data be a set of data with an output depth out_depth of 32, a convolutional kernel width ft_w of 3, and a convolutional kernel height ft_h of 3. Let the structure of the output data (feature map) outdata be: the depth is out_depth (the same as the input depth of the input data, which is 32), the width is out_width, and the height is out_height. In the convolutional calculation, there is a stride, let the stride be stride. The 3×3 input depth-direction data of the input data are:
[0244] [ax1,ax2,ax3,ax4,ax5,ax6,ax7,ax8,ax9,ax10,ax11,ax12,ax13,ax14,ax15,ax16,…ax32]
[0245] [bx1,bx2,bx3,bx4,bx5,bx6,bx7,bx8,bx9,bx10,bx11,bx12,bx13,bx14,bx15,bx16,…bx32]
[0246] [cx1,cx2,cx3,cx4,cx5,cx6,cx7,cx8,cx9,cx10,cx11,cx12,cx13,cx14,cx15,cx16,…cx32]
[0247] [dx1,dx2,dx3,dx4,dx5,dx6,dx7,dx8,dx9,dx10,dx11,dx12,dx13,dx14,dx15,dx16,…dx32]
[0248] [ex1,ex2,ex3,ex4,ex5,ex6,ex7,ex8,ex9,ex10,ex11,ex12,ex13,ex14,ex15,ex16,…ex32]
[0249] [fx1,fx2,fx3,fx4,fx5,fx6,fx7,fx8,fx9,fx10,fx11,fx12,fx13,fx14,fx15,fx16,…fx32]
[0250] [gx1,gx2,gx3,gx4,gx5,gx6,gx7,gx8,gx9,gx10,gx11,gx12,gx13,gx14,gx15,gx16,…gx32]
[0251] [hx1,hx2,hx3,hx4,hx5,hx6,hx7,hx8,hx9,hx10,hx11,hx12,hx13,hx14,hx15,hx16,…hx32]
[0252] [jx1,jx2,jx3,jx4,jx5,jx6,jx7,jx8,jx9,jx10,jx11,jx12,jx13,jx14,jx15,jx16,…jx32]-----(7)
[0253] The original data of the convolutional kernel is:
[0254] [a1,a2,…,a32;
[0255] b1,b2…,b32;
[0256] c1,c2,…,c32;
[0257] d1,d2,…,d32;
[0258] e1,e2,…,e32;
[0259] f1,f2,…,f32;
[0260] g1,g2,…,g32;
[0261] h1,h2,…,h32;
[0262] j1,j2,…,j32]; -----(8)
[0263] The optimized data of the convolutional kernel is:
[0264] [a1,b1,a2,b2,…,a16,b16;
[0265] c1,d1,c2,d2,…,c16,d16;
[0266] e1,f1,e2,f2,…,e16,f16;
[0267] g1,h1,g2,h2,…,g16,h16;
[0268] j1,0,j2,0,…,j16,0];
[0269] [a17,b17,a18,b18,…,a32,b32;
[0270] c17,d17,c18,d18,…,c32,d32;
[0271] e17, f17, e18, f18, …, e32, f32;
[0272] g17, h17, g18, h18, …, g32, h32;
[0273] j17, 0, j18, 0, …, j32, 0]; -----(9)
[0274] a) Conventional calculation method
[0275] Multiply formula (7) and formula (8) correspondingly and then sum them up, which is to calculate and obtain one data in the depth direction of the output result.
[0276] b) SIMD instruction optimization algorithm design
[0277] For the conventional algorithm, many instructions are used and there is also a lot of calculation repetition. Here, for the 3×3 depth data in the feature map for convolution, such as data (7), every two depths are grouped for cross-processing, which can form 4 groups. At the same time, there is one remaining depth data, and the remaining one depth data is combined with 0. At this time, 5 groups of depth data are formed, such as data (10). The data structure is:
[0278] [ax1, bx1, ax2, bx2, …, ax16, bx16, ax17, bx17, …, ax32, bx32]
[0279] [cx1, dx1, cx2, dx2, …, cx16, dx16, cx17, dx17, …, cx32, dx32]
[0280] [ex1, fx1, ex2, fx2, …, ex16, fx16, ex17, fx17, …, ex32, fx32]
[0281] [gx1, hx1, gx2, hx2, …, gx16, hx16, gx17, hx17, …, gx32, hx32]
[0282] [jx1, 0, jx2, 0, …, jx16, 0, jx17, 0, …, jx32, 0] -----(10)
[0283] Perform corresponding calculations on data (10) and the already processed convolution kernel data (9), multiply all data correspondingly and then sum them up. The result is exactly the same as the result after multiplying data (7) and data (8) and then summing them up.
[0284] SIMD is a type of vector calculation. The more numbers are calculated each time, the fewer instructions are used, and the more time-efficient the operation is. Here, mainly the SIMD instruction of multiplying and then adding adjacent numbers is used. This instruction can make the data become 16-bit after multiplying and adding 8-bit data. To use this instruction, the original feature map data needs to be processed, that is, the data (7) needs to be processed in the order of data (8) before use. Only in this way can the requirement of multiplying and then adding adjacent numbers of the instruction be met. Similarly, the convolution kernel data also needs to be processed in a cross order. Since this processing is time-consuming during the calculation, this processing is advanced to the convolution kernel storage, and the data order is adjusted and stored in the required order. The pseudo-code for the entire implementation is as follows:
[0285] SIMD type variable registers:
[0286] sum_0, sum_1;
[0287] in_value0, in_value1, in_value2, in_value3, in_value4, in_value5, in_value6, in_value7, in_value8, in_value9;
[0288] in_0, in_1, in_2, in_3, in_4, in_5, in_6, in_7, in_8, in_9; in_value, cvft_0, cvft_1, cvft_2, cvft_3, cvft_4, cvft_5, cvft_6, cvft_7, cvft_8, cvft_9;
[0289] stride is the step size;
[0290] in_data is the input feature map data;
[0291] input_height is the height of the input feature map data;
[0292] input_width is the width of the input feature map data;
[0293] input_depth is the depth of the input feature map data;
[0294] output_data is the output feature map data;
[0295] output_height is the height of the output feature map data;
[0296] output_width is the width of the output feature map data;
[0297] output_depth is the depth of the output feature map data (output_depth = input_depth); ft_data is the convolutional kernel data stored in the required order;
[0298] output_depth is the depth of the output feature map data;
[0299] Other parameters are pointers or regular data;
[0300] The first layer of loop: Determine whether h < output_height holds. If not, jump out of the loop. If it holds, perform the following calculations within the first layer of loop, and then increment out_h and re-determine whether h < output_height holds. Continue:
[0301] for(; out_h < output_height; out_h++){
[0302] int out_w = 0;
[0303] j = out_h * output_width;
[0304] int in_h = out_h * stride;
[0305] int8_t* in_ptr0 = in_data + in_h * input_width * input_depth;
[0306] The second layer of loop: The second layer of loop is included within the first layer of loop. Determine whether out_w < output_width holds. If not, jump out of the loop. If it holds, perform the following calculations within the second layer of loop, and then increment out_w and j, and re-determine whether out_w < output_width holds. Continue:
[0307] for(; out_w < output_width; out_w++, j++){
[0308] int in_w = out_w * stride;
[0309] int8_t* in_ptr1 = in_ptr0 + in_w * input_depth;
[0310] int8_t* out_ptr = output_data + j * output_depth;
[0311] ft_ptr = ft_data;
[0312] The third - level loop: The third - level loop is included within the second - level loop. dj = 0, and it is judged whether dj <=
[0313] output_depth - 16 holds. If not, the loop is exited. If it holds, the following calculations within the third loop are performed, and dj += 16, then it is re - judged whether dj <= output_depth - 16 holds, and continue:
[0314] for(dj = 0; dj <= output_depth - 16; dj += 16) {
[0315] Initialize sum_0 to 0: sum_0 = [0, 0, …, 0];
[0316] Initialize sum_1 to 0: sum_1 = [0, 0, …, 0];
[0317] Assignment:
[0318] in_locr = in_ptr;
[0319] in_locr2 = in_ptr + 1 * (input_width * input_depth);
[0320] in_locr3 = in_ptr + 2 * (input_width * input_depth);
[0321] Load 16 data from in_locr into the register in_value0: in_value0 = genenic_load(in_locr, 0);
[0322] Load 16 data from (in_locr + input_depth) into the register in_value1: in_value1 = genenic_load(in_locr + input_depth, 0);
[0323] Load 16 data from (in_locr + input_depth * 2) into the register in_value2: in_value2 = genenic_load(in_locr + input_depth * 2, 0);
[0324] Load 16 data from in_locr2 into the register in_value3: in_value3 = genenic_load(in_locr2, 0);
[0325] Load 16 data from (in_locr2 + input_depth) into register in_value4: in_value4 = ingenic_load(in_locr2 + input_depth, 0);
[0326] Load 16 data from in_locr2 into register in_value3: in_value5 = ingenic_load(in_locr2 + input_depth * 2, 0);
[0327] Load 16 data from in_locr3 into register in_value6: in_value6 = ingenic_load(in_locr3, 0);
[0328] Load 16 data from (in_locr3 + input_depth) into register in_value7: in_value7 = ingenic_load(in_locr3 + input_depth, 0);
[0329] Load 16 data from (in_locr3 + input_depth * 2) into register in_value8: in_value8 = ingenic_load(in_locr3 + input_depth * 2, 0);
[0330] Select the first 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_0: in_0 = ingenic_choise_h(in_value1, in_value0, cvt_ft_0);
[0331] Select the last 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_1: in_1 = ingenic_choise_h(in_value1, in_value0, cvt_ft_1);
[0332] Select the first 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_2: in_2 = ingenic_choise_h(in_value3, in_value2, cvt_ft_0);
[0333] Select the last 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_3: in_3 = ingenic_choise_h(in_value3, in_value2, cvt_ft_1);
[0334] Select the first 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_4: in_4 = ingenic_choise_h(in_value5, in_value4, cvt_ft_0);
[0335] Select the last 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_5: in_5 = ingenic_choise_h(in_value5, in_value4, cvt_ft_1);
[0336] Select the first 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_6: in_6 = ingenic_choise_h(in_value7, in_value6, cvt_ft_0);
[0337] Select the last 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_7: in_7 = ingenic_choise_h(in_value7, in_value6, cvt_ft_1);
[0338] Select the first 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_8. The 16 8-bit 0s are stored in zero: in_8 =
[0339] ingenic_choise_h(in_value8, zero, cvt_ft_0);
[0340] Select the last 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_9. The 16 8-bit 0s are stored in zero: in_9 =
[0341] ingenic_choise_h(in_value8, zero, cvt_ft_1);
[0342] Then copy the data, copying 16 8-bit data each time and loading them in sequence according to the order in ft_ptr. As follows:
[0343] The data in cvft_0 is: a1, b1, a2, b2, …, a8, b8;
[0344] The data in cvft_1 is: a9, b9, a10, b10, …, a16, b16;
[0345] The data in cvft_2 is: c1, d1, c2, d2, …, c8, d8;
[0346] The data in cvft_3 is: c9, d9, c10, d10, …, c16, d16;
[0347] ……
[0348] The data in cvft_8 is: j1, 0, j2, 0, …, j8, 0
[0349] The data in cvft_9 is: j9, 0, j10, 0, …, j16, 0
[0350] It can be expressed as follows:
[0351] cvft_0 = ingenic_load(ft_ptr, 0);
[0352] cvft_1 = ingenic_load(ft_ptr, 16);
[0353] cvft_2 = ingenic_load(ft_ptr, 32);
[0354] cvft_3 = ingenic_load(ft_ptr, 48);
[0355] cvft_4 = ingenic_load(ft_ptr, 64);
[0356] cvft_5 = ingenic_load(ft_ptr, 80);
[0357] cvft_6 = ingenic_load(ft_ptr, 96);
[0358] cvft_7 = ingenic_load(ft_ptr, 112);
[0359] cvft_8 = ingenic_load(ft_ptr, 128);
[0360] cvft_9 = ingenic_load(ft_ptr, 144);
[0361] ft_ptr += 160;
[0362] Calculation process of sum_0 and sum_1:
[0363] For in_0 to in_9 and cvft_0 to cvft_9, execute simd instructions in sequence. The result of multiplying and adding two registers is stored in the 128-bit registers sum_0 and sum_1. Each of sum_0 and sum_1 stores 8 16-bit data. The finally obtained sum_0 and sun_1 are the results generated by independent convolution calculation:
[0364] It can be expressed as follows:
[0365] sum_0 = ingenic_muladd_h(sum_0, in_0, cvft_0);
[0366] sum_1 = ingenic_muladd_h(sum_1, in_1, cvft_1);
[0367] sum_0 = ingenic_muladd_h(sum_0, in_2, cvft_2);
[0368] sum_1 = ingenic_muladd_h(sum_1, in_3, cvft_3);
[0369] sum_0 = ingenic_muladd_h(sum_0, in_4, cvft_4);
[0370] sum_1 = ingenic_muladd_h(sum_1, in_5, cvft_5);
[0371] sum_0 = ingenic_muladd_h(sum_0, in_6, cvft_6);
[0372] sum_1 = ingenic_muladd_h(sum_1, in_7, cvft_7);
[0373] sum_0 = ingenic_muladd_h(sum_0, in_8, cvft_8);
[0374] sum_1 = ingenic_muladd_h(sum_1, in_9, cvft_9);
[0375] / / -----
[0376] }
[0377] }
[0378] }
[0379] In the above pseudocode, in the calculation process of the said sum_0 and sum_1, the main instruction is to multiply and then add adjacent results. The calculation process is as follows:
[0380] The first iteration
[0381] sum_0 = [0, 0, …, 0]; initialized to 0;
[0382] sum_1 = [0, 0, …, 0]; initialized to 0;
[0383] In_0 = [ax1, bx1, ax2, bx2, …, ax8, bx8];
[0384] In_1 = [ax9, bx9, ax10, bx10, …, ax16, bx16];
[0385] cvft_0 = [a1, b1, a2, b2, …, a8, b8];
[0386] cvft_1 = [a9, b9, a10, b10, …, a16, b16];
[0387] The result after using the multiply-and-add-adjacent instruction:
[0388] sum_0 = [a1 * ax1 + b1 * bx1, a2 * ax2 + b2 * bx2, …, a8 * ax8 + b8 * bx8]
[0389] sun_1 = [a9 * ax9 + b9 * bx9, a10 * ax10 + b10 * bx10, …, a16 * ax16 + b16 * bx16];
[0390] In_2 = [cx1, dx1, cx2, dx2, …, cx8, dx8];
[0391] In_3 = [cx9, dx9, cx10, dx10, …, cx16, dx16];
[0392] cvft_2 = [c1, d1, c2, d2, …, c8, d8];
[0393] cvft_3 = [c9, d9, c10, d10, …, c16, d16];
[0394] Result after using the multiply-then-adjacent-add instruction:
[0395] sum_0 = [a1*ax1 + b1*bx1 + c1*cx1 + d1*dx1, a2*ax2 + b2*bx2 + c2*cx2 + d2*dx2, …, a8*ax8 + b8*bx8 + c8*cx8 + d8*dx8]
[0396] sun_1 = [a9*ax9 + b9*bx9 + c9*cx9 + d9*dx9, a10*ax10 + b10*bx10 + c10*cx10 + d10*dx10, …, a16*ax16 + b16*bx16 + c16*cx16 + d16*dx16];
[0397] Final iteration result:
[0398] sum_0 = [(a1*ax1 + b1*bx1 + c1*cx1 + d1*dx1 + … + j1*jx1), (a2*ax2 + b2*bx2 + c2*cx2 + d2*dx2 + j2*jx2), …, a8*ax8 + b8*bx8 + c8*cx8 + d8*dx8 + j8*jx8]
[0399] sun_1 = [(a9*ax9 + b9*bx9 + c9*cx9 + d9*dx9 + j9*jx9), (a10*ax10 + b10*bx10 + c10*cx10 + d10*dx10 + j10*jx10), …, (a16*ax16 + b16*bx16 + c16*cx16 + d16*dx16 + j16*jx16)];
[0400] Obviously, this is the convolution result of ordinary calculation. Since simd calculates 16 data at a time and the data reusability is very high, the time for loading data is reduced, thus greatly reducing the time.
[0401] Through the above simd instruction algorithm, the speed can be increased by nearly 40 times. The above algorithm is an algorithm with an arbitrary step size. If the step size is fixed, for example, it is 1, methods such as fixed window can be used for further optimization, and the running time can be further reduced.
[0402] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A design method for an independent convolution of 3×3 based on 4-bit to 6-bit, characterized in that, The method is to pre-transform the order of the data to be optimized and transformed, that is, during the storage of the convolution kernel data, the data of two adjacent depths are stored crosswise, and the data of the last depth is stored crosswise with 0; and during the subsequent convolution calculation process, the order of the original feature map data is adjusted, and then the simd instruction of multiplying and then adding adjacent ones is used, so that the data becomes 16-bit after multiplying and adding 8-bit data; the order of the convolution kernel data is adjusted and stored in the order required during use. The method further includes the following steps: S1, input data and store it. Among them, the processing of optimizing the storage method of the convolution kernel data: Since the subsequent convolution calculation is an independent convolution, after 9 pairs of data are multiplied and then accumulated together to generate one data; and there is only one instruction in the instruction set that can satisfy that the number of bits of the result of multiplying and then adding adjacent ones is 16 bits. According to the use of this instruction, it is necessary to store the data crosswise and add a 0; the data of two adjacent depths are stored crosswise, and the data of the last depth is stored crosswise with 0. S2, use the simd instruction to design the simd instruction optimization method to implement the convolution calculation: For the 3×3 depth data in the feature map for convolution, every two depths are grouped for cross-processing, which can form 4 groups. At the same time, there is one remaining depth data, and the remaining one depth data is combined with 0. At this time, 5 groups of depth data are formed. S3, perform corresponding calculations on the 5 groups of depth data formed in step S2 and the convolution kernel data that has been processed in step S1, multiply all the data correspondingly, and then accumulate.
2. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 1, characterized in that In step S1, the input data is stored in the order of depth first, width second, and height last; during the calculation, the spatial structure of the data is considered, and in the storage, it is a vector storage method.
3. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 1, characterized in that, The initial storage of the convolution kernel in step S1: Set a convolution kernel data-related information, the output depth out_dep is 4, the convolution kernel width is 3, and the convolution kernel height is 3; the continuous method of the data is continuous in the output depth direction, and then there are 3 groups in the output depth direction, where each direction has 4 data, and 12 data form a group of data in the width, and finally 3 groups of data in the width form the data in the height.
4. According to a 4-bit to 6-bit independent convolution 3×3 design method described in claim 1, the output depth used in the actual use of the method is a multiple of 16.
5. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 3, characterized in that Step S1 further includes: Set the input data indata as a group of data with an input depth in_depth of 32, a width in_width of 256, and a height in_height of 256; The convolution kernel data filter_data is a group of data with an output depth out_depth of 32, a convolution kernel width ft_w of 3, and a convolution kernel height ft_h of 3; Set the structure of the output data, that is, the feature map outdata: The depth is out_depth, which is the same as the input depth of the input data here, which is 32, the width is out_width, and the height is out_height; In convolution calculation, there is a stride, let the stride be stride: The 3×3 input data in the depth direction are: [ax1, ax2, ax3, ax4, ax5, ax6, ax7, ax8, ax9, ax10, ax11, ax12, ax13, ax14, ax15, ax16, … ax32] [bx1, bx2, bx3, bx4, bx5, bx6, bx7, bx8, bx9, bx10, bx11, bx12, bx13, bx14, bx15, bx16, … bx32] [cx1, cx2, cx3, cx4, cx5, cx6, cx7, cx8, cx9, cx10, cx11, cx12, cx13, cx14, cx15, cx16, … cx32] [dx1, dx2, dx3, dx4, dx5, dx6, dx7, dx8, dx9, dx10, dx11, dx12, dx13, dx14, dx15, dx16, … dx32] [ex1, ex2, ex3, ex4, ex5, ex6, ex7, ex8, ex9, ex10, ex11, ex12, ex13, ex14, ex15, ex16, … ex32] [fx1, fx2, fx3, fx4, fx5, fx6, fx7, fx8, fx9, fx10, fx11, fx12, fx13, fx14, fx15, fx16, … fx32] [gx1, gx2, gx3, gx4, gx5, gx6, gx7, gx8, gx9, gx10, gx11, gx12, gx13, gx14, gx15, gx16, … gx32] [hx1, hx2, hx3, hx4, hx5, hx6, hx7, hx8, hx9, hx10, hx11, hx12, hx13, hx14, hx15, hx16, … hx32] [jx1, jx2, jx3, jx4, jx5, jx6, jx7, jx8, jx9, jx10, jx11, jx12, jx13, jx14, jx15, jx16, … jx32] -----(7) The original data of the convolution kernel are: [a1, a2, …, a32; b1, b2 …, b32; c1, c2, …, c32; d1, d2, …, d32; e1, e2, …, e32; f1, f2, …, f32; g1, g2, …, g32; h1, h2, …, h32; j1, j2, …, j32]; -----(8) The optimized data of the convolution kernel are: [a1, b1, a2, b2, …, a16, b16; c1, d1, c2, d2, …, c16, d16; e1, f1, e2, f2, …, e16, f16; g1, h1, g2, h2, …, g16, h16; j1, 0, j2, 0, …, j16, 0]; [a17, b17, a18, b18, …, a32, b32; c17, d17, c18, d18, …, c32, d32; e17, f17, e18, f18, …, e32, f32; g17, h17, g18, h18, …, g32, h32; j17, 0, j18, 0, …, j32, 0]; -----(9).
6. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, characterized in that, In step S2, for the 3×3 depth data in the feature map for convolution, such as data (7); The composition of the 5 groups of depth data, such as data (10): The data structure is: [ax1, bx1, ax2, bx2, …, ax16, bx16, ax17, bx17, …, ax32, bx32] [cx1, dx1, cx2, dx2, …, cx16, dx16, cx17, dx17, …, cx32, dx32] [ex1, fx1, ex2, fx2, …, ex16, fx16, ex17, fx17, …, ex32, fx32] [gx1, hx1, gx2, hx2, …, gx16, hx16, gx17, hx17, …, gx32, hx32] [jx1, 0, jx2, 0, …, jx16, 0, jx17, 0, …, jx32, 0] -----(10).
7. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 1, characterized in that For step S3, advance it to the convolution kernel storage and adjust the data order to store it in the required order. The entire implementation is as follows: SIMD type variable registers: sum_0, sum_1; in_value0, in_value1, in_value2, in_value3, in_value4, in_value5, in_value6, in_value7, in_value8, in_value9; in_0, in_1, in_2, in_3, in_4, in_5, in_6, in_7, in_8, in_9; in_value, cvft_0, cvft_1, cvft_2, cvft_3, cvft_4, cvft_5, cvft_6, cvft_7, cvft_8, cvft_9; stride is the step size; in_data is the input feature map data; input_height is the height of the input feature map data; input_width is the width of the input feature map data; input_depth is the depth of the input feature map data; output_data is the output feature map data; output_height is the height of the output feature map data; output_width is the width of the output feature map data; output_depth is the depth of the output feature map data, output_depth = input_depth; ft_data is the convolution kernel data stored in the required order; output_depth is the depth of the output feature map data; Other parameters are pointers or regular data; The first layer of loop: Determine whether h < output_height holds. If not, jump out of the loop. If it holds, perform the following calculations within the first layer of loop, and then increment out_h and re-determine whether h < output_height holds. Continue: int out_w = 0; j = out_h * output_width; int in_h = out_h * stride; int8_t* in_ptr0 = in_data + in_h * input_width * input_depth; The second layer of loop: The second layer of loop is included within the first layer of loop. Determine whether out_w < output_width holds. If not, jump out of the loop. If it holds, perform the following calculations within the second layer of loop, and then increment out_w and j, and re-determine whether out_w < output_width holds. Continue: int in_w = out_w * stride; int8_t* in_ptr1 = in_ptr0 + in_w * input_depth; int8_t* out_ptr = output_data + j * output_depth; ft_ptr = ft_data; The third layer of loop: The third layer of loop is included within the second layer of loop. dj = 0. Determine whether dj <= output_depth – 16 holds. If not, jump out of the loop. If it holds, perform the following calculations within the third loop, and then increment dj by 16 and re-determine whether dj <= output_depth – 16 holds. Continue: sum_0 = [0, 0, …, 0]; Initialized to 0; sum_1 = [0, 0, …, 0]; Initialized to 0; in_locr = in_ptr; in_locr2 = in_ptr + 1 * (input_width * input_depth); in_locr3 = in_ptr + 2 * (input_width * input_depth); Load 16 data from in_locr into the register in_value0; Load 16 data from (in_locr + input_depth) into the register in_value1; Load 16 data from (in_locr + input_depth * 2) into the register in_value2; Load 16 data from in_locr2 into the register in_value3; Load 16 data from (in_locr2 + input_depth) into the register in_value4; Load 16 data from (in_locr2 + input_depth * 2) into the register in_value5; Load 16 data from in_locr3 into register in_value6; Load 16 data from (in_locr3 + input_depth) into register in_value7; load 16 data from (in_locr3 + input_depth * 2) into register in_value8; Select the first 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_0; Select the last 8 8-bit data from in_value1 and in_value0 respectively, store them crosswise, and place them in the 128-bit register in_1; Select the first 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_2; Select the last 8 8-bit data from in_value3 and in_value2 respectively, store them crosswise, and place them in the 128-bit register in_3; Select the first 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_4; Select the last 8 8-bit data from in_value5 and in_value4 respectively, store them crosswise, and place them in the 128-bit register in_5; Select the first 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_6; Select the last 8 8-bit data from in_value7 and in_value6 respectively, store them crosswise, and place them in the 128-bit register in_7; Select the first 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_8. Zero stores 16 8-bit 0s; Select the last 8 8-bit data from in_value8 and zero respectively, store them crosswise, and place them in the 128-bit register in_9. Zero stores 16 8-bit 0s; Next, copy the data. Each time, copy 16 8-bit data and load them in sequence according to the order in ft_ptr. Specifically as follows: The data in cvft_0 is: a1, b1, a2, b2, …, a8, b8; The data in cvft_1 is: a9, b9, a10, b10, …, a16, b16; The data in cvft_2 is: c1, d1, c2, d2, …, c8, d8; The data in cvft_3 is: c9, d9, c10, d10, …, c16, d16; …… The data in cvft_8 is: j1, 0, j2, 0, …, j8, 0 The data in cvft_9 is: j9, 0, j10, 0, …, j16, 0 The calculation process of sum_0 and sum_1: For in_0 to in_9 and cvft_0 to cvft_9, execute SIMD instructions in sequence. The result of multiplying and adding the two registers is stored in the 128-bit registers sum_0 and sum_1. Each of sum_0 and sum_1 stores 8 16-bit data. The finally obtained sum_0 and sum_1 are the results generated by the independent convolution calculation.
8. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 7, characterized in that The process of calculating sum_0 and sum_1 mainly involves adjacent addition after multiplication. The specific calculation process is as follows: The first iteration sum_0 = [0, 0, …, 0]; Initialized to 0; sum_1 = [0, 0, …, 0]; Initialized to 0; In_0 = [ax1, bx1, ax2, bx2, …, ax8, bx8]; In_1 = [ax9, bx9, ax10, bx10, …, ax16, bx16]; cvft_0 = [a1, b1, a2, b2, …, a8, b8]; cvft_1 = [a9, b9, a10, b10, …, a16, b16]; The result after using the adjacent addition instruction after multiplication: sum_0 = [a1*ax1 + b1*bx1, a2*ax2 + b2*bx2, …, a8*ax8 + b8*bx8] sun_1 = [a9*ax9 + b9*bx9, a10*ax10 + b10*bx10, …, a16*ax16 + b16*bx16]; In_2 = [cx1, dx1, cx2, dx2, …, cx8, dx8]; In_3 = [cx9, dx9, cx10, dx10, …, cx16, dx16]; cvft_2 = [c1, d1, c2, d2, …, c8, d8]; cvft_3 = [c9, d9, c10, d10, …, c16, d16]; The result after using the adjacent addition instruction after multiplication: sum_0 = [a1*ax1 + b1*bx1 + c1*cx1 + d1*dx1, a2*ax2 + b2*bx2 + c2*cx2 + d2*dx2, …, a8*ax8 + b8*bx8 + c8*cx8 + d8*dx8] sun_1 = [a9*ax9 + b9*bx9 + c9*cx9 + d9*dx9, a10*ax10 + b10*bx10 + c10*cx10 + d10*dx10, …, a16*ax16 + b16*bx16 + c16*cx16 + d16*dx16]; The final iteration result: sum_0 = [(a1*ax1 + b1*bx1 + c1*cx1 + d1*dx1 + … + j1*jx1), (a2*ax2 + b2*bx2 + c2*cx2 + d2*dx2 + j2*jx2), …, a8*ax8 + b8*bx8 + c8*cx8 + d8*dx8 + j8*jx8] sun_1 = [(a9 * ax9 + b9 * bx9 + c9 * cx9 + d9 * dx9 + j9 * jx9), (a10 * ax10 + b10 * bx10 + c10 * cx10 + d10 * dx10 + j10 * jx10), …, (a16 * ax16 + b16 * bx16 + c16 * cx16 + d16 * dx16 + j16 * jx16)], The resulting result is the convolution result of ordinary calculation.
9. A design method of an independent convolution 3×3 based on 4-bit to 6-bit, according to claim 5, characterized in that The method is a method using an arbitrary step size. If the step size is fixed, a method with a fixed window can also be used for optimization.
Citation Information
Patent Citations
Binary convolutional neural network processor and using method thereof
CN107153873A
Acceleration method of convolutional neural network, computer readable storage medium and application of method
CN111178505A