Design method based on floating point feature map format conversion

By using dma and simd instructions in image processing combined with the high-speed storage characteristics of oram, efficient conversion from DHWC32 format to HWC format is achieved, solving the problem of inefficiency in the prior art and improving processing speed.

CN120339031APending Publication Date: 2025-07-18INGENIC SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410069629.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the conversion from DHWC32 format to HWC format during image processing is inefficient, especially on Junzheng T40/T41/A1 chips, conventional processing methods are inefficient.

Method used

The data is stored from ddr to oram for format conversion through the dma instruction, and then stored from oram to ddr, combined with the Simd instruction for data processing, and the high-speed storage characteristics of oram are used to realize parallel processing.

Benefits of technology

It significantly improves the processing speed of data format conversion and doubles the efficiency, suitable for the minimum number of processing rows that oram can accommodate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339031A_ABST
    Figure CN120339031A_ABST
Patent Text Reader

Abstract

The invention provides a design method based on floating point feature map format conversion. The method comprises the following steps: S1, designing the number of rows processed each time; s2, converting a simd processing format; and S3, specific implementation steps based on Oram. The method comprises the following steps: firstly, carrying data from a ddr storage to an oam through a dma instruction, carrying out format conversion on the data in the oam, storing the converted data in the oam, and finally carrying the data from the oam storage to the ddr by using the dma instruction. Through the method, the processing speed is increased, and the time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a design method based on the format conversion of floating-point feature maps. Background Art

[0002] In the prior art, in the field of image processing, the reason for format conversion is that block computing is adopted in convolution computing, and it is more conducive to data loading and has higher efficiency to use the DHWC32 format as the input format and output format during the computing process. However, the final output result of the model requires data in the HWC format, so it is necessary to convert from the DHWC32 format to the HWC format.

[0003] In the prior art, for the conversion of all DHWC32 formats to HWC formats, for example, the situation in the current chips of Beijing Junzheng Integrated Circuit Co., Ltd. (abbreviation: Junzheng) models T40 / T41 / A1: during the format conversion, the conventional processing is to directly move data one by one, or to use simd processing for specific classification according to the actual number of channels, but the efficiency is very low.

[0004] In addition, the commonly used technical terms in the prior art include:

[0005] 1. Floating-point feature map: The result calculated by the model is floating-point data.

[0006] 2. Simd: An instruction for vector operations.

[0007] 3. Format conversion: Converting from one dimension to another dimension, or removing some redundant data at the same time.

[0008] 4. HWC format. Explanation: The outermost layer is the height H, followed by the width W, and the innermost dimension is C. C is the channel of the feature map.

[0009] 5. DHWC32 format. Explanation: The outermost dimension is D, then the height H, followed by the width W, and the innermost dimension is C32. C32 indicates that the minimum unit of this dimension is 32 data. If it is less than 32, it is filled with redundant data. This is a four-dimensional data format representation. If there is no redundant data in C32, then C32*D is the number of channels of the data after conversion from the HWC format.

[0010] 6. Oram. A high-speed storage space with a read / write speed 10 - 20 times that of the ddr bandwidth. Directly accessing the data in oram does not occupy the ddr bandwidth.

[0011] 7. Dma instruction: An instruction for moving data, which can realize the operation of moving data from ddr to oram. This instruction can run in parallel with the simd instruction. Summary of the Invention

[0012] To solve the above problems, the purpose of the present application is to design a method that first moves data from DDR storage to ORAM through DMA instructions, converts the data in ORAM, stores the converted data in ORAM, and finally uses DMA instructions to move the data from ORAM storage to DDR. By this method, the processing speed is accelerated and the time is reduced.

[0013] Specifically, the present invention provides a design method based on the format conversion of floating-point feature maps, and the method includes the following steps:

[0014] S1. Design of the number of rows processed each time; calculate in advance the minimum number of rows generated each time, use DMA instructions, and require that the pointer address and the data to be moved must be 64-byte aligned. By default, the initial pointer address is already 64-byte aligned; let the width of the input feature map be width, the height of the feature map be height, the D of the feature map be in_d, the input data DDR pointer be in_ptr, and the data type be floating-point; the number of output channels be out_chn, and the output data storage pointer be out_ptr;

[0015] S2. SIMD processing format conversion:

[0016] Let the pointer for loading ORAM into the SIMD register be in_ptr1, the ORAM pointer for saving the generated result be out_ptr1, and the data type be floating-point; the data storage format of in_ptr1 is also DHWC32, and the number of loaded rows is real_h where real_h <= deal_oram_h; the width of the input feature map is width, the height of the feature map is height, the D of the feature map is in_d, and the data type is floating-point; the number of output channels is out_chn, and the width and height of the input feature map are the same as the input;

[0017] The number of skipped data when jumping to the next line in the input feature map, expressed as

[0018] int in_line_size = width * 32;

[0019] The number of skipped data when jumping in the input feature map, expressed as

[0020] int in_once_size = deal_oram_h * in_line_size;

[0021] The number of skipped data when jumping to the next line in the output feature map, expressed as

[0022] int out_line_size = width * out_chn;

[0023] Initialize h = 0. When h < real_h, execute loop body 1; otherwise, jump out of loop body 1. After executing loop body 1, h++. It can be expressed as for(int h = 0; h < real_h; h++)

[0024] Start of loop body 1:

[0025] Pointer after jumping h lines, expressed as

[0026] float* in_ptr2 = in_ptr1 + h * in_line_size;

[0027] Pointer after jumping h lines, expressed as

[0028] float* out_ptr2 = out_ptr1 + h * out_line_size;

[0029] Initialize w = 0. When w < width, execute loop body 2; otherwise, jump out of loop body 2. After executing loop body 2, w++. It can be expressed as: for(int w = 0; w < width; w++);

[0030] Start of loop body 2:

[0031] Skip a pointer of 32, expressed as

[0032] float* in_ptr3 = in_ptr2 + w * 32;

[0033] Skip an output oc, expressed as

[0034] float* out_ptr3 = out_ptr2 + w * out_chn;

[0035] Initialize w = 0. When n < in_d, execute loop body 3; otherwise, jump out of loop body 3. After executing loop body 3, n++. It can be expressed as: for(int n = 0; n < in_d; n++);

[0036] Start of loop body 3:

[0037] SIMD instruction for loading data, expressed as

[0038] LA(o, VR0, 0, in_ptr3, 0);

[0039] LA(o, VR0, 1, in_ptr3, 32);

[0040] LA(o, VR1, 0, in_ptr3, 64);

[0041] LA(o, VR1, 1, in_ptr3, 96);

[0042] Skip an int in_once_size = deal_oram_h * in_line_size

[0043] Calculate the data address position for the next loop load

[0044] in_ptr3 += in_once_size;

[0045] The SIMD instruction for saving data, expressed as

[0046] SA(o, VR0, 0, out_ptr3, 0);

[0047] SA(o, VR0, 1, out_ptr3, 32);

[0048] SA(o, VR1, 0, out_ptr3, 64);

[0049] SA(o, VR1, 1, out_ptr3, 96);

[0050] Continuous storage, stored 32, pointer skipped 32

[0051] out_ptr3 += 32;

[0052] The loop body 3 ends;

[0053] The loop body 2 ends;

[0054] The loop body 1 ends;

[0055] S3. Steps for the ORAM-based implementation further include:

[0056] S3.1. Calculate the number of rows processed each time deal_oram_h, and determine whether deal_oram_h is greater than 0. If it is less than or equal to 0, it does not meet the current algorithm design requirements, and this method is exited;

[0057] S3.2. Calculate the number of times ih_loop that the input feature map height can loop deal_oram_h. Let the height of the feature map be height, expressed as ih_loop = (height + deal_oram_h - 1) / deal_oram_h; S3.3. According to deal_oram_h, divide the ORAM into four parts and calculate the pointer address of each block; The input part of the ORAM occupies two parts, and the pointers are ibuf_ping and ibug_pong respectively. The output part occupies two parts, and the pointers are obuf_ping and obug_pong respectively;

[0058] S3.4. Use the dma instruction to set the specific pointer and number of data in the input feature map to be moved from the initial pointer in_ddr_ptr of the ddr to the specific pointer in ibuf_ping, and then execute the data transfer. The data transfer can be executed synchronously with others; then in_ddr_ptr += in_once_size;

[0059] S3.5. Initialize ih_i = 0, ih < ih_loop;

[0060] S3.6. Determine whether the data transfer of the dma instruction in the input part is completed. If not, wait until it is completed, and then start from the new pointer in_ddr_ptr of the input feature map ddr. Use the dma instruction to set the specific pointer and number of data to be moved from the ddr to ibuf_pong, and execute the data transfer;

[0061] S3.7. Execute the simd processing format conversion, in_ptr1 = ibuf_ping, out_ptr1 = obuf_ping; S3.8. Determine whether the data transfer of the dma instruction in the output part is completed. If not, wait until it is completed; then use the dma instruction to set the pointer and number of data to be moved from ibuf_pong of the oram to the ddr, and execute the data transfer;

[0062] S3.9. in_ddr_ptr += in_once_size, out_ddr_ptr += width * out_chn * real_h; Swap ibuf_ping and ibuf_pong, and swap obuf_ping and ibuf_pong; where the number of skipped data in the input feature map in_once_size = deal_oram_h * in_line_size, the number of skipped data for the next line in the input feature map is in_line_size = width * 32, the width of the feature map is width, the channel depth of the output feature map is out_chn, and the actual data transfer height of the feature map is real_h;

[0063] S3.10. ih_i += 1. If ih < ih_loop, return to S3.6; otherwise, all conversions are completed;

[0064] S3.11. Determine whether the data transfer of the dma instruction in the output part is completed. If not, wait until it is completed; end.

[0065] The step S1 further includes:

[0066] S1.1, Calculate the least common multiple: The data generated in one processing must be a multiple of 64 bytes and is processed in multiples of the width. The requirements are as follows:

[0067] The least common multiple of width*(out_chn*sizeof(float)) and 64;

[0068] A floating point is 4 bytes, so there is:

[0069] The least common multiple of width*(out_chn*4) and 64; 4 is a common divisor in 64, so remove 4 from all, that is: the least common multiple of width*out_chn and 16; Let the least common multiple be dh, initialize dh = 16, and convert it into a selection formula as:

[0070] If width*out_chn % 16 == 0, then dh = 1, expressed as

[0071] if(width*out_chn % 16 == 0) dh = 1;

[0072] Otherwise, if width*out_chn % 8 == 0, then dh = 2, expressed as

[0073] elseif(width*out_chn % 8 == 0) dh = 2;

[0074] Otherwise, if width*out_chn % 8 == 0, then dh = 4, expressed as

[0075] elseif(width*out_chn % 4 == 0) dh = 4;

[0076] Otherwise, if width*out_chn % 2 == 0, then dh = 8, expressed as

[0077] elseif(width*out_chn % 2 == 0) dh = 8;

[0078] Otherwise, dh = 16, expressed as

[0079] else dh = 16;

[0080] S1.2, Calculate the minimum number of processing rows: Use the pingpang principle to process data; The oram for storing the input feature map is divided into two equal-sized parts, and the oram for storing the output feature map is also divided into two equal-sized parts; Here, the oram needs to be divided into four parts;

[0081] Usage of the pingpang principle: When the DMA instruction transfers data from the DDR to the first block of ORAM, the format conversion uses the second block of ORAM; similarly, when saving data to the third block of ORAM, the DMA instruction transfers data from the fourth block of ORAM to the DDR; in this way, the format conversion and data transfer can be effectively processed in parallel, improving the processing speed;

[0082] Let the minimum ORAM size occupied by the input data be in_oram, in units of floating point, with a row size of width * in_d * 32, and the minimum number of processed rows be dh. When processed using the pingpang principle, it needs to be multiplied by 2, so we have: in_oram = [(width * in_d * 32) * dh] * 2;

[0083] Let the minimum ORAM size occupied by the output data be out_oram, in units of floating point, with a row size of width * out_chn, and the minimum number of processed rows be dh. When processed using the pingpang principle, it needs to be multiplied by 2, so we have:

[0084] out_oram = [(width * out_chn) * dh] * 2;

[0085] Let the ORAM size be oramsize, in units of floating point, the starting pointer be oram_ptr, and the data type be floating point; let the maximum number of processed dh numbers based on dh be oram_dh_num;

[0086] oram_dh_num = oramsize / (in_oram + out_oram); we have

[0087] oram_dh_num = oramsize / ({[(width * in_d * 32) * dh] * 2 + [(width * out_chn) * dh] * 2)}

[0088] where oram_dh_num is rounded down;

[0089] Let the maximum number of rows processed at one time be deal_oram_h, which is expressed as

[0090] deal_oram_h = dh * oram_dh_num;

[0091] In the pingpang principle, it is necessary to load the first block of oram first before using the data in it. If the loaded data is very large, it will take a certain amount of time to load the data. Therefore, if oram_dh_num is relatively large, deal_oram_h will be relatively large. The deal_oram_h can be divided by 4, where 4 is empirical data, to reduce the amount of data loaded for the first time; otherwise, directly use deal_oram_h. If oram_dh_num >= 4, then deal_oram_h = (oram_dh_num / 4) * dh, which can be expressed as: if(oram_dh_num >= 4) deal_oram_h = (oram_dh_num / 4) * dh.

[0092] The specific implementation of step S3 is as follows:

[0093] The number of data skipped when jumping to the next line in the input feature map, expressed as

[0094] int in_line_size = width * 32;

[0095] The number of data in one block of HWC32 in the input feature map, expressed as

[0096] int in_block_size = in_line_size * height;

[0097] The number of data skipped when jumping to the next line in the input feature map, expressed as

[0098] int out_line_size = out_width * out_chn;

[0099] Let the initial pointer of oram be oram_addr, with the type of floating point.

[0100] Calculate the number of lines deal_oram_h processed each time:

[0101] First calculate the least common multiple, expressed as

[0102] int dh = 16;

[0103] If (out_chn * width) % 16 == 0, then dh = 1, expressed as

[0104] if((out_chn * width) % 16 == 0) dh = 1;

[0105] Otherwise, if (out_chn * width) % 8 == 0, then dh = 2, expressed as

[0106] else if ((out_chn * width) % 8 == 0) dh = 2;

[0107] Otherwise, if ((out_chn * width) % 4 == 0), then dh = 4, expressed as

[0108] else if ((out_chn * width) % 4 == 0) dh = 4;

[0109] Otherwise, if ((out_chn * width) % 2 == 0), then dh = 8, expressed as

[0110] else if ((out_chn * width) % 2 == 0) dh = 8;

[0111] Otherwise, dh = 16, expressed as else dh = 16;

[0112] Calculate the number of lines processed each time deal_oram_h, expressed as

[0113] int oram_dh_num =

[0114] ((oramsize) / (2 * dh * width * out_chn + 2 * dh * width * in_d * 32)));

[0115] int deal_oram_h = (oram_dh_num) * dh;

[0116] According to experience, when deal_oram_h is relatively large, it is divided into four parts for processing to reduce the time for the first data loading;

[0117] If oram_dh_num >= 4, then deal_oram_h = (oram_dh_num / 4) * dh, expressed as

[0118] if (oram_dh_num >= 4) deal_oram_h = (oram_dh_num / 4) * dh;

[0119] Calculate the number of times ih_loop that the input feature map height can loop deal_oram_h, expressed as

[0120] int ih_loop = (in_height + deal_oram_h - 1) / deal_oram_h;

[0121] Calculate the remaining height in the last loop;

[0122] int ih_tail = height % deal_oram_h;

[0123] If ih_tail == 0, then ih_tail = line_num, expressed as

[0124] if (ih_tail == 0) ih_tail = line_num;

[0125] The oram is divided into four parts. The first two parts are used for the input oram, and the last two parts are used for the output oram; the pointers are ibuf_ping, ibuf_pong, obuf_ping, obuf_pong respectively; the initial pointer of the oram is oram_addr;

[0126] uint32_t ibuf_ping = (uint32_t)oram_addr;

[0127] uint32_t ibuf_pong = ibuf_ping + deal_oram_h * width * in_d * 32;

[0128] uint32_t obuf_ping = ibuf_pong + deal_oram_h * width * in_d * 32;

[0129] uint32_t obuf_pong = obuf_ping + deal_oram_h * width * out_chn;

[0130] The number of data skipped in the input feature map when jumping over a data with a width of C32 and a height of deal_oram_h, expressed as

[0131] int in_once_size = deal_oram_h * in_line_size;

[0132] The number of data skipped in the remaining height, expressed as

[0133] int ibuf_tail_size = in_line_size * ih_tail;

[0134] The initial pointer of the ddr of the input feature map is in_ddr, and the initial pointer of the ddr of the output feature map is out_ddr, expressed as

[0135] float* in_batch_ddr = in_ddr;

[0136] float* out_batch_ddr = out_ddr;

[0137] In the first piece of oram, the number of data with a width of C32 and a height of deal_oram_h skipped by the jump is represented as

[0138] uint32_t dma_in_size = in_once_size;

[0139] Reassign the pointer,

[0140] uint32_t* ibuf_write_ptr = ibuf_ptr_ping;

[0141] float* in_read_ddr = in_batch_ddr;

[0142] Initialize ifp_i = 0. When it is judged that ifp_i < in_d, execute the loop body. After the loop body is executed, i++. It is represented as for(int ifp_i = 0; ifp_i < in_d; ifp_i++);

[0143] The loop body starts:

[0144] The two pointers set by the dma conversion and the quantity of each conversion are represented as

[0145] dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size);

[0146] The jump of each pointer is represented as

[0147] in_read_ddr += in_block_size;

[0148] ibuf_write_ptr += dma_in_size;

[0149] The loop body ends;

[0150] The dma instruction is executed and run, which is represented as

[0151] dma_run_IF(in_des_ptr);

[0152] int cur_line_num = deal_oram_h;

[0153] Initialize ih_i = 0. When it is judged that ih_i < ih_loop, execute the loop body. After executing the loop body, i++. It can be expressed as for(int ih_i = 0; ih_i < ih_loop; ih_i++);

[0154] Start of the ih_loop loop body:

[0155] Judge whether the data has been transferred completely. If not, wait. It can be expressed as

[0156] dma_wait_if;

[0157] Judge whether it is the last time. If it is not the last time of the first loading, assign the remaining height of the last time to cur_line_num. It can be expressed as

[0158] if(ih_i != 0 && ih_i == ih_loop - 1)

[0159] cur_line_num = ih_tail;

[0160] Recalculate the number of data to be skipped by jumping over a C32 width of width and height of cur_line_num

[0161] in_once_size = in_line_size * cur_line_num;

[0162] If ih_i != ih_loop - 1, which can be expressed as if(ih_i != ih_loop - 1), then execute

[0163] in_batch_ddr += in_once_size;

[0164] If ih_i == ih_loop - 2, which can be expressed as if(ih_i == ih_loop - 2), then execute

[0165] Calculate the data transferred in the last time, which is the data change of the dma pointer each time. It can be expressed as

[0166] dma_in_size = ibuf_tail_size;

[0167] Assign the pointer to another temporary pointer

[0168] in_read_ddr = in_batch_ddr;

[0169] ibuf_write_ptr = ibuf_ptr_pong;

[0170] Initialize ifp_i = 0. When it is judged that ifp_i < in_d, execute the loop body. After executing the loop body, ifp_i++. It can be expressed as for(int ifp_i = 0; ifp_i < in_d; ifp_i++ ifp_i++);

[0171] Start of the loop body:

[0172] Set two pointers for DMA conversion and the quantity for each conversion. It can be expressed as

[0173] dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size);

[0174] Jump of each pointer. It can be expressed as

[0175] in_read_ddr += in_block_size;

[0176] ibuf_write_ptr += dma_in_size;

[0177] End of the loop body;

[0178] Execute the DMA instruction. It can be expressed as

[0179] dma_run_IF(in_des_ptr);

[0180] Assign the pointer to another temporary pointer. It can be expressed as

[0181] float* in_ptr1 = (float*)ibuf_ptr_ping;

[0182] float* out_ptr1 = (float*)obuf_ptr_ping;

[0183] Use SIMD instructions to process format conversion;

[0184] Use the output DMA wait instruction to judge whether the data has been transferred completely. If not, wait until it is transferred completely. It can be expressed as out_wait;

[0185] The number of data transferred at a time. It can be expressed as

[0186] int dma_out_size = cur_line_num * width * out_once_size;

[0187] uint32_t* out_ddr_ptr = out_batch_ddr;

[0188] uint32_t* obuf_write_ptr = ibuf_ptr_pong;

[0189] The two pointers for DMA setting conversion and the quantity for each conversion are represented as

[0190] dma_node(obuf_write_ptr, out_ddr_ptr, dma_out_size);

[0191] The pointer for jumping and saving to DDR is represented as

[0192] out_batch_ddr += out_line_size * cur_line_num;

[0193] Swap the ibuf_ping and ibuf_pong pointers, represented as

[0194] swap(ibuf_ping, ibuf_pong);

[0195] Swap the obuf_ptr_ping and obuf_pong pointers, represented as

[0196] swap(obuf_ptr_ping, obuf_pong);

[0197] The ih_loop loop body ends;

[0198] Use the output DMA wait instruction to determine whether the data has been transferred completely. If not, wait until it is transferred completely, represented as out_wait.

[0199] The applicable scope of the described method is when the oram can accommodate the minimum number of processed lines.

[0200] Therefore, the advantages of this application are as follows: Through this method, the processing speed is accelerated and the time is reduced. The existing processing speed can be doubled. For example, on chips of the Junzheng T40 and T41 models, the speed can be doubled. Brief Description of the Drawings

[0201] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.

[0202] Figure 1 It is a flow schematic diagram of this method. Detailed Embodiment

[0203] In order to be able to more clearly understand the technical content and advantages of the present invention, the present invention will now be further described in detail in conjunction with the drawings.

[0204] In the prior art, data is directly read from the DDR memory into the registers of the SIMD, and then stored in the HWC format after being processed by the SIMD, resulting in a relatively long processing time. In the present invention, a design method is provided. First, the data is transferred from the DDR memory to the ORAM via a DMA instruction, the data in the ORAM is format-converted, the converted data is stored in the ORAM, and finally, the data is transferred from the ORAM memory to the DDR using a DMA instruction. This design method based on the format conversion of floating-point feature maps can double the processing speed. It is applicable when the ORAM can accommodate the minimum number of processing rows. As Figure 1 shown, it includes the following steps:

[0205] S1. Design of the number of rows processed each time:

[0206] Since there may be redundant numbers in the data in the DHWC32 format, it is necessary to calculate in advance the minimum number of rows of data generated each time. It further includes steps S1.1 and S1.2, where dh is the calculation result. Using a DMA instruction, it is required that the pointer address and the data to be transferred must be 64-byte aligned. By default, the initial pointer address is already 64-byte aligned.

[0207] Let the width of the input feature map be width, the height of the feature map be height, the D of the feature map be in_d, the input data DDR pointer be in_ptr, and the data type be floating point. The number of output channels is out_chn, and the output data storage pointer is out_ptr.

[0208] S1.1. Calculate the least common multiple.

[0209] The data generated in one processing must be a multiple of 64 bytes. The processing is carried out in multiples of the width. The following requirements apply:

[0210] The least common multiple of width*(out_chn*sizeof(float)) and 64.

[0211] A floating point is 4 bytes, so there is:

[0212] The least common multiple of width*(out_chn*4) and 64. Since 64 has a common divisor of 4, all 4s are removed, that is:

[0213] The least common multiple of width*out_chn and 16. Let the least common multiple be dh, and initialize dh = 16. The conversion to a selection formula is:

[0214] if(width*out_chn % 16 == 0) dh = 1;

[0215] elseif (width * out_chn % 8 == 0) dh = 2;

[0216] elseif (width * out_chn % 4 == 0) dh = 4;

[0217] elseif (width * out_chn % 2 == 0) dh = 8;

[0218] else dh = 16;

[0219] S1.2, calculate the minimum number of processing rows.

[0220] Use the pingpong principle to process data. The oram for storing the input feature map is divided into two equal-sized parts, and the oram for storing the output feature map is also divided into two equal-sized parts. Here, the oram needs to be divided into four parts. Usage of the pingpong principle: When the dma instruction transfers data from the ddr to the first oram block, the format conversion uses the second oram block; similarly, when saving data to the third oram block, the dma instruction transfers data from the fourth oram block to the ddr. This can effectively implement parallel processing of format conversion and data transfer, improving the processing speed.

[0221] Assume that the minimum oram size occupied by the input data is in_oram, with the unit of floating point, the size of one row is width * in_d * 32, the minimum number of processing rows is dh, and when using the pingpong principle for processing, it needs to be multiplied by 2, so there is:

[0222] in_oram = [(width * in_d * 32) * dh] * 2

[0223] Assume that the minimum oram size occupied by the output data is out_oram, with the unit of floating point, the size of one row is width * out_chn, the minimum number of processing rows is dh, and when using the pingpong principle for processing, it needs to be multiplied by 2, so there is:

[0224] out_oram = [(width * out_chn) * dh] * 2

[0225] Assume that the oram size is oramsize, with the unit of floating point, the starting pointer is oram_ptr, and the data type is floating point. Assume that the maximum number of processing dh numbers based on dh is oram_dh_num.

[0226] oram_dh_num = oramsize / (in_oram + out_oram)

[0227] There is

[0228] oram_dh_num = oramsize / ({[(width * in_d * 32) * dh] * 2 + [(width * out_chn) * dh] * 2})

[0229] where oram_dh_num is rounded down.

[0230] Let the maximum number of rows processed at one time be deal_oram_h.

[0231] deal_oram_h = dh * oram_dh_num

[0232] Since in the pingpong principle, the first block of oram needs to be loaded completely before its data can be used. If the loaded data is very large, resulting in a certain time occupied for loading the data, so when oram_dh_num is relatively large, deal_oram_h will be relatively large. deal_oram_h can be divided by 4 (4 is an empirical data) to reduce the amount of data loaded for the first time. Otherwise, directly use deal_oram_h.

[0233] if (oram_dh_num >= 4) deal_oram_h = (oram_dh_num / 4) * dh. S2. SIMD processing format conversion.

[0234] Let the pointer for loading oram into the SIMD register be in_ptr1, the oram pointer for saving the generated result be out_ptr1, and the data type be floating-point; the data storage format of in_ptr1 is also DHWC32, the number of loaded rows is real_h where real_h <= deal_oram_h; the width of the input feature map is width, the height of the feature map is height, the D of the feature map is in_d, and the data type is floating-point; the number of output channels is out_chn, and the width and height of the input feature map are the same as the input;

[0235] The number of data skipped when jumping to the next row in the input feature map is expressed as

[0236] int in_line_size = width * 32;

[0237] The number of data skipped when jumping in the input feature map is expressed as

[0238] int in_once_size = deal_oram_h * in_line_size;

[0239] The number of data skipped when jumping to the next row in the output feature map is expressed as

[0240] int out_line_size = width * out_chn;

[0241] Initialize h = 0. When h < real_h, execute loop body 1; otherwise, jump out of loop body 1. After executing loop body 1, h++. It can be expressed as: for(int h = 0; h < real_h; h++)

[0242] Start of loop body 1:

[0243] Pointer after jumping h lines, expressed as

[0244] float* in_ptr2 = in_ptr1 + h * in_line_size;

[0245] Pointer after jumping h lines, expressed as

[0246] float* out_ptr2 = out_ptr1 + h * out_line_size;

[0247] Initialize w = 0. When w < width, execute loop body 2; otherwise, jump out of loop body 2. After executing loop body 2, w++. It can be expressed as: for(int w = 0; w < width; w++);

[0248] Start of loop body 2:

[0249] Skip a pointer of 32, expressed as

[0250] float* in_ptr3 = in_ptr2 + w * 32;

[0251] Skip an output oc, expressed as

[0252] float* out_ptr3 = out_ptr2 + w * out_chn;

[0253] Initialize w = 0. When n < in_d, execute loop body 3; otherwise, jump out of loop body 3. After executing loop body 3, n++. It can be expressed as: for(int n = 0; n < in_d; n++);

[0254] Start of loop body 3:

[0255] SIMD instruction for loading data, expressed as

[0256] LA(o, VR0, 0, in_ptr3, 0);

[0257] LA(o, VR0, 1, in_ptr3, 32);

[0258] LA(o, VR1, 0, in_ptr3, 64);

[0259] LA(o, VR1, 1, in_ptr3, 96);

[0260] Skip an int in_once_size = deal_oram_h * in_line_size

[0261] Calculate the data loading address position for the next loop

[0262] in_ptr3 += in_once_size;

[0263] The SIMD instruction for saving data, expressed as

[0264] SA(o, VR0, 0, out_ptr3, 0);

[0265] SA(o, VR0, 1, out_ptr3, 32);

[0266] SA(o, VR1, 0, out_ptr3, 64);

[0267] SA(o, VR1, 1, out_ptr3, 96);

[0268] Continuous storage, storing 32, and the pointer skips 32

[0269] out_ptr3 += 32;

[0270] The loop body 3 ends;

[0271] The loop body 2 ends;

[0272] The loop body 1 ends;

[0273] S3. Steps for the ORAM-based implementation.

[0274] S3.1. Calculate the number of rows processed each time deal_oram_h, and determine whether deal_oram_h is greater than 0. If it is less than or equal to 0, it does not meet the requirements of the current algorithm design, and this method is exited.

[0275] S3.2. Calculate the number of times ih_loop that the input feature map height can loop deal_oram_h. Let the height of the feature map be height.

[0276] ih_loop = (height + deal_oram_h - 1) / deal_oram_h;

[0277] S3.3, Divide the ORAM into four parts according to deal_oram_h, and calculate the pointer address of each block. The input part of the ORAM occupies two parts, and the pointers are ibuf_ping and ibug_pong respectively. The output part occupies two parts, and the pointers are obuf_ping and obug_pong respectively.

[0278] S3.4, Use the dma instruction to set the specific pointer and number of data in the input feature map to be moved from the ddr initial pointer in_ddr_ptr to the specific pointer in ibuf_ping, and then execute the data transfer. The data transfer can be executed synchronously with others. Then in_ddr_ptr += in_once_size;

[0279] S3.6, Initialize ih_i = 0, ih < ih_loop;

[0280] S3.6, Determine whether the data transfer of the dma instruction in the input part is completed. If it is not completed, wait until it is completed. Then, starting from the new pointer in_ddr_ptr of the input feature map ddr, use the dma instruction to set the specific pointer and number of data to be moved from the ddr to ibuf_pong, and execute the data transfer.

[0281] S3.7, Execute the simd processing format conversion, in_ptr1 = ibuf_ping, out_ptr1 = obuf_ping. S3.8, Determine whether the data transfer of the dma instruction in the output part is completed. If it is not completed, wait until it is completed. Then use the dma instruction to set the pointer and number of data to be moved from ibuf_pong of the ORAM to the ddr, and execute the data transfer.

[0282] S3.9, in_ddr_ptr += in_once_size, out_ddr_ptr += width * out_chn * real_h; Swap ibuf_ping and ibuf_pong, and swap obuf_ping and ibuf_pong. Among them, the number of skipped data in the input feature map in_once_size = deal_oram_h * in_line_size, the number of skipped data for the next line in the input feature map is in_line_size = width * 32, the width of the feature map is width, the channel depth of the output feature map is out_chn, and the actual data transfer height of the feature map is real_h.

[0283] S3.10, ih_i += 1. If ih < ih_loop, return to S3.6, otherwise all conversions are completed. S3.11, Determine whether the data transfer of the dma instruction in the output part is completed. If it is not completed, wait until it is completed. End.

[0284] The specific implementation is as follows:

[0285] The number of data skipped when jumping to the next line in the input feature map, denoted as

[0286] int in_line_size = width * 32;

[0287] The number of data in one block of HWC32 in the input feature map, denoted as

[0288] int in_block_size = in_line_size * height;

[0289] The number of data skipped when jumping to the next line in the input feature map, denoted as

[0290] int out_line_size = out_width * out_chn;

[0291] Set the initial pointer of oram to oram_addr, with the type being floating-point

[0292] Calculate the number of rows processed each time, deal_oram_h:

[0293] First calculate the least common multiple, denoted as

[0294] int dh = 16;

[0295] If (out_chn * width) % 16 == 0, then dh = 1, denoted as

[0296] if ((out_chn * width) % 16 == 0) dh = 1;

[0297] Otherwise, if (out_chn * width) % 8 == 0, then dh = 2, denoted as

[0298] else if ((out_chn * width) % 8 == 0) dh = 2;

[0299] Otherwise, if (out_chn * width) % 4 == 0, then dh = 4, denoted as

[0300] else if ((out_chn * width) % 4 == 0) dh = 4;

[0301] Otherwise, if (out_chn * width) % 2 == 0, then dh = 8, denoted as

[0302] else if ((out_chn * width) % 2 == 0) dh = 8;

[0303] Otherwise, dh = 16, represented as else dh = 16;

[0304] Calculate the number of lines processed each time, deal_oram_h, represented as

[0305] int oram_dh_num =

[0306] ((oramsize) / (2*dh*width*out_chn + 2*dh*width*in_d*32)));

[0307] int deal_oram_h = (oram_dh_num)*dh;

[0308] According to experience, when deal_oram_h is relatively large, divide it into four parts for processing to reduce the time for the first data loading;

[0309] If oram_dh_num >= 4, then deal_oram_h = (oram_dh_num / 4)*dh, represented as

[0310] if(oram_dh_num >= 4) deal_oram_h = (oram_dh_num / 4)*dh;

[0311] Calculate the number of times ih_loop that the height of the input feature map can loop deal_oram_h, represented as

[0312] int ih_loop = (in_height + deal_oram_h - 1) / deal_oram_h;

[0313] Calculate the remaining height in the last loop;

[0314] int ih_tail = height % deal_oram_h;

[0315] If ih_tail == 0, then ih_tail = line_num, represented as

[0316] if(ih_tail == 0) ih_tail = line_num;

[0317] The oram is divided into four parts. The first two parts are used for the input oram, and the last two parts are used for the output oram; the pointers are ibuf_ping, ibuf_pong, obuf_ping, obuf_pong respectively; the initial pointer of the oram is oram_addr;

[0318] uint32_t ibuf_ping = (uint32_t)oram_addr;

[0319] uint32_t ibuf_pong = ibuf_ping + deal_oram_h * width * in_d * 32;

[0320] uint32_t obuf_ping = ibuf_pong + deal_oram_h * width * in_d * 32;

[0321] uint32_t obuf_pong = obuf_ping + deal_oram_h * width * out_chn;

[0322] The number of data skipped by jumping in the input feature map, with a width of C32 width and a height of deal_oram_h, is expressed as

[0323] int in_once_size = deal_oram_h * in_line_size;

[0324] The number of data skipped in the remaining height, expressed as

[0325] int ibuf_tail_size = in_line_size * ih_tail;

[0326] The initial pointer of the input feature map ddr is in_ddr, and the initial pointer of the output feature map ddr is out_ddr, expressed as

[0327] float* in_batch_ddr = in_ddr;

[0328] float* out_batch_ddr = out_ddr;

[0329] In the first block of oram, the number of data skipped by jumping, with a width of C32 width and a height of deal_oram_h, is expressed as

[0330] uint32_t dma_in_size = in_once_size;

[0331] Reassign the pointer,

[0332] uint32_t* ibuf_write_ptr = ibuf_ptr_ping;

[0333] float* in_read_ddr = in_batch_ddr;

[0334] Initialize ifp_i = 0. When it is judged that ifp_i < in_d, execute the loop body. After the loop body is executed, i++ is expressed as for(int ifp_i = 0; ifp_i < in_d; ifp_i++);

[0335] The loop body starts:

[0336] Set two pointers for dma conversion and the quantity for each conversion, which is expressed as

[0337] dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size);

[0338] The jump of each pointer is expressed as

[0339] in_read_ddr += in_block_size;

[0340] ibuf_write_ptr += dma_in_size;

[0341] The loop body ends;

[0342] The dma instruction is executed and run, which is expressed as

[0343] dma_run_IF(in_des_ptr);

[0344] int cur_line_num = deal_oram_h;

[0345] Initialize ih_i = 0. When it is judged that ih_i < ih_loop, execute the loop body. After the loop body is executed, i++ is expressed as for(int ih_i = 0; ih_i < ih_loop; ih_i++);

[0346] The ih_loop loop body starts:

[0347] Judge whether the data has been moved completely. If not, wait, which is expressed as

[0348] dma_wait_if;

[0349] Judge whether it is the last time. If it is not the last time of the first loading, assign the remaining height of the last time to cur_line_num, which is expressed as

[0350] if(ih_i != 0 && ih_i == ih_loop - 1)

[0351] cur_line_num = ih_tail;

[0352] Recalculate the jump to skip the number of data with a width of C32 and a height of cur_line_num

[0353] in_once_size = in_line_size * cur_line_num;

[0354] If ih_i!= ih_loop - 1, which is expressed as if(ih_i!= ih_loop - 1), then execute

[0355] in_batch_ddr += in_once_size;

[0356] If ih_i == ih_loop - 2, which is expressed as if(ih_i == ih_loop - 2), then execute

[0357] Calculate the number of data transferred in the last move, which represents the data change of the dma pointer each time, expressed as

[0358] dma_in_size = ibuf_tail_size;

[0359] Assign the pointer to another temporary pointer

[0360] in_read_ddr = in_batch_ddr;

[0361] ibuf_write_ptr = ibuf_ptr_pong;

[0362] Initialize ifp_i = 0. When judging ifp_i < in_d, execute the loop body. After executing the loop body, ifp_i++; which is expressed as for(int ifp_i = 0; ifp_i < in_d; ifp_i++ ifp_i++);

[0363] Start of the loop body:

[0364] The dma sets two pointers for conversion and the quantity of each conversion, expressed as

[0365] dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size);

[0366] The jump of the pointer each time, expressed as

[0367] in_read_ddr += in_block_size;

[0368] ibuf_write_ptr += dma_in_size;

[0369] End of loop body;

[0370] The DMA instruction is executed, denoted as

[0371] dma_run_IF(in_des_ptr);

[0372] Assign the pointer to another temporary pointer, denoted as

[0373] float*in_ptr1 = (float*)ibuf_ptr_ping;

[0374] float*out_ptr1 = (float*)obuf_ptr_ping;

[0375] Use SIMD instructions to process format conversion;

[0376] Use the output DMA wait instruction to determine if the data has been transferred completely. If not, wait until it is transferred completely, denoted as out_wait;

[0377] The number of data transferred at a time, denoted as

[0378] int dma_out_size = cur_line_num * width * out_once_size;

[0379] uint32_t*out_ddr_ptr = out_batch_ddr;

[0380] uint32_t*obuf_write_ptr = ibuf_ptr_pong;

[0381] Set the two pointers for DMA conversion and the quantity of each conversion, denoted as

[0382] dma_node(obuf_write_ptr,out_ddr_ptr,dma_out_size);

[0383] Jump to the pointer saved in DDR, denoted as

[0384] out_batch_ddr += out_line_size * cur_line_num;

[0385] Swap the ibuf_ping and ibuf_pong pointers, denoted as

[0386] swap(ibuf_ping,ibuf_pong);

[0387] Swap the obuf_ptr_ping and obuf_pong pointers, expressed as

[0388] swap(obuf_ptr_ping, obuf_pong);

[0389] The ih_loop loop body ends;

[0390] Use the output dma wait instruction to determine whether the data has been transferred. If not, wait until it is transferred, expressed as out_wait.

[0391] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A design method based on the conversion of floating-point feature map format, characterized in that, The method includes the following steps: S1. Design of the number of lines processed each time: Calculate in advance how many lines of data are required to be generated for each minimum process, and use the dma instruction. It is required that the pointer address and the data to be moved must be 64-byte aligned. By default, the initial pointer address is already 64-byte aligned; Let the width of the input feature map be width, the height of the feature map be height, the D of the feature map be in_d, the input data ddr pointer be in_ptr, and the data type be floating point; the number of output channels be out_chn, and the pointer for storing the output data be out_ptr; S2. simd processing format conversion: Let the pointer for loading oram into the simd register be in_ptr1, the oram pointer for saving the generated result be out_ptr1, and the data type be floating point; the data storage format of in_ptr1 is also DHWC32, and the number of loaded lines is real_h where real_h <= deal_oram_h; the width of the input feature map is width, the height of the feature map is height, the D of the feature map is in_d, and the data type is floating point; the number of output channels is out_chn, and the width and height of the input feature map are the same as the input; The number of skipped data when jumping to the next line in the input feature map, expressed as int in_line_size = width * 32; The number of skipped data when jumping in the input feature map, expressed as int in_once_size = deal_oram_h * in_line_size; The number of skipped data when jumping to the next line in the output feature map, expressed as int out_line_size = width * out_chn; Initialize h = 0. When h < real_h, execute loop body 1, otherwise jump out of loop body 1. After executing loop body 1, h++. Expressed as: for(int h = 0; h < real_h; h++) The start of loop body 1: The pointer after jumping h lines, expressed as float* in_ptr2 = in_ptr1 + h * in_line_size; The pointer after jumping h lines, expressed as float* out_ptr2 = out_ptr1 + h * out_line_size; Initialize w = 0. If w < width, execute loop body 2, otherwise jump out of loop body 2. After executing loop body 2, w++. Expressed as: for(int w = 0; w < width; w++); The start of loop body 2: The pointer skipping a 32, expressed as float* in_ptr3 = in_ptr2 + w * 32; Skipping one output oc, expressed as float* out_ptr3 = out_ptr2 + w * out_chn; Initialize w = 0. If n < in_d, execute loop body 3, otherwise jump out of loop body 3. After executing loop body 3, n++. Expressed as: for(int n = 0; n < in_d; n++); The start of loop body 3: The SIMD instruction for loading data, represented as LA(o, VR0, 0, in_ptr3, 0); LA(o, VR0, 1, in_ptr3, 32); LA(o, VR1, 0, in_ptr3, 64); LA(o, VR1, 1, in_ptr3, 96); Skip an int in_once_size = deal_oram_h * in_line_size Calculate the address position for loading data in the next loop in_ptr3 += in_once_size; The SIMD instruction for saving data, represented as SA(o, VR0, 0, out_ptr3, 0); SA(o, VR0, 1, out_ptr3, 32); SA(o, VR1, 0, out_ptr3, 64); SA(o, VR1, 1, out_ptr3, 96); Continuous storage, storing 32, and the pointer skips 32 out_ptr3 += 32; The end of loop body 3; The end of loop body 2; The end of loop body 1; S3. The implementation steps based on ORAM further include: S3.

1. Calculate the number of rows deal_oram_h processed each time, and judge whether deal_oram_h is greater than 0. If it is less than or equal to 0, it does not meet the requirements of the current algorithm design, and jump out of this method; S3.

2. Calculate the number of times ih_loop that the height of the input feature map can be looped by deal_oram_h. Let the height of the feature map be height, represented as ih_loop = (height + deal_oram_h - 1) / deal_oram_h; S3.

3. According to deal_oram_h, divide the ORAM into four parts and calculate the pointer address of each block; The input part of the ORAM occupies two parts, and the pointers are ibuf_ping and ibug_pong respectively. The output part occupies two parts, and the pointers are obuf_ping and obug_pong respectively; S3.

4. Use the DMA instruction to set the specific pointer and number of data in the input feature map from the initial pointer in_ddr_ptr of the DDR to the specific pointer in ibuf_ping, and then execute the data transfer. The data transfer can be executed synchronously with others; Then in_ddr_ptr += in_once_size; S3.

5. Initialize ih_i = 0, ih < ih_loop; S3.

6. Judge whether the data transfer of the DMA instruction in the input part is completed. If it is not completed, wait until it is completed, and then start from the new pointer in_ddr_ptr of the input feature map DDR, use the DMA instruction to set the specific pointer and number of data transferred from the DDR to ibuf_pong, and execute the data transfer; S3.

7. Perform SIMD processing format conversion, with in_ptr1 = ibuf_ping and out_ptr1 = obuf_ping; S3.

8. Determine whether the data transfer of the DMA instruction in the output part is completed. If not, wait until it is completed; then use the DMA instruction to set the pointer and number of data transferred from ibuf_pong in ORAM to DDR, and perform the data transfer. S3.

9. in_ddr_ptr += in_once_size, out_ddr_ptr += width * out_chn * real_h; Swap ibuf_ping and ibuf_pong, and swap obuf_ping and ibuf_pong; where the number of skipped data in the input feature map in_once_size = deal_oram_h * in_line_size, the number of skipped data when jumping to the next line in the input feature map in_line_size = width * 32, the width of the feature map is width, the channel depth of the output feature map is out_chn, and the actual data transfer height of the feature map is real_h. S3.

10. ih_i += 1. If ih < ih_loop, return to S3.6; otherwise, all conversions are completed. S3.

11. Determine whether the data transfer of the DMA instruction in the output part is completed. If not, wait until it is completed; end.

2. The design method based on the conversion of the floating-point feature map format according to claim 1, wherein The step S1 further includes: S1.

1. Calculate the least common multiple: The data generated in one processing must be a multiple of 64 bytes and is processed in multiples of the width; the requirements are as follows: The least common multiple of width * (out_chn * sizeof(float)) and 64. A floating point is 4 bytes, so there is: The least common multiple of width * (out_chn * 4) and 64; 4 is a common divisor in 64, so all 4s are removed, that is: the least common multiple of width * out_chn and 16; Let the least common multiple be dh, and initialize dh = 16. The conversion formula is: If width * out_chn % 16 == 0, then dh = 1, expressed as if (width * out_chn % 16 == 0) dh = 1; Otherwise, if width * out_chn % 8 == 0, then dh = 2, expressed as else if (width * out_chn % 8 == 0) dh = 2; Otherwise, if width * out_chn % 8 == 0, then dh = 4, expressed as else if (width * out_chn % 4 == 0) dh = 4; Otherwise, if width * out_chn % 2 == 0, then dh = 8, expressed as else if (width * out_chn % 2 == 0) dh = 8; Otherwise, dh = 16, expressed as else dh = 16; S1.

2. Calculate the minimum number of processed rows: Use the pingpang principle to process data; the oram for storing the input feature map is divided into two equal-sized parts, and the oram for storing the output feature map is also divided into two equal-sized parts; here, the oram needs to be divided into four parts; Usage of the pingpang principle: When the dma instruction moves data from the ddr to the first oram block, the format conversion uses the second oram; similarly, when saving data to the third oram block, the dma instruction moves data from the fourth oram block to the ddr; this can effectively achieve parallel processing of format conversion and data transfer, improving the processing speed; Assume that the minimum oram size occupied by the input data is in_oram, with the unit of floating point, the size of one row is width * in_d * 32, the minimum number of processed rows is dh, and using the pingpang principle for processing, it needs to be multiplied by 2, so we have: in_oram = [(width * in_d * 32) * dh] * 2; Assume that the minimum oram size occupied by the output data is out_oram, with the unit of floating point, the size of one row is width * out_chn, the minimum number of processed rows is dh, and using the pingpang principle for processing, it needs to be multiplied by 2, so we have: out_oram = [(width * out_chn) * dh] * 2; Assume that the oram size is oramsize, with the unit of floating point, the starting pointer is oram_ptr, and the data type is floating point; assume that the maximum number of processed dh based on dh is oram_dh_num; oram_dh_num = oramsize / (in_oram + out_oram); so oram_dh_num = oramsize / ({[(width * in_d * 32) * dh] * 2 + [(width * out_chn) * dh] * 2}) where oram_dh_num is rounded down; Assume that the maximum number of processed rows at one time is deal_oram_h, which is expressed as deal_oram_h = dh * oram_dh_num; Since in the pingpang principle, it is necessary to load the first oram block completely before using the data in it. If the loaded data is very large, it will take a certain amount of time to load the data. Therefore, if oram_dh_num is relatively large, deal_oram_h will be relatively large. The deal_oram_h can be divided by 4. Here, 4 is an empirical data to reduce the amount of data loaded for the first time; otherwise, directly use deal_oram_h; if oram_dh_num >= 4, then deal_oram_h = (oram_dh_num / 4) * dh, which is expressed as: if(oram_dh_num >= 4)deal_oram_h = (oram_dh_num / 4) * dh.

3. A design method based on the conversion of floating-point feature map format according to claim 1, characterized in that, The specific implementation of step S3 is as follows: The number of skipped data when jumping to the next row in the input feature map, which is expressed as int in_line_size = width * 32; The number of data in one block of HWC32 in the input feature map, denoted as int in_block_size = in_line_size * height; The number of data skipped when jumping to the next line in the input feature map, denoted as int out_line_size = out_width * out_chn; Let the initial pointer of oram be oram_addr, with the type of floating point, Calculate the number of lines deal_oram_h processed each time: First calculate the least common multiple, denoted as int dh = 16; If (out_chn * width) % 16 == 0, then dh = 1, denoted as if ((out_chn * width) % 16 == 0) dh = 1; Otherwise, if (out_chn * width) % 8 == 0, then dh = 2, denoted as else if ((out_chn * width) % 8 == 0) dh = 2; Otherwise, if (out_chn * width) % 4 == 0, then dh = 4, denoted as else if ((out_chn * width) % 4 == 0) dh = 4; Otherwise, if (out_chn * width) % 2 == 0, then dh = 8, denoted as else if ((out_chn * width) % 2 == 0) dh = 8; Otherwise, dh = 16, denoted as else dh = 16; Calculate the number of lines deal_oram_h processed each time, denoted as int oram_dh_num = ((oramsize) / (2 * dh * width * out_chn + 2 * dh * width * in_d * 32))); int deal_oram_h = (oram_dh_num) * dh; According to experience, when deal_oram_h is relatively large, it is divided into four parts for processing to reduce the time of loading data for the first time; If oram_dh_num >= 4, then deal_oram_h = (oram_dh_num / 4) * dh, denoted as if (oram_dh_num >= 4) deal_oram_h = (oram_dh_num / 4) * dh; Calculate the number of times ih_loop that the height of the input feature map can loop deal_oram_h, denoted as int ih_loop = (in_height + deal_oram_h - 1) / deal_oram_h; Calculate the remaining height in the last loop; int ih_tail = height % deal_oram_h; If ih_tail == 0, then ih_tail = line_num, denoted as if (ih_tail == 0) ih_tail = line_num; The ORAM is divided into four parts. The first two parts are used for the input ORAM, and the last two parts are used for the output ORAM. The pointers are ibuf_ping, ibuf_pong, obuf_ping, and obuf_pong respectively. The initial pointer of the ORAM is oram_addr; uint32_t ibuf_ping = (uint32_t)oram_addr; uint32_t ibuf_pong = ibuf_ping + deal_oram_h * width * in_d * 32; uint32_t obuf_ping = ibuf_pong + deal_oram_h * width * in_d * 32; uint32_t obuf_pong = obuf_ping + deal_oram_h * width * out_chn; The number of data skipped in the input feature map when jumping over a data with a width of C32 (width) and a height of deal_oram_h is expressed as int in_once_size = deal_oram_h * in_line_size; The number of data skipped in the remaining height is expressed as int ibuf_tail_size = in_line_size * ih_tail; The initial pointer of the DDR of the input feature map is in_ddr, and the initial pointer of the DDR of the output feature map is out_ddr, which is expressed as float* in_batch_ddr = in_ddr; float* out_batch_ddr = out_ddr; In the first ORAM block, the number of data skipped when jumping over a data with a width of C32 (width) and a height of deal_oram_h is expressed as uint32_t dma_in_size = in_once_size; Reassign the pointer, uint32_t* ibuf_write_ptr = ibuf_ptr_ping; float* in_read_ddr = in_batch_ddr; Initialize ifp_i = 0. When it is judged that ifp_i < in_d, execute the loop body. After the loop body is executed, i++. It is expressed as for(int ifp_i = 0; ifp_i < in_d; ifp_i++); The loop body starts: Set the two pointers for DMA conversion and the quantity for each conversion, which is expressed as dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size); The jump of each pointer is expressed as in_read_ddr += in_block_size; ibuf_write_ptr += dma_in_size; The loop body ends; Execute the DMA instruction, which is expressed as dma_run_IF(in_des_ptr); int cur_line_num = deal_oram_h; Initialize ih_i = 0. When it is judged that ih_i < ih_loop, execute the loop body, and after executing the loop body, i++. It can be expressed as for(int ih_i = 0; ih_i < ih_loop; ih_i++); The start of the ih_loop loop body: Judge whether the data has been transferred completely. If not, wait, which is expressed as dma_wait_if; Judge whether it is the last time. If it is not the last time of the first load, assign the remaining height of the last time to cur_line_num, which is expressed as if(ih_i != 0 && ih_i == ih_loop - 1) cur_line_num = ih_tail; Recalculate the number of data skipped by jumping over a C32 width of width height cur_line_num in_once_size = in_line_size * cur_line_num; If ih_i != ih_loop - 1, which is expressed as if(ih_i != ih_loop - 1), then execute in_batch_ddr += in_once_size; If ih_i == ih_loop - 2, which is expressed as if(ih_i == ih_loop - 2), then execute Calculate the data transfer for the last time, the data change of the dma pointer each time, which is expressed as dma_in_size = ibuf_tail_size; Assign the pointer to another temporary pointer in_read_ddr = in_batch_ddr; ibuf_write_ptr = ibuf_ptr_pong; Initialize ifp_i = 0. When it is judged that ifp_i < in_d, execute the loop body, and after executing the loop body, ifp_i++. It can be expressed as for(int ifp_i = 0; ifp_i < in_d; ifp_i++ifp_i++); The start of the loop body: Set the two pointers for dma conversion and the quantity for each conversion, which is expressed as dma_node(ibuf_write_ptr, in_read_ddr, dma_in_size); The jump of the pointer each time, which is expressed as in_read_ddr += in_block_size; ibuf_write_ptr += dma_in_size; The end of the loop body; Execute the dma instruction, which is expressed as dma_run_IF(in_des_ptr); Assign the pointer to another temporary pointer, which is expressed as float* in_ptr1 = (float*)ibuf_ptr_ping; float* out_ptr1 = (float*)obuf_ptr_ping; Use simd instructions to process format conversion; Use the output dma wait instruction to determine whether the data transfer is complete. If not, wait until it is complete, denoted as out_wait; The number of data transferred at one time, denoted as int dma_out_size = cur_line_num * width * out_once_size; uint32_t* out_ddr_ptr = out_batch_ddr; uint32_t* obuf_write_ptr = ibuf_ptr_pong; The two pointers for dma setting conversion and the quantity for each conversion, denoted as dma_node(obuf_write_ptr, out_ddr_ptr, dma_out_size); The pointer for jumping and saving to ddr, denoted as out_batch_ddr += out_line_size * cur_line_num; Swap the ibuf_ping and ibuf_pong pointers, denoted as swap(ibuf_ping, ibuf_pong); Swap the obuf_ptr_ping and obuf_pong pointers, denoted as swap(obuf_ptr_ping, obuf_pong); The ih_loop loop body ends; Use the output dma wait instruction to determine whether the data transfer is complete. If not, wait until it is complete, denoted as out_wait.

4. A design method based on the conversion of floating-point feature map format according to claim 1, characterized in that, The applicable range of the method is when oram can accommodate the minimum number of processing lines.

5. A design method based on the conversion of floating-point feature map format according to claim 1, characterized in that The method is applicable to the case of converting the DHWC32 format to the HWC format.