Design method based on Oram by pass non-redundancy calculation time in blocks

By optimizing the data loading and saving methods and using block convolution calculations, the redundant calculation problem when the convolution kernel width is greater than 1 is solved, efficient multi-layer convolution calculation is realized, and the calculation speed and efficiency are improved.

CN120339036APending Publication Date: 2025-07-18INGENIC SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410066608.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when the width of the convolution kernel is greater than 1, as the model level increases, the repeated redundant data increases sharply, resulting in excessive time consumption, and the ddr bandwidth bottleneck cannot be effectively solved. The existing oram by pass and chunking processing methods are inefficient in multi-layer computing.

Method used

By optimizing the loading and saving methods of data in the operator, using block convolutional calculations, using parallel processing of ddr and oram to avoid redundant calculations, specifically including: adding real data at each block height, the intermediate block both loading and saving duplicate part of the data into ddr, the last block only loads the data but not saves it, loading and saving is performed in parallel during the calculation process, and the ddr bandwidth is reasonably utilized.

Benefits of technology

It realizes the avoidance of redundant calculations in multi-layer convolutional calculations, improves the calculation speed, reduces the dependence on the ddr bandwidth, improves the calculation efficiency, and increases the speed by more than 40%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339036A_ABST
    Figure CN120339036A_ABST
Patent Text Reader

Abstract

The invention provides a design method based on Oram by pass non-redundancy calculation time in a block, the method is suitable for block convolution calculation, the method comprises the following two parts: realization of data from ddr to Oram in an operator and storage of the data in the ddr, and the method comprises the following steps: loading and using the data in the operator; according to the data storage in the ddr and the overall design method, an oam by pass and block processing method is fully exerted, no redundant calculation is generated, and the influence of the ddr bandwidth on calculation is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network models, and particularly relates to a design method based on the non-redundant calculation time of Oram bypass in blocks. Background Art

[0002] With the development of computer technology, especially the increasing popularity of AI applications, smart homes, smart cameras, face recognition, etc. have emerged in the consumer field. The neural network models involved in AI technology are also increasingly widely used. In the prior art, in the inference calculation of each layer of a model, the Oram is relatively small and cannot store the input, output, weights and other data of this layer. Therefore, the input and output data are stored in the DDR. During the calculation process, the DDR data is loaded into the Oram, and then the data is loaded from the Oram for relevant calculations. When the amount of calculation is relatively small, loading data from the DDR to the Oram becomes a bottleneck. Therefore, Oram bypass came into being. The input and output of each layer are in the Oram, without loading data from the DDR to the Oram, thus solving this situation. However, at the same time, another problem has emerged. The Oram is relatively small, and the input and output of each layer of many models are relatively large, resulting in it being unusable. At this time, the method of block processing emerged, dividing into very small blocks, so that the Oram can hold the input, output and weight data.

[0003] In addition, loading data from the DDR to the Oram, or saving Oram data to the DDR, and calculating using the data in the Oram can be processed in parallel.

[0004] However, the defects still existing in the prior art are as follows:

[0005] 1. In the existing Oram bypass and block processing, when the convolution kernel width is greater than 1, as the number of model layers increases, the repetitive redundant data increases sharply, resulting in more time than the situation without using Oram bypass and block processing.

[0006] 2. Only processing two layers cannot solve the DDR bandwidth bottleneck, and the received time effect is very small, without giving full play to the actual effect of Oram bypass.

[0007] In addition, the commonly used technical terms in the prior art include:

[0008] 1. Convolution kernel: The convolution kernel is a matrix used for image processing and is a parameter for performing operations with the original image. The convolution kernel is usually composed of a column matrix (for example, a 3*3 matrix), and each square in this area has a weight value. The matrix shape is generally 1×1, 3×3, 5×5, 7×7, 1×3, 3×1, 2×2, 1×5, 5×1,...

[0009] 2. Convolution: Place the center of the convolution kernel on the pixel to be calculated. Calculate the product of each element in the kernel and the pixel value of the image it covers one by one and sum them up. The resulting value is the new pixel value at this position, and this process is called convolution.

[0010] 3. Feature map: The result obtained after the input data is calculated by convolution is called a feature map. The result generated after the data passes through a fully connected layer is also called a feature map. The size of the feature map is generally expressed as length × width × depth, or 1×1 depth.

[0011] 4. DDR: A storage medium with a bandwidth speed more than ten times that of a common solid-state drive in the chip, and its size is generally 128MB or 256MB.

[0012] 5. ORAM by pass is a processing technology, which is a way of storing data where the input and output feature map data of each layer are in ORAM, or a way of storing data where the input or output feature map data of one layer is in ORAM. ORAM is a storage medium with a bandwidth speed dozens of times or even higher than that of ordinary DDR storage, and the medium material is the same as that of the cache. However, the space is relatively small, generally 512KB or 1MB.

[0013] 6. Block processing method. Divide the input of one layer into several blocks, for example, divide it into four or five blocks in terms of height. Calculate based on each block, and such a processing method is called a block processing method. Summary of the Invention

[0014] To solve the above problems, the purpose of the present application is as follows: This method makes full use of ORAM by pass and the block processing method through the loading and use of data in the operator, the preservation of data in DDR, and the overall design method, without generating any redundant calculations, and solves the impact of DDR bandwidth on calculations. The technical field to which the model in this method is specifically applied is in intelligent video surveillance, image detection and recognition, and image processing.

[0015] Specifically, the present invention provides a design method based on the non-redundant calculation time of Oram by pass in blocks, and the method is applicable to block convolution calculation, including: the implementation of using data in the operator from DDR to ORAM and the preservation of data in DDR, where,

[0016] (3) The implementation of using data in the operator from DDR to ORAM further includes:

[0017] 1), For each block in the chunk, only real data is added to the upper boundary of the height. If it is the first block, the upper boundary does not need to be added; the first block only saves the data of the repeated part to the ddr, and there is no process of loading data from the ddr to the oram; the last block only loads the data in the ddr and there is no process of saving to the ddr; each intermediate block needs to both load data from the ddr and save the data of the repeated part to the ddr; the repeated part is the part that is repeatedly calculated due to chunking in the prior art, and finally two lines are generated.

[0018] For the situation of each intermediate block, it further includes:

[0019] Assume the height of the oram for archiving input data is height, that is, 66, and 66 is the highest one among the heights of each block of the third-layer input.

[0020] Starting from the middle 33 of the oram height and calculating downward, while calculating, load the feature map from the ddr to the oram, that is, the process of (1) 2). The calculation process can execute the data loading in parallel. Only the first two lines of data required by this block of oram are saved in the ddr. The reason for loading the first two lines: Due to the scalability of 3x3, the height in each layer of calculation is reduced by one height, and since the upper layer gives one less height to the current layer, so there are a total of two heights. To make up for this loss, two heights of data, that is, two lines of data, are saved. The first two lines of this block of input data oram space are blank data, which are reserved for the space of loading data from the ddr to the oram and are also the part of repeated calculation.

[0021] Since there is enough time for the ddr to load data to the oram during the calculation time from the middle 33 lines to 66 lines, the loading of ddr data does not take time; save the last two lines of data generated by the calculation and saved in the oram to the ddr for use by the next block.

[0022] 2), During the execution of the calculation, load the data in the oram to the ddr and calculate downward from the starting height of the oram.

[0023] From the start of execution until the dividing line in the middle rows. Since the process of loading data from DDR to ORAM is parallel to the calculation, but the calculation requires data, directly needing the data to be loaded from DDR to ORAM will make it impossible to achieve parallelism, that is, waiting for the data to be loaded into ORAM before calculation can be performed. In ORAM, except for the first two rows which are the data required for calculation to be loaded, the data in other rows are all data required for calculation. Similarly, when saving ORAM data to DDR, if the calculation ends after generating the last two rows, these two rows of data have not been moved to DDR, resulting in the saving of ORAM data to DDR not being parallel to the calculation, and time is not saved. To enable parallelism between calculation and data loading, we adopt the most ideal starting position for calculation, starting from the middle, that is, after calculating to row 66, then starting from the first row to calculate. (4) The saving of data in DDR further includes:

[0024] 1) From the perspective of dividing a layer of feature map feature into blocks, the implementation process of saving data further includes: saving the data at the connection of each divided block in each layer. Here, only the division in the height H direction is considered. If the width is also divided into blocks, the same principle applies, and the division in the width direction is not processed here;

[0025] Save the part that will be recalculated in the next block generated at the bottom of each divided block in each layer, that is, save the two rows of data at the bottom of the divided block to DDR, and store the other part of the data in ORAM;

[0026] Here, it is not stored simultaneously. When executing the first block of consecutive layers, save the data that is recalculated between the first block and the second block to DDR; when executing the second block of consecutive layers, save the data that is recalculated between the second block and the third block to DDR; and so on; when continuously executing to the last block, the tail data in the height direction has no repetition and does not need to be saved;

[0027] 2), From the perspective of saving between layers, that is, the layer is divided into blocks inside the layer. Assuming it is divided into 4 blocks, inside the layer, it is the first block, the second block, the third block, and the fourth block; between layers, it is the i-th block of the first layer, the i-th block of the second layer, the i-th block of the third layer, …… The implementation process of saving data further includes:

[0028] When executing the first block from the first layer to the M-th layer, save the data that is recalculated between the first block and the second block in each layer. That is, save the two rows of repeated data of the first block in the first layer to DDR, and the stored data is ddr1-1; save the two rows of repeated data of the first block in the second layer to DDR, and the stored data is ddr1-2; save the two rows of repeated data of the first block in the third layer to DDR, and the stored data is ddr1-3; …… and so on; save the two rows of repeated data of the first block in the M-th layer to DDR, and the stored data is ddr1-M;

[0029] When executing the second block, calculate the second block of the first layer. Use the data ddr1-1 saved during the calculation of the first block of the first layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr2-1; calculate the second block of the second layer. Use the data ddr1-2 saved during the calculation of the first block of the second layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr2-2; calculate the second block of the third layer. Use the data ddr1-3 saved during the calculation of the first block of the third layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr2-3; and so on; calculate the second block of the Mth layer. Save the two rows of duplicate data to ddr, and the stored data is ddr2-M;

[0030] When executing the third block, calculate the third block of the first layer. Use the data ddr2-1 saved during the calculation of the second block of the first layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-1; calculate the third block of the second layer. Use the data ddr2-2 saved during the calculation of the second block of the second layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-2; calculate the third block of the third layer. Use the data ddr2-3 saved during the calculation of the second block of the third layer and re-save it into the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-3; and so on; save the two rows of duplicate data of the third block of the Mth layer to ddr, and the stored data is ddr1-M;

[0031] ……And so on. When executing the last block, denoted as the Nth block, calculate the Nth block of the first layer. Using the data ddr1-1 saved during the calculation of the (N-1)th block of the first layer, save it again to the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator; calculate the Nth block of the second layer. Using the data ddr1-2 saved during the calculation of the (N-1)th block of the second layer, save it again to the first two rows of the oram; calculate the Nth block of the third layer. Using the data ddr1-3 saved during the calculation of the (N-1)th block of the third layer, save it again to the first two rows of the oram; ……And so on; for the (N-1)th block of the Mth layer, using the data ddr1-M saved during the calculation of the (N-2)th block of the Mth layer, save it again to the first two rows of the oram; the Nth block is the tail of the last block height of each layer and is filled data, not real data, without repetition, so there is no need to save data to the ddr.

[0032] The block convolution calculation mentioned above refers to the block convolution calculation with an arbitrary convolution kernel height and an arbitrary stride.

[0033] The input and output data of each layer of the method described above are called the input feature map and the output feature map; assume that the convolution kernels of three consecutive layers are 3x3 and the strides are all 1. The implementation process is as follows: Use the oram by pass and block processing methods, and only consider block division in the height direction; let the height H be 256, and the pad type be same, that is, the input and output widths and heights are the same, divided into four blocks, use the oram bypass, and also use the ddr to store data.

[0034] In the above (1):

[0035] In the above 1), from line 33 to line 66, save the result saved to the oram to the ddr for use by the next block. Among them, the input feature occupies one block of the oram, and the output feature occupies another block of the oram. Here, only the oram used by the input feature is involved.

[0036] In the above 2), during the execution of this part of the calculation, it is also possible to prepare for loading weights to the oram for the next layer. If the oram is sufficient and space is allocated, it is also an optimization for deeper layers.

[0037] In the above (2):

[0038] In the implementation process of saving data from the perspective of block division of the feature map of one layer in the above 1):

[0039] In the case of dividing into 2 blocks: Save the 2 rows of data at the bottom of the first block to the ddr; that is, when calculating the first block, save this repeated part to the ddr;

[0040] Case of being divided into 4 blocks: For the parts generated by the first block, the second block, and the third block that will be recalculated in the next block, save the results of these parts in the ddr, and store the data of other parts in the oram; here, it is not stored simultaneously. Execute the first block of consecutive layers and save the first two rows of duplicate data in the ddr; execute the second block of consecutive layers and save the middle two rows of duplicate data in the ddr; execute the third block of consecutive layers and save the last two rows of duplicate data in the ddr; when continuously executing the fourth block, there are no duplicates in the tail data in the height direction and no need to save.

[0041] In the above (2):

[0042] In the implementation process of saving data from the perspective of inter-layer preservation in the above 2):

[0043] The blocks of the first layer and the blocks of the second layer are not parallel. This is because in the generation of each layer of the first block, it gradually decreases;

[0044] For the first block, the height situation of the input and output feature maps in each layer: In the first layer, the input height of the first block is height, and the output height is height - 1; in the second layer, the input height of the first block is height - 1, and the output height is height - 2; in the third layer, the input height of the first block is height - 2, and the output height is height - 3; for the second block, the height situation of the input and output feature maps in each layer: In the first layer, the input height of the second block is height + 2, and the output height is height; in the second layer, the input height of the second block is height + 2, and the output height is height; in the third layer, the input height of the second block is height + 2, and the output height is height; for the third block, the height situation of the input and output feature maps in each layer: In the first layer, the input height of the third block is height + 2, and the output height is height; in the second layer, the input height of the third block is height + 2, and the output height is height; in the third layer, the input height of the third block is height + 2, and the output height is height; for the fourth block, the height situation of the input and output feature maps in each layer: In the first layer, the input height of the fourth block is height + 1, and the output height is height; in the second layer, the input height of the fourth block is height, and the output height is height - 1; in the third layer, the input height of the fourth block is height - 1, and the output height is height - 2; actually, the first block generates with the loss of the last row;

[0045] Since the first block reduces one height in each layer, two heights are lost in the generation of the third layer; to recall this loss, add one to the height of each layer in the fourth block, and after correction, it is:

[0046] Fourth, the height of the input and output feature maps in each layer: The input height of the fourth block in the first layer is height + 1, and the output height is height; the input height of the fourth block in the second layer is height + 2. The increase in height is due to misalignment. When the height remains unchanged and misalignment occurs, in order to make the final output parallel, a value needs to be added to the input, that is, the input height is height + 2, and the output height is height + 1; the input of the fourth block in the third layer is two heights away from being parallel, that is, height + 3, and the output height is height + 2; thus eliminating the missing number of rows.

[0047] In the last block, the height needs to be calculated separately.

[0048] Therefore, the advantages of this application are as follows: avoiding redundant calculations, improving work efficiency, achieving unrestricted by the ddr bandwidth speed, enabling the use of oram by pass and block methods between dozens of layers or even more layers, thereby improving speed. It can achieve a speed increase of more than 40%. Brief Description of the Drawings

[0049] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.

[0050] Figure 1 It is a schematic diagram of calculating downward from the middle of the oram height in this method.

[0051] Figure 2 It is a schematic diagram of calculating downward from the starting height of the oram in this method.

[0052] Figure 3 It is a schematic diagram of the situation where each layer is divided into two blocks and the two bottom rows of data of each block are saved to the ddr in this method.

[0053] Figure 4 It is a schematic diagram of the situation where the two - block structure is upgraded to a four - block structure in this method.

[0054] Figure 5 It is a schematic diagram of the implementation process of the inter - layer saving situation in this method. Detailed Embodiment

[0055] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.

[0056] First, let's understand the analysis of the existing block method:

[0057] The input and output data of each layer are called the input feature map and the output feature map respectively. Now, taking three consecutive convolutional kernels with a size of 3x3 and a stride of 1 as an example, the implementation process is as follows. Using oram by pass and the block processing method, only consider blocking in the height direction. Let the height H be 256, and the pad type be same, that is, the input and output widths and heights are the same. Divided into four blocks, oram bypass can be used.

[0058] The output height of the first block of the third layer is 64. For the first block of the third layer, the upper boundary is filled and does not require a height value, while the lower boundary is the real value. Therefore, an additional height value is needed for the lower boundary. The input height of the first block of the third layer is the output height of the first block of the second layer. So, the output of the first block of the second layer needs to be a height of 64 + 1, that is, 65; similarly, the input height of the first block of the second layer needs to be 65 + 1, that is, 66. The output height of the first block of the first layer is the input height of the first block of the second layer, which is 66; the input height of the first block of the first layer is 66 + 1, that is, 67.

[0059] The height of the second block of the third layer becomes 64. For the second block of the third layer, both the upper and lower boundaries use real values. Therefore, one height value is needed for each of the upper and lower boundaries. The input height of the second block of the third layer is the output height of the second block of the second layer. So, the output of the second block of the second layer needs to be a height of 64 + 1 + 1, that is, 66; for the input height of the second block of the second layer, one height value is needed for each of the upper and lower boundaries during calculation. So, the input height of the second block of the second layer is 66 + 1 + 1, that is, 68. The output height of the second block of the first layer is the input height of the second block of the second layer, which is 68. When calculating the input height of the second block of the first layer, real values for one height value each of the upper and lower boundaries are also needed. Therefore, the input height of the second block of the first layer is 68 + 1 + 1, that is, 70.

[0060] The height of the third block of the third layer becomes 64. For the third block of the third layer, both the upper and lower boundaries use real values. Therefore, one height value is needed for each of the upper and lower boundaries. The input height of the third block of the third layer is the output height of the third block of the second layer. So, the output of the third block of the second layer needs to be a height of 64 + 1 + 1, that is, 66; for the input height of the third block of the second layer, one height value is needed for each of the upper and lower boundaries during calculation. So, the input height of the third block of the second layer is 66 + 1 + 1, that is, 68. The output height of the third block of the first layer is the input height of the third block of the second layer, which is 68. When calculating the input height of the third block of the first layer, real values for one height value each of the upper and lower boundaries are also needed. Therefore, the input height of the third block of the first layer is 68 + 1 + 1, that is, 70.

[0061] The height of the fourth block in the third layer becomes 64. The upper boundary of the fourth block in the third layer uses the real value, and the lower boundary is filled with data. Therefore, a height value is required for the upper boundary. The input height of the fourth block in the third layer is the output height of the fourth block in the second layer. So, the output of the fourth block in the second layer needs to be 64 + 1, that is, 65. For the input height of the fourth block in the second layer, a height value is required for the upper boundary during calculation, and the lower boundary is filled with data, so no height value is needed. Therefore, the input height of the fourth block in the second layer is 65 + 1, that is, 66. The output height of the fourth block in the first layer is 66, the input height of the fourth block in the first layer. When calculating, a real value of the height is required for the upper boundary, and the lower boundary is filled with data. So, the input height of the fourth block in the first layer is 66 + 1, that is, 67.

[0062] As can be seen from the above, up to the first layer, the calculated height is 67 + 70 + 70 + 67, that is, 274, but the actual value is only 256. Some intermediate parts will have repeated calculations. When the convolution kernel is relatively large or the stride is relatively large, this repeated phenomenon is more obvious, and the calculation time increases sharply, resulting in the inability to use this method.

[0063] To avoid repeated calculations, when using the oram by pass method, no ddr is used at all, completely wasting the ddr bandwidth. And the way to improve speed is to be able to meet the actual application, not to completely abandon it. From this perspective, partial ddr can be used. As long as the ddr data loading time is less than the calculation time, it will not cause any impact and can be used.

[0064] Therefore, this method proposes a design method based on the non-redundant calculation time of Oram by pass in the block, that is, the non-redundant block design method in oram by pass.

[0065] Continuing with the example of three consecutive convolutional layers with a kernel size of 3x3 and a stride of 1 for each layer, the process is implemented. Using the oram by pass and block processing methods, only consider dividing the blocks in the height direction. Assume the height is 256, and the pad type is same, that is, the input and output widths and heights are the same. Divide it into four blocks, use oram by pass, and also use ddr to store data.

[0066] The implementation idea is that when both the input and output are saved in oram, only a very small part of the weights use the bandwidth to the weight register. Generally, when the channel depth is relatively small and the feature map feature is relatively large, blocks are used, so the bandwidth is almost idle. Therefore, make full use of this part of the bandwidth to transfer the data in ddr to oram. That is, save the data in the intermediate repeated calculation part to ddr, and when using it, load the data in ddr to oram, so that the time bandwidth is hidden in the calculation.

[0067] (1) Implementation of using DDR to ORAM in the operator

[0068] 1), for each block, only the upper boundary of the height is added with real data. If it is the first block, the upper boundary does not need to be added;

[0069] The first block only saves the data of the repeated part (the part that is repeatedly calculated due to block division in the prior art, and finally two rows are generated) to the DDR, and there is no process of loading data from the DDR to the ORAM; the last block only loads the data in the DDR, and there is no process of saving to the DDR; each intermediate block needs to both load data from the DDR and save the data of the repeated part to the DDR;

[0070] The situation of each intermediate block:

[0071] Assume the height of the ORAM for archiving the input data is height, that is, 66. 66 is the highest one among the heights of each block in the third-layer input; the reason for choosing 66 is that since the final input is of equal height, for the input, only the highest one needs to be selected. The third-layer output input fully conforms to this situation, so the largest one is selected from 65 and 66, that is, 66. This is because the inputs of each block in the third layer are: the input height of the first block in the third layer is 65, the input height of the second block in the third layer is 66, the input height of the third block in the third layer is 66, and the input height of the fourth block in the third layer is 65. So the largest 66 is selected.

[0072] As Figure 1 shown, starting from the middle 33 of the ORAM height and calculating downward, while calculating, load the feature map from the DDR to the ORAM, that is, the same as the process in 2). The calculation process can execute the data loading in parallel. Only the first two rows of data required by this block of ORAM are saved in the DDR; the reason for loading the first two rows: due to the scalability of 3x3, the height of each layer of calculation is reduced by one height. And since the upper layer gives one less height to the current layer, so a total of two heights are reduced. To make up for this loss, the data of two heights, that is, two rows of data, are saved. The first two rows of the ORAM space of this input data are blank data, which are reserved for the space of loading data from the DDR to the ORAM, and are also the part of repeated calculation;

[0073] Since there is enough time during the calculation time from the middle 33 rows to 66 rows for the DDR to load data to the ORAM, so loading the DDR data does not take time; save the last two rows of data generated by the calculation and saved in the ORAM to the DDR for use by the next block. Here it is explained that the input feature occupies one block of ORAM, and the output feature occupies another block of ORAM. Here only the ORAM used for the input feature is mentioned, and it is also simply referred to as ORAM.

[0074] (2), starting from the starting height of the oram, calculate downward, as Figure 2 shown by the dashed arrow. During the execution of the calculation, the data in the oram is loaded into the ddr.

[0075] From the start of execution until the demarcation line of the middle row. Since the process of loading data from the ddr to the oram is parallel to the calculation, but the calculation requires data, directly waiting for the data loaded from the ddr to the oram will make parallelism impossible, that is, waiting for the data to be loaded into the oram before calculation can be performed. Except for the first two rows in the oram, the data in other rows are all data required for calculation. Similarly, when saving the oram data to the ddr, if the calculation ends after generating two rows at the end, these two rows of data have not been moved to the ddr, resulting in the oram data being saved to the ddr not being parallel to the calculation, and time is not saved. In order to achieve parallelism between calculation and data loading, we adopt the most ideal starting calculation position method, starting from the middle to calculate, that is, after calculating to row 66, then starting from the first row to calculate. During the execution of this part of the calculation, it is also possible to prepare for loading weights to the oram for the next layer. Of course, when the oram is sufficient and space is allocated, it is also an optimization for deeper layers. Here is just a mention of the idea of further optimization without any processing.

[0076] (2) Saving of ddr data

[0077] The data at the connection of each block saved here only considers the block division in the height H direction. If the width is also divided into blocks, the same principle applies, and no processing for width block division is done here. Save the two bottom rows of data of each layer block to the ddr, that is, when calculating the first block, save the Figure 3 two dashed lines in it to the ddr, which is the case of being divided into two blocks.

[0078] Upgraded from two blocks to four blocks, as Figure 4 shown. The two solid lines and four dashed lines in the middle are the parts generated by the first, second, and third blocks that will be recalculated in the next block. Save the results of these parts to the ddr, and store the data of other parts in the oram. But here it is not stored simultaneously. When executing the first block of consecutive layers, save the Figure 4 first two dashed line data in it to the ddr; when executing the second block of consecutive layers, save the data of the two middle solid lines to the ddr; when executing the third block of consecutive layers, save the data of the last two dashed lines to the ddr; when continuously executing the fourth block, the tail data in the height direction has no repetition and does not need to be saved.

[0079] The above is to save data from the perspective of dividing a feature map into blocks. In actual use, it is achieved from inter-layer saving. The following explains the implementation process from the situation of inter-layer saving.

[0080] AsFigure 5 As shown, when executing the first, second, and third layers of the first block, save the two dashed lines at the first red line adjacent to it. That is, the two dashed line data of the first layer of the first block are saved to the ddr, and the stored data is ddr1-1. The two dashed line data of the second layer of the first block are saved to the ddr, and the stored data is ddr1-2. The two dashed line data of the third layer of the first block are saved to the ddr, and the stored data is ddr1-2. Due to the scalability of 3x3, the height of each layer's calculation is reduced by one height. Also, since the upper layer gives one less height to the current layer, there are a total of two heights. To make up for this loss, save the data of two heights. That is, two solid lines or dashed lines. The two-line data of the first layer of the first block are saved to the ddr.

[0081] When executing the second block, calculate the first layer of the second block. Use the data ddr1-1 saved during the calculation of the first layer of the first block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr2-1. Calculate the second layer of the second block. Use the data ddr1-2 saved during the calculation of the second layer of the first block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr2-2. Calculate the third layer of the second block. Use the data ddr1-3 saved during the calculation of the third layer of the first block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr2-3.

[0082] When executing the third block, calculate the first layer of the third block. Use the data ddr2-1 saved during the calculation of the first layer of the second block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr1-1. Calculate the second layer of the third block. Use the data ddr2-2 saved during the calculation of the second layer of the second block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr1-2. Calculate the third layer of the third block. Use the data ddr2-3 saved during the calculation of the second layer of the second block and save it back to the first two rows of the oram. The specific implementation is to use the ddr-to-oram implementation in the (1) operator. When executing to generate the last two rows, save the data to ddr1-3.

[0083] When executing the fourth block, calculate the fourth block of the first layer. Use the data ddr1-1 saved during the calculation of the third block of the first layer and save it again to the first two rows of the oram. The specific implementation is the implementation of using ddr to oram in the (1) operator. Calculate the fourth block of the second layer. Use the data ddr1-2 saved during the calculation of the third block of the second layer and save it again to the first two rows of the oram. Calculate the fourth block of the third layer. Use the data ddr1-3 saved during the calculation of the third block of the third layer and save it again to the first two rows of the oram. The fourth block is the tail of the last block height of each layer and is filled data, not real data, without repetition, so there is no need to save data to the ddr.

[0084] In Figure 5 , the blocks of the first layer and the blocks of the second layer are not parallel. This is because in the generation of each layer of the first block, it gradually decreases.

[0085] For the first block, the height of the input and output feature maps in each layer is as follows: the input height of the first block of the first layer is height, and the output height is height-1; the input height of the first block of the second layer is height-1, and the output height is height-2; the input height of the first block of the third layer is height-2, and the output height is height-3;

[0086] For the second block, the height of the input and output feature maps in each layer is as follows: the input height of the second block of the first layer is height+2, and the output height is height; the input height of the second block of the second layer is height+2, and the output height is height; the input height of the second block of the third layer is height+2, and the output height is height;

[0087] For the third block, the height of the input and output feature maps in each layer is as follows: the input height of the third block of the first layer is height+2, and the output height is height; the input height of the third block of the second layer is height+2, and the output height is height; the input height of the third block of the third layer is height+2, and the output height is height;

[0088] For the fourth block, the height of the input and output feature maps in each layer is as follows: the input height of the fourth block of the first layer is height+1, and the output height is height; the input height of the fourth block of the second layer is height, and the output height is height-1; the input height of the fourth block of the third layer is height-1, and the output height is height-2. The first block actually generates a loss of the last row.

[0089] Since the height of each layer of the first block decreases by one, two heights are lost during the generation of the third layer. To recall this loss, the height of each layer in the fourth block is increased by one. After correction: For the fourth block, the height of the input and output feature maps in each layer: The input height of the fourth block in the first layer is height + 1, and the output height is height; the input height of the fourth block in the second layer is height + 2. Here, the height increases due to misalignment. As Figure 5 shown, when the height remains unchanged, misalignment occurs. For the final output to be parallel, a value needs to be added to the input, that is, the input height is height + 2, and the output height is height + 1; the input of the fourth block in the third layer is two heights away from being parallel, i.e., height + 3, and the output height is height + 2; thus eliminating the lost number of rows.

[0090] In the last block, the height needs to be calculated separately.

[0091] This design method can achieve oram_by_pass, that is, both the input and output are in the oram, and only a very small amount of duplicate data in the middle is saved to the ddr. Since only a very small part uses the ddr, through the logic of using the ddr to the oram in a reasonable operator, the time for loading data from the ddr to the oram and saving oram data to the ddr can be completely hidden in the calculation. Thus, oram_by_pass is truly achieved.

[0092] From the above example, it can be inferred to the block convolution calculation with any convolution kernel height and any stride.

[0093] This method can be applied to model fields involving chips and key fields of artificial intelligence, such as image recognition, image processing, target detection, etc. This method runs on a computer. It is precisely due to the use of this method that the excessive and redundant requirements in the design of hardware (chip, memory requirements) and software (bandwidth requirements) are reduced.

[0094] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Design method based on the non-redundant calculation time of Oram by pass in block division, characterized in that The method is applicable to block convolution calculation, including: the implementation of data transfer from DDR to ORAM in the operator and the storage of data in DDR. Among them, (1) The implementation of data transfer from DDR to ORAM in the operator further includes: 1), For each block in the block, only real data is added to the upper boundary of the height. If it is the first block, no data needs to be added to the upper boundary; the first block only saves the repeated part of the data to DDR, and there is no process of loading data from DDR to ORAM; the last block only loads the data in DDR and there is no process of saving data to DDR; for each intermediate block, it is necessary to load data from DDR and also save the repeated part of the data to DDR. The situation of each intermediate block further includes: Assume that the convolution kernels of three consecutive layers are 3x3 with a stride of 1. Let the height of the ORAM for storing the input data be height, that is, 66, and 66 is the highest among the heights of each block of the third-layer input. Start calculating downward from the middle 33 of the ORAM height. During the calculation, load the feature map from DDR to ORAM, that is, the same process as (1) 2) below. The calculation process can be executed in parallel with the data loading. Only the first two rows of data required for this block of ORAM are saved in DDR. The first two rows of this block of input data ORAM space are blank data, which are reserved for the space of loading data from DDR to ORAM and are also the part of repeated calculation. Since there is enough time for DDR to load data to ORAM during the calculation time from the middle 33 rows to 66 rows, the loading of DDR data does not take time; save the last two rows of data generated by the calculation, which are saved in ORAM, to DDR for use by the next block. 2), During the execution of the calculation, load the data in ORAM to DDR and start calculating downward from the starting height of ORAM. From the start of execution until the dividing line of the middle row, since the process of loading data from DDR to ORAM is parallel to the calculation, but the calculation requires data. If directly using the data to be loaded from DDR to ORAM, parallelism cannot be achieved, that is, waiting for the data to be loaded to ORAM before calculation can be performed; in ORAM, except for the first two rows which are the data required for calculation to be loaded, the data of other rows are all data required for calculation; similarly, when saving the data in ORAM to DDR, if the calculation ends after generating the last two rows, these two rows of data have not been moved to DDR, resulting in the data saving from ORAM to DDR not being parallel to the calculation and no time being saved. In order to achieve parallelism between calculation and data loading, the most ideal starting position for calculation is adopted, that is, start calculating from the middle. After calculating to 66 rows, start calculating from the first row again. (2) The storage of data in DDR further includes: 1), From the perspective of block division of each layer of feature map feature, the implementation process of data storage further includes: Save the data at the connection of each block of each layer. Here, only the block division in the height H direction is considered. If the width width is also divided into blocks, the same principle applies and the block division in the width direction is not processed here. Save the part generated at the bottom of each layer block that will be recalculated in the next block, i.e., the two rows of data at the bottom of the block, to the DDR, and store the other part of the data in the ORAM; Here, it is not stored simultaneously. For the first block of consecutive layers, save the data that is recalculated between the first block and the second block to the DDR; for the second block of consecutive layers, save the data that is recalculated between the second block and the third block to the DDR; and so on. When continuously executing until the last block, the tail data in the height direction has no repetition and does not need to be saved; 2), From the perspective of inter-layer saving, that is, the layer is divided into blocks. Assuming it is divided into 4 blocks, within the layer, there are the first block, the second block, the third block, and the fourth block. Then, between layers, it is the i-th block of the first layer, the i-th block of the second layer, the i-th block of the third layer, …… The implementation process of saving data further includes: When executing the first block for the first layer, the second layer, the third layer, ……, the M-th layer, save the data that is recalculated between the first block and the second block in each layer. That is, save the two rows of repeated data of the first block in the first layer to the DDR, and the stored data is ddr1-1; save the two rows of repeated data of the first block in the second layer to the DDR, and the stored data is ddr1-2; save the two rows of repeated data of the first block in the third layer to the DDR, and the stored data is ddr1-3; …… and so on. Save the two rows of repeated data of the first block in the M-th layer to the DDR, and the stored data is ddr1-M; When executing the second block, calculate the second block of the first layer, and use the data ddr1-1 saved when calculating the first block of the first layer to save it back to the first two rows of the ORAM. The specific implementation is to use the implementation of DDR to ORAM in the (1) operator. When executing to generate the last two rows, save the data to ddr2-1; calculate the second block of the second layer, and use the data ddr1-2 saved when calculating the first block of the second layer to save it back to the first two rows of the ORAM. The specific implementation is to use the implementation of DDR to ORAM in the (1) operator. When executing to generate the last two rows, save the data to ddr2-2; calculate the second block of the third layer, and use the data ddr1-3 saved when calculating the first block of the third layer to save it back to the first two rows of the ORAM. The specific implementation is to use the implementation of DDR to ORAM in the (1) operator. When executing to generate the last two rows, save the data to ddr2-3; …… and so on. Calculate the second block of the M-th layer, and save the two rows of repeated data to the DDR, and the stored data is ddr2-M; When executing the third block, calculate the third block of the first layer. Use the data ddr2-1 saved during the calculation of the second block of the first layer and re-save it to the first two rows of the oram. The specific implementation is to use the implementation of ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-1; calculate the third block of the second layer. Use the data ddr2-2 saved during the calculation of the second block of the second layer and re-save it to the first two rows of the oram. The specific implementation is to use the implementation of ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-2; calculate the third block of the third layer. Use the data ddr2-3 saved during the calculation of the second block of the second layer and re-save it to the first two rows of the oram. The specific implementation is to use the implementation of ddr to oram in the (1) operator. When executing to generate the last two rows, save the data to ddr1-3; and so on; Save the duplicate data of the last two rows of the Mth layer of the third block to ddr, and the stored data is ddr1-M; and so on; When executing the last block, denoted as the Nth block, calculate the Nth block of the first layer. When N is even, use the data ddr1-1 saved during the calculation of the (N - 1)th block of the first layer and re-save it to the first two rows of the oram. The specific implementation is to use the implementation of ddr to oram in the (1) operator; calculate the Nth block of the second layer. Use the data ddr1-2 saved during the calculation of the (N - 1)th block of the second layer and re-save it to the first two rows of the oram; calculate the Nth block of the third layer. Use the data ddr1-3 saved during the calculation of the (N - 1)th block of the third layer and re-save it to the first two rows of the oram; and so on; For the (N - 1)th block of the Mth layer, use the data ddr1-M saved during the calculation of the (N - 2)th block of the Mth layer and re-save it to the first two rows of the oram; The Nth block is the tail of the last block height of each layer, and the data is padding data, not real data, without duplicates, so there is no need to save data to ddr.

2. The design method based on the Oram by pass non-redundant calculation time in the block according to claim 1, characterized in that The block convolution calculation refers to the block convolution calculation with an arbitrary convolution kernel height and an arbitrary stride.

3. The design method based on the Oram by pass non-redundant calculation time in the block according to claim 2, wherein The input and output data of each layer of the method are called the input feature map and the output feature map; Assume that the convolution kernels of three consecutive layers are 3x3 and the strides are all 1. To implement its process: Use the oram by pass and block processing methods, and only consider block division in terms of height; Let the height H be 256 and the pad type be same, that is, the input and output widths and heights are the same. Divide it into four blocks and use oram bypass, and also use ddr to store data.

4. The design method based on the Oram by pass non-redundant calculation time in the block according to claim 3, characterized in that In the (1): During the calculation time from line 33 to line 66 in the 1), save the result saved to the oram to the ddr for use by the next block. Among them, the input feature occupies one block of the oram, and the output feature occupies another block of the oram. Here, only the oram used by the input feature is involved; In the 2), during the execution of this part of the calculation, it is also possible to prepare to load weights to the oram for the next layer. The oram is sufficient and space is allocated, which is also an optimization for deeper layers.

5. The design method based on the non-redundant calculation time of Oram by pass in the block according to claim 3, characterized in that, In (2): In the implementation process of saving data from the perspective of dividing a layer of feature map feature into blocks: In the case of dividing into 4 blocks: For the parts that will be recalculated in the next block generated by the first block, the second block, and the third block, save the results of these parts to the ddr, and store the data of other parts in the oram; here, they are not stored simultaneously. When executing the first block of consecutive layers, save the first two rows of repeated data to the ddr; when executing the second block of consecutive layers, save the middle two rows of repeated data to the ddr; when executing the third block of consecutive layers, save the last two rows of repeated data to the ddr; When continuously executing the fourth block, the tail data in the height direction has no repetition and does not need to be saved.

6. The design method based on the Oram by pass non-redundant calculation time in the block according to claim 3, characterized in that, In (2): In (2), from the perspective of inter-layer saving, that is, the layer is divided into blocks. Assuming it is divided into 4 blocks, within the layer, they are the first block, the second block, the third block, and the fourth block; between layers, they are the i-th block of the first layer, the i-th block of the second layer, the i-th block of the third layer, ……, in the implementation process of saving data: The blocks of each layer in the first layer and the second layer are not parallel because in the generation of each layer of the first block, it gradually decreases; For the first block, the height situations of the input and output feature maps in each layer: The input height of the first block in the first layer is height, and the output height is height - 1; the input height of the first block in the second layer is height - 1, and the output height is height - 2; the input height of the first block in the third layer is height - 2, and the output height is height - 3; For the second block, the height situations of the input and output feature maps in each layer: The input height of the second block in the first layer is height + 2, and the output height is height; the input height of the second block in the second layer is height + 2, and the output height is height; the input height of the second block in the third layer is height + 2, and the output height is height; For the third block, the height situations of the input and output feature maps in each layer: The input height of the third block in the first layer is height + 2, and the output height is height; the input height of the third block in the second layer is height + 2, and the output height is height; the input height of the third block in the third layer is height + 2, and the output height is height; For the fourth block, the height situations of the input and output feature maps in each layer: The input height of the fourth block in the first layer is height + 1, and the output height is height; the input height of the fourth block in the second layer is height, and the output height is height - 1; the input height of the fourth block in the third layer is height - 1, and the output height is height - 2; actually, the first block generates a loss of the last row; Since the first block reduces one height in each layer, resulting in a loss of two heights in the generation of the third layer; to recall this loss, add one to the height of each layer in the fourth block, and after correction, it is: Fourth block, height of input and output feature maps in each layer: The input height of the fourth block in the first layer is height + 1, and the output height is height; the input height of the fourth block in the second layer is height + 2. The increase in height here is due to misalignment. When the height remains unchanged, misalignment occurs. To make the final output parallel, a value needs to be added to the input, that is, the input height is height + 2, and the output height is height + 1; the input of the fourth block in the third layer is two heights away from being parallel, i.e., height + 3, and the output height is height + 2; thus eliminating the missing number of rows. In the last block, the height needs to be calculated separately.