System and method to process image data by a convolutional neural network
By splitting a CNN into sub-networks based on optimized subset configurations and data overlaps, the method addresses the high memory and computational demands of CNNs, reducing memory needs while maintaining performance and automating the optimization process.
Patent Information
- Application Number
- PCT/EP2024/082407
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-05
AI Technical Summary
Convolutional neural networks (CNNs) require significant memory and computational resources due to their high accuracy, leading to increased processing demands and memory consumption, which can impact performance and cost.
A method to analyze and split a CNN into sub-networks by determining subset configurations of input feature maps for each convolutional layer, calculating data overlaps, and optimizing memory and computational demands, thereby reducing the need for dedicated memory with a moderate impact on computational cost.
The proposed method effectively lowers the memory requirements for processing CNNs on dedicated hardware accelerators while maintaining performance, fully automating the process of finding the best compromise between memory saving and additional computational cost.
Smart Images

Figure EP2024082407_05062025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD TO PROCESS IMAGE DATA BY A CONVOLUTIONALNEURAL NETWORK
[0001] Neural networks represent computational architectures modeled after biological brains and are used to process complex input data such as ID data and 2D / 3D image data . In such computational architectures neurons or nodes may be interconnected and operate collectively to process the data . There are various examples of these networks such as convolutional neural networks , recurrent neural networks , deep belief networks , and restricted Boltzmann machines . All of these are applied in many applications such as computer vision, speech recognition, signal processing, and bioinformatics .
[0002] Speci fically, convolutional neural networks ( CNNs ) process data in multiple layers to model high-level abstraction of the data . Each layer may receive input data such as an input feature map ( e . g . , an 2D-image ) and generates output data such as an output feature map by processing the input through the layer . In greater detail , an output feature map may be generated by convolving an input feature map with a convolutional kernel . This kernel may be a convolution matrix, or a mask, which is a small matrix used for blurring, sharpening, embossing, edge detection, and more . The processing through a layer may then be accomplished by performing a convolution between the kernel and the input feature map .
[0003] In this context , initial layers of a CNN, which are also known as convolution layers , may be operative to extract low level features such as edges and / or gradients from input data . These initial layers may also be called feature extraction layers . Subsequent layers of the CNN, referred to as feature classi fication layers , may extract or detect progressively more complex features such as eyes , a nose , or the like . The featureclassification layers may also be referred to as "fully-connected layers."
[0004] For example, if a plurality of two-dimensional pictures of faces is provided as input data to a CNN, the CNN may learn a variety of characteristics of faces such as edges, curves, angles, dots, color contrasts, bright spots, dark spots, etc. These one or more features are learned at one or more convolutional layers of the CNN. Then, in one or more feature classification layers, the CNN may learn a variety of recognizable features of faces such as eyes, eyebrows, foreheads, hair, noses, mouths, cheeks, etc.; each of which is distinguishable from all of the other features. That is, the CNN may learn to recognize and distinguish an eye from an eyebrow or any other facial feature. In one or more third and then subsequent feature extraction layers, the CNN may learn entire faces and higher order characteristics such as race, gender, age, emotional state, etc. The CNN may even be taught in some cases to recognize the specific identity of a person.
[0005] Given the high accuracy of CNNs, they are instrumental for computer vision tasks, natural language processing, information retrieval, speech recognition fraud detection and more. However, the accuracy improvements of CNNs entail massive increases in processing requirements. As the accuracy of recognition increases, the number of computations required to evaluate a CNN also grows. In addition, extensive input data generate correspondingly large feature maps that consume extensive memory .
[0006] Indeed, for such accurate processing, the whole input data may be needed by the CNN. For inference, when processing images, such input data may be quite memory consuming and may have a significant impact on the memory need. Processing thewhole input data at the same time implies also that the working memory used in the di f ferent layers of the CNN is quite high . This high memory usage has a direct impact on the memory si zing and consequently, on its cost .
[0007] An obj ect of the present invention is therefore to provide an improved memory management for a CNN . According to embodiments , the above obj ect is achieved by the claimed matter according to the independent claims . Further developments are defined in the dependent claims .SUMMARY
[0008] According to embodiments a computer-implemented method to analyze a convolutional neural network and split the convolutional neural network ( CNN) comprising a plurality of convolutional layers is provided . The method comprising : determining, for each convolutional layer, di f ferent subset configurations of an input feature map to be processed by the respective convolutional layer, wherein the input feature map is divided into subsets of input feature data, each of the same si ze . For each determined subset configuration, the method further comprises : determining, for each convolutional layer, a data overlap to be added to each one of the subsets of input feature data, wherein the data overlap is determined from the input feature map, determining a memory- and computational demand required for processing the subsets of input feature data by each convolutional layer, selecting a subset configuration of the determined subset configurations depending on the determined memory- and computational demand, and splitting the convolutional network into di fferent sub-networks depending on the selected subset configuration to obtain a split convolutional network .
[0009] Splitting the CNN enables to lower the need of dedicated memory with a moderate impact on computational cost. In greater detail, the above method lowers the need of dedicated memory to process CNNs on a dedicated Al hardware accelerator and may be used for many existing CNNs. Furthermore, it fully automatizes the process of discovering the best compromise between memory saving and additional computational cost. Furthermore, it automatically modifies the CNN according to the best compromise found by processing the subset configurations. In addition, it automatizes the process of scheduling an optimized network on the hardware .
[0010] For each convolutional layer, the input feature map may comprise image data that is divided into subsets of input feature data based on pre-defined rules. The subsets of input feature data may be represented by stripes.
[0011] For each determined subset configuration, the convolutional layers that have the same number of subsets of input feature data may be grouped into one group.
[0012] This may allow splitting the CNN into pieces, each piece being a group of layers. This may further simplify the execution of a graph of the resulting CNN, and also optimizes the needed data overlaps.
[0013] Each convolutional layer may comprise a specific kernel- , padding-, dilation-, and stride configuration.
[0014] Furthermore, the operation determining, for each convolutional layer, the data overlap to be added to each one of the subsets of input feature data may further comprise determining, for each convolutional layer, a minimal data overlap to be added to each one of the subsets of input feature data such that theresult of the convolution operations by the convolutional layer matches the result of the convolution operations when proces sing the whole input feature map . The amount of minimal data overlap for a given convolutional layer may depend on the amount of data overlap of the subsequent convolutional layers and a kernel- , padding- , dilation- , and stride configuration of the given convolutional layer .
[0015] In addition, this step may comprise determining, for groups of convolutional layers having the same number of subsets of input feature data, a further data overlap to be added to the respective subsets of input feature data . The further data overlap may take into account redundant data due to padding .
[0016] Furthermore , this step may also comprise adding, for each convolutional layer, a fixed amount of data overlap to each one of the subsets of input feature data to account for di f ferent striding parameters .
[0017] Moreover, it may further comprise determining a data overlap to be added to the subsets of input feature data for each convolutional layer of a given group that accounts for suf ficient data overlap for a subsequent group .
[0018] These additional data overlaps may result in an accurate processing of the subsets of input feature data .
[0019] The operation determining a memory- and computational demand required for processing the subsets of input feature data by each convolutional layer may further comprise determining the memory demand in an On-Chip memory . In detail , the memory demand may depend on the si zes of the subsets of input feature data and their convolved outputs , their corresponding data overlaps , and a si ze o f the weights and biases for each one of theconvolutional layers .
[0020] The operation may further comprise determining a number of multiple accumulate operations (MACs ) that results from the convolution operations of one subset of input feature data and the determination of the respective data overlaps 18 multiplied by the number of subsets of input feature data for a given convolutional layer, and determining a number of clock cycles required for memory trans fers between a system memory and an On- chip memory .
[0021] The memory transfers may comprise : the trans fer of the weights for each convolutional layer from the system memory to the On-chip memory, the trans fer of the subsets of input feature data and the data overlaps from the system memory to the On-chip memory for each convolutional layer that starts a given group of convolutional layers , and the trans fer-back of the convolved outputs of the subsets of input feature data and the data overlaps of the last convolutional layer of a given group of convolutional layers from the On-Chip memory to the system memory .
[0022] The method may further comprise the following operation : splitting a convolutional layer according to output channels . In greater detail , each sub-network may correspond to a part of the convolutional layer processing one of the output channels . The processing of a convolutional layer may be split into several steps . In each step, the results for a given output channel may be computed . In such way, for each step, only the weights associated to this output channel need to be fetched in OCM memory .
[0023] The channel splitting may be performed on additional convolutional layers that do not have an input feature map divided into several subsets of input feature data .
[0024] The number of generated sub-networks may be re-grouped according to collocated output channels .
[0025] According to embodiments , a computer-implemented method for scheduling and performing inference using a split convolutional neural network obtained according to the method described above comprises : performing inference on groups of convolutional layers that correspond to di f ferent sub-networks , wherein, for a given group, each subset of input feature data is processed by the convolutional layers of the group and recomposed as input for a subsequent group, and parts of the data overlap of the subsets of input feature data are considered for recomposing .
[0026] Performing inference based on groups of convolutional layers allows automati zing the process of scheduling process on the underlying hardware .
[0027] Before performing inference on the given group, memory space at an On-chip memory for an input of the subsequent group may be allocated .
[0028] Furthermore , after processing the sub-networks , the processing of the subnetworks based on channel split may be performed .
[0029] After processing the sub-networks , inference may be performed on a last group that comprises no subsets of input feature data .
[0030] According to embodiments , a system for performing the above-described computer-implemented methods comprises : a central processor unit ( CPU) , a system memory, and a neural processing unit (NPU) comprising an On-Chip memory ( OCM) .
[0031] In addition, an apparatus according to embodiments is provided for executing the system .
[0032] The apparatus may comprise an image sensor according to embodiments that comprises a pixel array .BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings are included to provide a further understanding of embodiments of the invention and are incorporated in and constitute a part of this speci fication . The drawings illustrate the embodiments of the present invention and together with the description serve to explain the principles . Other embodiments of the invention and many of the intended advantages will be readily appreciated, as they become better understood by reference to the following detailed description . The elements of the drawings are not necessarily to scale relative to each other . Like reference numbers designate corresponding similar parts .Fig . 1A is a schematic diagram of a convolutional neural network according to examples .Fig . IB is a schematic diagram of an input feature map according to embodiments .Fig . 1C is a schematic diagram of a convolution operation according to embodiments .Fig . 2A is a schematic diagram o f a method to analyze the convolutional neural network and split the convolutional neural network according to embodiments .Fig . 2B is a schematic diagram of the method in combination with a system according to embodiments .Fig . 3A is a schematic diagram of an input feature map divided into subsets of input feature data according to embodiments .Fig . 3B is a schematic diagram of a subset of input feature data being processed by convolutional kernels according to embodiments .Fig . 3C is a schematic diagram of a subset of input feature data being processed by convolutional kernels according to embodiments .Fig . 3D is a schematic diagram of a subset of input feature data according to embodiments .Fig . 4 is a schematic diagram of a split convolutional network according to embodiments .Fig . 5A is a schematic diagram of a sequence for performing inference based on subsets / stripes STR according to embodiments .Fig . 5B is a diagram of a method for scheduling and performing inference using a spl it convolutional neural network obtained according to embodiments .Fig . 6 is a schematic diagram of a convolutional layer split by channels according to embodiments .Fig . 7 is a schematic diagram illustrating a system for performing image data proces sing by the convolutional neural network according to embodiments .Fig. 8 is a schematic diagram illustrating an apparatus including an image sensor according to embodiments.Fig. 9 is a schematic diagram illustrating a memory gain versus additional processing cost ratio.DETAILED DESCRIPTION
[0034] In the following detailed description reference is made to the accompanying drawings, which form a part hereof and in which are illustrated by way of illustration specific embodiments in which the invention may be practiced. In this regard, directional terminology such as "top", "bottom", "front", "back", "over", "on", "above", "leading", "trailing" etc. is used with reference to the orientation of the Figures being described. Since components of embodiments of the invention can be positioned in a number of different orientations, the directional terminology is used for purposes of illustration and is in no way limiting. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope defined by the claims.
[0035] The description of the embodiments is not limiting. In particular, elements of the embodiments described hereinafter may be combined with elements of different embodiments.
[0036] As used herein, the terms "having", "containing", "including", "comprising" and the like are open ended terms that indicate the presence of stated elements or features, but do not preclude additional elements or features. The articles "a", "an" and "the" are intended to include the plural as well as the singular, unless the context clearly indicates otherwise.
[0037] In the following, the term "convolutional layer" may be used and understood as "kernel," "convolutional kernel," "convolutional matrix," "mask," or "filter." The terms "input feature map" and "output feature map" may be used interchangeably with "input data," "input image," "output data," and "output image . "
[0038] It is further noted that the embodiments described herein are not limited to classification networks. The present disclosure may operate on „encoder+decoder" networks which output 2D output features and where feature classification layers are replaced by decoder layers.
[0039] Fig. 1A is a schematic diagram of a convolutional neural network (CNN) 10 according to examples, which may be used herein. The CNN 10 comprises an architecture that resembles the connectivity pattern of neurons in the human brain and is an example of a Deep Learning algorithm used for feature extraction and classification. The CNN 10 may take in an (initial) input feature map 12 such as a 2D input image and assigns learnable weights and biases (i.e., importance) to various aspects / ob j ects in the image. The CNN 10 may then be trained to be able to differentiate one aspect / ob j ect from the other.
[0040] Specifically, the CNN 10 may be able to capture spatial and dependencies in the input image through the application of convolutional kernels 14 that act as filters on the input image. These convolution operations may be performed in a plurality of convolutional layers ConvL as indicated in Fig. 1A. Thereby, a convolutional layer ConvL may receive a respective input feature map 12 and produces an output feature map that serves as input to a subsequent convolutional layer ConvL. When processing the input feature map 12 by the respective convolutional layer ConvL, a portion of the input feature map 12 is multiplied with thekernel 14, which shifts depending on a stride length over the input image. The kernels 14 are indicated by small rectangles in Fig. 1A hovering over respective portions or parts of the input feature maps 12.
[0041] A corresponding convolution operation (also called convolution therein) is illustrated in a simplified manner in Figs. IB to 1C. An input feature map 12 may be traversed by a kernel 14, whose movement is indicated in Fig. IB via a corresponding arrow. That is, the kernel 14 may move to the right with a certain stride value until it parses the complete width of the input feature map 12. Moving on, it may traverse down to the beginning (left) of the input image with the same stride value and repeats the process until the entire image is processed. In general, the stride value may comprise a parameter for processing the height and a parameter for processing the width of the input feature map 12. These parameters may be equal or different from each other. For example, when the stride parameters are set to (SIZE height = 1, SIZE width = 1) this indicates that the input image 12 may be traversed in units of one unit pixel 2 in the height and one unit pixel 2 in the width. Similarly, the kernel values may be indicated by (SIZE height, SIZE width) . In case of 3D input feature maps 12 there may also be a value for the depth. The input feature map 12 may comprise a plurality of unit pixels 2 having specific pixel values that may be elementwise multi- plicated with the kernel 14 (Hadamard Product) when it passes over .
[0042] Fig. 1C illustrates a convolution operation with a kernel 14 having the size of (SIZE height = 3, SIZE width = 3) and two different portions or parts of the input image map 12. In this example, dilation parameters, which will be explained in the following paragraphs are set to (SIZE height = 1, SIZE width = 1) . The results of the elementwise multiplication operationsbetween the kernel 14 and the parts of the input images are shown on the right side of Fig. 1C. As can be seen, the multiplied pixel values are added together. The pixel values in this example are arbitrary and only illustrated for explaining the convolution operation. Furthermore, the (initial) input feature map 12 shown in Figs. 1A to 1C is a 2D image. However, also 3D images or ID images may be processed by the CNN 10. Moreover, the size of the kernel 14 may be of any size and not limited to (3,3) . The same may apply to the dilation parameters, which may have different values.
[0043] There may also be a plurality of channels (not shown in Figs. 1A to 1C) comprising, for example, grayscale, Red-Green- Blue (RGB) , aHue Saturation Value (HSV) color scale, cyan-ma- genta-yellow- , and key (black) (CMYK) , and many more.
[0044] The stride parameters of the convolution may define the number of unit pixels 2 moved over between each convolution operation. For a 2D convolution as shown in Figs. 1A to 1C, there may be two values, each value for each dimension (i.e., width and height) as explained above. For stride parameters of (1,1) , all unit pixels 2 of the input feature map may be active pixels of the convolution. An active pixel may be a pixel in which the data is within the center of the kernel 14 for a convolution operation. These active pixels are indicated in Fig. 1C by a circle. For example, in case of stride values (1,1) each pixel will be an active pixel at one point, i.e., when the kernel 14 hovers over it such that it is located in the center. In contrast, for stride values of (2,2) every two unit pixels 2 in a row and every two rows in the input feature map 12 may be active pixels. The stride parameters (i.e., the stride values) for the height may be limited to a power of two (e.g., (SIZE height = , SIZE width) , (SIZE height = 2, SIZE width) , (SIZE height = 4,SIZE width) , ...) . The stride parameters for the width may be of any value.
[0045] A further parameter is the so-called dilation parameters, which is similar to the stride parameters. They indicate how much the kernel 14 is widened. There are usually spaces inserted between the kernel elements, and so the unit pixels 2 taken for the convolution operations may not necessarily be collocated. We will consider any value of the dilation parameters in this present disclosure.
[0046] The convolution operations as explained above may result in the extraction of features such as edges from an input feature map 12. As shown in Fig. 1A, the CNN comprises a plurality of convolutional layers ConvL . Conventionally, the first convolutional layer ConvL is responsible for capturing low-level features such as edges, color, gradient orientation, etc. With additional layers, the architecture may adapt to high-level features as well, resulting in a network that may have an understanding of images present in the (initial) input feature map 12.
[0047] Furthermore, the convolution operations performed on an input feature map 12 may result in a convolved feature which is reduced in dimensionality as compared to the input feature. This is achieved by applying "valid padding." In addition, it may be possible that the convolved feature remains the same in dimensionality. This done by applying "same padding." In greater detail, padding is the amount of data (e.g., zeros) added (or not) to the borders of the input feature map 12 such that there is enough data to perform the convolution operations on the borders of the input feature map 12. In the following, convolution operations are considered in which padding is set to "same," so that the output size (width, height) is the same as the inputsize (if the stride values are equal to (1,1) ) . In addition, convolution operations are considered, where the padding is set to "valid." In such a case, no data (zeros) is added on the borders and so the output size is smaller than the input size.
[0048] Each convolutional layer ConvL of the CNN 10 may comprise a specific kernel-, padding-, stride and dilation configuration.
[0049] Moreover, additional layers may be present in the CNN 10. For example, one or more "pooling layers" may be included (not shown in Fig. 1A) , which may be responsible for reducing the spatial size of a convolved feature. This may decrease the computational power required to process the data through dimensionality reduction. Furthermore, it may be useful for extracting dominant features which are rotational and positional invariant, thus maintaining the process of effectively training the model.
[0050] There are two types of pooling: max pooling and average pooling. Max pooling returns the maximum value from the portion of the image covered by the kernel 14. On the other hand, average pooling returns the average of all the values from the portion of the image covered by the kernel 14.
[0051] Furthermore, a fully-connected layer FCL may be included in the CNN 10, which may allow learning non-linear combinations of the high-level features represented by the resulting recomposed output feature maps, which result from the processing of the respective input feature maps 12 by the plurality of convolutional layers ConvL.
[0052] Before the recomposed output feature maps may be provided to the fully connected layer FCL, these may be further flattened into a column vector CV that is then fed to the fully-connected layer FCL as shown in Fig. 1A. Over a series of epochs, the modelmay be able to distinguish between dominating and certain low- level features in images and classify them using the "Softmax Classification technique." The result may then be provided as output OUT .
[0053] The above processing may be performed in two parts, the feature extraction FE and feature classification FC part as indicated in Fig. 1A.
[0054] Common CNNs have a large number of parameters and a large number of feature maps that require memory. For example, if large input feature maps are employed the memory usage and memory management can be a bottleneck to reach an appropriate performance. Extensive input data generates correspondingly large feature maps that consume extensive memory. On the other hand, in the CNN processing, all of the input feature maps should be processed simultaneously, and all of them are required at the same time to have efficient processing. This high memory usage directly impacts the memory sizing and corresponding costs.
[0055] According to embodiments, a method is proposed that analyzes the CNN 10 and splits the CNN 10 as described above into several sub-networks 20A, 20B, 20C as indicated in Figs. 2A and 2B. Fig. 2A shows the corresponding operations performed by the method to obtain a split CNN 10. Fig. 2B shows the method in combination with an underlying system 100 according to embodiments, which will be explained in reference to Fig. 7. In detail, Fig. 2B shows schematically how the single operations S110 to S150 may be performed to obtain the sub-networks 20A, 20B, 20C by splitting the CNN 10. In Fig. 2B three sub-networks are illustrated. The number of sub-networks 20A, 20B, 20C is not limited, and there may be more or less sub-networks 20A, 20B, 20C than tree.
[0056] According to further embodiments, the operations S110 to S150 may be performed offline on a computer.
[0057] In operation S110, for each convolutional layer ConvL, different subset configurations of an input feature map 12 to be processed by a respective convolutional layer ConvL are determined. Thereby, the input feature map 12 is divided into subsets STR of input feature data, each of the same size. In other words, it is identified into how many subsets STR the input feature map 12 may be divided into. For 2D image data (as assumed in the following description) these subsets may be referred to as stripes. According to further embodiments, other subsets may be feasible. For example, for 3D image data, these subsets may have a cuboid structure. This division or splitting is illustrated in Fig. 3A for an input feature map 12 that is representative for other input feature maps 12 to be processed by the respective convolutional layers ConvL of the CNN 10 and will be explained later on.
[0058] Then, for each determined subset configuration, the following operations are performed.
[0059] In operation S120, for each convolutional layer ConvL, a data overlap 18 to be added to each one of the subsets STR of input feature data is determined. The data overlap 18 is determined from the input feature map 12. More specifically, the data overlap 18 is required to correctly perform convolution operations based on the subsets STR of input feature data instead of the whole input feature map 12. This will be explained in more detail with regard to Figs. 3B to 3D.
[0060] In operation S130, a memory- and computational demand required for processing the subsets STR of input feature data by each convolutional layer ConvL is determined. For example, foreach subset configuration the memory and processing cost to perform convolutional operations of the corresponding sub- sets / stripes STR including the data overlaps 18 may be identified. It is then determined which configuration provides the best trade-off regarding memory and processing cost.
[0061] Furthermore, in operation S140, a subset configuration of the determined subset configurations is selected depending on the determined memory- and computational demand. In greater detail, the subset configuration having the best trade-off between memory consumption and processing cost is selected.
[0062] In operation S150, the convolutional network 10 is split into different sub-networks 20A, 20B, 20C depending on the selected subset configuration to obtain a split convolutional network.
[0063] Splitting the CNN 10 according to the selected subset configuration may result in a plurality of sub-networks 20A, 20B, 20C with an overall optimized memory usage and moderate impact on computational costs. As a result, the sub-networks 20A, 20B, 20C perform inference requiring less memory than the original (un-split) CNN as they operate on subsets e.g., stripes STR instead of whole input feature maps 12.
[0064] It is noted that the different sub-networks 20A, 20B, 20C resulting from the CNN 10 split may not be the same. There may be a first sub-network 20A, another sub-network 20B, and for instance, a further sub-network 20C.
[0065] For example, the first sub-network 20A may be executed iteratively on each subset and the results may be stored in a corresponding memory until all the data necessary to execute the next sub-network 20B is computed.
[0066] Each sub-network 20A, 20B, 20C may represent a group of convolutional layers ConvL sharing the same subset configuration . This will be explained in detail later on .Determining subset configurations :
[0067] In the following, the identi fication of the subset configurations according to the operation S 110 is explained in de- tail . For each convolutional layer ConvL, the input feature map12 may comprise image data that are divided into subsets STR of input feature data based on pre-defined rules . The subsets STR of input feature data may be represented by stripes . The examples described below refer to 2D image data for simplicity reasons . It is noted that also ID and 3D image data may be processed .
[0068] For each convolutional layer ConvL, there may be only a given number of subset configurations . More speci fically, there may be a limited number of how many subsets / stripes STR the input feature map 12 of a respective convolutional layer ConvL may be divided into . In the present disclosure , only a split on the height of the input feature map 12 is considered . For example , the possible height of a subset / stripe STR may be of a power of two and at least two . This may limit the complexity of a computer algorithm and may be compatible with the possible stride configuration for the height , which is also limited to a power of two as explained above . Another division on the width of the input feature map 12 may be added too .
[0069] Moreover, at least for one convolutional layer ConvL, the number of subsets / stripes STR of the convolutional layer ConvL may be a multiple of the number of subsets / stripes STR of the subsequent convolutional layer ConvL . This means that the deeper the processing advances into the CNN 10 , fewer subsets / stripesSTR may be present. This is since the input feature maps 12 of the convolutional layers ConvL may be smaller, and so the less they need to be divided or split. Each resulting subset / s tripe STR may comprise of a plurality of unit pixels 2 over which convolution operations are performed.
[0070] Fig. 3A illustrates three possibilities to separate an input feature map 12 (to be processed by a respective convolutional layer ConvL) into several subsets STR of input feature data of the same size. On the left side of Fig. 3A a first subset configuration is illustrated, which includes eight sub- sets / stripes STR of input feature data. The middle illustrates a second subset configuration with four subsets / stripes STR. And on the right side of Fig. 3A, a third subset configuration is shown with two subsets / stripes STR. More subset configurations may be possible, and these indicated in Fig. 3A are only by way of example. In detail, for every convolutional layer ConvL, different subset configurations for their respective input feature maps 12 may be identified.
[0071] Furthermore, to embrace all input feature map 12 sizes, it is possible that the height of the input feature map 12 is not limited to be a multiple of the height of the subset / stripe STR. In such a case, the last subset / stripe STR may be taken so that the last line of the input feature map 12 corresponds to the last line of the last subset / stripe STR. There may be duplicated data at the beginning of this subset / stripe STR, which may also be present in the previous subset / stripe STR. This duplicated data may be removed when recomposing the intermediate feature at the end of the processing by the subsets / stripes as explained later on.
[0072] During the operation S110, it may also be necessary to consider the other layers of the CNN 10 apart from theconvolutional layers ConvL . For example, in case one layer needs the full input feature map 12 (such as a "GlobalAveragePooling layer") , the division or split of the input feature map 12 may not be possible anymore. In such a case, all subset configurations for input feature maps 12 of subsequent convolutional layers ConvL may be set to one, i.e., they are not divided.
[0073] In addition, subset configurations may be filtered according to a number of active feature maps at the same time in a graph. If, for instance, a subset configuration changes between two convolutional layers ConvL, and there may be another active feature map in parallel that will be processed later in the graph, this configuration may be invalidated. In this case, the subset configuration having a minimal number of subsets / stripes STR may be taken for these two convolutional layers ConvL to correctly process the data, and such a configuration may be in a list of all possible subset configurations explored.
[0074] In greater detail, Fig. 1A shows only one active feature at a time. In other words, there may be only one path of data processing from beginning to the end. However, there are more complicated networks with multiple data branches / paths to which the method is applied to. For instance, an output feature 'OUTn' from a layer n may be the input of layer n+1 and be retained to be combined later in an ADD layer n+m where 'OUTn' is added with 'OUTn+m-1' . At the level of layer n+m-1, there may be two active features, 'OUTn' that must be retained and kept valid for the next layer and 'OUTn+m-2' which is processed by layer n+m-1. Such architectures are known as networks with skip connections. One example is the ResNet network family. An 'active feature' used herein may thus refer to either the input of the current layer or another input feature coming from a previous layer used in a subsequent layer in the graph.
[0075] After every possible subset configuration is determined, the operations S120 to S150 are performed, which are described in detail in the following paragraphs.Determining data overlaps:
[0076] Turning to operation S120, for each determined subset configuration, the data overlap 18 to be added to each one of the subsets STR of input feature data is determined. That is, for each determined subset configuration, i.e., for every specific subset / stripe STR of a subset configuration, the data overlaps 18 are computed. This may be done in four steps.
[0077] The first step relates to determining the minimal amount of data overlap 18 required for each subset / stripe STR of a given subset configuration. The second step may relate to the determination of the "real" data overlap 18 by propagating useless rows due to the padding. The third step may correspond to a harmonization of the data overlaps 18 regarding the stride values. And in the fourth step, a harmonization with regards to groups of convolutional layers ConvL may be performed.First step:
[0078] More specifically, the data overlap 18 or minimal data overlap 18 may result in accurate convolution operations that are performed based on a subset STR of input feature data. The left side of Fig. 3B shows an exemplary subset / stripe STR with added data overlap 18 resulting from a previous subset / stripe STR of the input feature map 12. The data overlap 18 may also be called co-located data and is added to the corresponding subset / stripe STR to get a convolution result that may resemble the convolution result obtained when processing the whole input feature map 12 instead of the subsets / stripes STR. The unit pixels2 are represented by respective small rectangles with the active pixels being highlighted by small circles. The kernel 14 that passes over these unit pixels 2 for performing respective convolution operations is shown by a large rectangle and has a kernel size, for instance, of (3,3) . The stride parameter (i.e., the stride values) is set to (2,2) , the dilation parameter is set to (1,1) and the padding to "same." Furthermore, in this example, a height of the subset / stripe STR is taken to be of eight unit pixels 2. The right side of Fig. 3B shows schematically the resulting subset / stripe STR after the convolution operations took place.
[0079] As can be seen at the left side of Fig. 3B, when processing the top row of unit pixels 2 at the upper border of the sub- set / stripe STR, the kernel 14 requires one additional row of unit pixels 2 from the previous subset / stripe STR in addition to the row of unit pixels 2 following the top row to perform convolution operations.
[0080] Moreover, a further row of unit pixels 2 is required, because of the stride values of (2,2) . Due to these striding values, the top row, the third row, and so on, will be processed for the convolution operations. These lines are also called active lines. However, the second row, the fourth row, and so on would be left out. Since these rows are also required to get a convolution result resembling the one performed on the whole input feature map 12, an additional row of unit pixels 2 is added to the data overlap 18 to perform the convolution operations including these rows.
[0081] For the last row of the subset / stripe STR shown in Fig. 3B, no additional data is needed. Because of the stride values of (2,2) , the last input row for the convolution operations is the penultimate row, and so the last row will be there for theinput of the kernel with a size of (3,3) .
[0082] The output feature map or output feature subset / stripe (i.e., the resulting subset / stripe after the convolution operations are performed) , as shown on the right side of Fig. 3B, includes one line of trash data 22 because zero-padding is applied. This trash data 22 is going to be removed at a later time. In case the padding is set to "valid, " this line of trash data would not have been present.
[0083] The data overlap 18 as explained above is determined for the first convolutional layer ConvL of the CNN 10. However, the CNN 10 comprises a plurality of convolutional layers ConvL and for processing subsets / stripes STR deeper in the CNN 10 these data overlaps 18 need to be backpropagated .
[0084] Fig. 3C shows an example of processing a subset / stripe STR by two convolutional layers with stride values (1, 1) , a kernel size of (3, 3) , a dilation parameter set to (1, 1) and the padding set to "same." The second convolutional layer ConvL may need a data overlap 18 of one row at the top and one row at the bottom of the subset / stripe STR to perform convolutional operations based on the kernel size (3,3) .
[0085] In greater detail, to obtain convolution results that may resemble those when processing the whole input feature map 12, the data overlaps 18 of the subset / stripe STR processed by the second convolutional layer ConvL may need additional data overlap 18 rows in the first convolutional layer ConvL. The consequence is that another data overlap 18 needs to be added to the subset / stripe STR. In total, two rows of unit pixels 2 on the top and two rows on the bottom of the subset / stripe STR to be processed by the first convolutional layer ConvL may beconsidered to perform convolution operations by the two convolutional layers ConvL in cascade .
[0086] Given the plurality of convolutional layers ConvL of the CNN 10 , a small data overlap 18 in the last convolutional layers ConvL may generate a huge number of data overlaps 18 in the first convolutional layer ConvL due to this propagation of data overlaps 18 .
[0087] This backpropagation of data overlaps 18 also applies , i f padding is set to "valid, " but in that case there will be no trash data 22 as shown in Fig . 3C .
[0088] In view of the above , the operation S 120 may comprise ( as the first step ) determining, for each convolutional layer ConvL, a minimal data overlap 18 to be added to each one of the subsets STR of input feature data such that the result of the convolution operations by the convolutional layer ConvL matches the result of the convolution operation when processing the whole input feature map 12 . The amount of minimal data overlap 18 for a given convolutional layer ConvL may depend on the amount of data overlap 18 of the subsequent convolutional layers ConvL, and a kernel- , padding- , dilation, and stride configuration of the given convolutional layer ConvL, as shown for example in Fig . 3C .
[0089] As follows from above , the (minimal ) data overlap 18 needed for a given convolutional layer ConvL depends on all data overlaps 18 of the subsequent or next convolutional layers ConvL in the CNN 10 . These data overlap 18 may be computed by scanning the convolutional layers ConvL in reverse order . Thereby the data overlap 18 according to the data overlap 18 of the next convolutional layer ConvL and the kernel- , padding- , dilation, and stride configuration of the current convolutional layer ConvL may be considered .
[0090] Below is a short excerpt of program code for calculating the (minimal) data overlaps 18. The following parameters may be considered :• overlap (n) (top) , the top data overlap 18 for convolutional layer ConvL with number n,• overlap (n) (bottom) , the bottom data overlap 18 for convolutional layer ConvL with number n,• stride (n) , the stride parameter for the height of the input feature and the convolutional layer ConvL with number n,• padding (n) , the padding parameter in pixels for convolutional layer ConvL with number n (this could be different for top and bottom, but for simplicity they are set to be equal, which may be the case for real use cases) ,• overlap (m) , which is zero for top and bottom if m is bigger or equal to the first convolutional layer ConvL number where split configuration is set to 1 (no more splitting or division) ,• (int) , which is the rounding to the nearest below integer.
[0091] In greater detail, the formula applied to determine the minimal data overlaps 18 when the padding is set to "same" is as follows : overlap (n) (top) = (overlap (n+1 ) (top) + (int) ( (padding(n) + stride (n) - 1) / stride (n) ) ) * stride (n) overlap (n) (bottom) = (overlap (n+1 ) (bottom) +(int) (padding (n) / stride (n) ) ) * stride (n)
[0092] This formula may be different for the top and bottom data overlaps 18. It may be applied iteratively from the last to the first convolutional layer ConvL and may be implemented for all padding configurations (the example algorithm below is provided in Python, but any other programming language may be applied) :
[0093] As described above, to obtain processing results based on subsets / stripes STR that may have the same result as the original processing using the complete input feature map 12, some minimal data overlap 18 on the borders of each subset / stripe STR may be added. This additional data comes from the previous or the next subset / stripe STR of a given input feature map 12.
[0094] However, for the first subset / stripe STR of a given input feature map 12, there is no previous subset / stripe STR.
[0095] That is, in addition to the minimal data overlap 18 that accounts for correct convolutional results when processing sub- sets / stripes STR, also the borders of the input feature map 12 need to be taken into consideration.
[0096] One technique might be to add data overlaps 18 filled with zeros. However, this technique produces incorrect results. For example, for convolution operations where padding is set to "same," the CNN 10 may automatically add some zeros for the top and bottom padding. When considering two convolutional layers ConvL with stride-, dilation- and kernel configurations asdescribed with regard to Fig. 3C above, i.e., with stride values of (1,1) , dilation values of (1,1) and a kernel size of (3,3) , then a data overlap 18 of two rows at the top and two rows at the bottom is required. If these are filled with zeros, the result of the convolutional operations performed by the first convolutional layer ConvL via the kernel 14 will be the same as normal processing. But the overlapping row resulting after the processing by the first convolutional layer ConvL will not be filled with zeros anymore. This is due to the adding of biases by the first convolutional layer ConvL, which will not be zeros. Consequently, the result of the processing by the second convolutional layer ConvL will not be the same as normal processing.
[0097] To imitate the calculation of normal processing, no data overlap 18 is added on top of the first subset / stripe STR. Therefore, benefit is taken from the inherent padding of the convolutional layers ConvL in case the padding is set to "same." The same is for the case when the padding is set to "valid." But, as described above, the input size of all subsets / stripes STR (for a respective subset configuration) processed by a convolutional layer ConvL must be the same. Therefore, a data overlap 18 based on data of the next subset / stripe STR is taken at the bottom of the first subset / stripe STR. This additional data overlap 18 does not affect the convolution results, and the resulting outputs are discarded when recomposing the output feature at the end of the processing by subsets / stripes STR. This additional data overlap 18 may also have been filled with zeros.
[0098] The same technique with regard to the top subset / stripe STR may be used for the last subset / stripe STR of the input feature map 12. A data overlap 18 may be taken so that the last row of the input feature map 12 is taken as the last row of the subset / stripe STR.
[0099] This is illustrated in Fig. 3D, which shows the data overlap 18 selections according to subset / stripe STR location. The upper panel of Fig. 3D illustrates the data overlap 18 selection for a subset / stripe STR located in the middle of an input feature map 12. The calculation of this minimal data overlap 18 has been explained in detail above. The middle and lower panel of Fig. 3D illustrate how a data overlap 18 is determined for subsets / stripes STR that are the first and the last ones of the input feature map 12, respectively.
[0100] The right side of each panel in Fig. 3D shows the resulting subsets / stripes STR with data overlap 18.Second step:
[0101] As will be explained in detail later on, for each determined subset configuration, the convolutional layers ConvL that have the same number of subsets STR of input feature data may be grouped into one group. For each convolutional layer ConvL of a given group, some redundant additional rows of unit pixels 2 may be produced if the padding is set to "same." These rows of redundant data cannot be removed from the processing as the convolutional layers ConvL of the given group are executed at the same time.
[0102] This "real" data overlap 18 may be determined by starting from the first convolutional layer ConvL of a given group and computing the data overlap sizes for each next convolutional layer ConvL of the group. In fact, the only parameter that may affect the data overlaps 18 at this step for the next convolutional layers ConvL is the current stride parameter in case the padding is set to "same." In case the padding is set to "valid," no redundant rows are present.
[0103] Thus, the operation 120 may comprise (as the second step) : determining, for groups of convolutional layers ConvL having the same number of subsets STR of input feature data, a further data overlap 18 to be added to the respective subsets STR of input feature data. The further data overlap 18 may take into account redundant data due to padding.
[0104] The following shows an excerpt of an example of program code for calculating this data overlap 18 (also in Python) :Third step:
[0105] Furthermore, as explained above, when the stride parameter for the height (i.e., the stride value) is not equal to (SIZE height = 1) , an additional row (or additional rows, if it is greater than 2) of unit pixels 2 may be added so that the first active row of a subset / stripe STR (after the top data overlap18) will be an active row for the convolution operations. Therefore, a data overlap 18 may be added in each convolutional layer ConvL. Thus, the operation S120 may comprise (as the third step) : adding, for each convolutional layer ConvL, a fixed amount of data overlap 18 to each one of the subsets STR of input feature data to account for different striding parameters.
[0106] This may be achieved by incrementing iteratively by one the top and bottom data overlap 18 of the first convolutional layer ConvL until the data overlaps 18 of all convolutional layers ConvL of the groups are correct.
[0107] The following illustrates an excerpt of an example of program code for this calculation (also in Python) :Fourth step:
[0108] Finally, there may still be a need for a correction. In the previous processing's the data overlaps 18 may have been increased within each group. Therefore, a situation might arise, where the data overlap 18 for the first convolutional layer ConvL of a group n is bigger than the data overlap 18 implied by the convolution operations of the last convolutional layer ConvL of the group n-1.
[0109] To address such a situation, the operation S120 may further comprise (as the fourth step) : determining S120 a data overlap 18 to be added to the subsets STR of input feature data for each convolutional layer ConvL of a given group that accounts for sufficient data overlap 18 for a subsequent group.
[0110] This may be achieved by scanning the data overlap 18 of each first convolutional layer ConvL of each group and increasing the data overlap 18 of the previous group such that enough data overlap 18 may be provided to next group.
[0111] The following illustrates an excerpt of an example of program code for this calculation (also in Python) :
[0112] After the required data overlaps 18 have been determined for each subset configuration based on steps one through four, the required memory and processing costs are derived as follows.Determining memory and computational demand:
[0113] In the following, the focus is set to operation S130. After the previous operation S120, the number of subset configurations and corresponding data overlaps 18 for the respective subsets / stripes STR are determined. With this information it is possible to compute the memory needed in an On-Chip memory (OCM) 50 of an underlying system 100 according to embodiments, which is used to execute the CNN 10. This system 100 will be further explained in combination with Fig. 7. The operation S130 may thus comprise: determining S130 the memory demand in the OCM 50. Said memory demand depends on the sizes of the subsets STR of input feature data and their convolved outputs, their corresponding data overlaps 18, and a size of the weights and biases for each one of the convolutional layers ConvL .
[0114] Furthermore, it is also possible to compute the processing cost. The processing cost may be expressed in number of clock cycles and may be the sum of a number of multiple accumulateoperations (MACS ) and the number of clock cycles needed for the memory trans fers between a system memory 60 of the system 100 and the OCM 50 . According to embodiments , one MAC may be performed in one clock cycle ( e . g . , a clock used for a neural processing unit (NPU) 30 of the system 100 ) .
[0115] That is , operation S 130 may further comprise : determining S 130 a number of multiple accumulate operations (MACS ) that results from the convolution operations of one subset STR of input feature data and the determination of the respective data overlaps 18 multiplied by the number of subsets STR of input feature data for a given convolutional layer ConvL, and determining S 130 a number of clock cycles required for memory trans fers between the system memory 60 of the system 100 and the OCM 50 .
[0116] These memory trans fers may comprise :• the trans fer of the weights for each convolutional layer ConvL from the system memory 60 to the OCM 50 ,• the trans fer of the subsets STR of input feature data and the data overlaps 18 from the system memory 60 to the OCM 50 for each convolutional layer ConvL that starts a given group of convolutional layers ConvL, and• the trans fer-back of the convolved outputs of the subsets STR of input feature data and the data overlaps 18 of the last convolution layer ConvL of a given group of convolutional layers ConvL from the OCM 50 to the system memory 60 .
[0117] In addition, it is noted that all layers of the CNN 10 may be considered for determining the memory- and computational demand . I f there are further layers that are not convolutional , it is important to consider the memory needed for input feature maps 12 and output feature maps for these layers .
[0118] It is further noted that for all layers of the CNN 10, there might be other active feature maps not used in the respective layer but later in the graph. If that is the case, the memory needs to be added for such feature maps considering the current subset configuration. Once this is done for all layers, the total MACs processing cost needed by all layers for this subset configuration, and thus, the maximum memory needed in OCM 50 (which is the maximum memory needed for a given layer in the CNN 10) may be determined.
[0119] The required system memory 60 may depend on the weights of the CNN 10, which are stored and fetched from the system memory 60.
[0120] The required system memory 60 may also depend on:• the size of the subset STR input feature data. This size depends on the subset configuration and the underlying data overlaps 18. Considered is the subset / stripe STR currently processed, in addition to, x subsets / stripes STR to get the top data overlap 18, plus y subsets / stripes 18 to get the bottom data overlap 18 (the system 100 acquires the input by multiple of subsets / stripes STR) .• the size of the resulting output feature maps of the CNN 10 (without subset / stripe STR consideration) .• the sizes of the active subsets / stripes STR to store the results for each group of convolutional layers ConvL .Subset configuration choice:
[0121] After the memory- and computational demand for each subset configuration is determined, the method progresses to operation S140. In operation S140, a subset configuration of the determined subset configurations is selected depending on the determined memory- and computational demand.
[0122] In greater detail, in operation S140, it is empirically decided to invalidate subset configurations where the processing cost is more than a predetermined value, e.g., two times larger than the original CNN 10 without dividing input feature maps 12 into subsets / stripes STR. This invalidation setting may be easily changed and is only by way of example. For example, if a smaller value is taken, it will lower the increase of processing cost but the gain the OCM 50 will be lower as well, as shown by results of use cases described later on in combination with Fig. 9.
[0123] Once this condition is verified, the subset configuration with the lowest OCM memory size is chosen. According to embodiments, the system memory 60 is not considered as it varies only marginally depending on the subset configurations. Furthermore, the focus is to lower the OCM 50.Generation of a split convolutional network according to embodiments :
[0124] Once the desired subset configuration is chosen, the next operation is to split the CNN 10 into several pieces or subnetworks 20A, 20B. 20C. This is done in operation S150.
[0125] In addition, as already described above, for each determined subset configuration, the convolutional layers ConvL that have the same number of subsets STR of input feature data may be grouped into one group. That is, all convolutional layers ConvL with the same split or subset / stripe STR number are regrouped into a group. More specifically, for a given subset configuration, groups of convolutional layers ConvL may correspond to groups where each layer has the same number of subsets / stripes STR. This grouping may simplify the execution of the graph ofthe resulting CNN 10. This may also allow removing useless data overlaps 18 between these groups, as explained in previous paragraphs regarding the second step for determining the data overlaps 18.
[0126] For the selected input feature subset configuration, the CNN 10 may be split into sub-networks 20A, 20B, 20C according to the different groups of convolutional layers ConvL . Each group of convolutional layers ConvL may represent a different subnetwork 20A, 20B, 20C of the CNN 10.
[0127] In other words, each piece or sub-network 20A, 20B, 20C may be a group of convolutional layers ConvL where all layers have the number of subsets / stripes STR.
[0128] The resulting sub-networks 20A, 20B. 20C, which are based on the selected subset configuration provide the best trade-off between memory demand and computational cost.
[0129] An example of such a split CNN 10 is illustrated schematically in Fig. 4. Fig. 4 shows an (initial) input feature map 12 divided into eight subsets / stripes STR that are going to be processed by groups of convolutional layers ConvL as described above. In Fig. 4 the data overlaps 18 of the respective subsets / stripes STR are not illustrated.
[0130] As can be seen in Fig. 4, the sub-networks 20A, 20B comprise each a group of convolutional layers ConvL, in the illustrated example each may comprise two convolutional layers ConvL, that may process the respective input subsets / stripes STR. In greater detail, in Fig. 4, the first group of convolutional layers 20A may process eight subsets / stripes STR whereas the second group 20B may process four subsets / stripes STR. Thenumber of sub-networks 20A, 20B is not limited, and there might be more sub-networks 20A, 20B.
[0131] It may further be noted that the method described above may be performed on any CNN.Performing inference using a split network according to embodi- ments :
[0132] Once the sub-networks 20A, 20B, 20C resulting from the splitting of the original CNN 10 are obtained, inference is performed on the targeted system.
[0133] Fig. 5A shows a simplified sequence for performing inference based on subsets / stripes STR according to embodiments.
[0134] In the first part (e.g., the feature extraction part as described with reference to Fig. 1A) , the inference may be derived in several steps. Each step may operate on a subset / stripe STR of the input feature map 12. For instance, an (initial) input feature map 12 or input image of 224x224 unit pixels 2 may be processed by applying a selected subset configuration (indicated by a first arrow in Fig. 5A) where corresponding subsets / stripes STR may include eight unit pixels 2 in height. Thus, these subsets / stripes STR may be of the size of 224x8 unit pixels 2 and are then processed (as indicated by a second arrow in Fig. 5A) by the sub-networks 20A, 20B, 20C (not shown in Fig. 5A) resulting from splitting the CNN 10 as described above. At the end of the first part, the resulting output subsets / stripes STR are recomposed (indicated by two small arrows in Fig. 5A) to get the same intermediate result as processing the whole input feature map 12 done by the original CNN 10.
[0135] Then the second part (i.e., the feature classification part FC of the CNN 10 as described with regard to Fig. 1A) is executed on this intermediate result to obtain the output feature of the CNN 10.
[0136] It is noted that the second part is not limited to feature classification. It may also refer to an "encoder" part of the network in addition to an optional "decoder" part.
[0137] Furthermore, it is noted that in Fig. 5A, to avoid overcomplexity, the number of subsets / stripes STR is always the same for the first part.
[0138] The second part may be optional if, for instance, the CNN 10 is composed exclusively of convolution layers ConvL (apart from transpose convolutions) and if the output feature sub- set / stripe STR may be 2D image data. In such a case, it may be possible to split the CNN 10 according to subsets / stripes STR.
[0139] When processing input feature maps 12 via subsets / stripes STR, it is not necessary to get the whole input feature map 12 in a buffer for processing. Instead, the subset / stripe selection and processing may be done on the fly. This allows saving some memory used for storing the input feature map 12 of the CNN 10. Furthermore, the processing may start during capturing 2D image data .
[0140] According to embodiments, a method is provided for scheduling and performing inference using a split convolutional neural network obtained according to the method described above. The method comprises the following operations as shown in Fig. 5B.
[0141] In operation S210, inference on groups of convolutional layers ConvL that correspond to different sub-networks 20A, 20B, 20C is carried out.
[0142] Then, for a given group of convolutional layers ConvL, operation S220 is performed. In operation S220, each subset STR of input feature data is processed by the convolutional layers ConvL of the group and recomposed as input for a subsequent group. In greater detail, only parts of the data overlap 18 of the subsets STR of input feature data are considered for recomposing. After recomposing, the resulting output subset / stripe serves as input into the next group of convolutional layers ConvL (i.e., the next sub-network 20A, 20B, 20C) .
[0143] For example, given a sub-network 20A, 20B, 20C, which corresponds to a specific group of layers, an input may correspond to a subset / stripe STR with additional data overlap 18 at the top and at the bottom of this subset / stripe STR. Once the output subset / stripe STR of said sub-network 20A, 20B, 20C is obtained, it must be recomposed into an input subset / stripe STR for the next group of layers. For that, the following may be taken into account.
[0144] Output data overlaps 18 may be partially duplicated, for example, when considering two consecutive subsets / stripes STR. Therefore, only portions of these data overlaps 18 may be written into the input subset / stripe STR of the next group. For instance, if a ratio between the number of subsets / stripes STR of the current group and the number of subsets / stripes STR of the next group ( k=split ( group (n-m) ) / split ( group (n-m+1 ) ) is two, the top data overlap 18 of the first subset / stripe STR and the bottom data overlap 18 of the second subset / stripe STR may be considered as input into the output subset / stripe STR.
[0145] In addition, as explained above, there may be some trash data 22 present due to the padding. Therefore, some rows of the data overlap 18 may be skipped or removed.
[0146] Furthermore, as mentioned before, the first and the last subsets / stripes STR of the input feature map 12 may be treated separately from the rest of the subsets / stripes STR. For the top subset / stripe STR, the processing may be started at the top of the input feature map 12. For subsequent subsets / stripes STR until a multiple of the ratio k, the processing starts from the top of the input feature map 12, which is shifted to the height of the input subset / stripe STR. At the output, when recomposing the subset / stripe STR for the next group, for the first subset / stripe STR the top output feature may be taken. For the subsets / stripes STR until a multiple of the ratio k, the output feature may be shifted so that there is a continuity of data when recomposing it with the output feature of the first subset / stripe STR. This may be done for the subsequent groups until a group is reached that does not process subsets / stripes STR (e.g., a layer with a number of subsets / stripes STR set to one) .
[0147] A similar processing is done for the last subsets / stripes STR of the input feature map 12 by aligning the bottom of the last subset / stripe STR to the bottom of the input feature map 12.
[0148] This method may provide a proper looping over the subnetworks 20A, 20B, 20C, which depends on the number of subsets / stripes STR. Furthermore, the input / output feature maps of the sub-networks 20A, 20B, 20C are efficiently managed, since the data overlaps 18 corresponding to the subsets / stripes STR including the special cases when a subset / stripe STR is the first or the last subset / stripe STR are considered in the processing.
[0149] The main scheduling of the inference, for which a short excerpt of an example program code is shown below, may be summarized as follows:• Imbricating loops to perform inference on all groups where the input feature maps 12 are divided into subsets / stripes (e.g., splitting is set above one) . For simplicity, there may be n groups like this. For a given loop m (0 being the first main loop, n being the last) , inference may be performed for the sub-network 20A, 20B, 20C containing layers of the group n-m.• The total loop count depends on the number of subsets / stripes STR. For a given loop m, it may be given by the ratio k=split ( group (n-m) ) / split (group (n-m+1 ) ) ,• Before each loop, given a loop m and so a group n-m, allocating input of the sub-network 20A, 20B, 20C corresponding to group n-m+1. This buffer may be directly filled by the output of the sub-network 20A, 20B, 20C corresponding to group n-m.• After looping, performing inference on the last group if this last group has no subsets / stripes STR (group n + 1) .
[0150] Provided below is an example of network inference for the CNN 10 according to embodiments, where the data overlaps 18 of the first group may include eight unit pixels 2 on the top and two unit pixels 2 on the bottom and the subset / stripe STR height without data overlap may include 16 unit pixels (in this example, which is a Python prototype, all input features maps 12 of the network CNN 10 may be in input_data) :Processing by channels:
[0151] A wide range of CNNs are based, at least in their first part, on encoder-like architecture where the size of the intermediate features decreases, whereas the size of the weightsincreases when going deeper into the network. At one point in the network, the size of the weights needed in a corresponding OCM 50 may be larger than the size of the intermediate features.
[0152] For processing by subsets / stripes STR as explained above, this may impact the memory optimization, which may decrease when going deeper into the network.
[0153] According to embodiments, another technique may be provided, which is called processing by channels, to address this issue. The main idea is to split the processing of a convolutional layer ConvL into several steps. In each step, the results for a given output channel may be computed. In this way, for each step, only the weights associated to this output channel may need to be fetched into the OCM 50 (given that the weights for each channel in a convolutional layer ConvL are distinct together) .
[0154] Fig. 6 shows a schematic diagram of a convolutional layer ConvL split by channels 26 according to embodiments.
[0155] Thus, a convolutional layer of the CNN 10 as described above may be split according to output channels 26. In this case, each sub-network 20' A, 20' B may correspond to a part of the convolutional layer ConvL processing one of the output channels 26, as shown in Fig. 6. The channel split is not limited to two sub-networks 20'A, 20'B. There may be more.
[0156] To simplify the whole process, the channel splitting may be applied only on last layers of the CNN 10, where there is no split or division by subsets / stripes STR to decorrelate them. In other words, the channel splitting may be performed on additional convolutional layers ConvL that do not have an input feature map 12 divided into several subsets STR of input feature data. Thechannel split may be performed on input channels 26' as shown inFig. 6.
[0157] For further simplification and reduction of the memory transfers between system memory 60 and OCM 50, it may be possible to limit the number of generated sub-networks 20' A, 20' B by regrouping collocated output channels 26, given that there is no need to insulate each output channel 26 to be able to gain enough OCM. In other words, the number of generated sub-networks 20'A, 20' B may be regrouped according to collocated output channels 26. Regrouping channels 26 may limit the need of fetching the input feature from the system memory 60 to the OCM 50 each time the sub-network 20' A, 20' B is invoked for each output channel. For example, the split may be by a factor of two, four, eight, and so on, according to the resulting gain on the OCM 50.
[0158] As shown in Fig. 6, a convolutional layer ConvL according to embodiments may take in four input channels 26' and may output two channels 26. These may be split in two small sub-networks 2 O' A, 20' B each producing one output channel 26. At the end of processing, the two output channels may be merged in order to get the same result as the original CNN 10. It is noted that each inference based on the sub-networks 20' A, 20' B is done sequentially on the NPU 30 so that only the weights of a given sub-network 20A, 20B, 20C are fetched at a time in the OCM 50.
[0159] When performing the inference using a split network as described above, it may further include an additional operation where after processing the sub-networks 20A, 20B, 20C, the processing of the subnetworks 20' A, 20' B based on the channel split is performed.
[0160] Then, after processing all of the sub-networks 20A, 20B, 20C, 20 ' A, 20 ' B, inference is performed on a last group that comprises no subsets STR of input feature data .Underlying system architecture :
[0161] As shortly mentioned above , Fig . 7 shows the underlying system 100 providing a hardware architecture for processing the CNN 10 . In detail , in the present disclosure neural network inference is done on a dedicated accelerated hardware , called a Neural Processing Unit (NPU) 30 . In such an architecture , the inference scheduling is generally done on a general-purpose central processing unit ( CPU) 40 , whereas the core operators of the CNN 10 are executed on the NPU 30 . Furthermore , in such an architecture , two types of memories may be distinguished . The first one is the system memory 60 , accessible by the CPU 40 and the NPU 30 to get the input feature maps 12 and the network weights . Another kind of memory is also needed, the memory used exclusively by the NPU 30 as a working memory during its processing . Herein, it is called On Chip Memory ( OCM) 50 . This memory and the NPU 30 are included on the same hardware module 70 that needs to be as small as possible for cost reason, as opposed to the general-purpose CPU 40 and the system memory 60 that may be included in the larger system 100 . The system memory 60 is usually larger as it deals with other processing in addition to the Al processing . So , the OCM 50 is generally smaller than the system memory 60 , and the memory optimi zation technique described herein has the focus to lower the usage of the OCM 50 . Moreover, the methods described above may allow to start processing earlier during pixels acquisition, which may improve the latency of the total processing ( image acquisition plus neural network inference ) .
[0162] Fig. 7 shows the system 100 in more detail. The system 100 allows performing the methods described above and comprises the CPU 40, the system memory 60, and the NPU 30. The NPU 30 further comprises the OCM 50 and may be included in the same hardware module 70. In addition, Fig.7 shows also other peripherals 80. The memory transfers are indicated by double arrows. Single arrows may indicate other operations such as control operations .
[0163] Fig. 8 is a schematic diagram illustrating an apparatus 1000 comprising the system 100.
[0164] According to embodiments the apparatus 1000 may further comprise an image sensor 200 comprising a pixel array.Results for a fixed processing cost increase ratio:
[0165] The methods as described above have been experimented on using two different CNNs 10. For this experiment the allowed increase of processing cost may be fixed to a factor of two.
[0166] The first use case is a very simple network based on mobilenet vl backbone, which detects if a face is present in an image .
[0167] The second use case is a more complex network based on mobilenet v3 backbone that performs a face segmentation on an image .
[0168] As described above, a program in Python may be implemented that explores all possible subset configurations in terms of processing by subsets / stripes STR and processing by channels 26. For each configuration, subset / stripe STR data overlaps 18 are computed and an evaluation of the cost in terms of memory andprocessing is performed. At the end, the program selects the configuration where there is a best trade-off between gain on the OCM 50 and increase of processing cost.
[0169] The results for the first use case may be summarized as follows .
[0170] As outlined in the table above, the gain on the OCM 50 is quite significant (% 3.7) , but the system memory 60 remains stable. The processing cost is increased by a factor of nearly two, i.e., the limit that has been configured for this experiment. A smaller ratio refers to better results.
[0171] It is noted that in this configuration also processing by channels 26 has been considered. When including processing by channels 26, the memory cost of the OCM 50 further decreases from 90880 bytes to 80784 bytes.
[0172] The results for the second use case may be summarized as follows :
[0173] From the results listed above, it is apparent that the gain on the OCM 50 is less than in the first use case. This is due to the fact that "Global Average Pooling layers" are included in the middle of the CNN 10, which prevents applying processing by subsets / stripes STR for deeper layers in the CNN 10. A smaller ratio refers to better results.Exploration on processing cost:
[0174] An experiment has been performed to study the impact of the processing cost increase factor on the memory gain. This experiment has been performed in the first use case, but not in the second use case where the possible values for this factor are limited (from 1.0 to 1.02) .
[0175] The underlying program code has been performed for a varying processing cost factor from 1 to 3. For each configuration, the real additional processing cost (that will always be lower than the configured one) and the memory gain on the system memory 60 and the OCM 50 have been determined.
[0176] The results can be seen in Fig. 9, which illustrates a memory gain versus additional processing cost ratio according to embodiments .
[0177] From Fig. 9, it follows that the system memory 60 usage remains stable, and its variation does not really follow any logic according to processing cost. The best compromise between memory gains and additional processing cost is achieved for a factor of 1.3. In such a configuration, the gain on the OCM 50 is good (divided by 3) and a small gain for the system memory 60(92% vs original) is obtained. For a larger processing cost ratio, the further improvement on memory gains may be more limited: the OCM 50 usage decreases a little while additional processing cost increases a lot.Validation :
[0178] As explained above, the CNN 10 is split into several pieces or sub-networks 20A, 20B, 20C according to the optimal subset configurations, and a program based on Python is developed to perform inference with such sub-networks 20A, 20B, 20C. In addition, the inference based on the sub-networks 20' A. 20' B has been explained.
[0179] In greater detail, the example program excerpts in Python provided in previous paragraphs may be implemented such that it explores all possible subset configurations in terms of processing by subsets / stripes STR and processing by channels 26. For each configuration, subset / stripe STR data overlaps 18 are computed and an evaluation of the cost in terms of memory and processing may be performed. At the end, the program may select the configuration where there is a best trade-off between gain on the OCM 50 and increase of processing cost.
[0180] Furthermore, its noted that the determination of the subset configurations and the selection of the one that provides the best trade-off may be performed offline on a computer.
[0181] To shortly summarize, the method described in the present disclosure may provide the following technical contributions:• it exposes an entire end-to-end detailed method from offline preprocessing to network inference on hardware,• it describes in detail the method to compute necessary data overlaps 18 between stripes STR according to networkarchitecture ,• it deals with the necessary trade-of f between memory optimi zation and additional computational cost ,• it includes a companion method, called processing by channels 26 , used to optimi ze further the memory,• it includes a novel technique to trans form the input network for optimi zed inference , and• it includes network scheduling for execution on the system 100 .
[0182] The method described in the present disclosure may have the following technical ef fects :• the need of dedicated memory to process CNNs on dedicated Al hardware may be lowered,• many existing CNNs may be processed by the novel techniques described therein,• these techniques may further ful ly automati ze the process of discovering the best compromise between memory saving and additional computational cost ,• the latency of execution may be reduced i f input data is acquired on-the- fly and network inference is done in parallel of this acquisition,• they may automatically modify the CNN according to the best compromise found, and they may automati ze the process of scheduling the optimi zed network on the hardware .LIST OF REFERENCES2 Unit pixel10 Convolutional network12 Input feature map14 Convolutional kernel18 Data overlap20A, 20B, 20C Sub networks2 O ' A, 20 ' B Sub networks22 Trash data26 Output channel26 ' Input channel30 Neural processing unit50 On Chip Memory60 System Memory70 Hardware module100 System200 Image sensor1000 ApparatusS 110 OperationS 120 OperationS 130 OperationS 140 OperationS 150 OperationS210 OperationS220 OperationSTR Subsets / stripesConvL Convolutional layerFCL Fully connected layerCV Column vectorOUT OutputFE Feature extraction partFC Feature classi fication part
Claims
CLAIMS1. A computer-implemented method to analyze a convolutional neural network and split the convolutional neural network (10) comprising a plurality of convolutional layers (ConvL) , the method comprising: determining ( S 110 ) , for each convolutional layer (ConvL) , different subset configurations of an input feature map (12) to be processed by the respective convolutional layer (ConvL) , wherein the input feature map (12) is divided into subsets (STR) of input feature data, each of the same size, wherein, for each determined subset configuration, the method further comprises: determining (S120) , for each convolutional layer (ConvL) , a data overlap (18) to be added to each one of the subsets (STR) of input feature data, wherein the data overlap (18) is determined from the input feature map (12) ; determining (S130) a memory- and computational demand required for processing the subsets (STR) of input feature data by each convolutional layer (ConvL) ; selecting (S140) a subset configuration of the determined subset configurations depending on the determined memory- and computational demand; and splitting (S150) the convolutional network (10) into different sub-networks (20A, 20B, 20C) depending on the selected subset configuration to obtain a split convolutional network .
2. The method according to claim 1, wherein, for each convolutional layer (ConvL) , the input feature map (12) comprises image data that is divided into subsets (STR) of input feature data based on pre-defined rules, wherein the subsets (STR) of input feature data are represented by stripes (STR) .
3. The method according to claim 1 or 2, wherein, for each determined subset configuration, the convolutional layers (ConvL) that have the same number of subsets (STR) of input feature data are grouped into one group.
4. The method according to any one of the preceding claims, wherein, for the selected input feature subset configuration, the convolutional network (10) is split into sub-net- works (20A, 20B, 20C) according to the different groups of convolutional layers (ConvL) , wherein each group of convolutional layers (ConvL) represents a different sub-network (20A, 20B, 20C) of the convolutional network (10) .
5. The method according to anyone of the preceding claims, wherein each convolutional layer (ConvL) comprises a specific kernel-, padding-, dilation-, and stride configuration.
6. The method according to anyone of the preceding claims, wherein determining, for each convolutional layer (ConvL) , the data overlap (18) to be added to each one of the subsets (STR) of input feature data further comprises: determining (120) , for each convolutional layer (ConvL) , a minimal data overlap (18) to be added to each one of the subsets (STR) of input feature data such that the result of the convolution operations by the convolutional layer (ConvL) matches the result of the convolution operations when processing the whole input feature map (12) , wherein the amount of minimal data overlap (18) for a given convolutional layer (ConvL) depends on the amount of data overlap (18) of the subsequent convolutional layers (ConvL) and a kernel-, padding-, dilation-, and stride configuration of the given convolutional layer.
7. The method according to claim 6, wherein determining (S120) , for each convolutional layer (ConvL) , the data overlap (18) to be added to each one of the subsets (STR) of input feature data further comprises: determining (S120) , for groups of convolutional layers (ConvL) having the same number of subsets (STR) of input feature data, a further data overlap (18) to be added to the respective subsets (STR) of input feature data, wherein the further data overlap (18) takes into account redundant data due to padding.
8. The method according to claims 6 or 7, wherein determining (S120) , for each convolutional layer (ConvL) , the data overlap (18) to be added to each one of the subsets (STR) of input feature data further comprises: adding (S120) , for each convolutional layer (ConvL) , a fixed amount of data overlap (18) to each one of the subsets (STR) of input feature data to account for different striding parameters .
9. The method according to claims 6 to 8, wherein determining (S120) , for each convolutional layer (ConvL) , the data overlap (18) to be added to each one of the subsets (STR) of input feature data further comprises: determining (S120) a data overlap (18) to be added to the subsets (STR) of input feature data for each convolutional layer (ConvL) of a given group that accounts for sufficient data overlap (18) for a subsequent group.
10. The method according to anyone of the preceding claims, wherein determining (S130) a memory- and computational demand required for processing the subsets (STR) of input feature data by each convolutional layer (ConvL) further comprises:determining (S130) the memory demand in an On-Chip memory, wherein the memory demand depends on the sizes of the subsets (STR) of input feature data and their convolved outputs, their corresponding data overlaps (18) , and a size of the weights and biases for each one of the convolutional layers (ConvL) .
11. The method according to anyone of the preceding claims, wherein determining (S130) a memory- and computational demand required for processing the subsets (STR) of input feature data by each convolutional layer (ConvL) further comprises: determining (S130) a number of multiple accumulate operations, MACs, that results from the convolution operations of one subset (STR) of input feature data and the determination of the respective data overlaps (18) multiplied by the number of subsets (STR) of input feature data for a given convolutional layer (ConvL) ; and determining (S130) a number of clock cycles required for memory transfers between a system memory (60) and an On- chip memory (50) .
12. The method according to claim 11, wherein the memory transfers comprise: the transfer of the weights for each convolutional layer (ConvL) from the system memory (60) to the On-chip memory ( 50 ) ; the transfer of the subsets (STR) of input feature data and the data overlaps (18) from the system memory (60) to the On-chip memory (50) for each convolutional layer (ConvL) that starts a given group of convolutional layers (ConvL) ; and the transfer-back of the convolved outputs of the subsets (STR) of input feature data and the data overlaps (18) of the last convolutional layer (ConvL) of a given group ofconvolutional layers (ConvL) from the On-Chip memory (50) to the system memory (60) .
13. The method according to any one of the preceding claims, further comprising: splitting a convolutional layer (ConvL) according to output channels (26) , wherein each sub-network (20'A, 20'B) corresponds to a part of the convolutional layer (ConvL) processing one of the output channels (26) .
14. The method according to claim 13, wherein the channel splitting is performed on additional convolutional layers (ConvL) that do not have an input feature map (12) divided into several subsets (STR) of input feature data.
15. The method according to claim 13 or 14, wherein the number of generated sub-networks (20' A, 20' B) is regrouped according to collocated output channels.
16. A computer-implemented method for scheduling and performing inference using a split convolutional neural network (10) obtained according to a method of any one of the preceding claims 1 to 16, the method comprising: performing inference on groups of convolutional layers (ConvL) that correspond to different sub-networks (20A, 20B, 20C) , wherein, for a given group, each subset (STR) of input feature data is processed by the convolutional layers (ConvL) of the group and recomposed as input for a subsequent group, and parts of the data overlap (18) of the subsets (STR) of input feature data are considered for recomposing.
17. The method according to claim 16, wherein before performing inference on the given group, memory space at an On-chip memory (50) for an input of the subsequent group is allocated .
18. The method according to claims 16 or 17, wherein after processing the sub-networks (20A, 20B, 20C) , the processing of the subnetworks (20'A, 20'b) based on channel split is performed .
19. The method according to claims 18, wherein after processing the sub-networks (20A, 20B, 20C, 20' A, 20'B) , inference is performed on a last group that comprises no subsets (STR) of input feature data.
20. A system (100) for performing a computer-implemented method according to any one of claims 1 to 15 and a computer- implemented method according to any one of claims 16 to 19, the system (100) comprising: a central processor unit, CPU (40) ; a system memory (60) ; and a neural processing unit, NPU (30) , comprising an On- Chip memory, OCM (50) .
21. An apparatus comprising a system (100) according to claim 20.
22. The apparatus (1000) according to claim 21, further comprising an image sensor (200) comprising a pixel array.
Citation Information
Patent Citations
Implementation of a neural network in multicore hardware
US20220121914A1
Hardware accelerator optimized group convolution based neural network models
WO2023059335A1
A system and method for evaluating convolutional neural networks
WO2023062443A1