Method and apparatus for convolution operation
By rearranging and padding input data and filters to match fixed channel sizes, the method addresses efficiency issues in CNN hardware accelerators, ensuring effective convolution operations regardless of channel size mismatches.
Patent Information
- Application Number
- PCT/KR2023/021822
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-03
AI Technical Summary
Existing hardware accelerators for convolutional neural networks (CNNs) face efficiency issues when the input or output channel sizes of convolution operations do not match the fixed sizes they are designed for, leading to reduced performance.
A method and apparatus that rearranges and pads input data and filters to match the fixed channel sizes of hardware accelerators, using techniques such as channel rearrangement, duplication, and padding to perform convolution operations efficiently.
Enables efficient convolution operations even when input or output channel sizes are smaller than the hardware accelerator's fixed sizes, without altering the hardware accelerator's channel sizes.
Smart Images

Figure KR2023021822_03072025_PF_FP_ABST
Abstract
Description
Method and device for convolution operation
[0001] The disclosed embodiments relate to convolution operation techniques.
[0002] Recently, learning models based on convolutional neural networks (CNNs) have been applied to various fields such as image classification, speech recognition, and natural language processing, demonstrating excellent performance.
[0003] CNNs typically consist of deep neural networks, which increases the computational complexity. Therefore, methods utilizing hardware accelerators to accelerate the convolutional operations performed in CNNs have been proposed. However, these hardware accelerators typically have fixed input or output channel sizes. Therefore, if the input or output channel sizes of the convolutional operation performed using the hardware accelerator are smaller than those of the hardware accelerator, the efficiency of the convolutional operation deteriorates.
[0004] The disclosed embodiments are intended to provide a method and apparatus for convolution operation.
[0005] According to one embodiment, a convolution operation method comprises: when the number of channels of first input data is less than the number of preset input channels, rearranging elements of each channel of the first input data to generate second input data whose width or height is reduced by 1 / p (where p is a ratio of the number of preset input channels to the number of channels of the first input data) and whose number of channels is increased by p times compared to the first input data; padding each channel of the second input data with a preset value to generate third input data whose width or height is increased by p-1 compared to the first input data; duplicating p of each channel of one or more first filters whose number of channels is N / p (where N is the number of preset input channels) to generate one or more second filters whose number of channels is p·N; padding each channel of the one or more second filters with the preset value to generate one or more third filters whose width or height is K+p-1; And a step of performing a convolution operation with a stride of S (where S is a preset stride) on the third input data using the one or more third filters.
[0006] The step of generating the second input data may include generating p channels with a width or height reduced by 1 / p for each channel of the first input data, and among the p channels generated for the i-th channel of the first input data (where i is an integer 1≤i≤N / p), the j-th channel (where j is an integer 1≤j≤p) may include elements extracted from the j-th element at intervals of p in the width or height direction from the i-th channel.
[0007] The step of generating the third input data may generate the third input data by padding the preset value so that p-1 of the preset value is included before and after each element included in each channel of the second input data in the width or height direction for each channel of the second input data.
[0008] The one or more second filters include p channels identical to specific channels of the one or more first filters, and the step of generating the one or more third filters comprises: a first channel of the p channels is successively padded with p-1 preset values at the front end in the width or height direction, and n of the p channels (where n is 1) <n≤p인 정수) 번째 채널은 상기 첫 번째 채널을 상기 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 상기 하나 이상의 제3 필터를 생성할 수 있다.
[0009] The step of performing the above convolution operation can perform the above convolution operation using a hardware accelerator that performs a convolution operation in which the number of input channels is fixed to a preset number of input channels.
[0010] According to one embodiment, a convolution operation method includes, when the number of one or more first filters is less than a preset number of output channels, replicating one or more first filters q times (where q is a ratio of the number of output channels to the number of one or more first filters) to generate q·F second filters (where F is the number of one or more first filters); padding each channel of the q·F second filters with a preset value to generate q·F third filters, the width or height of which is increased by (q-1)·S (where S is a preset stride) more than that of the q·F second filters; and performing a convolution operation with a stride of q·S on input data using the q·F third filters.
[0011] The q·F third filters include q filters in which a specific filter among the one or more first filters is padded with the preset value, and the step of generating the q·F third filters includes: a first filter among the q filters is padded with (q-1)·S preset values sequentially at the rear end in the width or height direction in each channel of the specific filter, and m of the q filters (wherein m is 1) <m≤q인 정수) 번째 필터는 상기 첫 번째 필터의 각 채널을 상기 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 상기 q·F개의 제3 필터를 생성할 수 있다.
[0012] The step of performing the above convolution operation can perform the convolution operation using a hardware accelerator that performs a convolution operation in which the number of output channels is fixed to a preset number of output channels.
[0013] A convolutional computing device according to one embodiment comprises: one or more processors; And a memory storing one or more programs executed by the one or more processors, wherein the one or more processors, when the number of channels of the first input data is less than the preset number of input channels, rearrange the elements of each channel of the first input data to generate second input data whose width or height is reduced by 1 / p (where p is the ratio of the number of preset input channels to the number of channels of the first input data) times and whose number of channels is increased by p times compared to the first input data, pad each channel of the second input data with a preset value to generate third input data whose width or height is increased by p-1 compared to the first input data, duplicate each channel of one or more first filters whose number of channels is N / p (where N is the number of preset input channels) p times to generate one or more second filters whose number of channels is p·N, pad each channel of the one or more second filters with the preset value to generate one or more third filters whose width or height is K+p-1, and use the one or more third filters to generate the third input data. A convolution operation is performed on the data with a stride of S (where S is a preset stride).
[0014] The one or more processors generate p channels with a width or height reduced by 1 / p for each channel of the first input data, and among the p channels generated for the i-th channel of the first input data (where i is an integer 1≤i≤N / p), the j-th channel (where j is an integer 1≤j≤p) may include elements extracted from the j-th element at intervals of p in the width or height direction from the i-th channel.
[0015] The one or more processors may generate the third input data by padding the preset value so that p-1 of the preset values are included before and after each element included in each channel of the second input data in the width or height direction for each channel of the second input data.
[0016] The one or more second filters include p channels identical to specific channels of the one or more first filters, and the one or more processors sequentially pad a first channel of the p channels with p-1 preset values at the front end in the width or height direction, and n of the p channels, where n is 1. <n≤p인 정수) 번째 채널은 상기 첫 번째 채널을 상기 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 상기 하나 이상의 제3 필터를 생성할 수 있다.
[0017] The one or more processors may perform the convolution operation using a hardware accelerator that performs the convolution operation with the number of input channels fixed to a preset number of input channels.
[0018] According to one embodiment, a convolution operation device includes: one or more processors; and a memory storing one or more programs executed by the one or more processors, wherein the one or more processors, when the number of one or more first filters is less than a preset number of output channels, replicate the one or more first filters q times (where q is a ratio of the number of output channels to the number of the one or more first filters) to generate q·F second filters (where F is the number of the one or more first filters), pad each channel of the q·F second filters with a preset value to generate q·F third filters, the width or height of which is increased by (q-1)·S (where S is a preset stride) more than that of the q·F second filters, and perform a convolution operation with a stride of q·S on input data using the q·F third filters.
[0019] The q·F third filters include q filters in which a specific filter among the one or more first filters is padded with the preset value, and the one or more processors are configured such that a first filter among the q filters is padded with (q-1)·S preset values sequentially at the rear end in the width or height direction for each channel of the specific filter, and m of the q filters (wherein m is 1) <m≤q인 정수) 번째 필터는 상기 첫 번째 필터의 각 채널을 상기 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 상기 q·F개의 제3 필터를 생성할 수 있다.
[0020] The one or more processors may perform the convolution operation using a hardware accelerator that performs the convolution operation with the number of output channels fixed to a preset number of output channels.
[0021] According to the disclosed embodiments, efficient convolution operations are possible even when the input channel size or output channel size of the convolution operation is smaller than a preset size.
[0022] In addition, according to the disclosed embodiments, when using a hardware accelerator for a convolution operation, even if the input channel size or the output channel size of the convolution operation to be performed is smaller than the input or output channel size of the hardware accelerator, the convolution operation using the hardware accelerator is possible without changing the input or output channel size of the hardware accelerator.
[0023] Fig. 1 is a configuration diagram of a convolution operation device according to one embodiment.
[0024] FIGS. 2 to 5 are drawings for illustrative purposes only, illustrating an example of generating second input data and third input data according to one embodiment.
[0025] FIGS. 6 to 9 are drawings for illustrative purposes only, illustrating other examples of generation of second input data and third input data according to one embodiment.
[0026] FIG. 10 and FIG. 11 are drawings for illustrative purposes explaining an example of generating a second filter and a third filter according to one embodiment.
[0027] FIG. 12 and FIG. 13 are drawings for illustrative purposes to explain another example of generating a second filter and a third filter according to one embodiment.
[0028] Fig. 14 is a configuration diagram of a convolution operation device according to another embodiment.
[0029] FIG. 15 and FIG. 16 are drawings for illustrative purposes only, illustrating an example of the generation of a second filter and a third filter according to one embodiment.
[0030] FIG. 17 and FIG. 18 are drawings for illustrative purposes only, illustrating another example of generation of a second filter and a third filter according to one embodiment.
[0031] Fig. 19 is a flowchart of a convolution operation method according to one embodiment.
[0032] Fig. 20 is a flowchart of a convolution operation method according to another embodiment.
[0033] FIG. 21 is a block diagram illustrating a computing environment including a computing device according to one embodiment.
[0034] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, devices, and / or systems described herein. However, these are merely examples and the present invention is not limited thereto.
[0035] In describing embodiments of the present invention, if a detailed description of a known technology related to the present invention is judged to unnecessarily obscure the gist of the present invention, the detailed description will be omitted. In addition, the terms described below are terms defined in consideration of their functions in the present invention, and this may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification. The terminology used in the detailed description is only for the purpose of describing embodiments of the present invention and should not be limited in any way. Unless clearly used otherwise, the singular form includes the plural form. In this description, expressions such as "comprises" or "having" are intended to indicate certain features, numbers, steps, operations, elements, parts or combinations thereof, and should not be construed to exclude the presence or possibility of one or more other features, numbers, steps, operations, elements, parts or combinations thereof other than those described.
[0036] Fig. 1 is a configuration diagram of a convolution operation device according to one embodiment.
[0037] Referring to FIG. 1, a convolution operation device (100) according to one embodiment includes an input data conversion unit (110), a filter conversion unit (120), and a convolution operation unit (130).
[0038] In one embodiment, the input data conversion unit (110), the filter conversion unit (120), and the convolution operation unit (130) may be implemented using one or more physically separate devices, or may be implemented by one or more processors or a combination of one or more processors and software, and may not be clearly distinguished in specific operations, unlike the illustrated example.
[0039] The input data conversion unit (110) converts the first input data and the first filter to have the same number of channels as the preset number of input channels when the number of channels of the first input data and the first filter input for the convolution operation is less than the preset number of input channels.
[0040] Specifically, when the number of channels of the first input data is less than the preset number of input channels, the input data conversion unit (110) rearranges the elements of each channel of the first input data to generate second input data whose width or height is reduced by 1 / p (where p is the ratio of the preset number of input channels to the number of channels of the first input data) times and whose number of channels is increased by p times compared to the first input data.
[0041] At this time, according to one embodiment, the second input data may include p channels generated by rearranging elements of the i-th channel of the first input data (where i is an integer such that 1 ≤ i ≤ N / p, and N is a preset number of input channels), and among the p channels, the j-th channel (where j is an integer such that 1 ≤ j ≤ p) may include elements extracted from the j-th element in the width or height direction at intervals of p from the i-th channel.
[0042] Meanwhile, the input data conversion unit (110) pads each channel of the second input data with a preset value to generate third input data whose width or height is increased by p-1 compared to the first input data. At this time, the preset value may be, for example, 0, but is not necessarily limited thereto and may be set to a different value depending on the embodiment.
[0043] Specifically, according to one embodiment, the input data conversion unit (110) may generate third input data by padding each channel of the second input data with a preset value such that p-1 preset values are included before and after each element included in each channel of the second input data in the width or height direction. At this time, the preset value may be 0, but is not necessarily limited thereto, and may be set differently depending on the embodiment.
[0044] Specifically, FIGS. 2 to 5 are drawings for illustratively explaining an example of generating second input data and third input data according to one embodiment.
[0045] In FIGS. 2 to 5, for convenience of explanation, it is assumed that the preset number of input channels (N) is 8, the number of channels (M) of the first input data (210) is 4, and each channel (211, 212, 213, 214) of the first input data is one-dimensional data with a width (W) of 10. However, the preset number of input channels (N) and the first input data (210) are not necessarily limited to the illustrated example.
[0046] Referring to FIG. 2, since the ratio (p) of the preset input channel number (N) to the channel number (M) of the first input data (210) is 2 (i.e., p=N / M), the input data conversion unit (110) can generate two channels whose widths are reduced by 1 / 2 (i.e., 1 / p) from each channel (211, 212, 213, 214) of the first input data (210) to generate second input data (220) whose channel number is the same as the preset input channel number (N).
[0047] Specifically, the input data conversion unit (110) can generate a first channel (221) and a second channel (222) of the second input data (220) from a first channel (211) of the first input data (210), generate a third channel (223) and a fourth channel (224) of the second input data (220) from a second channel (212) of the first input data (210), generate a fifth channel (225) and a sixth channel (226) of the second input data (220) from the third channel (213) of the first input data (210), and generate a seventh channel (227) and an eighth channel (228) of the second input data (220) from the fourth channel (214) of the first input data (210).
[0048] At this time, as in the example illustrated in FIG. 3, the first channel (221) of the two channels (221, 222) of the second input data (220) generated from the first channel (211) of the first input data (210) may include five elements extracted at intervals of 1 in the width direction from the first element of the first channel (211) of the first input data (210). In addition, the second channel (222) of the two channels (221, 222) of the second input data (220) generated from the first channel (211) of the first input data (210) may include five elements extracted at intervals of 1 in the width direction from the second element of the first channel (211) of the first input data (210).
[0049] Similarly, as in the example illustrated in FIG. 4, the first channel (223) of the two channels (223, 224) of the second input data (220) generated from the second channel (212) of the first input data (210) may include five elements extracted at intervals of 1 in the width direction from the first element of the second channel (212) of the first input data (210). In addition, the second channel (224) of the two channels (223, 224) of the second input data (220) generated from the second channel (212) of the first input data (210) may include five elements extracted at intervals of 1 in the width direction from the second element of the second channel (212) of the first input data (210).
[0050] Meanwhile, referring to FIG. 5, the input data conversion unit (110) can generate third input data (230) including eight channels (231, 232, 233, 234, 235, 236, 237, 238) having a width (W'') of 11 (i.e., W''=W+p-1) by performing padding so that each element included in each channel (221, 222, 223, 224, 225, 226, 227, 228) of the second input data (220) includes 1 (i.e., p-1) of the preset value 0 in the width direction.
[0051] Meanwhile, FIGS. 6 to 9 are drawings for illustratively explaining other examples of generation of second input data and third input data according to one embodiment.
[0052] In FIGS. 6 to 9, for convenience of explanation, it is assumed that the preset number of input channels (N) is 8, the number of channels (M) of the first input data (210) is 2, and each channel (611, 612) of the first input data (610) is one-dimensional data with a width (W) of 12. However, the preset number of input channels (N) and the first input data (610) are not necessarily limited to the illustrated examples.
[0053] Referring to FIG. 6, since the ratio (p) of the preset input channel number (N) to the channel number (M) of the first input data (610) is 4 (i.e., p=N / M), the input data conversion unit (110) can generate four channels with a width (W') reduced by 1 / 4 (i.e., 1 / p) from each channel (611, 612) of the first input data (610) to generate second input data (620) including the same number of channels as the preset input channel number.
[0054] Specifically, the input data conversion unit (110) can generate a first channel (621), a second channel (622), a third channel (623), and a fourth channel (624) of the second input data (620) from a first channel (611) of the first input data (610), and can generate a fifth channel (625), a sixth channel (626), a seventh channel (627), and an eighth channel (628) of the second input data (620) from a second channel (612) of the first input data (610).
[0055] At this time, as in the example shown in FIG. 7, among the four channels (621, 622, 623, 624) of the second input data (620) generated from the first channel (611) of the first input data (610), the first channel (621) may include three elements extracted at three intervals in the width direction from the first element of the first channel (611) of the first input data (610), the second channel (622) may include three elements extracted at three intervals in the width direction from the second element of the first channel (611) of the first input data (610), the third channel (623) may include three elements extracted at three intervals in the width direction from the third element of the first channel (611) of the first input data (610), and the fourth channel (624) may include three elements extracted at three intervals in the width direction from the fourth element of the first channel (611) of the first input data (610). It can contain three elements extracted at intervals of three from the element in the width direction.
[0056] Similarly, as in the example illustrated in FIG. 8, among the four channels (625, 626, 627, 628) of the second input data (620) generated from the second channel (612) of the first input data (610), the first channel (625) may include three elements extracted at three intervals in the width direction from the first element of the second channel (612) of the first input data (610), the second channel (626) may include three elements extracted at three intervals in the width direction from the second element of the second channel (612) of the first input data (610), the third channel (627) may include three elements extracted at three intervals in the width direction from the third element of the second channel (612) of the first input data (610), and the fourth channel (628) may include three elements extracted at three intervals in the width direction from the fourth element of the second channel (612) of the first input data (610). It can contain three elements extracted at intervals of three from the element in the width direction.
[0057] Meanwhile, referring to FIG. 9, the input data conversion unit (110) can generate third input data (630) including eight channels (631, 632, 633, 634, 635, 636, 637, 638) having a width (W'') of 15 (i.e., W''=W+p-1) by performing padding so that three (i.e., p-1) 0s, which are preset values, are included in the width direction before and after each element included in each channel (621, 622, 623, 624, 625, 626, 627, 628) of the second input data (620).
[0058] Again, referring to FIG. 1, the filter transformation unit (120) duplicates p of each channel of one or more first filters having a channel count of N / p (i.e., M) to generate one or more second filters having a channel count of p·N, and pads each channel of the generated one or more second filters with a preset value to generate one or more third filters having a width or height of K+p-1.
[0059] At this time, according to one embodiment, one or more second filters may each include p channels identical to specific channels of one or more first filters, and the filter transformer (120) sequentially pads the first channel of the p channels with p-1 preset values at the front end in the width or height direction, and n of the p channels (where n is 1) <n≤p인 정수) 번째 채널은 p개의 채널 중 첫 번째 채널을 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 하나 이상의 제3 필터를 생성할 수 있다.
[0060] Specifically, FIGS. 10 and 11 are drawings for illustratively explaining an example of generating a second filter and a third filter according to one embodiment.
[0061] In the examples illustrated in FIGS. 10 and 11, for convenience of explanation, it is assumed that the first filter (1110) is composed of one filter with a channel number (M) of 4 and a width (K) of 3, that the padding preset value is 0, and that the preset number of input channels (N) is 8. However, the first filter (1110) and the preset value are not necessarily limited to the illustrated examples.
[0062] Referring to FIG. 10, the second filter (1120) may include eight channels (1121, 1122, 1123, 1123, 1124, 1125, 1126, 1127, 1128) generated by replicating two (i.e., p=N / M) of each channel (1111, 1112, 1113, 1114) of the first filter (1110).
[0063] Specifically, the first channel (1121) and the second channel (1122) of the second filter (1120) may be identical to the first channel (1111) of the first filter (1110), and the third channel (1123) and the fourth channel (1124) of the second filter (1120) may be identical to the second channel (1112) of the first filter (1110).
[0064] Additionally, the fifth channel (1125) and the sixth channel (1126) of the second filter (1120) may be identical to the third channel (1113) of the first filter (1110), and the seventh channel (1127) and the eighth channel (1128) of the second filter (1120) may be identical to the fourth channel (1114) of the first filter (1110).
[0065] Meanwhile, referring to FIG. 11, the filter conversion unit (120) can generate a third filter (1130) having a width (K') of 4 (i.e., K+p-1) by padding 0 to each channel (1111, 1112, 1113, 1114) of the second filter (1120).
[0066] Specifically, the first channel (1131) of the third filter (1130) may be in the form of the first channel (1121) of the second filter (1120) being padded with 0 at the front end in the width direction, and the second channel (1132) of the third filter (1130) may be in the form of the first channel (1131) of the third filter (1130) being cyclically shifted by 1 in the opposite direction in the width direction.
[0067] In addition, the third channel (1133) of the third filter (1130) may be in the form of the third channel (1123) of the second filter (1120) padded with 0 at the front end in the width direction, and the fourth channel (1134) of the third filter (1130) may be in the form of the third channel (1133) of the third filter (1130) being cyclically shifted by 1 in the opposite direction in the width direction.
[0068] Additionally, the fifth channel (1135) of the third filter (1130) may be in the form of the fifth channel (1125) of the second filter (1120) padded with 0 at the front end in the width direction, and the sixth channel (1136) of the third filter (1130) may be in the form of the fifth channel (1135) of the third filter (1130) being cyclically shifted by 1 in the opposite direction in the width direction.
[0069] The seventh channel (1137) of the third filter (1130) may be in the form of the seventh channel (1127) of the second filter (1120) padded with 0 at the front end in the width direction, and the eighth channel (1138) of the third filter (1130) may be in the form of the seventh channel (1137) of the third filter (1130) being cyclically shifted by 1 in the opposite direction in the width direction.
[0070] FIG. 12 and FIG. 13 are drawings for illustrative purposes only, illustrating another example of generation of a second filter and a third filter according to one embodiment.
[0071] In the examples illustrated in FIGS. 12 and 13, for convenience of explanation, it is assumed that the first filter (1210) is composed of one filter with a channel number (M) of 2 and a width (K) of 3, that the padding preset value is 0, and that the preset number of input channels (N) is 8. However, the first filter (1100) and the preset value are not necessarily limited to the illustrated examples.
[0072] Referring to FIG. 12, the second filter (1220) may include eight channels (1221, 1222, 1223, 1224, 1225, 1226, 1227, 1228) generated by replicating each channel (1211, 1212) of the first filter (1210) four times (i.e., p=N / M).
[0073] Specifically, the first channel (1221), the second channel (1222), the third channel (1223), and the fourth channel (1224) of the second filter (1220) may be identical to the first channel (1211) of the first filter (1210), and the fifth channel (1225), the sixth channel (1226), the seventh channel (1227), and the eighth channel (1228) of the second filter (1220) may be identical to the second channel (1212) of the first filter (1210), respectively.
[0074] Meanwhile, referring to FIG. 13, the filter conversion unit (120) can generate a third filter (1230) having a width (K') of 6 (i.e., K+p-1) by padding 0 to each channel of the second filter (1220).
[0075] Specifically, the first channel (1231) of the third filter (1230) may be in the form of 3 (i.e., p-1) 0s consecutively padded at the front end in the width direction of the first channel (1221) of the second filter (1220). In addition, the second channel (1232) of the third filter (1230) may be in the form of the first channel (1231) of the third filter (1230) being cyclically shifted by 1 in the opposite direction in the width direction, the third channel (1233) of the third filter (1230) may be in the form of the first channel (1231) of the third filter (1230) being cyclically shifted by 2 in the opposite direction in the width direction, and the fourth channel (1234) of the third filter (1230) may be in the form of the first channel (1231) of the third filter (1230) being cyclically shifted by 3 in the opposite direction in the width direction.
[0076] In addition, the fifth channel (1235) of the third filter (1230) may be in the form of three (i.e., p-1) 0s consecutively padded at the front end in the width direction of the fifth channel (1225) of the second filter (1220). In addition, the sixth channel (1236) of the third filter (1230) may be in the form of the fifth channel (1235) of the third filter (1230) being cyclically shifted by 1 in the opposite direction in the width direction, the seventh channel (1237) of the third filter (1230) may be in the form of the fifth channel (1235) of the third filter (1230) being cyclically shifted by 2 in the opposite direction in the width direction, and the eighth channel (1238) of the third filter (1230) may be in the form of the fifth channel (1235) of the third filter (1230) being cyclically shifted by 3 in the opposite direction in the width direction.
[0077] Referring again to FIG. 1, the convolution operation unit (130) performs a convolution operation with a stride of S (where S is a preset stride) on the third input data using the third filter.
[0078] At this time, according to one embodiment, the convolution operation unit (130) can perform the convolution operation using a hardware accelerator that performs the convolution operation with the number of input channels fixed to N.
[0079] As a specific example, if the first input data (210) and the first filter (1110) are the same as the examples shown in FIG. 2 and FIG. 10, respectively, and S=1, the convolution operation unit (130) can perform a convolution operation with a stride of 1 (i.e., S) on the third input (230) shown in FIG. 5 using the third filter (1130) shown in FIG. 11.
[0080] As another example, if the first input data (610) and the first filter (1210) are the same as the examples illustrated in FIG. 6 and FIG. 12, respectively, and S=1, the convolution operation unit (130) can perform a convolution operation with a stride of 1 (i.e., S) on the third input (630) illustrated in FIG. 9 using the third filter (1230) illustrated in FIG. 13.
[0081] Fig. 14 is a configuration diagram of a convolution operation device according to another embodiment.
[0082] Referring to FIG. 14, a convolution operation device (1400) according to one embodiment includes a filter transformation unit (1410) and a convolution operation unit (1420).
[0083] In one embodiment, the filter transformer (1410) and the convolution operation unit (1420) may be implemented using one or more physically separate devices, or may be implemented by one or more processors or a combination of one or more processors and software, and may not be clearly distinguished in specific operations, unlike the illustrated example.
[0084] The filter conversion unit (1410) converts the filters so that the number of filters is equal to the preset number of output channels when the number of filters for the convolution operation is less than the preset number of output channels.
[0085] Specifically, the filter conversion unit (1410) generates q·F second filters by replicating q of the first filters (where q is the ratio of the number of output channels (P) to the number of first filters (F)) when the number of one or more first filters (F) is less than the preset number of output channels (P). Thereafter, the filter conversion unit (1410) pads each channel of the generated q·F second filters with a preset value to generate q·F third filters whose width or height is increased by (q-1)·S (where S is a preset stride) compared to the second filters.
[0086] At this time, according to one embodiment, the q·F third filters may include q filters in which a preset value is padded to a specific filter among one or more first filters, and the filter transformation unit (1410) sequentially pads (q-1)·S preset values at the rear end of each channel of the specific filter in the width or height direction of the first filter among the q filters, and m (wherein m is 1) of the q filters <m≤q인 정수) 번째 필터는 첫 번째 필터의 각 채널을 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 q·F개의 제3 필터를 생성할 수 있다.
[0087] Specifically, FIGS. 15 and 16 are drawings for illustratively explaining an example of generating a second filter and a third filter according to one embodiment.
[0088] In the examples illustrated in FIGS. 15 and 16, for convenience of explanation, it is assumed that the first filter (1510) is composed of one (i.e., F=1) filter having a channel number (N) of 4 and a width (K) of 3, a padding preset value of 0, a preset output channel number (P) of 2, and a preset stride (S) of 1. However, the first filter (1510), the preset value, the preset output channel number (P), and the preset stride (S) are not necessarily limited to the illustrated examples.
[0089] Referring to FIG. 15, the second filter (1520) may include two (i.e., q·F) filters (1521, 1522) generated by replicating the first filter (1510). That is, the two filters (1521, 1522) included in the second filter (1520) are each identical to the first filter (1510).
[0090] Meanwhile, referring to FIG. 16, the filter conversion unit (1410) can generate a third filter (1530) including two filters (1531, 1532) each having a width (K') of 4 (i.e., K+(q-1)·S) by padding 1 (i.e., (q-1)·S) 0s to each filter (1521, 1522) included in the second filter (1520).
[0091] Specifically, the first filter (1531) of the two filters (1531, 1532) included in the third filter (1530) may be in a form in which each channel (1521-1, 1521-2, 1521-3, 1521-4) of the first filter (1521) of the two filters (1521, 1522) included in the second filter (1520) is padded with 0 at the rear in the width direction, and the second filter (1532) of the two filters (1531, 1532) included in the third filter (1530) may be in a form in which each channel (1531-1, 1531-2, 1531-3, 1531-4) of the first filter (1531) is cyclically shifted by 1 (i.e., S) in the width direction.
[0092] FIG. 17 and FIG. 18 are drawings for illustrative purposes only, illustrating another example of generation of a second filter and a third filter according to one embodiment.
[0093] In the examples illustrated in FIGS. 17 and 18, for convenience of explanation, it is assumed that the first filter (1710) is composed of two (i.e., F=2) filters (1711, 1712) having a channel number (N) of 4 and a width (K) of 3, a padding preset value is 0, a preset output channel number (P) is 4, and a preset stride (S) is 2. However, the first filter (1510), the preset value, the preset output channel number (P), and the preset stride (S) are not necessarily limited to the illustrated examples.
[0094] Referring to FIG. 17, the second filter (1720) may include four (i.e., q·F) filters, including two filters (1721, 1722) generated by duplicating the first filter (1711) among the two filters (1711, 1712) included in the first filter (1710) and two filters (1723, 1724) generated by duplicating the second filter (1712) among the two filters (1711, 1712) included in the first filter (1710). That is, among the four filters (1721, 1722, 1723, 1724) included in the second filter (1720), the first filter (1721) and the second filter (1722) are identical to the first filter (1711) among the two filters (1711, 1712) included in the first filter (1710), and among the four filters (1721, 1722, 1723, 1724) included in the second filter (1720), the third filter (1723) and the fourth filter (1724) are identical to the second filter (1712) among the two filters (1711, 1712) included in the first filter (1710).
[0095] Meanwhile, referring to FIG. 18, the filter conversion unit (1410) can generate a third filter (1730) including four filters (1731, 1732, 1733, 1734) each having a width (K') of 5 (i.e., K+(q-1)·S) by padding each filter (1721, 1722, 1723, 1724) included in the second filter (1720) with 2 (i.e., (q-1)·S) 0s.
[0096] Specifically, the first filter (1731) among the four filters (1731, 1732, 1733, 1734) included in the third filter (1730) may be padded with two 0s at the rear end in the width direction in each channel (1721-1, 1721-2, 1721-3, 1721-4) of the first filter (1721) among the four filters (1721, 1722, 1723, 1724) included in the second filter (1720), and the second filter (1732) among the four filters (1731, 1732, 1733, 1734) included in the third filter (1730) may be padded with two 0s at the rear end in the width direction in each channel (1731-1, 1731-2, 1731-3, 1731-4) may be in the form of a circular shift of 2 (i.e., S) in the width direction.
[0097] In addition, the third filter (1733) among the four filters (1731, 1732, 1733, 1734) included in the third filter (1730) may be padded with two 0s at the rear end in the width direction in each channel (1723-1, 1723-2, 1723-3, 1723-4) of the third filter (1723) among the four filters (1721, 1722, 1723, 1724) included in the second filter (1720), and the fourth filter (1734) among the four filters (1731, 1732, 1733, 1734) included in the third filter (1730) may be padded with two 0s at the rear end in the width direction in each channel (1733-1, 1733-2, 1733-3, 1733-4) may be in the form of a circular shift of 2 (i.e., S) in the width direction.
[0098] Referring again to FIG. 14, the convolution operation unit (1420) performs a convolution operation with a stride of q·S on the input data using the third filter.
[0099] At this time, according to one embodiment, the convolution operation unit (1420) may perform a convolution operation using a hardware accelerator that performs a convolution operation in which the number of output channels is fixed to P.
[0100] As a specific example, if the first filter (1510) is the same as the example illustrated in FIG. 15, and P=2 and S=1, the convolution operation unit (1410) can perform a convolution operation with a stride of 2 (i.e., q·S) on input data having a channel count of N using the third filter (1530) illustrated in FIG. 16.
[0101] As another example, if the first filter (1710) is the same as the example illustrated in FIG. 17, and P=4 and S=2, the convolution operation unit (1410) can perform a convolution operation with a stride of 4 (i.e., q·S) on input data having a channel count of N using the third filter (1730) illustrated in FIG. 18.
[0102] Fig. 19 is a flowchart of a convolution operation method according to one embodiment.
[0103] The method illustrated in FIG. 19 can be performed, for example, by the convolution operation device (100) illustrated in FIG. 1.
[0104] Referring to FIG. 19, the convolution operation device (100) determines whether the number of channels (M) of the first input data for the convolution operation is less than the preset number of input channels (N) (1910).
[0105] If the number of channels (M) of the first input data is less than the preset number of input channels (N), the convolution operation device (100) rearranges the elements of each channel of the first input data to generate second input data whose width or height is reduced by 1 / p (where p=N / M) times and whose number of channels is increased by p times compared to the first input data (1920).
[0106] At this time, according to one embodiment, the convolution operation device (100) can generate p channels with a width or height reduced by 1 / p for each channel of the first input data, and among the p channels generated for the i-th channel of the first input data (where i is an integer 1≤i≤N / p), the j-th channel (where j is an integer 1≤j≤p) can include elements extracted from the j-th element in the width or height direction at intervals of p in the i-th channel.
[0107] Thereafter, the convolution operation device (100) pads each channel of the second input data with a preset value to generate third input data whose width or height is increased by p-1 compared to the first input data (1930).
[0108] At this time, according to one embodiment, the convolution operation device (100) can generate third input data by padding preset values so that p-1 preset values are included in the width or height direction before and after each element included in each channel of the second input data.
[0109] Thereafter, the convolution operation unit (100) replicates p of each channel of one or more first filters having N / p (i.e., M) channels to generate one or more second filters having p·N channels (1940).
[0110] Thereafter, the convolution operation unit (100) pads each channel of one or more generated second filters with preset values to generate one or more third filters having a width or height of K+p-1 (1950).
[0111] At this time, according to one embodiment, one or more second filters may include p channels identical to specific channels of one or more first filters, and the convolution operation device (100) sequentially pads the first channel of the p channels with p-1 preset values at the front end in the width or height direction, and n of the p channels (where n is 1) <n≤p인 정수) 번째 채널이 p개의 채널 중 첫 번째 채널을 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 하나 이상의 제3 필터를 생성할 수 있다.
[0112] Thereafter, the convolution operation device (100) performs a convolution operation with a stride of S (where S is a preset stride) on the third input data using one or more third filters generated (1960).
[0113] At this time, according to one embodiment, the convolution operation device (100) can perform a convolution operation using a hardware accelerator that performs a convolution operation in which the number of input channels is fixed to N.
[0114] Meanwhile, in the flowchart illustrated in FIG. 19, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into sub-steps and performed, or one or more steps not illustrated may be added and performed.
[0115] Fig. 20 is a flowchart of a convolution operation method according to another embodiment.
[0116] The method illustrated in FIG. 20 can be performed, for example, by a convolution operation device (1400) illustrated in FIG. 14.
[0117] Referring to FIG. 20, a convolution operation device (1400) determines whether the number (F) of one or more first filters for convolution operation is less than the preset number of output channels (P) (2010).
[0118] If the number of one or more first filters (F) is less than the preset number of output channels (P), the convolution operation unit (1400) replicates one or more first filters q times (where q = P / F) to generate q·F second filters (2020).
[0119] Thereafter, the convolution operation unit (1400) pads each channel of the q·F second filters with a preset value to generate q·F third filters whose width or height is increased by (q-1)·S (where S is a preset stride) compared to the q·F second filters (2030).
[0120] At this time, according to one embodiment, the q·F third filters may include q filters in which a preset value is padded to a specific filter among one or more first filters, and the convolution operation unit (1400) may sequentially pad the first filter among the q filters with (q-1)·S preset values at the rear end in the width or height direction of each channel of the specific filter, and m of the q filters (wherein m is 1) <m≤q인 정수) 번째 필터는 첫 번째 필터의 각 채널을 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 q·F개의 제3 필터를 생성할 수 있다.
[0121] Thereafter, the convolution operation device (1400) performs a convolution operation with a stride of q·S on the input data using the generated q·F third filters (2040).
[0122] At this time, according to one embodiment, the convolution operation device (100) can perform a convolution operation using a hardware accelerator that performs a convolution operation in which the number of output channels is fixed to P.
[0123] Meanwhile, in the flowchart illustrated in FIG. 20, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into sub-steps and performed, or one or more steps not illustrated may be added and performed.
[0124] Figure 21 is a block diagram illustrating a computing environment including a computing device according to one embodiment. In the illustrated embodiment, each component may have different functions and capabilities other than those described below, and may include additional components other than those described below.
[0125] The illustrated computing environment (10) includes a computing device (12). The computing device (12) may be one or more components included in a convolution operation device (100, 1400) according to one embodiment.
[0126] A computing device (12) includes one or more processors (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) may cause the computing device (12) to operate according to the exemplary embodiments mentioned above. For example, the processor (14) may execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, and the computer-executable instructions, when executed by the processor (14), may be configured to cause the computing device (12) to perform operations according to the exemplary embodiments. Meanwhile, according to one embodiment, the one or more processors (14) may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit, but are not necessarily limited thereto.
[0127] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by the processor (14). In one embodiment, the computer-readable storage medium (16) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that can be accessed by the computing device (12) and store desired information, or a suitable combination thereof.
[0128] A communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and computer-readable storage media (16).
[0129] The computing device (12) may also include one or more input / output interfaces (22) and one or more network communication interfaces (26) that provide interfaces for one or more input / output devices (24). The input / output interfaces (22) and the network communication interfaces (26) are connected to the communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) via the input / output interfaces (22). Exemplary input / output devices (24) may include input devices such as pointing devices (such as a mouse or a trackpad), a keyboard, a touch input device (such as a touchpad or a touchscreen), a voice or sound input device, various types of sensor devices and / or photographing devices, and / or output devices such as display devices, printers, speakers and / or network cards. The exemplary input / output devices (24) may be included within the computing device (12) as a component constituting the computing device (12), or may be connected to the computing device (12) as a separate device distinct from the computing device (12).
[0130] Meanwhile, embodiments of the present invention may include a program for performing the methods described herein on a computer, and a computer-readable recording medium including the program. The computer-readable recording medium may include program commands, local data files, local data structures, etc., alone or in combination. The medium may be specially designed and configured for the present invention, or may be one commonly used in the field of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of the program may include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0131] While representative embodiments of the present invention have been described in detail above, those skilled in the art will appreciate that various modifications to the above-described embodiments are possible without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be determined not only by the claims set forth below but also by equivalents thereof.
Claims
1. One or more processors, and A method performed on a computing device having a memory storing one or more programs executed by one or more processors, A step of rearranging elements of each channel of the first input data to generate second input data whose width or height is reduced by 1 / p (wherein p is a ratio of the number of preset input channels to the number of channels of the first input data) times and whose number of channels is increased by p times compared to the first input data, when the number of channels of the first input data is less than the preset number of input channels; A step of padding each channel of the second input data with a preset value to generate third input data whose width or height is increased by p-1 compared to the first input data; A step of generating one or more second filters having p·N channels by replicating p of each channel of one or more first filters having N / p channels (where N is the preset number of input channels); generating one or more third filters having a width or height of K+p-1 by padding each channel of the one or more second filters with the preset value; and A convolution operation method, comprising the step of performing a convolution operation with a stride of S (wherein S is a preset stride) on the third input data using the one or more third filters.
2. In claim 1, The step of generating the second input data generates p channels with a width or height reduced by 1 / p times for each channel of the first input data, A convolution operation method, wherein the jth channel (jth channel, where j is an integer where 1≤j≤p) among the p channels generated for the ith channel (where i is an integer where 1≤i≤N / p) of the first input data includes elements extracted from the jth element at intervals of p in the width or height direction from the ith channel.
3. In claim 1, A convolution operation method in which the step of generating the third input data generates the third input data by padding the preset value so that p-1 preset values are included before and after each element included in each channel of the second input data in the width or height direction for each channel of the second input data.
4. In claim 1, wherein said one or more second filters include p channels identical to specific channels of said one or more first filters, The step of generating the one or more third filters comprises: sequentially padding the first channel of the p channels with p-1 preset values at the front end in the width or height direction, and n of the p channels (where n is 1). <n≤p인 정수) 번째 채널은 상기 첫 번째 채널을 상기 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 상기 하나 이상의 제3 필터를 생성하는, 합성곱 연산 방법.
5. In claim 1, A convolution operation method, wherein the step of performing the above convolution operation performs the convolution operation using a hardware accelerator that performs a convolution operation in which the number of input channels is fixed to a preset number of input channels.
6. One or more processors, and A method performed on a computing device having a memory storing one or more programs executed by one or more processors, If the number of one or more first filters is less than the preset number of output channels, a step of duplicating the one or more first filters q times (where q is the ratio of the number of output channels to the number of the one or more first filters) to generate q F second filters (where F is the number of the one or more first filters); A step of padding each channel of the q·F second filters with a preset value to generate q·F third filters whose width or height is increased by (q-1)·S (where S is a preset stride) more than that of the q·F second filters; and A convolution operation method, comprising the step of performing a convolution operation with a stride of q·S on input data using the q·F third filters.
7. In claim 6, The q·F third filters include q filters in which a specific filter among the one or more first filters is padded with the preset value, The step of generating the q·F third filters comprises: a first filter among the q filters is sequentially padded with (q-1)·S preset values at the end in the width or height direction of each channel of the specific filter, and m of the q filters (where m is 1) <m≤q인 정수) 번째 필터는 상기 첫 번째 필터의 각 채널을 상기 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 상기 q·F개의 제3 필터를 생성하는, 합성곱 연산 방법.
8. In claim 6, A convolution operation method, wherein the step of performing the above convolution operation performs the convolution operation using a hardware accelerator that performs a convolution operation in which the number of output channels is fixed to a preset number of output channels.
9. One or more processors; and A memory storing one or more programs executed by said one or more processors, One or more of the above processors, If the number of channels of the first input data is less than the preset number of input channels, the elements of each channel of the first input data are rearranged to generate second input data whose width or height is reduced by 1 / p (where p is the ratio of the number of channels of the first input data to the number of preset input channels) times and whose number of channels is increased by p times, Generate third input data whose width or height is increased by p-1 compared to the first input data by padding each channel of the second input data with a preset value, One or more second filters having p·N channels are generated by replicating p of each channel of one or more first filters having N / p channels (where N is the preset number of input channels), Generating one or more third filters having a width or height of K+p-1 by padding each channel of the one or more second filters with the preset value, A convolution operation device that performs a convolution operation with a stride of S (where S is a preset stride) on the third input data using one or more of the third filters.
10. In claim 9, The one or more processors generate p channels, each of which has a width or height reduced by 1 / p times, for each channel of the first input data, A convolution operation device, wherein the jth channel (jth integer where 1≤j≤p) among the p channels generated for the ith channel (where i is an integer where 1≤i≤N / p) of the first input data includes elements extracted from the jth element at p intervals in the width or height direction from the ith channel.
11. In claim 9, A convolution operation unit in which the one or more processors generate the third input data by padding the preset value so that p-1 preset values are included before and after each element included in each channel of the second input data in the width or height direction for each channel of the second input data.
12. In claim 9, wherein said one or more second filters include p channels identical to specific channels of said one or more first filters, The one or more processors are configured such that the first channel of the p channels is sequentially padded with p-1 preset values at the front end in the width or height direction, and n of the p channels, where n is 1. <n≤p인 정수) 번째 채널은 상기 첫 번째 채널을 상기 폭 또는 높이 방향의 반대 방향으로 n-1만큼 순환 시프트한 형태가 되도록 상기 하나 이상의 제3 필터를 생성하는, 합성곱 연산 장치.
13. In claim 9, A convolution operation device in which the one or more processors perform the convolution operation using a hardware accelerator that performs the convolution operation in which the number of input channels is fixed to a preset number of input channels.
14. One or more processors; and A memory storing one or more programs executed by said one or more processors, One or more of the above processors, If the number of one or more first filters is less than the preset number of output channels, the one or more first filters are duplicated q times (where q is the ratio of the number of output channels to the number of the one or more first filters) to generate q·F second filters (where F is the number of the one or more first filters). By padding each channel of the above q·F second filters with a preset value, q·F third filters are generated whose width or height is increased by (q-1)·S (where S is a preset stride) more than the q·F second filters, A convolution operation unit that performs a convolution operation with a stride of q·S on input data using the q·F third filters.
15. In claim 14, The q·F third filters include q filters in which a specific filter among the one or more first filters is padded with the preset value, The one or more processors are configured such that the first filter among the q filters is sequentially padded with (q-1)·S preset values at the end in the width or height direction for each channel of the specific filter, and m of the q filters, where m is 1. <m≤q인 정수) 번째 필터는 상기 첫 번째 필터의 각 채널을 상기 폭 또는 높이 방향으로 S만큼 순환 시프트한 형태가 되도록 상기 q·F개의 제3 필터를 생성하는, 합성곱 연산 장치.
16. In claim 14, A convolution operation unit in which the one or more processors perform the convolution operation using a hardware accelerator that performs the convolution operation in which the number of output channels is fixed to a preset number of output channels.
Citation Information
Patent Citations
METHOD AND SYSTEM FOR reducing computations in a neural network
KR1020160143505A
Cocktail medium-added composition for in vitro culture and activation of human natural killer cells
KR1020230113055A
System and method for colposcopy by using a deep learning model
KR1020230131687A
Method for convolution operation redution and system for performing the same
KR102034659B1
Method for providing transportation service which supports earning points of passenger and apparatuses using the same
KR102249082B1