Method and apparatus for convolution operation using kernel shape control

By generating sub-kernels and using a mask to exclude weights, the convolution operation method effectively compresses the weight matrix, addressing the challenge of reduced inference accuracy in existing SDK mapping techniques.

WO2025116294A1PCT designated stage expired Publication Date: 2025-06-05RES & BUSINESS FOUND SUNGKYUNKWAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016234
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-10-24
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing convolution operation methods, particularly when using the SDK mapping technique, struggle to effectively compress the weight matrix without relying on pattern-based pruning or other pruning methods, which can lead to reduced inference accuracy.

Method used

The method involves generating a base kernel by adding padding to the original kernel, then creating smaller sub-kernels from the base kernel. A mask is used to determine which weights to exclude from the convolution operation, allowing for effective weight matrix compression without pruning.

Benefits of technology

This approach enables efficient weight matrix compression and improves model accuracy when using the SDK mapping technique, reducing hardware waste and energy consumption while maintaining inference accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016234_05062025_PF_FP_ABST
    Figure KR2024016234_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A convolution operation method using kernel shape control according to a first aspect of the present invention includes the steps of: generating a base kernel by adding padding to a kernel including a weight; generating one or more sub-kernels smaller than the base kernel by using a partial area of the base kernel; and performing a convolution operation on the basis of the one or more sub-kernels and an input feature matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Convolution operation method and device using kernel shape control

[0001] The present invention relates to a convolution operation method and device using kernel shape control.

[0002] Processing in Memory (PIM) is a computing architecture where computations are performed in memory. It is attracting attention as a suitable computing architecture for real-time inference using deep learning models on edge or mobile devices. When computations are performed in PIM, weights remain in memory, inputs are passed, computations are performed in memory, and only the resulting values ​​are output.

[0003] SDK (shift and duplicate kernel), a cutting-edge mapping technique, shifts and duplicates kernel input data one by one across columns that don't share the same weights, thereby generating multiple output values ​​in a single cycle. Because SDK reuses input data and kernel weights, it uses memory more efficiently and reduces computational complexity compared to existing mapping methods.

[0004] Model lightweighting, or pruning, refers to a method of weight reduction that retains important parameters while pruning unimportant ones during model training. Pattern-based pruning, one of the model lightweighting techniques, applies a set pattern to all kernels to prune weights.

[0005] Pattern-based pruning is a method of removing weights using patterns. The more pruning is performed, the more the weight matrix can be compressed. However, at the same time, the inference accuracy decreases, and the compression ratio varies depending on which pattern is used.

[0006] In addition, pattern-based pruning can compress the size of the weight matrix, but it is difficult to apply it to the SDK mapping method, so there is a problem that the weight matrix cannot be effectively compressed when the SDK mapping method is used.

[0007] The problem to be solved by the present invention is to provide a convolution operation method and device using kernel shape control that can provide a weight matrix compression effect without using pattern-based pruning or other pruning methods when using a mapping technique.

[0008] However, the problems to be solved by the present invention are not limited to those mentioned above, and other problems to be solved that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.

[0009] A convolution operation method using kernel shape control according to a first aspect of the present invention includes a step of generating a base kernel by adding padding to a kernel including a weight, a step of generating one or more sub kernels smaller in size than the base kernel by using a portion of the base kernel, and a step of performing a convolution operation based on the one or more sub kernels and an input feature matrix.

[0010] In the step of performing the above convolution operation, the convolution operation can be performed simultaneously based on each sub-kernel and the input feature matrix.

[0011] All weight values ​​within the above padding can be 0.

[0012] The above kernel may be a two-dimensional matrix expressed as M by N (M and N are natural numbers). In addition, the base kernel may be a two-dimensional matrix expressed as (M+2) by (N+2).

[0013] The one or more sub-kernels may be a two-dimensional matrix expressed as (M+1) by (N+1). In this case, they may be generated by removing at least one of the first row and the (M+1)th row of the base kernel or at least one of the first column and the (N+1)th column of the base kernel.

[0014] A convolution operation method using kernel shape control according to the first aspect of the present invention may further include a step of generating a mask, which is a two-dimensional matrix expressed as (M+1) by (N+1), and a step of compressing one or more sub-kernels by determining weights to be excluded from the convolution operation through the mask.

[0015] The above mask may be composed of a number of non-zero components existing at the same position of the one or more sub-kernels, and a ratio of the number of non-zero components calculated for each position of a two-dimensional matrix expressed as (M+1) by (N+1) is equal to or less than a predetermined ratio, and 0 otherwise.

[0016] The weights excluded from the above convolution operation may be weights within each sub-kernel that exist at the same location as the location of 0 within the mask.

[0017] In the step of performing the convolution operation, the convolution operation can be performed based on the compressed one or more sub-kernels and the input feature matrix.

[0018] The above input feature matrix may be a two-dimensional matrix expressed as (M+1) by (N+1). In this case, the convolution operation method using kernel shape control according to the first aspect of the present invention may further include a step of compressing the input feature matrix by determining input features to be excluded from the convolution operation through the mask.

[0019] In the step of performing the convolution operation, the convolution operation can be performed based on the compressed one or more sub-kernels and the compressed input feature matrix.

[0020] A convolution operation device according to a second aspect of the present invention includes at least one memory capable of storing computer-executable instructions, and a processor that generates a base kernel by adding padding to a kernel including weights by executing the instructions, generates one or more sub-kernels smaller in size than the base kernel by using a portion of the base kernel, and performs a convolution operation based on the one or more sub-kernels and an input feature matrix.

[0021] A non-transitory computer-readable recording medium storing computer-executable instructions according to a third aspect of the present invention, wherein the computer-executable instructions, when executed by a processor, cause the processor to perform a method comprising the steps of: generating a base kernel by adding padding to a kernel including weights; generating one or more sub-kernels smaller in size than the base kernel by using a portion of the base kernel; and performing a convolution operation based on the one or more sub-kernels and an input feature matrix.

[0022] A computer program stored in a non-transitory computer-readable recording medium according to a fourth aspect of the present invention, wherein the computer program includes instructions for causing the processor to perform a method, the method comprising the steps of: generating a base kernel by adding padding to a kernel including weights; generating one or more sub-kernels smaller in size than the base kernel by using a portion of the base kernel; and performing a convolution operation based on the one or more sub-kernels and an input feature matrix, when executed by the processor.

[0023] According to the present invention, when using a mapping technique, it is possible to effectively compress a weight matrix and improve the accuracy of a model without using pattern-based pruning or other pruning methods.

[0024] The effects that can be obtained from the present invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0025] FIG. 1 is a flowchart exemplarily showing a convolution operation method using kernel shape control according to the first aspect of the present invention.

[0026] FIG. 2 is a block diagram exemplarily showing a convolution operation device according to the second aspect of the present invention.

[0027] Figure 3 is a block diagram exemplifying the function of a convolution operation program using kernel shape control.

[0028] Figure 4 is an example diagram showing a convolutional layer mapping method.

[0029] Figure 5 is an example diagram showing the application of pattern-based pruning when using the SDK mapping method.

[0030] FIG. 6 is an exemplary diagram showing how to omit weights from operations by applying a mask in a convolution operation method using kernel shape control according to the present invention.

[0031] FIG. 7 is an exemplary diagram showing a method for determining a row to omit weights in a convolution operation method using kernel shape control according to the present invention.

[0032] Figure 8 is an exemplary diagram showing components of a convolution operation method using kernel shape control according to the present invention.

[0033] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined solely by the scope of the claims.

[0034] When describing embodiments of the present invention, detailed descriptions of known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the invention. Furthermore, the terms described below are defined in light of their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0035] The terms used in this specification will be briefly explained, and the present invention will be described in detail.

[0036] The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present invention. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this invention should not be defined simply as names, but rather based on their inherent meanings and the overall content of the present invention.

[0037] When a part of a specification is said to 'include' a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise stated.

[0038] Also, the term 'part' used in the specification means a software or hardware component such as an FPGA or ASIC, and the 'part' performs certain functions. However, the 'part' is not limited to software or hardware. The 'part' may be configured to reside on an addressable storage medium or may be configured to play one or more processors. Thus, as an example, the 'part' includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and 'parts' may be combined into a smaller number of components and 'parts' or further separated into additional components and 'parts'.

[0039] Below, with reference to the attached drawings, an embodiment of the present invention is described in detail so that a person having ordinary skill in the art to which the present invention pertains can easily practice it.

[0040] FIG. 1 is a flowchart exemplarily showing a convolution operation method using kernel shape control according to the first aspect of the present invention.

[0041] Hereinafter, the convolution operation method using kernel shape control will be described on the assumption that it is performed by a convolution operation device.

[0042] In addition, for convenience of explanation, the present specification describes the convolution operation method using kernel shape control according to the present invention as being applied to PIM (processing-in-memory), but is not limited thereto. In other words, the present invention can be applied to any computing device to which the convolution operation method is applied.

[0043] As shown in FIG. 1, a convolution operation method using kernel shape control according to a first aspect of the present invention includes a step (S100) of generating a base kernel by adding padding to a kernel including a weight, a step (S110) of generating one or more sub kernels smaller in size than the base kernel by using a portion of the base kernel, and a step (S120) of performing a convolution operation based on the one or more sub kernels and an input feature matrix.

[0044] An input feature matrix may refer to a matrix representing features corresponding to signals including images, voices, natural language, etc.

[0045] Weights are used in the convolution operation of the inference process using a deep learning model and may be predetermined during the training process of the deep learning model.

[0046] A kernel is a numerical matrix used in a convolution operation, and may be designed to detect features of an input feature matrix. A kernel may contain one or more weights.

[0047] For a convolution operation, the output of the convolution operation can be calculated by multiplying the corresponding parts of the kernel and the input feature matrix and then adding the results of the multiplication.

[0048] Padding can mean creating a new matrix by filling the outer edges of a matrix with a specified number of predefined values. The predefined values ​​can be 0. Applying padding can prevent the output feature map size from shrinking.

[0049] A base kernel can be a collection of one or more subkernels. The weights contained in the base kernel can be determined by the subkernels and the column-dependent mask. The size of the base kernel can be larger than the kernel size. The base kernel can be created by adding padding to the outer edges of the kernel. For example, the base kernel can be created by adding padding with a value of 0 to the upper, lower, left, and right edges of the kernel.

[0050] A subkernel can represent a kernel mapped to each column of a PIM array. A subkernel can be generated using a portion of the base kernel. For example, a subkernel can be obtained by slicing rows or columns of the base kernel. Each subkernel can perform operations independent of the input features. Using subkernels, we can implement the feature of the SDK mapping method, where the same kernel is mapped to each column.

[0051] A column-dependent mask (CDM) can be applied individually to the same kernel to exclude certain elements mapped to the memory cells of each row from the computation, thereby omitting the computation of the corresponding column. By using a column-dependent mask, an independent pattern mask can be applied to the SDK of each kernel mapped to the PIM array. In other words, the column-dependent mask can be used to independently determine the weights that do not participate in the computation for each kernel. Furthermore, like a kernel, a column-dependent mask can be applied to the input feature matrix to determine which input features do not participate in the computation.

[0052] The fields in which the above-described input feature matrix, kernel, mask, etc. are used are examples and are not limited thereto. That is, the present invention can be applied to all fields in which convolution operations are used.

[0053] FIG. 2 is a block diagram exemplarily showing a convolution operation device according to the second aspect of the present invention.

[0054] As shown in FIG. 2, the convolution operation device (200) may include an input unit (210), an output unit (220), a processor (230), a memory (240), and a communication unit (260).

[0055] Hereinafter, for the convenience of explanation, the convolution operation device (200) is described as an example in which an input unit (210), an output unit (220), a processor (230), a memory (240), and a communication unit (260) are included, but the present invention is not limited thereto. That is, each unit configuration can interact with the convolution operation device (200) from outside the convolution operation device (200).

[0056] The input unit (210) may be a hardware device that can directly input commands, information, etc. used to control the convolution operation device (200) through a user interface (e.g., keyboard, touch input, voice input, etc.).

[0057] In one embodiment, the input unit (210) may receive information required for a convolution operation from a user. Specifically, the user may input information including an input feature matrix, a kernel including weights, information related to a base kernel, information related to a sub-kernel, information related to a column-dependent mask, information related to a deep learning model, and convolution operation conditions through the input unit (210).

[0058] The output unit (220) can provide information including an input feature matrix, a kernel including weights, information related to a base kernel, information related to a sub-kernel, information related to a heat-dependent mask, information related to a deep learning model, convolution operation conditions, and convolution operation results to a user as visual information through an interface or display device.

[0059] The processor (230) can control the overall operation of the convolution operation device (200) to perform the present invention.

[0060] The processor (230) can load the convolution operation program (250) using kernel shape control and the information necessary for executing the convolution operation program (250) using kernel shape control from the memory (240) to execute the convolution operation program (250) using kernel shape control.

[0061] The processor (230) can control to store data received from an external device through the communication unit (260) in the memory (240). In addition, the processor (230) can control to transmit information including an input feature matrix, a kernel including weights, information related to a base kernel, information related to a sub-kernel, information related to a column-dependent mask, information related to a deep learning model, convolution operation conditions, and a convolution operation result to an external device through the communication unit (260).

[0062] The processor (230) may refer to a processing device such as a microprocessor, a central processing unit (CPU), a graphic processing unit (GPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a micro controller unit (MCU), but is not limited to the above-described embodiment.

[0063] The memory (240) can store information necessary for executing a convolution operation program (250) using kernel shape control and a convolution operation program (250) using kernel shape control. In addition, the memory (240) can also store processing results by the processor (230).

[0064] A convolution operation program (250) using kernel shape control may mean software including commands programmed to perform a method according to the present invention.

[0065] The memory (240) can store information including an input feature matrix, a kernel including weights, information related to a base kernel, information related to a sub-kernel, information related to a column-dependent mask, information related to a deep learning model, convolution operation conditions, and convolution operation results. In addition, the memory (240) can store information received from an external device via a communication unit (260).

[0066] Memory (240) may refer to a computer-readable recording medium, such as a hard disk, a magnetic media such as a floppy disk and a magnetic tape, an optical media such as a CD-ROM or a DVD, a magneto-optical media such as a floptical disk, a hardware device specifically configured to store and execute program instructions such as a flash memory, but is not limited to the above-described embodiment.

[0067] The communication unit (260) may be a wireless communication module capable of performing wireless communication by adopting a communication method such as CDMA, GSM, W-CDMA, TD-SCDMA, WiBro, LTE, EPC, 5G, wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD (Wi-Fi Direct), UWB (Ultra Wide Band), infrared communication (IrDA; infrared data association), BLE (Bluetooth Low Energy), or NFC (Near Field Communication), but is not limited to the above-described embodiment.

[0068] In addition, information input and output through the input unit (210) and output unit (220), information stored in the memory (240), and information transmitted and received through the communication unit (260) include all information related to the present invention, and are not limited to the above-described embodiment.

[0069] The function or operation of the convolution operation program (250) using kernel shape control will be examined in detail with reference to FIG. 3.

[0070] Figure 3 is a block diagram exemplifying the function of a convolution operation program using kernel shape control.

[0071] As shown in Fig. 3, the convolution operation program (250) using kernel shape control may include a base kernel generation unit (310), a sub kernel generation unit (320), a mask generation unit (330), a matrix compression unit (340), and a convolution operation execution unit (350). The base kernel generation unit (310), the sub kernel generation unit (320), the mask generation unit (330), the matrix compression unit (340), and the convolution operation execution unit (350) are exemplary divisions of the functions of the convolution operation program (250) using kernel shape control, and are not limited thereto.

[0072] According to an embodiment, the functions of the base kernel generation unit (310), the sub kernel generation unit (320), the mask generation unit (330), the matrix compression unit (340), and the convolution operation execution unit (350) can be merged / separated and implemented as a series of commands included in one program.

[0073] The base kernel generation unit (310), sub kernel generation unit (320), mask generation unit (330), matrix compression unit (340), and convolution operation execution unit (350) may be implemented by a processor (230), and may mean a data processing device built into hardware having a physically structured circuit to perform a function expressed by a code or command included in a convolution operation program (250) using kernel shape control stored in a memory (240).

[0074] The base kernel generation unit (310) can generate a base kernel by adding padding to a kernel containing weights. A kernel is a numeric matrix used in a convolution operation and may be designed to detect characteristics of an input feature matrix. A kernel may include one or more weights.

[0075] Padding can mean creating a new matrix by filling the outer edges of a matrix with a specified number of predefined values. The predefined values ​​can be 0. That is, all weight values ​​within the padding can be 0. Applying padding can prevent the output feature map size from decreasing.

[0076] A base kernel can be a collection of one or more subkernels. The weights contained in the base kernel can be determined by the subkernels and the column-dependent mask. The size of the base kernel can be larger than the kernel size. The base kernel can be created by adding padding to the outer edges of the kernel. For example, the base kernel can be created by adding padding with a value of 0 to the upper, lower, left, and right edges of the kernel.

[0077] The kernel can be a two-dimensional matrix expressed as M by N (M and N are natural numbers).

[0078] The base kernel can be a two-dimensional matrix expressed as (M+2) by (N+2).

[0079] The subkernel generation unit (320) can generate one or more subkernels smaller than the base kernel by using a portion of the base kernel. The subkernels may represent kernels mapped to each column of the PIM array. The subkernels may be generated by using a portion of the base kernel. For example, the subkernels may be obtained by slicing rows or columns of the base kernel. Each subkernel can perform operations independent of the input features. By using the subkernels, the feature of the SDK mapping method in which the same kernel is mapped to each column can be implemented.

[0080] A subkernel can be a two-dimensional matrix expressed as (M+1) by (N+1).

[0081] A subkernel may be generated by removing at least one of the first row and the (M+1)th row of the base kernel or at least one of the first column and the (N+1)th column of the base kernel.

[0082] The mask generation unit (330) can generate a column dependent mask (CDM).

[0083] Column-dependent masks can be applied individually to the same kernel to exclude only certain elements mapped to memory cells in each row from computation, thereby omitting computations on the corresponding columns. By using column-dependent masks, an independent pattern mask can be applied to the SDK of each kernel mapped to the PIM array. In other words, column-dependent masks can be used to independently determine weights that do not participate in computation for each kernel. Furthermore, column-dependent masks can be applied to the input feature matrix, similar to kernels, to determine which input features do not participate in computation.

[0084] The mask generation unit (330) can generate a mask that is a two-dimensional matrix expressed as (M+1) by (N+1).

[0085] The mask generated by the mask generation unit (330) may be composed of a number of non-zero components existing at the same position of one or more sub-kernels, and a ratio of the number of non-zero components calculated for each position of a two-dimensional matrix expressed as (M+1) by (N+1) is equal to or less than a predetermined ratio, and 0 otherwise. A detailed description related to this will be given in Fig. 7.

[0086] The matrix compression unit (340) can compress one or more sub-kernels by determining weights to be excluded from the convolution operation through a mask. At this time, the weights to be excluded from the convolution operation may be weights within each sub-kernel that exist at the same location as the location of 0 within the mask.

[0087] The matrix compression unit (340) can compress the input feature matrix by determining input features to be excluded from the convolution operation through a mask.

[0088] The convolution operation performing unit (350) can perform a convolution operation based on one or more sub-kernels and an input feature matrix.

[0089] An input feature matrix may refer to a matrix representing features corresponding to signals including images, voices, natural language, etc. The input feature matrix may be a two-dimensional matrix expressed as (M+1) by (N+1).

[0090] The convolution operation performing unit (350) can calculate the output of the convolution operation by multiplying the corresponding parts of the kernel and the input feature matrix and then adding the multiplication results for the convolution operation.

[0091] The convolution operation performing unit (350) can perform convolution operations simultaneously based on each sub-kernel and input feature matrix.

[0092] The convolution operation performing unit (350) can perform a convolution operation based on one or more sub-kernels and an input feature matrix compressed by the matrix compression unit (340).

[0093] The convolution operation performing unit (350) can perform a convolution operation based on one or more compressed sub-kernels and a compressed input feature matrix.

[0094] Figure 4 is an example diagram showing a convolutional layer mapping method.

[0095] Figure 4(a) illustrates the image-to-column (im2col) method, a common mapping method, which maps each kernel to each column of the PIM array in a fixed-size format. Input features of the same size as the kernel are input for each cycle. However, since this method's memory usage is determined by the array and weight matrix sizes, unused memory cells may be generated in some cases, resulting in unnecessary energy consumption and increased latency.

[0096] Figure 4(b) illustrates the shift and duplicate kernel (SDK) mapping method. This method allows for obtaining multiple output values ​​in a single cycle by reassigning columns that do not share the same weights. The SDK mapping method forms a parallel window (PW), a set of input feature windows that share a portion of the kernel, to compute multiple operands in a single cycle. Because it reuses input feature and kernel weights, it utilizes memory more efficiently and reduces computational complexity than the im2col mapping method.

[0097] Figure 5 is an example diagram showing the application of pattern-based pruning when using the SDK mapping method.

[0098] Pattern-based pruning is a technique that applies the same pattern to all kernels based on a predetermined pattern to prune the weights. For example, in the case of the 8-entry of Fig. 5, when using the SDK mapping method, if the first row is to be omitted in the operation, the pattern-based pruning technique removes the 'a' from the kernel, so the 'a' in all rows can be removed. This example can be applied equally to 'a', 'h', and 'I' in the 6-entry of Fig. 5, and 'a', 'e', ​​'f', 'h', and 'I' in the 4-entry.

[0099] The N-entry value shown in Figure 5 represents the number of weights within the kernel that are not actually pruned. For example, 6-entry means that 3 weights are pruned and only 6 weights are used for calculations. Patterns with fewer entries use fewer weights for calculations, which can lead to a loss of inference accuracy.

[0100] While pattern-based pruning can effectively compress the weight matrix by performing row-skipping, the SDK mapping method suffers from the problem of ineffective row-skipping due to the unsorted nature of weight mapping to reuse input features. This means that for a given pattern, pruning efficiency can vary for each kernel, and even depending on the pattern used. In some cases, pruning may not compress the weight matrix, resulting in no hardware benefit.

[0101] Row-skipping is one of the compression techniques for weight matrices. It is a method of reducing the size of a weight matrix by not mapping the weights of a row if all weight values ​​mapped to that row are 0.

[0102] FIG. 6 is an exemplary diagram showing how to omit weights from operations by applying a mask in a convolution operation method using kernel shape control according to the present invention.

[0103] As shown in Fig. 6, the convolution operation method using kernel shape control according to the present invention can omit only the 'a' weight of the first row from the operation by applying a mask to the first row without removing the 'a' weight as in the existing pruning.

[0104] FIG. 7 is an exemplary diagram showing a method for determining a row to omit weights in a convolution operation method using kernel shape control according to the present invention.

[0105] According to one embodiment of the present invention, the weights excluded from the convolution operation may be weights within each sub-kernel that exist at the same location as the location of 0 within the mask. At this time, the mask may be configured to calculate the number of non-zero components existing at the same location of one or more sub-kernels, and if the ratio of the number of non-zero components calculated for each location is less than or equal to a predetermined ratio, it is 0, and otherwise, it is 1. Here, the predetermined ratio may be determined using a threshold.

[0106] Weights that exist at the same location in each subkernel can be mapped to the same row when mapped for operation. Therefore, if the mask calculates the number of non-zero components existing at the same location in one or more subkernels, it can mean that it calculates the number of non-zero components among the weights located in the same row in the array mapped for operation. In other words, the mask can exclude rows from the convolution operation by selecting rows in the array mapped for operation in which the proportion of non-zero components among the weights located in the same row is smaller than a threshold.

[0107] In one embodiment of the present invention, a threshold value may be used to determine rows to be excluded from a convolution operation, as shown in FIG. 7.

[0108] The mask can use row utilization and a threshold to determine which weights in a row to omit. That is, rows with many zeros are considered to have low row utilization, and row-skipping can minimize loss of inference accuracy.

[0109] The weight matrix can be compressed by row-skipping, which omits weights in rows with row utilization below a threshold.

[0110] In Fig. 7, for example, when the threshold is 25%, rows 1, 4, 13, and 16 with row utilization of 25% or less (i.e., rows with weights of 3 or more out of 4 being 0) can be omitted from the calculation.

[0111] Figure 8 is an exemplary diagram showing components of a convolution operation method using kernel shape control according to the present invention.

[0112] The kernel is a two-dimensional matrix expressed as M by N, and the base kernel can be a two-dimensional matrix expressed as (M+2) by (N+2) by adding padding with 0 values ​​to the upper, lower, left, and right outer corners of the kernel. In addition, the sub kernel is a two-dimensional matrix expressed as (M+1) by (N+1), and can be generated by removing at least one of the first row and the (M+1)th row of the base kernel or at least one of the first column and the (N+1)th column of the base kernel.

[0113] A parallel window may contain input features as part of the input feature matrix.

[0114] For convolution operations, input features within parallel windows and weights within each subkernel can be mapped according to a PIM array. In this case, input features within parallel windows and weights within each subkernel can be mapped along the same column.

[0115] A column-dependent mask (CDM) can be applied individually to the same kernel to exclude specific elements mapped to the memory cells of each row from computation, thereby skipping the computation of the corresponding column. By using a column-dependent mask, an independent pattern mask can be applied to the SDK of each kernel mapped to the PIM array. In other words, the column-dependent mask can be used to independently determine the weights that do not participate in the computation for each kernel. Furthermore, the column-dependent mask can be applied to the input feature matrix, similar to a kernel, to determine which input features do not participate in the computation. The column-dependent mask can be a two-dimensional matrix with the same size as the subkernel or parallel window. The column-dependent mask can determine which weights in which rows to skip using row utilization and a threshold. In other words, rows with many zeros are considered to have low row utilization and can be skipped to minimize inference accuracy loss.

[0116] As described above, the present invention effectively compresses the mapped weight matrix by targeting only a defined portion, taking into account the operating method and structure of the PIM array. This allows for efficient computation without hardware waste. This reduces computation cycles, improving computation speed and reducing energy consumption.

[0117] Furthermore, the present invention effectively compresses weight matrices and improves model accuracy without using pattern-based pruning or other pruning methods when using a mapping technique. Specifically, while conventional pattern-based pruning methods eliminate weight elements, the present invention omits them. Consequently, important weight elements that influence inference accuracy can be preserved, minimizing loss of inference accuracy and improving memory utilization.

[0118] In addition, compared to existing pattern-based pruning methods that remove weights, the present invention improves the weight matrix compression ratio by about 36% and the array utilization by about 39% while maintaining the accuracy loss at a level similar to that of existing networks through weight omission.

[0119] Therefore, according to the present invention, a weight matrix can be effectively compressed without using pruning in an SDK mapping method, and can be a fast and hardware-efficient mapping method for real-time inference of a deep learning model in the PIM field.

[0120] The embodiments of the present invention described above may be implemented through various means. For example, the embodiments of the present invention may be implemented using hardware, firmware, software, or a combination thereof.

[0121] The combination of each block of the block diagram and each step of the flowchart attached to the present invention may be performed by computer program instructions. These computer program instructions may be installed in an encoding processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, so that the instructions executed by the encoding processor of the computer or other programmable data processing equipment create a means for performing the functions described in each block of the block diagram or each step of the flowchart. These computer program instructions may also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to implement the functions in a specific manner, so that the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes an instruction means for performing the functions described in each block of the block diagram or each step of the flowchart. Since the computer program instructions can also be installed on a computer or other programmable data processing device, a series of operational steps are performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for executing the functions described in each block of the block diagram and each step of the flowchart can also provide steps for executing the functions described in each block of the block diagram and each step of the flowchart.

[0122] Additionally, each block or step may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specific logical function(s). In some embodiments, the functions mentioned in the blocks or steps may occur out of order. For example, two blocks or steps depicted in succession may actually be performed substantially simultaneously, or the blocks or steps may sometimes be performed in reverse order depending on the corresponding function.

[0123] The above description is merely an illustrative illustration of the technical idea of ​​the present invention, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential quality of the present invention. Therefore, the embodiments disclosed in the present invention are intended to illustrate, rather than limit, the technical idea of ​​the present invention, and the scope of the technical idea of ​​the present invention is not limited by these embodiments. The scope of protection of the present invention should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.

Claims

1. A method for controlling a kernel shape for a convolution operation performed by a convolution operation device, A step of generating a base kernel by adding padding to a kernel containing weights; A step of generating one or more sub kernels smaller in size than the base kernel by using a portion of the base kernel; and Comprising a step of performing a convolution operation based on one or more of the sub-kernels and an input feature matrix, A convolution operation method using kernel shape control.

2. In paragraph 1, In the step of performing the above convolution operation, Performing the convolution operation simultaneously based on each sub-kernel and the input feature matrix, A convolution operation method using kernel shape control.

3. In paragraph 1, All weight values ​​within the above padding are 0, A convolution operation method using kernel shape control.

4. In paragraph 1, The above kernel is, It is a two-dimensional matrix expressed as M by N (M and N are natural numbers), The above base kernel is, A two-dimensional matrix expressed as (M+2) by (N+2), A convolution operation method using kernel shape control.

5. In paragraph 4, One or more of the above subkernels, A two-dimensional matrix expressed as (M+1) by (N+1), generated by removing at least one of the first row and the (M+1)th row of the base kernel or at least one of the first column and the (N+1)th column of the base kernel. A convolution operation method using kernel shape control.

6. In paragraph 4, A step of generating a mask, which is a two-dimensional matrix expressed as (M+1) by (N+1); and Further comprising a step of compressing the one or more sub-kernels by determining weights to be excluded from the convolution operation through the mask. A convolution operation method using kernel shape control.

7. In paragraph 6, The above mask, The number of non-zero components existing at the same position of one or more of the above sub-kernels is calculated, and if the ratio of the number of non-zero components calculated for each position of a two-dimensional matrix expressed as (M+1) by (N+1) is less than or equal to a predetermined ratio, it is 0, otherwise it is 1. A convolution operation method using kernel shape control.

8. In paragraph 6, The weights excluded from the above convolution operation are: The weight in each sub-kernel that exists at the same location as the location of 0 in the above mask, A convolution operation method using kernel shape control.

9. In paragraph 6, In the step of performing the above convolution operation, Performing the convolution operation based on the compressed one or more sub-kernels and the input feature matrix, A convolution operation method using kernel shape control.

10. In paragraph 6, The above input feature matrix is, It is a two-dimensional matrix expressed as (M+1) by (N+1), Further comprising a step of compressing the input feature matrix by determining input features to be excluded from the convolution operation through the mask. A convolution operation method using kernel shape control.

11. In paragraph 10, In the step of performing the above convolution operation, Performing the convolution operation based on the compressed one or more sub-kernels and the compressed input feature matrix, A convolution operation method using kernel shape control.

12. At least one memory capable of storing computer-executable instructions; and By executing the above command, A processor comprising: a processor configured to generate a base kernel by adding padding to a kernel including weights; generate one or more sub-kernels smaller in size than the base kernel by using a portion of the base kernel; and perform a convolution operation based on the one or more sub-kernels and an input feature matrix. A convolution operation unit using kernel shape control.

13. A non-transitory computer-readable recording medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, A step of generating a base kernel by adding padding to a kernel containing weights; A step of generating one or more sub-kernels smaller in size than the base kernel by using a portion of the base kernel; and Causing the processor to perform a method including performing a convolution operation based on the one or more sub-kernels and the input feature matrix; Computer readable recording medium.

Citation Information

Patent Citations

  • Convolution neural network system and operation method thereof

    KR1020180052063A

  • Robot arm for education

    KR1020210096887A

  • Genetic health monitoring system to provide customized diet exercise information using personal genetic information

    KR1020220003353A

  • Frame interpolation via adaptive convolution and adaptive separable convolution

    US20200012940A1

  • KR20200049366A