Electronic device for gating mechanism and method thereof

KR103012689B1Inactive Publication Date: 2026-09-02KOREA ELECTRONICS TECH INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020240178658
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2026-09-02
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure 112024134478872-PAT00012_ABST
    Figure 112024134478872-PAT00012_ABST
Patent Text Reader

Abstract

An electronic device for performing a gating mechanism according to an embodiment of the present invention comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element in which the quantized value of the gate vector is 0 and can omit the output calculation of the second linear layer for the element based on the index.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an electronic device and method for performing a gating mechanism. Background Technology

[0002] With the recent launch of chatGPT developed by OpenAI, competition among Large Language Models (LLMs) is intensifying. Transformer-based language models have seen performance improvements in various aspects, including training methods, data collection, and network structures; consequently, new technologies are being developed daily through diverse initiatives.

[0003] With the release of many LLMs, modified transformer-based network layers have been researched to improve performance. However, due to the extensive computational power required for LLM development, designing large-scale LLMs is essential for performance; consequently, current LLMs can only run on large servers.

[0004] Accordingly, proprietary LLMs with improved performance are being proactively launched by companies such as OpenAI's GPT, Google's PaLM and Bard, and Meta's Llama1 and Llama2. In particular, Meta has facilitated research and development in various fields by providing pre-trained models through open source.

[0005] Meanwhile, activation functions are an essential component of neural networks (NNs). Activation functions introduce non-linear characteristics into the network, significantly impacting the performance of artificial intelligence; consequently, problems can only be solved by using non-linear activation functions. Accordingly, various activation functions such as the sigmoid function, hyperbolic tangent (tanh) function, ReLU (Rectified Linear Unit), and swish have been continuously developed. Since the introduction of ReLU, activation functions have been developed for learning purposes to ensure differentiability while maintaining a shape similar to ReLU. These functions exhibit the characteristic of generating values ​​close to zero for negative inputs and generating input values ​​for positive inputs.

[0006] Various methods have been introduced to reduce the computational cost of model training and inference, one of which involves utilizing these properties of activation functions. Activation functions can be designed as a component of the Gated Linear Unit (GLU), which is used in the feed-forward network portion of models such as Transformers. Feed-forward layers typically consist of two linear transformations and a non-linear activation function, and the non-linearity can be controlled through the GLU. This allows for the selective control of information flow, enabling the model to focus on more important information. Therefore, the GLU can be effectively used as a tool to improve efficiency and training performance in LLMs such as Transformers. The problem to be solved

[0007] The objective of the present invention is to provide an electronic device and method that can save hardware resources and reduce power consumption by reducing unnecessary computations when running an artificial neural network model.

[0008] The objective of the present invention is to provide an electronic device and method capable of saving memory bandwidth and increasing the data processing speed of the entire system when running an artificial neural network model. means of solving the problem

[0009] An electronic device for performing a gating mechanism according to an embodiment of the present invention comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element in which the quantized value of the gate vector is 0 and can omit the output calculation of the second linear layer for the element based on the index.

[0010] The above controller can index elements among the gate vectors in which the quantized value is 0.

[0011] The above controller may omit the linear transformation of at least one element corresponding to the index among the input vectors input to the second linear layer.

[0012] The controller above may omit the multiplication operation of the output vector of the second linear layer for the element and the gate vector based on the index above.

[0013] The above controller may omit the output operation of the second linear layer for the element on a channel basis and the multiplication operation of the output vector of the second linear layer and the gate vector.

[0014] The above activation function may be any one of the sigmoid function, hyperbolic tangent (tanh) function, ReLU (Rectified Linear Unit), and swish function.

[0015] An electronic device for performing a gating mechanism according to an embodiment of the present invention comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element among the gate vectors in which the quantized value is 0, and based on the index, can omit the multiplication operation of the output vector of the second linear layer and the gate vector for said element.

[0016] The above controller may omit the multiplication operation of the output vector of the second linear layer and the gate vector for the element on a channel basis.

[0017] An electronic device for performing a gating mechanism according to an embodiment of the present invention comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller may include a controller that extracts an index of an element among the gate vectors whose quantized value is 0 and omits an input operation of the third linear layer for the element based on the index.

[0018] The above controller may omit the input operation of the third linear layer for the element, the output operation of the second linear layer, and the multiplication operation of the output vector of the second linear layer and the gate vector on a channel-by-channel basis.

[0019] A method for performing a gating mechanism by an electronic device according to an embodiment of the present invention comprises: a step of linearly transforming an input vector input to a first linear layer of an artificial neural network model and then performing a gate vector using an activation function; a step of linearly transforming an input vector of a second linear layer; a step of elementally multiplying an output vector of the second linear layer and the gate vector; and a step of inputting the elementally multiplied value to a third linear layer, wherein the step of performing the gate vector includes a step of extracting an index of an element among the gate vectors in which the quantized value is 0, and based on the index, at least one of the output operation of the second linear layer for the element, the multiplication operation of the output vector of the second linear layer and the gate vector, and the input operation of the third linear layer may be omitted. Effects of the invention

[0020] According to one embodiment of the present invention, unnecessary operations caused by zero value input can be reduced to save hardware resources and reduce power consumption.

[0021] According to one embodiment of the present invention, memory bandwidth can be saved by skipping operations at the channel level, and the data processing speed of the entire system can be increased. Brief explanation of the drawing

[0022] Figure 1 is a drawing illustrating a Transformer model. Figure 2 is a drawing illustrating an improved transformer model. Figure 3 is a diagram illustrating a gating mechanism. FIG. 4 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present invention. FIG. 5 is a drawing illustrating a gating mechanism according to a first embodiment of the present invention. FIG. 6 is a drawing illustrating a gating mechanism according to a second embodiment of the present invention. FIG. 7 is a drawing illustrating a gating mechanism according to a third embodiment of the present invention. FIG. 8 is a diagram illustrating the operation flowchart of an electronic device according to one embodiment of the present invention. Specific details for implementing the invention

[0023] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below, together with the accompanying drawings, is intended to describe exemplary embodiments of the present invention and is not intended to represent the only embodiment in which the present invention can be practiced. In order to clearly explain the present invention in the drawings, parts unrelated to the description may be omitted, and the same reference numerals may be used for identical or similar components throughout the specification.

[0024] The words and terms used in this specification and claims are not limited to their ordinary or dictionary meanings, but should be interpreted in a meaning and concept consistent with the technical spirit of the invention in accordance with the principles by which the inventor defines terms and concepts to best describe his invention.

[0025] Therefore, the embodiments described in this specification and the configurations illustrated in the drawings correspond to preferred embodiments of the present invention and do not represent all technical concepts of the present invention; thus, various equivalents and modifications that may replace such configurations may exist at the time of filing the present invention.

[0026] In this specification, terms such as “comprising” or “having” are intended to describe the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should not be understood as precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0027] Figure 1 is a drawing illustrating a Transformer model.

[0028] The Transformer model is a very important deep learning architecture in the field of natural language processing (NLP) that uses the encoder-decoder structure of the existing Sequence-to-Sequence (seq2seq). However, unlike seq2seq, the Transformer model uses only attention instead of a recurrent neural network (RNN).

[0029] The biggest drawback of conventional RNN models is that computational efficiency decreases as the network is continuously updated iteratively. However, using the Transformer model eliminates the need for RNNs, which not only improves computational efficiency but also enhances performance through superior parallelization and shorter training times.

[0030] The encoder consists of multiple layers, and each layer is composed of Multi-Head Attention and a Feed Forward Network (FFN).

[0031] Attention refers to a method of determining which words to focus on based on context, and understanding the context through this is key. Unlike RNNs or CNNs, Transformers do not have a sequential structure, so they provide information about the order of input data through positional encoding.

[0032] Multi-head attention allows all words in an input sequence to learn relationships by referencing each other, and can focus attention in various ways using multiple heads.

[0033] Add & Norm is a layer that enhances learning stability by applying residual connections and normalization to each sub-layer of the layer.

[0034] A feedforward network is a fully connected layer that processes vectors at each location independently, and it has a structure where input data flows in only one direction. Feedforward networks are primarily implemented in the form of a Multi-Layer Perceptron (MLP) and consist of an input layer, one or more hidden layers, and an output layer.

[0035] In the Transformer architecture, the feed-forward network plays the role of transforming information by processing the given input through a linear transformation and then applying an activation function.

[0036] The decoder is also composed of multiple layers, and each layer consists of Masked Multi-Head Attention, Multi-Head Attention, and a Feed Forward Network (FFN).

[0037] Unlike the encoder, the decoder performs multi-head attention using masking to ensure that only past information is used when predicting the next word. Multi-head attention learns the relationship between the input sentence and the generated sentence by referring to the encoder's output. Similar to the encoder, the decoder also uses residual linking and regularization to improve learning stability.

[0038] Finally, the decoder outputs the probability of the next word for each position through a linear layer and a softmax layer.

[0039] As such, Transformers are advantageous for parallel processing and can learn much faster and more efficiently than RNN-based models. This makes them highly effective for various sequential data processing tasks, such as natural language processing.

[0040] Figure 2 is a drawing illustrating an improved transformer model.

[0041] Llama (Large Language Model Meta AI), developed by Meta, is a natural language processing model based on the Transformer architecture, but there are some differences from the Transformer.

[0042] One of the major changes in the Llama model compared to the Transformer is the normalization step. The Transformer model uses layer normalization after each multi-head attention block. Through normalization, values ​​are aligned around 0 with a standard deviation of 1, following a normal distribution. On the other hand, the Llama model uses a different normalization variation called Root Mean Square (RMS) normalization.

[0043] Furthermore, Transformer models prioritize the position of the input sequence and utilize positional encoding, with positional encoding vectors used throughout the training process. In contrast, Llama uses rotary positional embedding. This method introduces rotation operations into the positional encoding process, enabling the model to learn dynamic positional representations during training instead of relying on pre-calculated static positional encoding vectors.

[0044] Another major change in Llama in the Transformer model is that it uses activation functions with gated linear unit structures instead of standard activation functions like ReLU in the feedforward network.

[0045] Gated linear units reduce the amount of information a model needs to process by passing only necessary information, and can impart non-linear characteristics to enable the model to learn more complex representations. Furthermore, due to their simple structure, they facilitate parallelization, allowing them to maintain efficiency even in large-scale computations. Gated linear units can be combined with other existing activation functions; examples include SwiGLU, which combines the Swish function with a gated linear unit, and ReGLU, which combines it with ReLU.

[0046] Figure 3 is a diagram illustrating a gating mechanism.

[0047] The Gated Linear Unit (GLU) operates by introducing a gating mechanism during the neural network's learning process to determine the importance of input data, thereby blocking unnecessary information and passing only important information to the next layer. In this way, the network can select which words or features are important for predicting the next word.

[0048] A gate linear unit can be defined as shown in the following mathematical formula 1.

[0049]

[0050] x is the input vector, W1 and W2 are weight matrices, b1 and b2 are bias vectors, σ is the activation function, means the product of elements.

[0051] For example, a gated linear unit better regulates information flow by using gates that control the relationship between input and output. The gated linear unit linearly transforms the input data and then controls the gates through an activation function to selectively pass the results. Using this method, information can be transmitted more smoothly while maintaining non-linearity as the input is coupled with the gate.

[0052] As illustrated in FIG. 3, the structure of the feedforward network (300) can be composed of three parts. The three parts are one element-wise multiplication layer, three linear layers, and one activation function. Specifically, the same input vector is linearly transformed in the first linear layer (310) and the second linear layer (320), respectively, and the gate vector is computed by applying the activation function (330) to the linearly transformed vector of the first linear layer (310). The output of the activation function (330) and the output of the second linear layer (320) are element-wise multiplied (340) and input to the third linear layer (350). This can replace the MLP block of the transformer.

[0053] Gated linear units are used in the feed-forward network structure of transformers or models based on them, and can improve the performance of feed-forward networks by applying various activation functions in addition to the sigmoid function.

[0054] FIG. 4 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present invention.

[0055] An electronic device (100) according to one embodiment of the present invention is a device that performs a gating mechanism in a feedforward network of an artificial neural network model and can be implemented as a computer, server, smartphone, tablet PC, smart pad, laptop, etc.

[0056] In this case, the feedforward network requires significant computation due to linear layers and element-wise multiplication. However, the activation function used in the middle of the feedforward network converts most negative numbers into values ​​close to zero, which become zero upon quantization. By skipping unnecessary calculations for a large number of elements that have zero values, the amount of computation can be reduced and performance improved.

[0057] The present invention proposes a process that reduces computational complexity by utilizing the property that a large number of zero values ​​occur during the gating mechanism process.

[0058] Hereinafter, the configuration and operation of an electronic device (100) according to one embodiment of the present invention will be described in detail with reference to the drawings.

[0059] An electronic device (100) according to one embodiment of the present invention may include an input unit (110), a communication unit (120), a display unit (130), a memory (140), and a controller (150).

[0060] The input unit (110) generates input data in response to user input of the electronic device (100). User input can be applied without restriction if it is user input necessary to perform the gating mechanism.

[0061] The input unit (110) includes at least one input means. The input unit (110) may include a keyboard, a key pad, a dome switch, a touch panel, a touch key, a mouse, a menu button, etc.

[0062] The communication unit (120) can communicate with various components, such as memory (140), connected to the electronic device (100) through an external device such as a server or a system bus.

[0063] The communication unit (120) can support one or more communication standards for wired or wireless communication with various components connected to the electronic device (100), and the communication unit (120) can be implemented for each connected device, such as a host interface connected to a host device and a memory interface connected to an external memory.

[0064] The communication unit (120) may be configured by adopting at least one of a serial transmission standard, such as PCIe (Peripheral Component Interconnection Express), USB (Universal Serial Bus), SATA (Serial AT Attachment), or SAS (Serial Attached SCSI). Additionally, the communication unit (120) may be configured by adopting at least one of a parallel transmission standard, such as SCSI (Small Computer System Interface), ATA (AT Attachment), or PATA (Parallel AT Attachment).

[0065] In addition, the communication unit (120) can perform wireless communication such as 5G (5th generation communication), LTE-A (Long Term Evolution-Advanced), LTE (Long Term Evolution), Wi-Fi (Wireless Fidelity), Bluetooth, or serial communication such as Ethernet, LAN (Local Area Network), WAN (Wide Area Network), power line communication, etc.

[0066] The display unit (130) displays display data according to the operation of the electronic device (100). The display unit (130) includes a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a micro electro mechanical systems (MEMS) display, and an electronic paper display. The display unit (130) can be combined with the input unit (110) to be implemented as a touch screen.

[0067] The memory (140) can store software, firmware, and data for controlling the electronic device (100). The neural network involves large-scale matrix operations, which significantly affect memory bandwidth and data transfer speed. The electronic device (100) utilizes various memory layers and data buffers to optimize this data flow.

[0068] The memory (140) may include a memory with non-volatile properties that can store data (information) regardless of whether power is provided, and a memory with volatile properties in which data to be processed by the electronic device (100) is loaded and data cannot be stored if power is not provided.

[0069] The memory (140) stores operation programs of the electronic device (100). The memory (140) includes storage with non-volatile properties that can preserve data (information) regardless of whether power is provided, and memory with volatile properties in which data to be processed by the controller (150) is loaded and data cannot be preserved if power is not provided. Storage includes flash memory, hard-disc drive (HDD), solid-state drive (SSD), and ROM (Read Only Memory), and memory includes buffer and RAM (Random Access Memory).

[0070] The controller (150) can control at least one other component (e.g., hardware or software component) of the electronic device (100) by executing software such as a program, and can perform various data processing or operations.

[0071] The controller (150) may include a hardware component for processing data based on one or more instructions as an application processor (AP). The controller (150) may have a multi-core processor structure such as a dual core, quad core, or hexa core.

[0072] The controller (150) can create a neural network (NN), train a neural network, perform operations based on received input data, generate result information based on the operation, or retrain the neural network. The neural network may include various types of neural network models such as CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), R-CNN (Region with Convolutional Neural Network), S-DNN (Stacking-based Deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), RPN (Region Proposal Network), DBN (Deep Belief Network), RBM (Restricted Boltzmann Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory), and transformer, but is not limited thereto.

[0073] A controller (150) according to one embodiment of the present invention can linearly transform an input vector input to a first linear layer of an artificial neural network model and then calculate a gate vector using an activation function, linearly transform an input vector of a second linear layer, and input an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer.

[0074] At this time, the controller (150) can extract the index of an element whose quantized value is 0 among the gate vectors, and based on the index, omit the output operation of the second linear layer for the element.

[0075] As another example, the controller (150) may omit the multiplication operation of the output vector of the second linear layer and the gate vector for the element based on the index.

[0076] As another example, the controller (150) may omit the input operation of the third linear layer for the element based on the index.

[0077] Figures 5 to 7 below propose an efficient data flow and hardware structure that can skip unnecessary operations with zero values ​​as inputs occurring during GLU operations by utilizing the characteristics of a gating mechanism.

[0078] FIG. 5 is a drawing illustrating a gating mechanism according to a first embodiment of the present invention.

[0079] A controller (150) according to one embodiment of the present invention can calculate a gate vector using an activation function (520) after linearly transforming an input vector input to a first linear layer (510) of an artificial neural network model.

[0080] At this time, the first linear layer (510), the second linear layer (530), the third linear layer (560), the activation function (520), and the element-wise multiplier (550) can be applied to the feed-forward network (500) of the artificial neural network model.

[0081] The activation function (520) can be any one of the sigmoid function, hyperbolic tangent (tanh) function, ReLU (Rectified Linear Unit) and swish function.

[0082] The sigmoid function can be defined by the following mathematical equation 2.

[0083]

[0084] The sigmoid function restricts the output to between 0 and 1, allowing it to be interpreted as a probability.

[0085] The hyperbolic tangent (tanh) function can be defined by the following mathematical equation 3.

[0086]

[0087] The hyperbolic tangent (tanh) function was proposed as an alternative to the sigmoid, with an output range between -1 and 1 and a center of 0.

[0088] ReLU (Rectified Linear Unit) can be defined by the following mathematical equation 4.

[0089]

[0090] ReLU (Rectified Linear Unit) was introduced to solve the vanishing gradient problem and improve computational efficiency.

[0091] Leaky ReLU (Rectified Linear Unit) can be defined by the following mathematical equation 5.

[0092]

[0093] Leaky ReLU (Rectified Linear Unit) has a slight slope for negative values.

[0094] The Swish function can be defined by the following mathematical formula 6.

[0095]

[0096] σ is the sigmoid function, and β is a learnable parameter.

[0097] The controller (150) can extract the index (540) of the element whose quantized value is 0 among the gate vectors.

[0098] An activation function such as Swish generates values ​​close to zero for elements of an input vector that has undergone a linear transformation, and these values ​​are quantized to zero during the quantization process for inference. The controller (150) can index elements of the gate vector whose quantized value is zero.

[0099] The controller (150) may omit the output operation of the second linear layer (530) for elements whose quantized value is 0 based on the index (540). That is, the controller (150) may omit the linear transformation of at least one element corresponding to the index (540) among the input vectors input to the second linear layer (530).

[0100] A pruning technique can be applied during the process of omitting the output operation of the second linear layer (530). A pruning technique is a method of removing unnecessary elements to lighten and optimize an artificial neural network model. Among pruning techniques, there is a structured pruning technique that removes unnecessary elements in specific units while maintaining the structure of the model. While general pruning techniques reduce the number of parameters by removing individual weights of the model, structured pruning techniques simplify the model based on larger structural units. Structured pruning techniques remove parameters in large units, such as channels, filters, and layers, from the neural network. This allows for the removal of unnecessary parts without changing the dimensions of the network, making optimization on hardware easier.

[0101] When using structured pruning techniques, unnecessary channels or filters are eliminated by the network, reducing computational load and memory usage while accelerating model inference. In particular, efficiency is enhanced by enabling parallel processing optimization on hardware such as Graphics Processing Units (GPUs) and Neural Processing Units (NPUs).

[0102] The controller (150) may omit the output operation of the second linear layer (530) for elements whose quantized value is 0 on a channel-by-channel basis.

[0103] The controller (150) can input the output of the second linear layer (530), in which unnecessary operations are omitted, and the output of the activation function (520) to the third linear layer (560) by performing element-wise multiplication (550). At this time, the controller (150) can omit the multiplication operation between the output vector of the second linear layer (530) and the gate vector for elements whose quantized value is 0 based on the index. Similarly, the controller (150) can omit the multiplication operation between the output vector of the second linear layer (530) and the gate vector for elements whose quantized value is 0 on a channel-by-channel basis.

[0104] The zero value output from the activation function makes the output zero regardless of the input on the other side in element-wise multiplication. Therefore, there is no need to perform element-wise multiplication operations for the corresponding pixel, and there is no need to perform operations for the corresponding pixel in the second linear layer (530). Additionally, because a large number of zeros are input to the third linear layer (560) due to the zero value output from the activation function, there are operations that can be skipped in the third linear layer (560) as well.

[0105] In summary, the second linear layer (530) can skip operations for unnecessary outputs, and the third linear layer (560) can skip operations using zeros as inputs. To use this efficiently, the method of storing the weights of each layer in memory must be different. General memory works well when sequential reading is performed. Therefore, for efficient data movement, weight data for the second linear layer (530) must be stored per output channel, and weight data for the third linear layer (560) must be stored per input channel.

[0106] According to one embodiment of the present invention, since the amount of zeros varies depending on the input value and the layer position within the entire network, the amount of computation can be significantly reduced by using this structure to change the amount of zeros.

[0107] According to one embodiment of the present invention, unnecessary operations caused by zero value input can be reduced to save hardware resources and reduce power consumption.

[0108] According to one embodiment of the present invention, memory bandwidth can be saved by skipping operations at the channel level, and the data processing speed of the entire system can be increased.

[0109] To actually use this method, it must be implemented in hardware to efficiently skip zeros. In future work, we will implement the part that accelerates LLM inference using a field programmable gate array (FPGA).

[0110] FIG. 6 is a drawing illustrating a gating mechanism according to a second embodiment of the present invention.

[0111] A controller (150) according to one embodiment of the present invention can calculate a gate vector using an activation function (620) after linearly transforming an input vector input to a first linear layer (610) of an artificial neural network model.

[0112] At this time, the first linear layer (610), the second linear layer (630), the third linear layer (660), the activation function (620), and the element-wise multiplier (650) can be applied to the feed-forward network (600) of the artificial neural network model.

[0113] The activation function (620) may be any one of the sigmoid function, hyperbolic tangent (tanh) function, ReLU (Rectified Linear Unit), and swish function. The specific details of each activation function (620) are as described with reference to FIG. 5.

[0114] The controller (150) can extract the index (640) of an element in the gate vector whose quantized value is 0. The controller (150) can index an element in the gate vector whose quantized value is 0.

[0115] The controller (150) may omit the multiplication operation between the output vector of the second linear layer (630) and the gate vector for elements with a quantized value of 0 based on the index (640). The controller (150) may omit the multiplication operation between the output vector of the second linear layer (630) and the gate vector for elements with a quantized value of 0 on a channel-by-channel basis. A pruning technique may be applied during the process of omitting the multiplication operation between the output vector of the second linear layer (630) and the gate vector, as described above with reference to FIG. 5.

[0116] According to one embodiment of the present invention, unnecessary operations caused by zero value input can be reduced to save hardware resources and reduce power consumption.

[0117] According to one embodiment of the present invention, memory bandwidth can be saved by skipping operations at the channel level, and the data processing speed of the entire system can be increased.

[0118] FIG. 7 is a drawing illustrating a gating mechanism according to a third embodiment of the present invention.

[0119] The controller (150) can calculate a gate vector using an activation function after linearly transforming the input vector input to the first linear layer (710) of the artificial neural network model.

[0120] At this time, the first linear layer (710), the second linear layer (730), the third linear layer (760), the activation function (720), and the element-wise multiplier (740) can be applied to the feed-forward network (700) of the artificial neural network model.

[0121] The activation function (720) may be any one of the sigmoid function, hyperbolic tangent (tanh) function, ReLU (Rectified Linear Unit), and swish function. The specific details of each activation function (720) are as described with reference to FIG. 5.

[0122] The controller (150) can extract the index (750) of an element in the gate vector whose quantized value is 0. Based on the index (750), the controller (150) can omit the input operation for the element whose quantized value is 0 in the third linear layer (760). In this case, the input operation refers to an operation of inputting the multiplication value of the output vector of the second linear layer (730) and the gate vector for the element whose quantized value is 0. A pruning technique may be applied during the process of omitting the input operation for the element whose quantized value is 0 in the third linear layer (760), and the pruning technique is as described above with reference to FIG. 5.

[0123] According to one embodiment of the present invention, unnecessary operations caused by zero value input can be reduced to save hardware resources and reduce power consumption.

[0124] According to one embodiment of the present invention, memory bandwidth can be saved by skipping operations at the channel level, and the data processing speed of the entire system can be increased.

[0125] FIG. 8 is a diagram illustrating the operation flowchart of an electronic device according to an embodiment of the present invention. FIG. 8 borrows from the description previously made with reference to FIG. 5 to FIG. 7, and therefore, specific explanations regarding overlapping content are omitted.

[0126] The controller (150) can calculate a gate vector using an activation function after linearly transforming the input vector input to the first linear layer of the artificial neural network model (S10).

[0127] The controller (150) can extract the index of an element in the gate vector whose quantized value is 0 (S20).

[0128] The controller (150) can linearly transform the input vector of the second linear layer (S30). At this time, the controller (150) may omit the linear transformation of at least one element corresponding to an index among the input vectors input to the second linear layer.

[0129] The controller (150) can perform element-wise multiplication of the output vector of the second linear layer and the gate vector (S40). At this time, the controller (150) may omit the multiplication operation of the output vector of the second linear layer and the gate vector for elements whose quantized value is 0 based on the index.

[0130] The controller (150) can input the element-multiplied value into the third linear layer (S50). At this time, the controller (150) can omit the input operation of the third linear layer for elements whose quantized value is 0 based on the index.

[0131] Table 1 shows the reduction in computational load according to the ratio of quantized values ​​being 0 according to the activation function. Referring to this, even if the ratio of quantized values ​​(0) is 20%, the computation can be reduced by 13%.

[0132] Zero value ratio Computational reduction ratio 10% 6.6% 20% 13.3% 30% 20.0% 40% 26.6% 50% 33.3% 60% 40.0% Explanation of the symbols

[0133] 100: Electronic device 110: Input section 120: Communications Department 130: Display unit 140: Memory 150: Controller

Claims

Claim 1 An electronic device for performing a gating mechanism comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element whose quantized value is 0 among the gate vectors, omits a linear transformation of at least one element corresponding to the index among the input vectors input to the second linear layer, and omits a multiplication operation of the output vector of the second linear layer and the gate vector for the element based on the index. Claim 2 In claim 1, the controller is an electronic device that indexes elements of the gate vectors in which the quantized value is 0. Claim 3 delete Claim 4 delete Claim 5 In claim 1, the controller is an electronic device that omits the output operation of the second linear layer for the element on a channel basis and the multiplication operation of the output vector of the second linear layer and the gate vector. Claim 6 An electronic device according to claim 1, wherein the activation function is any one of a sigmoid function, a hyperbolic tangent (tanh) function, a ReLU (Rectified Linear Unit) and a swish function. Claim 7 An electronic device for performing a gating mechanism comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element among the gate vectors in which the quantized value is 0, and based on the index, omits the multiplication operation of the output vector of the second linear layer and the gate vector for said element. Claim 8 In claim 7, the controller is an electronic device that indexes elements of the gate vectors in which the quantized value is 0. Claim 9 In claim 7, the controller is an electronic device that omits the multiplication operation of the output vector of the second linear layer and the gate vector for the element on a channel basis. Claim 10 An electronic device according to claim 7, wherein the activation function is any one of a sigmoid function, a hyperbolic tangent (tanh) function, a ReLU (Rectified Linear Unit), and a swish function. Claim 11 An electronic device for performing a gating mechanism comprises a controller that linearly transforms an input vector input to a first linear layer of an artificial neural network model and then calculates a gate vector using an activation function, linearly transforms an input vector of a second linear layer, and inputs an element-wise multiplication value of the output vector of the second linear layer and the gate vector to a third linear layer, wherein the controller extracts an index of an element among the gate vectors in which the quantized value is 0, and based on the index, performs a multiplication operation of the output vector of the second linear layer and the gate vector for the element, and then omits the input operation of the third linear layer. Claim 12 In claim 11, the controller is an electronic device that indexes elements of the gate vectors in which the quantized value is 0. Claim 13 In paragraph 11, the controller is an electronic device that omits the linear transformation of at least one element corresponding to the index among the input vectors input to the second linear layer. Claim 14 In claim 11, the controller is an electronic device that omits the multiplication operation of the output vector of the second linear layer and the gate vector for the element based on the index. Claim 15 In paragraph 14, the controller is an electronic device that omits the input operation of the third linear layer for the element, the output operation of the second linear layer, and the multiplication operation of the output vector of the second linear layer and the gate vector on a channel-by-channel basis. Claim 16 An electronic device according to claim 11, wherein the activation function is any one of a sigmoid function, a hyperbolic tangent (tanh) function, a ReLU (Rectified Linear Unit), and a swish function. Claim 17 A method for performing a gating mechanism performed by an electronic device, comprising: a step of linearly transforming an input vector input to a first linear layer of an artificial neural network model and then performing a gate vector operation using an activation function; a step of linearly transforming an input vector of a second linear layer; a step of elementally multiplying an output vector of the second linear layer and the gate vector; and a step of inputting the elementally multiplied value to a third linear layer, wherein the step of performing the gate vector operation includes a step of extracting an index of an element in which the quantized value of the gate vector is 0, and wherein at least one of the linear transformation of at least one element corresponding to the index among the input vector input to the second linear layer, the multiplication operation of the output vector of the second linear layer and the gate vector, and the input operation of the third linear layer after the multiplication operation of the output vector of the second linear layer and the gate vector is omitted.

Citation Information

Patent Citations

  • Accelerator device for multimode activation function

    KR1020230143041A

  • Computing-In-Memory and Zero Skip Operation Method Therefor

    KR1020240170232A