Artificial intelligence neural network accelerator and method for transformer neural network

The AI neural network accelerator addresses Transformer neural network power consumption and memory access issues by employing a hybrid model with reduced weight values and implicit generation, enhancing energy efficiency and reducing memory access.

JP2025105535AActive Publication Date: 2025-07-10KOREA ADVANCED INST OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024226255
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2024-12-23
Publication Date
2025-07-10
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Transformer neural networks face challenges with high power consumption due to large weight values and extensive external memory access, making them unsuitable for mobile devices.

Method used

An artificial intelligence neural network accelerator that reduces weight values and external memory access by using a hybrid model with a primary prediction step using a reduced model and a secondary step using a full model only when accuracy falls below a threshold, along with implicit weight value generation and on-chip network transmission.

Benefits of technology

The accelerator achieves high energy efficiency by reducing weight values and memory access, maintaining accuracy through a mixed network structure and implicit weight generation, with significant reductions in power consumption and memory access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105535000001_ABST
    Figure 2025105535000001_ABST
Patent Text Reader

Abstract

To provide a highly energy-efficient mobile large language model accelerator.SOLUTION: An AI neural network accelerator 100 includes a computing unit configured to predict output tokens in units of input tokens and including a plurality of transformer operation cores operating based on a transformer model using n weights, and a controller configured to control operation of each of the transformer operation cores after determining a size of the transformer model based on the number of weights. The controller performs control to perform primary prediction in which each of the plurality of transformer operation cores predicts output tokens using a reduced transformer model using m weights (m<n) in response to acquisition of the input tokens, and further perform secondary prediction of predicting output tokens for the input tokens using a transformer model before reduction only when prediction accuracy of a result of the prediction is less than or equal to a preset threshold value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence neural network accelerator and a method thereof, and more particularly, to an artificial intelligence neural network accelerator and a method thereof for a transformer neural network that is difficult to reuse weight values, requires a large amount of weight values, has a large amount of external memory access, and consequently has a high power consumption.

Background Art

[0002] A Transformer neural network is a neural network that tracks relationships within sequential data such as words in a sentence, learns context and meaning, and replaces a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN).

[0003] Such a Transformer neural network does not need to construct a large labeled training dataset by mathematically searching for patterns between elements, is suitable for parallel processing, and due to its fast execution speed, is widely used in image classification and large language models, and is also expected to be used in mobile systems that provide real-time responses.

[0004] However, such a Transformer neural network has a problem that it is difficult to reuse weight values, requires a large amount of weight values, has a large amount of external memory access, and as a result, power consumption increases.

[0005] Therefore, in order to solve the above problems, various techniques (References 1 to 3 (Non-Patent Documents 1-3)) have been proposed to increase hardware utilization and reduce power consumption. However, the system power consumption and response time of these transformer processors are still not suitable for mobile devices. For example, large language models such as GPT-2 have many weight values (400M to 700M), and external memory access consumes 68% of the total power.

[0006] In addition, in order to alleviate the bottleneck phenomenon of external memory access, a transformer accelerator (Reference 4 (Non-Patent Document 4)) applying pruning has been proposed. Although this can increase the sparsity of weight values, it can only be applied to simple operations such as predicting the next word (e.g., language modeling), and has the problem that high sparsity cannot be achieved in advanced operations such as language translation, question answering, and summarization.

[0007] Therefore, in order to reduce external memory access for accelerating large language models for mobile devices with high energy efficiency, there is a need for a new method that can compress weight values.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

[0009] The present invention has been made to solve the above-described problems, and by reducing the required amount of weight values for accelerating an artificial intelligence neural network, it reduces the external memory access amount, and as a result, by reducing the power consumption, it attempts to provide an artificial intelligence neural network accelerator and a method thereof that can accelerate a large-scale language model for mobile devices with high energy efficiency.

[0010] Further, the present invention first accelerates a reduced model in which the required amount of basic weight values necessary for accelerating the basic model is reduced at a predetermined ratio, and only in a special case where the prediction accuracy for the result is below a preset critical value, it further performs a secondary acceleration step of accelerating the basic model. As a result, it attempts to provide an artificial intelligence neural network accelerator and a method thereof that can reduce the required amount of weight values for accelerating the artificial intelligence neural network.

[0011] Furthermore, the present invention includes weight value embedding logic generated in advance as a result of learning each weight value of the Transformer neural network matched to the kernel position information of the Transformer neural network. After generating implicit weight values for accelerating the artificial intelligence neural network while receiving only the kernel position information from the external memory, by supplying the implicit weight values using an on-chip network, it attempts to provide an artificial intelligence neural network accelerator and a method thereof that can reduce the external memory access amount for receiving weight values.

[0012] Also, the present invention attempts to provide an artificial intelligence neural network accelerator and a method thereof that can shorten the time required to receive the kernel position information by going through a process of decompressing the kernel position information received from the external memory in a symbol-compressed manner, and as a result, can shorten the external memory access time.

Means for Solving the Problems

[0013] To solve the above problems, the artificial intelligence neural network accelerator provided by the present invention is an artificial intelligence neural network accelerator for accelerating a Transformer neural network. The operation unit includes a number of Transformer operation cores that operate based on a Transformer model that predicts output tokens in units of input tokens and uses n (where n is a natural number) weight values; and a controller that controls the operation of each of the Transformer operation cores after determining the size of the Transformer model according to the number of the weight values. The controller is configured to control each of the number of Transformer operation cores to perform a primary prediction of predicting an output token using a reduced Transformer model (hereinafter referred to as a "reduced model") that uses m (where m is a natural number and m < n) weight values in response to obtaining the input token, and to perform a secondary prediction of predicting an output token for the input token using the Transformer model before reduction (hereinafter referred to as a "basic model") only when the prediction accuracy of the primary prediction result is equal to or lower than a preset threshold value.

[0014] Preferably, the artificial intelligence neural network accelerator includes weight value embedding logic that is pre-generated as a result of learning by matching each weight value of the Transformer neural network with axb kernel position information of the Transformer neural network, and further includes a weight value generator that generates implicit weight values based on the position information of the kernel input from an external memory. The controller can provide the implicit weight values via an on-chip network in response to requests from the number of Transformer operation cores.

[0015] Preferably, the weight value generator includes a decompression unit that decompresses the compression of the position information of the kernel input in a symbol-compressed state from the external memory; and the weight value embedding logic, and applies the position information of the kernel decompressed by the decompression unit to the weight value embedding logic to generate an implicit weight value corresponding to the position of the decompressed kernel; an implicit weight value generation unit.

[0016] Preferably, the implicit weight value generation unit includes a two-dimensional MAC array that performs multiplication and accumulation operations for generating the implicit weight value using the position information of the decompressed kernel; and weight value embedding logic that selects a weight value embedding corresponding to the position information of the kernel and transmits it to the two-dimensional MAC array.

[0017] Preferably, the weight value generator can further include a transformer weight value memory that stores the implicit weight value; and an on-chip network switch that transmits the implicit weight value via an on-chip network in response to a request of at least one of the multiple transformer operation cores.

[0018] On the one hand, the artificial intelligence neural network acceleration method provided by the present invention includes a number of Transformer operation cores operating based on a Transformer model using n (where n is a natural number) weight values. In the artificial intelligence neural network acceleration method using an artificial intelligence neural network accelerator for accelerating a Transformer neural network, the artificial intelligence neural network accelerator performs a primary prediction step of predicting an output token using a reduced Transformer model (hereinafter referred to as a "reduced model") using m (where m is a natural number and m < n) weight values in response to obtaining an input token; a prediction accuracy calculation step of calculating the prediction accuracy for the prediction result of the primary prediction step; and a secondary prediction step in which the artificial intelligence neural network accelerator further performs a secondary prediction of predicting an output token for the input token using the Transformer model before reduction (hereinafter referred to as a "basic model") when the prediction accuracy is less than or equal to a preset threshold value.

[0019] Preferably, the method includes an implicit weight value generation step in which the artificial intelligence neural network accelerator has a weight value embedding logic generated in advance as a result of learning by matching each weight value of an existing Transformer neural network with the axb kernel position information of the Transformer neural network, and generates an implicit weight value based on the position information of the kernel input from an external memory and the weight value embedding logic; an implicit weight value storage step in which the artificial intelligence neural network accelerator stores the implicit weight value; and an implicit weight value transmission step in which the artificial intelligence neural network accelerator transmits the implicit weight value via an on-chip network. Each of the primary prediction step and the secondary prediction step can predict an output token for the input token using the implicit weight value.

[0020] Preferably, the implicit weight value generation stage further includes a decompression stage for decompressing the position information of the kernel input in a compressed state from the external memory, and applying the position information of the kernel decompressed in the decompression stage to the weight value embedding logic to generate an implicit weight value corresponding to the position of the decompressed kernel.

Advantages of the Invention

[0021] The artificial intelligence neural network accelerator and method thereof of the present invention as described above can accelerate a large-scale language model for mobile devices with high energy efficiency by reducing the amount of external memory access by reducing the required amount of weight values for accelerating the artificial intelligence neural network, and as a result, reducing power consumption.

[0022] In addition, the present invention first accelerates a reduced model in which the required amount of basic weight values necessary for accelerating the basic model is reduced by a predetermined ratio, and further performs a secondary acceleration stage of accelerating the basic model only in a special case where the prediction accuracy for the result is below a preset critical value. As a result, the present invention has the effect of reducing the required amount of weight values for accelerating the artificial intelligence neural network.

[0023] In addition, the present invention is provided with weight value embedding logic generated in advance as a result of learning each weight value of the transformer neural network matched to the kernel position information of the transformer neural network, and generates an implicit weight value for accelerating the artificial intelligence neural network while receiving only the kernel position information from the external memory, and then supplies the implicit weight value using an on-chip network. As a result, the present invention has the effect of reducing the amount of external memory access for receiving weight values.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0025] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings, and will be described in detail so that those having ordinary knowledge in the technical field to which the present invention pertains can easily implement the present invention. However, the present invention can be embodied in various different forms and is not limited to the embodiments described herein. On the other hand, in the drawings, parts not related to the description are omitted in order to clearly describe the present invention, and similar parts throughout the specification are given similar drawing reference numerals. Also, even when detailed descriptions are omitted, descriptions of parts that can be easily understood by those skilled in the art are omitted.

[0026] Throughout the specification and the claims, when one part is said to include one component, this means that, unless otherwise stated to the contrary, it does not exclude other components, but may further include other components.

[0027] FIG. 1 and FIG. 2 are diagrams for explaining the structure of an artificial intelligence neural network accelerator according to an embodiment of the present invention. FIG. 1 shows a schematic block of the artificial intelligence neural network accelerator of the present invention, and FIG. 2 shows the arithmetic unit and weight value generator included in the artificial intelligence neural network accelerator of the present invention in more detail.

[0028] Referring to FIGS. 1 and 2, an artificial intelligence neural network accelerator 100 according to an embodiment of the present invention includes a plurality of gateways 10, 20, a token input unit 110, an arithmetic unit 120, a controller 130, and a weight value generator 200. In particular, in the example of FIG. 2, 48 transformer arithmetic cores 121, two weight value generators 200, and an artificial intelligence neural network accelerator 100 including a top controller, ID SIMD, and an on-chip network are illustrated.

[0029] At this time, the transformer operation core 121 may be configured to include eight multipliers and accumulators (for 8-bit inputs and 8-bit weighted values), an input loader, a controller, an on-chip network switch, and input / weighted value / output memories in order to perform matrix multiplications required for each decoder block of the transformer and generate output tokens each time.

[0030] Each gateway 10, 20 can connect an external memory (not shown) and the artificial intelligence neural network accelerator 100. Each gateway 10, 20 can be used to transmit each weighted value stored in the external memory (not shown) to the artificial intelligence neural network accelerator 100 and transmit each processing result generated by the artificial intelligence neural network accelerator 100 to the external memory (not shown).

[0031] The token input unit 110 receives tokens. That is, the token input unit 110 receives one token each time from an input sentence composed of a large number of tokens (token) and transmits it to the operation unit 120.

[0032] The operation unit 120 predicts output tokens in units of input tokens input via the token input unit 110. For this purpose, the operation unit 120 can include a large number of transformer operation cores 121, and each of the transformer operation cores 121 includes a transformer model.

[0033] At this time, the size of the transformer model is determined by the number of required weighted values, and has the characteristic that the larger the size of the transformer model, the greater the power consumption. That is, the transformer model has the characteristic that the larger the required amount of weighted values, the larger its size. Due to such characteristics, when reducing the required amount of weighted values of an arbitrary transformer model (hereinafter referred to as the "basic model") using n (where n is a natural number) weighted values, the size of the basic model can be reduced, and as a result, the power consumption can be greatly reduced.

[0034] Therefore, the present invention uses such features to reduce a large number of tokens among the entire input tokens to m (where m is a natural number such that m < n) in terms of the required amount of weight values, and predicts using a transformer model with its size reduced (hereinafter referred to as "reduced model"), and as a result, reduces the required amount of weight values for accelerating the artificial intelligence neural network, thereby reducing the external memory access amount and ultimately reducing the power consumption.

[0035] For this purpose, the arithmetic unit 120 receives the control of a controller 130 described later.

[0036] After determining the size of the transformer model based on the number of weight values, the controller 130 can control the operations of each of the transformer arithmetic cores 121. In particular, the controller 130 can control each of the multiple transformer arithmetic cores 121 such that after each of the multiple transformer arithmetic cores 121 makes a primary prediction for the input tokens using the reduced model, calculates the prediction accuracy for the result of the primary prediction, and further makes a secondary prediction using the basic model only when the prediction accuracy is equal to or lower than a preset threshold value (for example, a prediction accuracy of 60%).

[0037] At this time, the prediction accuracy can be calculated using various known techniques. For example, in order to calculate the prediction accuracy, the controller 130 can manage to preferentially calculate a small model and calculate a large model as appropriate.

[0038] For example, the controller 130 causes each of the transformer operation cores 121 to predict an output token linearly using a reduced model in which the required amount of weighted values of the basic model is reduced to n / 10 (where n is a natural number) in response to the acquisition of the input token, and reduces the required amount of weighted values by omitting the secondary prediction process when the prediction accuracy exceeds the threshold value. The processing procedure of such a controller 130 is illustrated schematically in FIG. 3.

[0039] FIG. 3 is a diagram for explaining an acceleration process by a hybrid structure of an artificial intelligence neural network accelerator according to an embodiment of the present invention, and for each of the input tokens acquired by the controller 130, an output token prediction is preferentially performed using a reduced model with a low required amount of weighted values. As a result, only for a predetermined number of input tokens with low prediction accuracy (i.e., difficult to classify), an additional output token prediction is performed using a basic model with a high required amount of weighted values, which is illustrated schematically.

[0040] Referring to FIG. 3, for four input tokens ( <bos>For each of (the / man / walks), the token probabilities of each output token (the / man / walks / runs) predicted using the reduced model (i.e., the transformer model with a weight requirement of "N / 10") 30 are illustrated in the drawing. That is, since the token probabilities (i.e., prediction accuracies) (P) of each output token (the / man / walks) except the rightmost output token (runs) exceed the threshold value (Th), for only the rightmost output token (runs), an additional output token is predicted using the basic model (i.e., the transformer model with a weight requirement of "N") 40, and as a result, it can be seen that an output token (across) with a token probability (P) exceeding the threshold value (Th) is predicted.

[0041] Thus, the present invention predicts output tokens for all of the input tokens using a reduced model with a low weight requirement, and only for a small number of input tokens with low prediction accuracy, additional output tokens are predicted by a basic model with a high weight requirement. Compared with the conventional technique of predicting output tokens for all input tokens by the basic model, the present invention has the feature that the weight requirement can be significantly reduced.

[0042] Thus, first, only the calculation of the small model is performed, and when the prediction probability of a specific token exceeds a predefined threshold value, the calculation of the large model is skipped. When the method of the present invention is applied to language modeling using GPT-2, there is an effect that external memory access can be reduced by 39%.

[0043] Referring to FIGS. 1 and 2 again, the weight generator 200 generates weights provided to each of the transformer operation cores 121.

[0044] The weight value generator 200 generates implicit weight values using an "artificial intelligence neural network" that is trained to implicitly remember the weight values of a transformer neural network, and can receive only the axb kernel position information of the transformer neural network as input and output the implicit weight values corresponding to that position. At this time, the "artificial intelligence neural network" is a neural network different from the transformer neural network, and a multi-layer perceptron, which is usually used to generate the weight values of the transformer neural network, can be trained and used. In this way, the artificial intelligence neural network is trained, and the process of moving data will be described later with reference to FIGS. 5 and 6.

[0045] To generate the implicit weight values, the weight value generator 200 can include an entropy decoding unit 210, an implicit weight value generation unit 220, a transformer weight memory 230, and an on-chip network switch 240.

[0046] The entropy decoding unit 210 decodes the compression of the kernel position information input in a compressed state from an external memory. For this purpose, the entropy decoding unit 210 includes an index memory (IDX_MEM) 211, an MSB memory (MSB_MEM) 212, a sign memory (Sign_MEM) 213, an LSB memory (LSB_MEM) 214, and a router 215 as a configuration for decoding the compression of normal compressed data.

[0047] The index memory (IDX_MEM) 211 stores the addresses of data whose MSB is not filled with consecutive sign bits. The MSB memory (MSB_MEM) 212 stores the MSB partial values of each data (i.e., uncompressed data) where the MSB-side data is not filled with sign extension bits. The sign memory (Sign_MEM) 213 stores the sign bits compressed to 1 bit. The LSB memory 214 stores the LSB data. The router 215 decompresses the data compressed by sign compression according to the respective data information of the index memory (IDX_MEM) 211, the MSB memory (MSB_MEM) 212, the sign memory (Sign_MEM) 213, and the LSB memory 214, and transfers this to an 8-bit queue.

[0048] The implicit weight generation unit 220 includes a two-dimensional MAC array (2D MAC Array) 221 that performs multiplication and accumulation operations for generating implicit weights using the position information of the kernel decompressed by the sign decompression unit 210, and weight embedding logic 222 that is pre-generated as a result of learning by matching each weight value of the transformer neural network with the axb kernel position information of the transformer neural network.

[0049] When the position information of the kernel decompressed by the sign decompression unit 210 is input to the implicit weight generation unit 220, the weight embedding logic 222 can detect the weight embedding corresponding to the position information and generate an implicit weight.

[0050] For this purpose, the weight embedding logic 222 selects the weight embedding corresponding to the position information of the kernel decompressed by the sign decompression unit 210 and transfers it to the two-dimensional MAC array (2D MAC Array) 221 to generate an implicit weight.

[0051] The transformer weight memory 230 stores the implicit weights generated by the implicit weight generation unit 220.

[0052] In response to a request from at least one of a number of transformer operation cores 121, the on-chip network switch 240 transmits the implicit weighted values stored in the transformer weighted value memory 230 to the transformer operation core 121 via the on-chip network. For this purpose, the on-chip network switch 240 can be controlled by the controller 130. That is, the controller 130 can control the operation of the on-chip network switch 240 to provide the implicit weighted values via the on-chip network in response to requests from a number of transformer operation cores 121.

[0053] The processing process of such a weighted value generator 200 is schematically shown in FIG. 4.

[0054] FIG. 4 is a diagram for explaining a process for an artificial intelligence neural network accelerator according to an embodiment of the present invention to generate implicit weight values. Referring to FIG. 4, when sign-compressed weight values with Start IDX being 0 and End IDX being 7 are input to the sign decompression unit 210 to generate the first eight weight values of the neural network, the sign decompression unit 210 receives an address value from the IDX MEM 211 and checks whether there are weight values with MSBs filled with consecutive sign bits among the weight values from 0 to 7. In the example of FIG. 4, since the addresses of the data where the MSBs are not filled with consecutive sign bits are 1 and 6, 1 and 6 are stored in the IDX MEM 211, and the sign decompression unit 210 reads the address values of 1 and 6. In this case, the sign decompression unit 210 determines that the MSBs of the weight values of 0, 2, 3, 5, and 7 are all filled with sign extension bits, while the weight values of 1 and 6 are not. The weight values of 0, 2, 3, 5, and 7 fill the MSB-side data while calling 1-bit compressed sign bits from the Sign MEM 213, and the weight values of 1 and 6 fill the MSB-side data while calling data from the MSB MEM 212. This is to completely supply the MSB-side data to the implicit weight value generation unit 220. At this time, the MSB data router 215 performs the process of calling two pieces of data required for 1 and 6 from the MSB MEM 212. On the other hand, whether to fill the MSB part while calling data from the MSB MEM 212 for each weight value or to fill the MSB while calling data from the LSB MEM 214 changes according to the situation, which is determined by the multiplexer logic connected to the rear end of the MSB data router 215.

[0055] The data decompressed through such a process (i.e., the position information of the decompressed kernel) is stored in an 8-bit queue and then transmitted to a two-dimensional MAC array (2D MAC Array) 221, and an implicit weighted value is generated by detecting the weighted value embedding corresponding to the position information from the weighted value embedding logic 222.

[0056] The implicitly weighted value generated in this way is used in the decoder block operations (not shown) that form the transformer network, specifically, it can be used in attention or feed-forward operations.

[0057] FIG. 5 is a diagram for explaining the process of training a neural network for generating an implicitly weighted value according to an embodiment of the present invention, and FIG. 6 is a diagram for explaining compression and decompression for using the implicitly weighted value according to an embodiment of the present invention.

[0058] First, referring to FIG. 5, in order to train the artificial intelligence neural network, first, the embedding value 310 of a specific a x b kernel based on the position information of the weights of the Transformer neural network is transmitted to the artificial intelligence neural network 320 composed of a multi-layer perceptron. Then, a specific M x N kernel is generated through the operation of the artificial intelligence neural network 320, and a new Transformer neural network 330 can be generated by collecting all these kernels. In such a training process, it is most important how similar the generated new Transformer neural network 330 can be made to the existing Transformer neural network 340. Therefore, for this purpose, the difference (loss) between the results generated by the new Transformer neural network 330 and the existing Transformer neural network 340 is calculated and substituted into the loss function. Then, training is performed on the artificial intelligence neural network 320 that generates the Transformer network through backpropagation. And the artificial intelligence neural network 320 generated through such a training process can be used to reduce external memory access by being used as the weight embedding logic 222 of the implicit weight generation unit 220.

[0059] Referring to FIG. 6, as the movement of data for the use of implicit weight values, when the chip is externally compressed, while the M x N transformer neural network weight values at specific positions are input into the artificial intelligence neural network 320 and an embedding is generated, only the embedding of a small size is stored in the DRAM instead of the M x N transformer neural network weight values, and in the chip, by loading only the embedding, the external memory access can be reduced. When the chip is internally decompressed, the embedding based on the position information of the kernel to be processed is input into the artificial intelligence neural network 320, and the artificial intelligence neural network 320 generates the implicit weight values of the transformer neural network through operations. The implicit weight values of the transformer neural network generated at this time are used for language model work.

[0060] FIGS. 7 and 8 are process flowcharts for explaining an artificial intelligence neural network acceleration method according to an embodiment of the present invention. Hereinafter, with reference to FIGS. 1 to 8, an artificial intelligence neural network acceleration method according to an embodiment of the present invention will be explained.

[0061] First, in step S110, the artificial intelligence neural network accelerator 100 of the present invention generates implicit weight values. That is, in step S110, the artificial intelligence neural network accelerator 100 is provided with weight value embedding logic that has been learned by matching each weight value of an existing transformer neural network with the a x b kernel position information of the transformer neural network, and generates implicit weight values based on the position information of the kernel input from the external memory and the weight value embedding logic.

[0062] For this purpose, in step S111, the implicit weight value generation unit 220 stores weight value embedding logic, and in steps S112 and S113, the compression of the position information of the kernel input in a symbol-compressed state from the external memory is released.

[0063] In step S114, the implicit weight value generation unit 220 applies the position information of the kernel decompressed in step S113 to the weight value embedding logic to generate an implicit weight value corresponding to the position of the decompressed kernel. That is, in step S114, the implicit weight value generation unit 220 generates an implicit weight value based on the decompressed kernel position and the corresponding weight value embedding logic.

[0064] In step S115, the transformer weight memory 230 stores the implicit weight value.

[0065] In steps S120 and S130, in response to the acquisition of input tokens, the transformer arithmetic core 121 performs a primary prediction of predicting output tokens using a reduced transformer model (hereinafter referred to as the "reduced model") that uses m (where m is a natural number such that m < n) weight values. For this purpose, in step S130, the transformer arithmetic core 121 is under the control of the controller 130.

[0066] In step S140, the controller 130 calculates the prediction accuracy for the primary prediction result of step S130.

[0067] Also, in step S150 and step S160, the controller 130 compares the prediction accuracy with a preset threshold value. When the prediction accuracy is equal to or lower than the preset threshold value, the controller 130 further controls the transformer operation core 121 to perform a secondary prediction of predicting an output token for the input token using the transformer model before reduction (hereinafter referred to as the "basic model"). That is, in step S160, the transformer operation core 121 performs the secondary prediction while receiving the control of the controller 130.

[0068] For this purpose, in step S130 and step S160, the transformer operation core 121 may further include an implicit weight value transmission step in which the transformer weight value memory 230 requests an implicit weight value, and in response, the on-chip network switch 240 transmits the implicit weight value via the on-chip network.

[0069] Also, in step S130 and step S160, the transformer operation core 121 predicts an output token for the input token using the implicit weight value.

[0070] In this way, the present invention has the effect of reducing external memory access by performing some operations inside the accelerator instead of calling all weight values for accelerating the transformer neural network from the outside. That is, the present invention can increase energy efficiency while maintaining the accuracy of transformer inference through a mixed structure of large and small artificial intelligence neural networks and implicit weight value generation.

[0071] As an example, in the case of language modeling using GPT-2, first, only the calculation of the small model is performed, and when the prediction probability of a specific token exceeds a predefined threshold, the calculation of the large model can be skipped through the method of the present invention, and the external memory access can be reduced by 39%. Also, by generating implicit weighting values using only the position of the kernel as input, the external memory access can be reduced by 60%, and by applying symbol compression technology to compress the artificial intelligence neural network for generating implicit weighting values, the external memory access can be reduced to 74%.

[0072] As another example, in the case of language translation using mT5, the external memory access can be reduced by 42% through a mixed structure of large and small networks, the external memory access can be reduced by 67% through implicit weighting value generation, and the external memory access can be reduced to 78% through symbol compression technology.

[0073] Also, in the case of summarization using T5, the external memory access can be reduced by 59% through a mixed structure of large and small networks, the external memory access can be reduced by 71% through implicit weighting value generation, and the external memory access can be reduced to 78% through symbol compression technology.

[0074] Finally, in the case of language translation using FSMT, the external memory access can be reduced by 48% through a mixed structure of large and small networks, the external memory access can be reduced by 72% through implicit weighting value generation, and the external memory access can be reduced to 81% through symbol compression technology.

[0075] The effects of the present invention are illustrated schematically in FIG. 9.

[0076] FIG. 9 is a graph for explaining the effects when the artificial intelligence neural network accelerator and its method according to an embodiment of the present invention are applied to various transformer models (for example, GPT-2, mT5, T5, FSMT). FIG. 9(a) shows the reduction amount of the weighted value loading energy of each of the models when the present invention is applied. FIG. 9(b) shows the difference in accuracy when the present invention is applied to each of the models compared with the prior art. FIG. 9(c) shows the difference in the weighted value amount when the present invention is applied to each of the models compared with the conventional technology.

[0077] Referring to FIG. 9(a), when the present invention is applied to GPT-2, the weighted value loading energy is reduced by 71%. When the present invention is applied to mT5, the weighted value loading energy is reduced by 73%. When the present invention is applied to T5, the weighted value loading energy is reduced by 76%. When the present invention is applied to FSMT, the weighted value loading energy is reduced by 76%.

[0078] Also, referring to FIG. 9(b), when the present invention is applied to GPT-2, the accuracy is reduced by 1.26. When the present invention is applied to mT5, the accuracy is reduced by 0.59. When the present invention is applied to T5, the accuracy is reduced by 0.96. When the present invention is applied to FSMT, the accuracy is reduced by 1.29.

[0079] On the other hand, referring to FIG. 9(c), when the present invention is applied to GPT-2, mT5, T5, and FSMT, it can be seen that the weighted value amount is significantly reduced.

[0080] Thus, the present invention shows a weighted value compression rate of 74% to 81%, a reduction rate of weighted value loading energy of 71% to 76%, while showing an accuracy of -0.52 to -1.29 compared with the prior art.

[0081] Thus, the artificial intelligence neural network accelerator and its method of the present invention can accelerate a large-scale language model for mobile devices with high energy efficiency by reducing the amount of external memory access by reducing the required amount of weighted values for accelerating the artificial intelligence neural network, and as a result, reducing power consumption.

[0082] Further, the present invention first accelerates a reduced model in which the required amount of basic weighted values necessary for accelerating the basic model is reduced by a predetermined ratio, and further performs a secondary acceleration step of accelerating the basic model only in a special case where the prediction accuracy for the result is equal to or lower than a preset critical value. As a result, the present invention has the feature that the required amount of weighted values for accelerating the artificial intelligence neural network can be reduced.

[0083] In addition, the present invention includes weighted value embedding logic generated in advance as a result of learning each weighted value of the Transformer neural network matched to the kernel position information of the Transformer neural network. After generating implicit weighted values for accelerating the artificial intelligence neural network while receiving only the kernel position information from the external memory, the present invention supplies the implicit weighted values using an on-chip network, thereby reducing the amount of external memory access for receiving the weighted values.

[0084] In the above, embodiments of the present invention have been described. However, the scope of rights of the present invention is not limited thereto, and the present invention can be easily modified by those having ordinary knowledge in the technical field to which the present invention pertains from the embodiments, and includes all changes and modifications within the scope recognized as equivalent.

Description of Reference Numerals

[0085] 100 Artificial intelligence neural network accelerator 110 Token input unit 120 Arithmetic unit 121 Transformer arithmetic core 130 Controller 200 Weight Value Generator 210 Symbol Decompression Unit 211 IDX_MEM 212 MSB_MEM 213 Sign_MEM 214 LSB_MEM 215 Router 220 Implicit Weight Value Generator 221 2D MAC Array 222 Weight Value Embedding Logic 230 Transformer Weight Memory 240 On-chip Network Switch< / bos>

Claims

1. In an artificial intelligence neural network accelerator for accelerating a transformer neural network, an arithmetic unit including a number of transformer arithmetic cores that operate based on a transformer model that predicts output tokens in units of input tokens and uses n (where n is a natural number) weight values; and a controller that controls the operation of each of the transformer arithmetic cores after determining the size of the transformer model according to the number of the weight values; wherein the controller controls each of the number of transformer arithmetic cores to perform a primary prediction of predicting an output token using a reduced transformer model (hereinafter referred to as a "reduced model") that uses m (where m is a natural number and m < n) weight values in response to acquisition of the input token, and further perform a secondary prediction of predicting an output token for the input token using the transformer model before being reduced (hereinafter referred to as a "basic model") only when the prediction accuracy of the primary prediction result is equal to or lower than a preset critical value. The artificial intelligence neural network accelerator is characterized by the above.

2. It is provided with weight value embedding logic generated in advance as a result of learning by matching each weight value of the transformer neural network with axb kernel position information of the transformer neural network, and further includes a weight value generator that generates implicit weight values based on the position information of the kernel input from an external memory, wherein the controller provides the implicit weight values via an on-chip network in response to a request from the number of transformer arithmetic cores. The artificial intelligence neural network accelerator according to Claim 1.

3. The weight value generator includes an entropy decoding unit that decodes the compression of the position information of the kernel input in a state of being entropy compressed from the external memory; and an implicit weight value generation unit that is provided with the weight value embedding logic and applies the position information of the kernel decompressed by the entropy decoding unit to the weight value embedding logic to generate implicit weight values corresponding to the positions of the decompressed kernels. The artificial intelligence neural network accelerator according to Claim 2.

4. The implicit weight value generation unit is a two-dimensional MAC array that performs multiplication and accumulation operations for generating the implicit weight value using the position information of the decompressed kernel; and weight value embedding logic that selects a weight value embedding corresponding to the position information of the kernel and transmits it to the two-dimensional MAC array; and The artificial intelligence neural network accelerator according to claim 3.

5. The weight value generator is a transformer weight memory for storing the implicit weight value; and an on-chip network switch that transmits the implicit weight value via an on-chip network in response to a request of at least one of the multiple transformer operation cores; and The artificial intelligence neural network accelerator according to claim 3.

6. In an artificial intelligence neural network acceleration method using an artificial intelligence neural network accelerator provided with a number of transformer operation cores operating based on a transformer model using n (where n is a natural number) weight values to accelerate a transformer neural network, a primary prediction step in which the artificial intelligence neural network accelerator performs a primary prediction of predicting an output token using a reduced transformer model (hereinafter referred to as a "reduced model") using m (where m is a natural number less than n) weight values in response to acquisition of an input token; a prediction accuracy calculation step in which the artificial intelligence neural network accelerator calculates a prediction accuracy for the prediction result of the primary prediction step; and a secondary prediction step in which the artificial intelligence neural network accelerator further performs a secondary prediction of predicting an output token for the input token using the transformer model before reduction (hereinafter referred to as a "basic model") when the prediction accuracy is less than or equal to a preset threshold value; An artificial intelligence neural network acceleration method characterized by the above.

7. The artificial intelligence neural network accelerator includes pre-generated weight value embedding logic as a result of learning by matching each weight value of an existing transformer neural network with the axb kernel position information of the transformer neural network, and an implicit weight value generation stage that generates an implicit weight value based on the position information of the kernel input from an external memory and the weight value embedding logic; an implicit weight value storage stage in which the artificial intelligence neural network accelerator stores the implicit weight value; and, an implicit weight value transmission stage in which the artificial intelligence neural network accelerator transmits the implicit weight value via an on-chip network; and further includes, Each of the primary prediction stage and the secondary prediction stage, predicts an output token for the input token using the implicit weight value The artificial intelligence neural network acceleration method according to claim 6.

8. The implicit weight value generation stage, further includes an encoding decompression stage that decompresses the position information of the kernel input in a state of being encoded and compressed from the external memory, applies the position information of the kernel decompressed in the encoding decompression stage to the weight value embedding logic to generate an implicit weight value corresponding to the position of the decompressed kernel The artificial intelligence neural network acceleration method according to claim 7.

Citation Information

Patent Citations

  • Method, device and medium for solving heterogeneous federated learning by using hypernetwork

    CN115860135A

  • Auxiliary model for predicting new model parameters

    US20220147818A1

  • Processing method, processing system, and processing program

    WO2022113175A1