Accelerator configured to execute artificial intelligence calculation, method for operating accelerator, and artificial intelligence system including accelerator

The accelerator addresses the challenge of high data requirements in AI systems by quantizing data to low precision, improving performance and reducing costs through optimized memory usage.

JP2025114508APending Publication Date: 2025-08-05SAMSUNG ELECTRONICS CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025008121
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-13
Filing Date
2025-01-21
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing artificial intelligence systems face challenges in optimizing calculations to reduce cost and improve performance due to the large amount of data required for neural network operations, leading to increased memory bandwidth and power consumption.

Method used

An accelerator is designed with a quantizer to convert high-precision data to low-precision data, reducing the amount of data stored in memory and minimizing memory access requirements, thereby improving performance and reducing costs.

Benefits of technology

The accelerator effectively reduces memory bandwidth and capacity needs, enhancing system performance and lowering implementation costs by quantizing data during AI operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114508000001_ABST
    Figure 2025114508000001_ABST
Patent Text Reader

Abstract

To provide an accelerator configured to execute artificial intelligence calculation having a reduced cost and improved performance, a method for operating the accelerator, and an artificial intelligence system including the accelerator.SOLUTION: According to the present invention, an accelerator configured to execute artificial intelligence calculation, includes: a processing unit configured to execute first calculation for first active data and first weight data loaded from a memory to generate first result data; and a quantizer configured to execute quantization for the first result data to generate first output data. The first active data, the first weight data and the first output data are a low accuracy type, the first result data is a high accuracy type, and the first output data is stored in the memory.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an artificial intelligence system, and more particularly to an accelerator configured to perform artificial intelligence operations, a method of operating the accelerator, and an artificial intelligence system including the accelerator. [Background technology]

[0002] Artificial intelligence is a branch of computer science that has recently been widely applied in areas as diverse as natural language understanding, natural language translation, robotics, artificial vision, problem solving, learning, knowledge acquisition, and cognitive science.

[0003] Artificial intelligence is implemented based on various algorithms. For example, a neural network is a complex network of nodes and synapses repeatedly connected. As data moves from the current node to the next node, various signal processing operations can occur depending on the corresponding synapses. These signal processing operations are called layers. In other words, a neural network can include various layers that are intricately connected to each other. The various layers in a neural network require a large amount of calculations, and various methods for optimizing this are being researched. In other words, because a single layer requires a large amount of calculations, even a small improvement in calculation optimization can have a significant impact on the speed, efficiency, power consumption, etc. of a multi-layer system. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] US Patent Application Publication No. 2022 / 0044114 [Patent Document 2] U.S. Patent No. 11,461,614 [Patent Document 3] U.S. Patent No. 11,556,772 [Patent Document 4] Chinese Patent Publication No. 116050464 [Patent Document 5] Chinese Patent Publication No. 116432712 [Patent Document 6] Chinese Patent Publication No. 116502691 Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention has been made in view of the above-mentioned prior art, and an object of the present invention is to provide an accelerator configured to perform artificial intelligence calculations with reduced cost and improved performance, a method for operating the accelerator, and an artificial intelligence system including the accelerator. [Means for solving the problem]

[0006] According to one embodiment of the present invention, an accelerator configured to perform artificial intelligence (AI) operations includes a processing unit configured to perform a first operation on first activation data and first weight data loaded from a memory to generate first result data, and a quantizer configured to perform quantization on the first result data to generate first output data, wherein the first activation data, the first weight data, and the first output data are of low precision type, and the first result data is of high precision type, and the first output data is stored in the memory.

[0007] According to one embodiment of the present invention, a method for operating an accelerator configured to perform artificial intelligence (AI) operations includes loading first activation data and first weight data from a memory, performing a first operation on the first activation data and the first weight data to generate first result data, performing quantization on the first result data to generate first output data, and storing the first output data in the memory, wherein the first activation data, the first weight data, and the first output data are of a low-precision type, and the first result data is of a high-precision type.

[0008] According to one embodiment of the present invention, an artificial intelligence system includes a memory configured to store first activation data and first weight data, an accelerator configured to load the first activation data and the first weight data from the memory, perform a first operation on the first activation data and the first weight data to generate first result data, and perform quantization on the first result data to generate first output data, and a CPU (Central Processing Unit) configured to control the memory and the accelerator, wherein the first activation data, the first weight data, and the first output data are of low precision type, and the first result data is of high precision type, and the first output data is stored in the memory. [Effects of the Invention]

[0009] According to the present invention, an accelerator can perform operations on an artificial intelligence model. The accelerator may include a quantizer configured to perform quantization on various data generated during the accelerator's learning or inference. This reduces the amount of data (e.g., activations, weights, etc.) generated during the accelerator's learning or inference, thereby reducing the required bandwidth and capacity of a memory configured to store or load the various data. Therefore, an accelerator configured to perform artificial intelligence operations with reduced cost and improved performance, a method for operating the accelerator, and an artificial intelligence system including the accelerator are provided. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram illustrating a system configured to process an artificial intelligence model. [Figure 2] 1 is a block diagram illustrating a system according to one embodiment of the present invention. [Figure 3] FIG. 3 is a conceptual diagram for explaining deep learning performed by the accelerator of FIG. 2. [Figure 4] 3 is a diagram for explaining the concept of a MAC operation performed in the accelerator of FIG. 2. FIG. [Figure 5] 3 is a diagram for explaining a quantization operation performed by a quantizer of the accelerator of FIG. 2. FIG. [Figure 6] FIG. 3 is a block diagram illustrating the accelerator of FIG. 2. [Figure 7] FIG. 7 is a block diagram showing the quantizer of FIG. 6. [Figure 8] 8 is a block diagram illustrating one of the quantization cores of FIG. 7. [Figure 9] FIG. 9 is a block diagram showing the conversion circuit of FIG. 8. [Figure 10] 3 is a flowchart showing the operation of the accelerator of FIG. 2. [Figure 11A]8 is a diagram for explaining the quantization operation performed by the quantizer of FIG. 7. FIG. [Figure 11B] 8 is a diagram for explaining the quantization operation performed by the quantizer of FIG. 7. FIG. [Figure 11C] 8 is a diagram for explaining the quantization operation performed by the quantizer of FIG. 7. FIG. [Figure 12] FIG. 3 is a block diagram illustrating the accelerator of FIG. 2. [Figure 13] FIG. 2 illustrates the structure of an accelerator according to one embodiment of the present invention. [Figure 14] FIG. 2 illustrates the structure of an accelerator according to one embodiment of the present invention. [Figure 15] FIG. 2 is a block diagram showing the configuration of an accelerator according to an embodiment of the present invention. [Figure 16] 1 is a block diagram illustrating a system according to one embodiment of the present invention. [Figure 17] 1 is a block diagram illustrating a system according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following, embodiments of the present invention will be described clearly and in detail so as to enable those skilled in the art to easily practice the present invention.

[0012] Terms such as "unit," "module," and the like used in the detailed description or drawings, or functional blocks illustrated in the drawings, may be implemented in the form of hardware, software, or a combination thereof configured to perform specific functions. As an example, a "computing module" may be a hardware circuit configured to perform the corresponding functions or operations described herein. Furthermore, a processing circuit may include a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a programmable logic unit, a microprocessor, an application-specific integrated circuit (ASIC), or the like, or may include active and / or passive electrical components such as transistors, resistors, capacitors, and / or electronic circuits having one or more of the above-mentioned components.

[0013] FIG. 1 is a block diagram illustrating a system configured to process an artificial intelligence model. Referring to FIG. 1, the system 10 may include a central processing unit (CPU) 11, a memory 12, and an accelerator 13. The CPU 11 can control the general operation of the system. The memory 12 is configured to store various information necessary for the system to operate. The accelerator 13 can perform various learning or inference operations using the artificial intelligence model stored in the memory 12 under the control of the CPU 11.

[0014] In one embodiment, the accelerator 13 performs learning or inference on the artificial intelligence model using the artificial intelligence model (e.g., weights) stored in the memory 12. For example, the accelerator 13 performs repeated multiplication and addition operations on inputs (e.g., activations and weights) to perform learning or inference on the artificial intelligence model. At this time, the accelerator's inputs (e.g., activations and weights), calculation intermediate values, or calculation results are stored in the memory 12, and the accelerator 13 repeatedly accesses the memory 12 to perform learning or inference. In this case, if the size of the data stored in the memory 12 is large, the memory 12 requires a large bandwidth and a large capacity. However, if the size of the data transferred between the accelerator 13 and the memory increases, the calculation speed of the accelerator 13 may decrease or the power consumed by the memory 12 may increase due to limitations in the memory bandwidth.

[0015] In one embodiment, the artificial intelligence model (or weights) stored in memory 12 may be quantized by CPU 11 after learning is complete. Quantization may refer to the operation of converting relatively high-precision data into relatively low-precision data. For example, the first data may have a value expressed in floating point −32 (FP32). In this case, quantization is performed on the first data to convert the first data into a value expressed in integer −8 (Integer-8). Quantization is described in more detail with reference to FIG. 5. When the artificial intelligence model (or weights) is quantized, the amount of data associated with the artificial intelligence model may be reduced.

[0016] As described above, the artificial intelligence model (or weights) stored in the memory 12 is quantized by the CPU 11 and can occupy a relatively small amount of space. This can improve the speed at which the accelerator 13 accesses the artificial intelligence model (or weights) or reduce power consumption for accessing the memory 12. However, when the accelerator 13 performs inference on the artificial intelligence model, the calculation result or intermediate value output from the accelerator 13 may be data with relatively high accuracy. The calculation result or intermediate value is used as input to the next layer of the accelerator. That is, because the accelerator 13 repeatedly accesses relatively high-accuracy or large-capacity data from the memory 12, a high bandwidth and large capacity are still required for the memory 12, and power consumption by the memory 12 increases. Therefore, the cost of implementing the system 10 increases or the performance of the system 10 decreases.

[0017] FIG. 2 is a block diagram illustrating a system according to an embodiment of the present invention. FIG. 3 is a conceptual diagram illustrating deep learning performed by the accelerator of FIG. 2. Referring to FIGS. 2 and 3, system 100 may include memory 101, controller 102, and accelerator 1000. In one embodiment, system 100 may be dedicated hardware configured to perform processing of an artificial intelligence model. For example, system 100 may be a graphics processing unit (GPU), a neural processing unit (NPU), or other dedicated hardware. In one embodiment, system 100 may be included in an application processor (AP) or a mobile device.

[0018] In one embodiment, the artificial intelligence model driven by system 100 is generated through machine learning. Machine learning may include various learning methods such as supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning, but the scope of the present invention is not limited thereto.

[0019] In one embodiment, the artificial intelligence model is generated or trained through one or a combination of at least two or more of various neural networks, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, etc. The artificial intelligence model may include multiple neural network layers, each of which may be configured to perform an artificial intelligence operation based on a trained model or weights. In one embodiment, for example, system 10 may be applied to smartphones, tablet devices, smart TVs, AR devices, IoT devices, autonomous vehicles, robots, medical devices, drones, advanced driver assistance systems (ADAS), image display devices, data processing servers, measurement devices, etc., which use neural networks to perform speech recognition, image recognition, image classification, and image processing, and may be implemented in one of various types of electronic devices.

[0020] Hereinafter, the term "artificial intelligence operation" will be used to facilitate the description of embodiments of the present invention. "Artificial intelligence operation" may refer to various operations performed within system 100 to train an artificial intelligence model or infer some result. As an example, the artificial intelligence operation may include multiply-accumulate (MAC) operations performed at various layers of the artificial intelligence model.

[0021] For example, as shown in FIG. 3, the system 100 may operate based on a deep neural network. The deep neural network may include an input layer IL, a hidden layer HL, and an output layer OL. The input layer IL may be configured to receive input data or input features F1-F4. The input features F1-F4 received through the input layer IL are operated on (e.g., weighted multiplication) based on corresponding weights, and the results are transmitted to the next layer (e.g., hidden layer HL). The hidden layer HL is configured to perform weighted multiplication on various inputs and accumulate the results. The results of the hidden layer HL are provided to the output layer OL. The output layer OL may perform multiplication or accumulation on the results of the hidden layer HL and output inference results IFR1 and IFR2. As described above, the deep neural network may repeatedly perform multiplication and accumulation operations (i.e., MAC operations) on the input data and weights. However, the scope of the present invention is not limited in this respect, and the system 100 may be configured to perform various operations.

[0022] The memory 101 is configured to store various data, weights, parameters, etc. required for the artificial intelligence calculations of the system 100. For example, the memory 101 may store an artificial intelligence model for the artificial intelligence calculations of the system 100. The artificial intelligence model may include various weight information. In one embodiment, the memory 101 may be a dynamic random access memory (DRAM). However, the scope of the present invention is not limited in this respect, and the memory 101 may include various types of memory such as an SRAM, a PRAM, an MRAM, a RRAM, an FRAM (registered trademark), a flash memory, etc.

[0023] The accelerator 1000 is configured to perform an artificial intelligence operation using data, weights, or parameters stored in the memory 101. In one embodiment, the accelerator 1000 may include multiple processing elements (PEs) for the artificial intelligence operation. Each of the multiple processing elements may be configured to perform a multiply-accumulate (MAC) operation on the data, weights, or parameters stored in the memory 101. An artificial intelligence model may be trained or a specific result may be inferred based on the operation results of the multiple processing elements.

[0024] The controller 102 is configured to control the memory 101 and the accelerator 1000. In one embodiment, the controller 102 may be a CPU (central processing unit) configured to control the general operation of the system 100.

[0025] In one embodiment, as described with reference to FIG. 1, the artificial intelligence model (or weights) stored in memory 101 may be quantized by controller 102 (or CPU). Meanwhile, accelerator 1000 may perform calculations based on a data type with relatively high precision to improve the accuracy of artificial intelligence inference. In this case, the calculation results or calculation intermediate values by accelerator 1000 will be of a data type with relatively high precision. In this case, as described with reference to FIG. 1, the calculation results or calculation intermediate values of accelerator 1000 are stored in memory 101, which may result in a decrease in memory access speed or an increase in power consumption.

[0026] The accelerator 1000 according to an embodiment of the present invention may include a quantizer 1100. The quantizer 1100 may be configured to perform quantization on an intermediate value or a result of an operation of the accelerator 1000. That is, the accelerator 1000 may perform quantization on an intermediate value or a result of an operation generated during an AI operation or in an inference process. In this case, the amount of intermediate value or result stored in the memory 101 is reduced, thereby reducing the bandwidth and capacity required for the memory 101. Alternatively, the power consumption of the memory 101 may be reduced. Therefore, the implementation cost of the system 100 may be reduced, or the performance of the system 100 may be improved.

[0027] For example, inference in an AI system for a large-scale language model such as the recent Ghat-GPT requires large amounts of data to be loaded from or stored in memory. In this case, if memory resources are limited, a bottleneck in memory access may occur, resulting in a decrease in overall performance of the AI system. According to the present invention, a quantizer included in the accelerator quantizes operation result data or subsequent activation data within the accelerator, thereby reducing the total amount of data stored in or loaded from memory. This prevents bottlenecks in limited memory resources, providing an accelerator or AI system with improved performance and reduced costs.

[0028] Fig. 4 is a diagram illustrating the concept of a MAC operation performed in the accelerator of Fig. 2. In one embodiment, the MAC operation may include one multiplication operation and one accumulation operation on two pieces of data.

[0029] For example, as shown in FIG. 4, multiplication of an N-bit weight by an N-bit activation is performed. In this case, the multiplication result has a size of 2N bits. Then, as the multiplication results are accumulated, the accumulated result may have a size of (2N+M) bits. As the size of the accumulated result increases, the memory capacity required to define the accumulated result increases, so the size of the accumulated result needs to be reduced. Thus, the accumulated result is normalized with N-bit output data. The normalized output data may be stored again in memory 101. The output data stored in memory 101 is used as the activation for the next MAC operation.

[0030] In one embodiment, the data for the MAC operation may have various data types. For example, the data for the MAC operation may be an integer type. Alternatively, the data for the MAC operation may be a floating-point type. The floating-point type is a method of expressing data in the form of a sign, a fraction, and an exponent. Floating-point types include 32-bit single precision and 64-bit double precision. Depending on the type and size of the data, the accuracy of the operation result, the area and power consumption of the hardware structure may vary. Therefore, the type and size of the data may be determined in various ways depending on the purpose of the system 100.

[0031] In one embodiment, it is assumed that the MAC operation is performed based on floating-point type. In this case, input data (e.g., weights and activations) and output data (i.e., MAC operation results) have floating-point type. As described above, floating-point type has relatively high precision, but requires a relatively large number of bits to represent one piece of information. In this case, power consumption increases or a relatively long time is required when loading input data from memory 101 or storing output data to memory 101.

[0032] According to an embodiment of the present invention, the accelerator 1000 includes a quantizer 1100 that performs quantization on output data to represent the output data with a relatively small number of bits. The output data with a relatively small number of bits is stored in the memory 101, and the output data stored in the memory 101 can be loaded into the accelerator 1100 as the activity for the next MAC operation. That is, the quantizer 1100 of the accelerator 1000 can quantize an intermediate value or an operation result generated during the AI operation process of the accelerator 1000, thereby reducing the size of information stored in the memory 101. This can improve the overall performance of the system 100 or reduce power consumption.

[0033] 5 is a diagram illustrating a quantization operation performed by the quantizer of the accelerator of FIG. 2. Referring to FIG. 5, data may be represented in various ways. For example, data may be represented as a high precision type (HP). In one embodiment, the high precision type may refer to a floating-point type that represents data in the form of a sign, a fraction, and an exponent. Alternatively, data may be represented as a low precision type (LP). In one embodiment, the low precision type LP may refer to an integer type that represents data in integer format. In one embodiment, the low precision type LP may represent a single piece of data using a relatively smaller number of bits but may have a relatively lower precision than the high precision type HP.

[0034] In one embodiment, the high-precision type HP may include data types represented by a relatively large number of bits and having a relatively high precision, such as the Brain Floating Point Format (BF16) type, the half-precision IEEE Floating Point Format (FP16) type, the single-precision floating-point format (FP32) type, the double-precision floating-point format (FP64) type, etc. The low-precision type LP may include data types represented by a relatively small number of bits and having a relatively low precision, such as INT4, INT8, INT16, etc. In one embodiment, the low-precision type LP may be realized by a combination of integer data types such as the INT4 type, the INT8 type, the INT16 type, etc. and floating-point data types such as the BF16 type, the FP16 type, the FP32 type, etc., but represented by a relatively small number of bits compared to the high-precision type HP.

[0035] The quantizer 1100 can convert or downsize high-precision type data to low-precision type data. For example, the quantizer 1100 can convert high-precision type HP data to low-precision type LP data. In this case, the number of bits required to represent the data can be reduced. In this case, the amount of data stored in or loaded from the memory 101 can be reduced, thereby reducing memory bandwidth and memory capacity requirements.

[0036] Figure 6 is a block diagram illustrating the accelerator of Figure 2. Referring to Figures 2 and 6, the accelerator 1000 may include a quantizer 1100 and a processing unit 1200.

[0037] The processing unit 1200 can load the activation data ACT and the weight data WT stored in the memory 101 and perform an AI calculation on the loaded activation data ACT and the weight data WT. For example, the processing unit 1200 can repeatedly perform a MAC calculation on the activation data ACT and the weight data WT and output calculation result data RST.

[0038] In one embodiment, the processing unit 1200 may perform an AI calculation based on a high-precision HP. For example, the processing unit 1200 may perform a MAC calculation on the activity data ACT and the weight data WT based on an FP16 type, which is a type of high-precision HP. In this case, the calculation result RST output by the processing unit 1200 has a high-precision HP type.

[0039] The quantizer 1100 can quantize the operation result data RST of the processing unit 1200 to generate output data OUT. For example, as described above, the operation result data RST can be of high precision type HP. The quantizer 1100 can quantize the operation result data RST of high precision type HP to generate output data OUT of low precision type LP. The output data OUT of low precision type LP may be stored in the memory 101.

[0040] In one embodiment, accelerator 1000 iteratively performs artificial intelligence operations through multiple layers, as described with reference to Figure 4. The output data OUT stored in memory 101 is used as input (e.g., activity data ACT) for the artificial intelligence operation for the next layer of accelerator 1000. In this case, because the output data OUT stored in memory 101 is of low-precision type LP, the activity data ACT input to processing unit 1200 of accelerator 1000 for the artificial intelligence operation for the next layer will also be of low-precision type LP.

[0041] As described above, the quantizer 1100 of the accelerator 1000 can quantize the operation result data RST of the processing unit 1200 to generate output data OUT. The output data OUT is stored in the memory 101, and the output data OUT stored in the memory 101 can be used as input (i.e., activation data ACT) to the next layer of the accelerator 1000. In this case, the data stored in or loaded from the memory 101 is low-precision type LP with a relatively small capacity, so that the bandwidth and capacity required for the memory 101 can be reduced.

[0042] Figure 7 is a block diagram showing the quantizer of Figure 6. Referring to Figures 6 and 7, the quantizer 1100 may include a round robin switch 1110, a plurality of quantization cores 1120 to 112n, and a control logic circuit 1130.

[0043] The round robin switch 1110 may receive the calculation result RST as an input INPUT from the processing unit 1200. In one embodiment, the calculation result RST (i.e., the input INPUT) may be a high-precision type HP. The round robin switch 1110 may sequentially provide the calculation result RST to the plurality of quantization cores 1120-112n in a round robin manner. For example, each of the plurality of quantization cores 1120-112n may perform quantization on a predetermined number (e.g., k) of data. In this case, the round robin switch 1110 may provide k calculation results RST to each of the plurality of quantization cores 1120-112n in a round robin manner.

[0044] Each of the quantization cores 1120 to 112n may perform a quantization operation on input data. For example, each of the quantization cores 1120 to 112n may include various arithmetic modules or arithmetic units for performing the quantization operation on the input data. Each of the quantization cores 1120 to 112n may perform the quantization operation on the input data using the arithmetic modules or arithmetic units and generate output data OUT. The output data OUT may be of low precision type LP. The generated output data OUT may be provided to the round robin switch 1110. The round robin switch 1110 may output the output data OUT to the memory 101.

[0045] The control logic circuit 1130 is configured to control each of the multiple quantization cores 1120 to 112n. For example, each of the multiple quantization cores 1120 to 112n can operate in parallel or independently. The control logic circuit 1130 can control the operation timing of each of the multiple quantization cores 1120 to 112n.

[0046] Alternatively, each of the quantization cores 1120-112n may perform quantization based on various algorithms. The operation modules and operation orders performed by each of the quantization cores 1120-112n may differ depending on the quantization algorithm performed by each of the quantization cores 1120-112n. The control logic circuit 1130 may individually control the operation modules of each of the quantization cores 1120-112n depending on the quantization algorithm performed by each of the quantization cores 1120-112n.

[0047] Fig. 8 is a block diagram showing one quantization core among the multiple quantization cores in Fig. 7. For convenience of explanation, the 0th quantization core 1120 will be described with reference to Fig. 8, but the scope of the present invention is not limited thereto, and the other quantization cores 1121 to 112n may also have a structure similar to that of the 0th quantization core 1120.

[0048] 8, the quantization core 1120 may include an input reformatter 1120a, an output reformatter 1120b, and a conversion circuit 1120c. The input reformatter 1120a may be configured to receive input data INPUT from the round robin switch 1110 and change the format of the received input data INPUT. In one embodiment, the input data INPUT may be result data RST generated from the processing unit 1200. The input data INPUT may include multiple pieces of result data RST. The number of pieces of result data RST included in one piece of input data INPUT may be the number of pieces of data that can be quantized through one quantization in the quantization core 1120.

[0049] The input reformatter 1120a may include a first-in-first-out unit FIFO, a transposition unit TRSP, a scalar-vector duplicator REPC, and a first register RGST1. The first-in-first-out unit FIFO is configured to perform first-in-first-out on the input data INPUT received from the round robin switch 1110. The transposition unit TRSP may be configured to calculate a permutation of the input data INPUT. For example, if the input data INPUT is in the form of a vector consisting of one row and four columns, the transposition unit TRSP can perform a transposition on the input data INPUT to generate transposed data consisting of four rows and one column. The scalar-to-vector duplicator REPC can duplicate the input data INPUT, which is a scalar value, and convert it into a vector value. The first register RGST1 is configured to store a value, data, or vector generated by the input reformatter 1120a or an intermediate value generated by the conversion circuit 1120c. Each of the first-in-first-out unit FIFO, the transpose unit TRSP, the scalar-to-vector duplicator REPC, and the first register RGST1 can communicate with the other components of the input reformatter 1120a. For example, each of the first-in-first-out unit FIFO, the transpose unit TRSP, the scalar-to-vector duplicator REPC, and the first register RGST1 can perform unidirectional, bidirectional, and / or broadcast communication with other components in a serial and / or parallel manner to send and receive information, data, and / or commands. Information may be encoded in various formats, such as analog and / or digital formats.

[0050] The output reformatter 1120b is configured to store the results converted by the conversion circuit 1120c (i.e., quantized result data) and output the converted results (i.e., output data OUT) to the round robin switch 1110. In one embodiment, the output data OUT may include multiple quantized data. The multiple quantized data may be data obtained by quantizing multiple pieces of result data RST, respectively.

[0051] The output reformatter 1120b may include a second register RGST2 and an address selector ADDR. The second register RGST2 may be configured to store output data OUT. The address selector ADDR may select or control an address of the second register RGST2 so that the output data OUT stored in the second register RGST2 is output to the round robin switch 1110. The second register RGST2 may communicate with the address selector ADDR. For example, each of the second register RGST2 and the address selector ADDR may perform unidirectional, bidirectional, and / or broadcast communication with other components in a serial and / or parallel manner to transmit and receive information, data, and / or commands. The information may be encoded in various formats, such as an analog format and / or a digital format.

[0052] The conversion circuit 1120c receives input data INPUT or modified input data INPUT from the input reformatter 1120a and performs various operations on the modified input data INPUT. The various operations may include a quantization operation on the input data INPUT or modified input data INPUT. Intermediate data generated during the operation of the conversion circuit 1120c is stored in a first register RGST1 of the input reformatter 1120a. Once quantization is complete by the conversion circuit 1120c, the conversion circuit 1120c may provide output data OUT to the output reformatter 1120b.

[0053] In one embodiment, the transform circuit 1120c may include various computation modules to support various quantization algorithms. The transform circuit 1120c may execute the various computation modules under the control of the control logic circuit 1130.

[0054] Figure 9 is a block diagram showing the transform circuit of Figure 8. Referring to Figures 2, 8, and 9, the transform circuit 1120c may include various operation modules to support various quantization algorithms. For example, the transform circuit 1120c may include a sign operation module 1120c-1, a scalar operation module 1120c-2, a vector-scalar operation module 1120c-3, and a vector-vector operation module 1120c-4.

[0055] The sign operation module 1120c-1 is configured to manage the sign of the input data INPUT (or data received from the input reformatter 1120a) or perform sign-related operations. The sign operation module 1120c-1 may include a sign extraction unit SIGN-EXT, a sign inversion unit SIGN-INV, and an absolute value unit ABS.

[0056] The sign extraction unit SIGN-EXT may be configured to extract the sign of the input data INPUT. For example, the input data INPUT may include multiple pieces of result data RST. The sign extraction unit SIGN-EXT can extract the sign of each of the multiple pieces of result data RST included in the input data INPUT and generate code data corresponding to the extracted sign. In one embodiment, the code data may have the same format (e.g., vector or scalar) as the input data INPUT. For example, if the input data INPUT has a vector format of [2.66 1.05 -0.07 0.65], the code data may have a vector format of [1 1 -1 1]. Each of the above-mentioned code data is represented by one bit (i.e., 1 if the corresponding data is positive and 0 if negative).

[0057] The sign inverter SIGN-INV may be configured to invert the sign of the input data INPUT. For example, the sign inverter SIGN-INV may invert the sign of each of the plurality of result data RST included in the input data INPUT to generate inverted data. For example, if the input data INPUT has a vector format of [2.66 1.05 -0.07 0.65], the inverted data may have a vector format of [-2.66 -1.05 0.07 -0.65].

[0058] The absolute value unit ABS may be configured to extract an absolute value of the input data INPUT. For example, the absolute value unit ABS may extract the absolute value of each of the multiple result data RST included in the input data INPUT to generate absolute value data. As an example, if the input data INPUT has a vector format of [2.66 1.05 -0.07 0.65], the absolute value data may have a vector format of [2.66 1.05 0.07 0.65].

[0059] The scalar operation module 1120c-2 is configured to perform a scalar operation on the input data INPUT. For example, the input data INPUT may have a vector format. In this case, the scalar operation module 1120c-2 may be configured to perform a scalar operation on one piece of data included in the input data INPUT. The scalar operation module 1120c-2 may include a reciprocal part RCP and a precision control part PRC.

[0060] The reciprocal part RCP can calculate the reciprocal of one piece of data included in the input data INPUT. The precision control part PRC can change the precision of one piece of data included in the input data INPUT.

[0061] The vector-scalar operation module 1120c-3 can perform a vector-scalar operation on the input data INPUT. For example, the input data INPUT may have a vector format. In this case, the vector-scalar operation module 1120c-3 can perform a vector-scalar operation on the input data INPUT to generate data in a scalar format.

[0062] The vector-scalar operation module 1120c-3 may include an addition tree unit ADD1, a minimum value search unit MIN, and a maximum value search unit MAX. The addition tree unit ADD1 may perform addition on multiple data included in the input data INPUT and output the sum data. For example, if the input data INPUT is [2.66 1.05 -0.07 0.65], the sum data may be 4.29. The minimum value search unit MIN may search for the minimum value among the multiple data included in the input data INPUT and output the minimum value data. For example, if the input data INPUT is [2.66 1.05 -0.07 0.65], the minimum value data may be -0.07. The maximum value search unit MAX may search for the maximum value among the multiple data included in the input data INPUT and output the maximum value data. For example, if the input data INPUT is [2.66 1.05 -0.07 0.65], the maximum value data may be 2.66.

[0063] The vector-vector operation module 1120c-4 can perform a vector-vector operation on the input data INPUT. For example, the input data INPUT may have a vector format. The vector-vector operation module 1120c-4 can perform a vector-vector operation on the input data INPUT and other data in vector format to generate data in vector format.

[0064] The vector-vector operation module 1120c-4 may include an adder ADD2, a multiplier MU, and a shifter SFT. The adder ADD2 may perform an addition operation on two vector data and output the sum data. For example, if the two vector data are [2.66 1.05 -0.07 0.65] and [1.11 1.11 -1.11 1.11], the sum data may be [3.77 2.16 -1.18 1.76].

[0065] The multiplication unit MUL may perform a multiplication operation on two vector data. In this case, the multiplication unit MUL may perform an inner product (scalar product) operation or an outer product (vector product) operation on two vector data. Alternatively, the multiplication unit MUL may perform a scalar product operation on one vector data. For example, if one vector data is [1 1 -1 1] and is multiplied by 1.11, the multiplied data may be [1.11 1.11 -1.11 1.11]. The shifter unit SFT may be configured to perform a shift operation on the vector data.

[0066] As described above, the conversion circuit 1120c can perform various operations for quantization of the input data INPUT. Each component provided in the code management module 1120c-1, the scalar operation module 1120c-2, the vector-scalar operation module 1120c-3, and the vector-vector operation module 1120c-4 can communicate with each other component provided in the code management module 1120c-1, the scalar operation module 1120c-2, the vector-scalar operation module 1120c-3, and the vector-vector operation module 1120c-4. For example, the sign extractor SIGN-EXT, the sign inverter SIGN-INV_, and / or the absolute value unit ABS can perform unidirectional, bidirectional, and / or broadcast communication with the components included in the scalar operation module 1120c-2, the vector-scalar operation module 1120c-3, and / or the vector-vector operation module 1120c-4 to transmit and receive information, data, and / or commands. In one embodiment, the quantization core 1120 may perform quantization based on various quantization algorithms, and the types or order of operation modules executed in the quantization core 1120 may vary depending on the quantization algorithm being performed. The control logic circuit 1130 may control each operation module of the quantization core 1120 to suit the quantization algorithm being performed.

[0067] Figure 10 is a flowchart illustrating the operation of the accelerator of Figure 2. For convenience of explanation, an embodiment in which one artificial intelligence operation is performed by the accelerator 1000 will be described with reference to Figure 10. That is, the accelerator 1000 can repeatedly or in parallel perform the operations of the flowchart of Figure 10 for inference on specific data, and the result data or output data generated by the operations of the flowchart of Figure 10 is used as input (i.e., activation data ACT) in the next operation process.

[0068] 2, 6, and 10, in step S110, the accelerator 1000 may load the activity data ACT and the weights WT. For example, the accelerator 1000 may load the activity data ACT and the weights WT from the memory 101.

[0069] In step S120, the accelerator 1000 may perform inverse quantization on the activity data ACT and the weights WT. For example, the activity data ACT and the weights WT loaded from the memory 101 may be of low-precision type LP. Meanwhile, the processing unit 1200 of the accelerator 1000 may perform a MAC operation based on high-precision type HP. In this case, the activity data ACT and the weights WT loaded from the memory 101 may be converted to high-precision type HP. In one embodiment, the high-precision type HP may be of BF16 type, FP16 type, or FP32 type.

[0070] In one embodiment, the operation of step S120 is omitted according to the calculation algorithm of the processing unit 1200. For example, as described below, when the quantizer 1100 performs quantization based on a binary-coding based quantization (BCQ) algorithm and the processing unit 1200 performs calculations based on non-general matrix to matrix multiplication for binary-coding based quantized neural networks (BiQGEMM), the calculations can be performed without separate conversion or inverse quantization of the activity data ACT and the weight data WT.

[0071] In step S130, the accelerator 1000 may perform an operation on the activation data ACT and the weight data WT. For example, the processing unit 1200 of the accelerator 1000 may perform a MAC operation on the activation data ACT and the weight data WT to generate result data RST. In one embodiment, the processing unit 1200 performs the MAC operation based on a high-precision type HP, and therefore the result data RST generated by the processing unit 1200 may be of a high-precision type HP.

[0072] In step S140, the accelerator 1000 may perform quantization on the result data RST of the operation. For example, the processing unit 1200 may perform an operation based on a high-precision type HP. In this case, the result data RST calculated by the processing unit 1200 may be of a high-precision type HP. The quantizer 1100 of the accelerator 1000 may quantize the result data RST of a high-precision type HP and convert it into output data OUT of a low-precision type LP. In one embodiment, the result data RST may be of a BF16 type, an FP16 type, or an FP32 type, and the output data OUT may be of an INT8 type or a combination of an INT8 and an FP16 type. In this case, the overall size or capacity of the output data OUT may be smaller than the overall size or capacity of the result data RST.

[0073] In step S150, the accelerator 1000 may store the quantized output data OUT in the memory 101. In one embodiment, the output data OUT stored in the memory 101 is used as input (i.e., activation data ACT) for the next calculation operation of the accelerator 1000.

[0074] As described above, the accelerator 1000 according to an embodiment of the present invention can perform quantization on data (e.g., activity data ACT, weight data WT, or calculation result data RST) generated during the learning or inference process. In this case, the size or capacity of data stored in or loaded from the memory 101 is reduced, thereby reducing the required bandwidth or capacity for the memory 101 and reducing the power consumption of the memory 101. Therefore, an accelerator 1000 with reduced cost and improved performance is provided.

[0075] 11A to 11C are diagrams illustrating the quantization operation performed by the quantizer of FIG. 7. For convenience of explanation, it is assumed that the quantizer 1100 performs non-uniform quantization on 16 FP16 type data. The quantization techniques described with reference to FIGS. 11A to 11C are merely examples, and the scope of the present invention is not limited thereto. The quantizer 1100 of the accelerator 1000 according to the present invention can perform quantization based on various quantization algorithms.

[0076] 2, 7, 8, 9, 11A, 11B, and 11C, 16 pieces of result data are generated by the calculation operation of the processing unit 1200. Each of the 16 pieces of result data may be of FP16 type. The 16 pieces of result data are input to the quantizer 1100. The round robin switch 1110 of the quantizer 1100 may provide the 16 pieces of result data to the 0th quantization core 1220 as the 0th input data INPUT0.

[0077] First, referring to Figure 11A, the 0th input data INPUT0 including 16 pieces of result data is arranged in a vector format as shown in Figure 11A by the input reformatter 1120a. For example, the scalar-to-vector duplicator REPC of the input reformatter 1120a can arrange the 16 pieces of result data in a vector format like the 0th input data INPUT0 shown in Figure 11A. The 0th input data INPUT0 may be temporarily stored in a first register RGST1 of the input reformatter 1120a.

[0078] The conversion circuit 1120c may then perform a quantization operation on the 0th input data INPUT0. As an example, the conversion circuit 1120c may perform an absolute value operation and an average operation on the 0th input data INPUT0 row by row to generate the 0th average data a0. For example, the absolute value unit ABS of the sign operation module 1120c-1 of the conversion circuit 1120c may perform an absolute value operation on the 0th input data INPUT0 to generate absolute value data. The addition unit ADD1 of the vector-scalar operation module 1120c-3 of the conversion circuit 1120c may then perform an addition operation on the absolute value data row by row to generate sum data. The multiplication unit MUL of the vector-vector operation module 1120c-3 of the conversion circuit 1120c may then perform a division operation on the sum data by the number of elements in each row of the 0th input data INPUT0 to generate an average value for each row. The average value for each row is stored in the first register RGST1 of the input reformatter 1120a. The scalar-to-vector duplicator REPC of the input reformatter 1120a duplicates the average value of each row to generate the 0th average data a0 in vector format, which is stored in the first register RGST1 of the input reformatter 1120a.

[0079] In one embodiment, the zeroth average data a0 may include four pieces of data. At this time, each of the four pieces of data included in the zeroth average data a0 may be temporarily stored in the first register RGST1 of the input reformatter 1120a in an FP16 type format. That is, the zeroth average data a0 may have a size of 16×4=64 bits.

[0080] The conversion circuit 1120c performs sign extraction on the 0th input data INPUT0 and can generate the 0th code data b0. For example, the sign extraction unit SIGN-EXT of the sign operation module 1120c-1 of the conversion circuit 1120c can perform sign extraction on the 0th input data INPUT0 and generate the 0th code data b0 in vector format.

[0081] In one embodiment, the 0th code data b0 may include 16 pieces of data. At this time, since the 16 pieces of data of the 0th code data b0 represent positive or negative numbers, each of the 16 pieces of data is temporarily stored as 1 bit (i.e., 1 or 0) in the first register RGST1 of the input reformatter 1120a. That is, the 0th code data b0 may have a size of 16×1=16 bits.

[0082] The conversion circuit 1120c can multiply the 0th average data a0 and the 0th code data b0 to generate the a-th intermediate data INTa. For example, the multiplication unit MUL of the vector-vector operation module 1120c-4 of the conversion circuit 1120c can perform scalar multiplication on the elements (elements) of the 0th average data a0 and the rows of the 0th code data b0 to generate the a-th intermediate data INTa. That is, since the 0th intermediate data INTa is expressed as the product of the 0th average data a0 and the b-th code data b0, when the 0th average data a0 and the b-th code data b0 are stored in the first register RGST1 of the input reformatter 1120a, the a-th intermediate data INTa is generated.

[0083] 11B, the conversion circuit 1120c can perform a subtraction operation on the 0th input data INPUT0 and the 0th intermediate data INTa to generate the first intermediate data INT1. For example, the sign inversion unit SIGN-INV of the sign operation module 1120c-1 of the conversion circuit 1120c can perform a sign inversion on the 0th intermediate data INTa to generate inverted data. Then, the addition unit ADD2 of the vector-vector operation module 1120c-4 of the conversion circuit 1120c can perform an addition on the 0th input data INPUT0 and the inverted data to generate the first intermediate data INT1.

[0084] Thereafter, the conversion circuit 1120c may perform row-wise absolute value calculation and average calculation on the first intermediate data INT1 to generate first average data a1, and perform sign extraction on the first intermediate data INT1 to generate first sign data b1.

[0085] 11A, row-by-row absolute value and average operations are performed by the absolute value unit ABS of the sign operation module 1120c-1, the adder unit ADD1 of the vector-scalar operation module 1120c-3, the multiplier unit MUL of the vector-vector operation module 1120c-3, etc., and the sign extraction operation is performed by the sign extraction unit SIGN-EXT of the sign operation module 1120c-1, a detailed description of which will be omitted.

[0086] 11A, the first average data a1 may include four pieces of data, each of which is temporarily stored in the first register RGST1 of the input reformatter 1120a in the form of an FP16 type. That is, the first average data a1 may have a size of 16×4=64 bits.

[0087] 11A, the first code data b1 may include 16 pieces of data, and since the 16 pieces of data represent positive or negative numbers, each of the 16 pieces of data is temporarily stored as one bit (i.e., 1 or 0) in the first register RGST1 of the input reformatter 1120a. That is, the first code data b1 may have a size of 16×1=16 bits.

[0088] The conversion circuit 1120c may generate the bth intermediate data INTb by calculating the sum of the multiplication result of the 0th average data a0 and the 0th code data (i.e., a0 × b0) and the multiplication result of the first average data a1 and the first code data b1 (i.e., a1 × b1). As described with reference to FIG. 11A, the multiplication of each data is performed by the multiplication unit MUL of the vector-vector operation module. The addition operation of the multiplication result (i.e., a0 × b0 + a1 × b1) is performed by the addition unit ADD2 of the vector-vector operation module 1120c-4.

[0089] 11B, since the b-th intermediate data INTb is expressed as a0×b0+a1×b1, when the 0-th average data a0, the 0-th code data b1, the 1st average data a1, and the 1st code data b2 are stored in the first register RGST1 of the input reformatter 1120a, the b-th intermediate data INTb is generated. That is, the b-th intermediate data INTb is expressed using 16×4+1×16+16×4+1×16=160 bits.

[0090] 11B, the conversion circuit 1120c may perform a subtraction operation on the 0th input data INPUT0 and the bth intermediate data INTb to generate second intermediate data INT2. Thereafter, the conversion circuit 1120c may perform a row-wise absolute value operation and an average operation on the second intermediate data INT2 to generate second average data a2, and perform sign extraction on the second intermediate data INT2 to generate second code data b2.

[0091] As described with reference to Figures 11A and 11B, the subtraction operation is performed by the sign inversion unit SIGN-INV of the sign operation module 1120c-1 and the addition unit ADD2 of the vector-vector operation module 1120c-4, the row-wise absolute value operation and average operation are performed by the absolute value unit ABS of the sign operation module 1120c-1, the addition unit ADD1 of the vector-scalar operation module 1120c-3, the multiplication unit MUL of the vector-vector operation module 1120c-3, etc., and the sign extraction operation is performed by the sign extraction unit SIGN-EXT of the sign operation module 1120c-1, and detailed description thereof will be omitted.

[0092] The conversion circuit 1120c can generate output data OUT by multiplying the 0th average data a0 by the 0th code data b0 (i.e., a0 x b0), multiplying the first average data a1 by the first code data b1 (i.e., a1 x b1), and adding (i.e., a0 x b0 + a1 x b1 + a2 x b2) the multiplication of the second average data a2 by the second code data b2 (i.e., a2 x b2).

[0093] 11C, the output data OUT is expressed as a0×b0+a1×b1+a2×b2. That is, the output data OUT is represented by the 0th average data a0, the 0th code data b0, the first average data a1, the first code data b1, the second average data a2, and the second code data b2, which requires a total capacity of 3×(16×4+1×16)=240 bits.

[0094] In one embodiment, the more the above-described operations are repeatedly performed, the more the error for the 0th input data INPUT0 can be reduced. For example, the mean square error MSEa for the 0th input data of the a-th intermediate data INTa may be greater than the mean square error MSEb for the 0th input data of the b-th intermediate data INTb, and the mean square error MSEb for the 0th input data of the b-th intermediate data INTb may be greater than the mean square error MSEc for the 0th input data of the output data OUT.

[0095] In one embodiment, the quantization described above causes some errors in the output data OUT compared to the 0th input data INPUT0 (i.e., the original data), but reduces the overall data capacity. For example, the 0th input data INPUT0 includes 16 FP16 data. That is, the 0th input data INPUT0 has a capacity of 16×16=256 bits. On the other hand, after the quantization operation described above, the output data OUT has a capacity of 3×(16×4+1×16)=240 bits. Therefore, as the calculation result or activation data ACT generated during the inference process of the accelerator 1100 is quantized, the bandwidth and capacity required for the memory 101 can be reduced.

[0096] 11A to 11C may be binary coding-based quantization (BCQ). In this case, output data OUT is expressed by multiplication and addition of average data and code data. Then, when the output data OUT is used as input for the next artificial intelligence operation, it is simply calculated by BiQGEMM (non-GENERAL Matrix to Matrix multiplication for Binary-coding based Quantized neural networks). In one embodiment, the operation result by BiQGEMM may have a high precision type HP (e.g., BF16, FP16, FP32, etc.), and the quantizer 1100 may perform quantization on the operation result in the same manner as described above.

[0097] Fig. 12 is a block diagram showing the accelerator of Fig. 2. Referring to Fig. 2 and Fig. 12, the accelerator 1000 may include a quantizer 1100, a general-purpose buffer unit 1300, a plurality of processing units PE11 to PE44, and an accumulator 1400.

[0098] The general-purpose buffer unit 1300 may be configured to store various data, weights, or parameters for artificial intelligence operations performed by the accelerator 1000. In one embodiment, the various data, weights, or parameters stored in the general-purpose buffer unit 1300 are provided from the memory 101 of FIG. 2 or obtained from the operation results of the multiple processing units PE11 to PE44 of the accelerator 1000. In one embodiment, the information stored in the general-purpose buffer unit 1300 may have a low precision type LP.

[0099] The processing units PE11 to PE44 may perform an AI operation or a MAC operation based on the data provided from the general-purpose buffer unit 1300. For example, each of the processing units PE11 to PE44 may receive activation data ACT and weight data WT from the general-purpose buffer unit 1300, perform an AI operation or a MAC operation on the received activation data ACT and weight data WT, and output partial sum data PSUM.

[0100] The accumulator 1400 may be configured to accumulate the partial sum data PSUM of each of the multiple processing units PE11 to PE44. The output of the accumulator 1400 may be provided to the quantizer 1100 as result data RST.

[0101] The quantizer 1100 may generate output data OUT by performing quantization on the result data RST received from the accumulator 1400. The output data OUT is stored in the general-purpose buffer unit 1300. In one embodiment, the output data OUT stored in the general-purpose buffer unit 1300 is reused as activation data ACT for the processing units PE11 to PE44.

[0102] In one embodiment, the processing units PE11 to PE44 and the accumulator 1400 may perform operations based on a high-precision type HP (e.g., BF16, FP16, FP32, etc.). In this case, the partial sum data PSUM output from the processing units PE11 to PE44 and the result data RST output from the accumulator 1400 should have a high-precision type HP. Meanwhile, the quantizer 1100 quantizes the result data RST to generate output data OUT, which is stored in the general-purpose buffer unit 1300. That is, the quantizer 1100 quantizes intermediate data generated by the inference operation of the accelerator 1000. In this case, a relatively small amount of data is stored in or output from the general-purpose buffer unit 1300, thereby reducing the required capacity and bandwidth of the general-purpose buffer unit 1300 and reducing the power consumption of the general-purpose buffer unit 1300.

[0103] 13 is a diagram illustrating an accelerator structure according to an embodiment of the present invention. For the sake of simplicity and convenience of explanation, only some components relevant to an embodiment of the present invention are shown in FIG.

[0104] 13, the accelerator 2000 may include a memory 2001, a quantizer 2100, and a processing unit 2200. The memory 2001 may be configured to store various data, weights, parameters, etc. required for artificial intelligence calculations.

[0105] The quantizer 2100 may be located between the memory 2001 and the processing unit 2200. The quantizer 2100 may perform inverse quantization on data received from the memory 2001 or quantization on data received from the processing unit 2200. For example, the processing unit 2200 may perform an artificial intelligence operation based on high precision type HP. That is, the processing unit 2200 may perform a MAC operation on high precision type HP data and output resulting high precision type HP data. Meanwhile, data stored in the memory 2001 may have low precision type LP. Thus, the quantizer 2100 may quantize high precision type HP data received from the processing unit 2200 to low precision type LP. Alternatively, the quantizer 2100 may inverse quantize low precision type LP data received from the memory 2001 to high precision type HP.

[0106] In one embodiment, the quantizer 2100 may have a structure similar to the quantizer 1100 described with reference to FIGS. 7-9, and may be controlled in various ways according to various quantization or inverse quantization algorithms.

[0107] 14 is a diagram showing the structure of an accelerator according to an embodiment of the present invention. For the sake of simplicity and convenience of explanation, only some components related to an embodiment of the present invention are shown in FIG.

[0108] 14, the accelerator 3000 may include a memory 3001, a quantizer 3100, a processing unit 3200, and an inverse quantizer 3200. The memory 3001 and the processing unit 3200 are similar to those described with reference to FIG. 13, and therefore detailed description thereof will be omitted.

[0109] The quantizer 3100 can quantize the high precision type HP operation result data RST generated from the processing unit 3200 to generate low precision type LP output data OUT. The output data OUT is stored in the memory 3001. The inverse quantizer 3300 can inverse quantize the low precision type LP data received from the memory 3001 to generate high precision type HP data (e.g., activation data or weights, etc.). The generated high precision type data is provided to the processing unit 3200.

[0110] Although some structures of the accelerators 2000 and 3000 have been described with reference to FIGS. 13 and 14, the scope of the present invention is not limited thereto. In one embodiment, low-precision type L data is generated by a quantizer of the accelerator. At this time, a processing unit of the accelerator may perform an operation on the low-precision type data and generate high-precision type result data. For example, the quantizer may generate low-precision type data based on BCQ, and the processing unit of the accelerator may perform an operation on the low-precision type data based on BiQGEMM (non-GEneral Matrix to Matrix multiplication for Binary-coding based Quantized neural networks). In this case, separate inverse quantization may not be performed within the accelerator.

[0111] Fig. 15 is a block diagram showing the structure of an accelerator according to an embodiment of the present invention. Referring to Fig. 15, an accelerator 4000 may include a memory 4001, a quantizer 4100, and a processing unit 4200. The memory 4001, the quantizer 4100, and the processing unit 4200 are similar to those described with reference to Fig. 2, and therefore detailed description thereof will be omitted.

[0112] In one embodiment, the accelerator 4000 may perform operations on multiple layers to perform learning or inference on an artificial intelligence model. Different operation methods may be applied depending on the operating characteristics or reliability of each of the multiple layers. For example, some of the multiple layers may require a relatively large amount of operation, while other layers may require a relatively small amount of operation. Alternatively, calculation accuracy may be important for some of the multiple layers, while calculation speed may be important for other layers. The accelerator 4000 may quantize the operation result RST and store it in the memory 4001, or may omit quantization and store the operation result RST in the memory 4001, depending on the characteristics of the subsequent layer.

[0113] For example, if a subsequent layer operated by the accelerator 4000 requires a large amount of calculation and a fast calculation speed, the quantizer 4100 can quantize the calculation result RST and store it in the memory 4001. In this case, the size of the data stored in the memory 4001 is reduced, allowing for fast memory access. On the other hand, if a subsequent layer operated by the accelerator 4000 requires a small amount of calculation and accurate calculation, the accelerator 4000 can omit quantization and store the result data RST (i.e., high-precision type HP) in the memory 4001. In this case, the accuracy of the data stored in the memory 4001 is high, allowing for relatively accurate calculation.

[0114] Fig. 16 is a block diagram showing a system according to one embodiment of the present invention. Referring to Fig. 16, a system 5000 may include a first accelerator 5110, a second accelerator 5120, a memory 5200, and a controller 5300. The memory 5200 and the controller 5300 have been described with reference to Fig. 2, and therefore a detailed description thereof will be omitted.

[0115] The system 5000 may be dedicated hardware configured to perform processing of artificial intelligence models. In one embodiment, the first accelerator 5110 and the second accelerator 5120 of the system 5000 may perform operations in parallel to process large artificial intelligence models. For example, the first accelerator 5110 and the second accelerator 5120 may process large artificial intelligence models in parallel or independently through data parallelism, model parallelism, or tensor parallelism. In one embodiment, the first accelerator 5110 and the second accelerator 5120 may operate based on the operating or computing methods described with reference to FIGS. 1-15. For example, the first accelerator 5110 may include a first quantizer 5111, and the second accelerator 5210 may include a second quantizer 5121. Each of the first quantizer 5111 and the second quantizer 5121 may be the quantizer described with reference to FIGS. 1 to 15 or may operate based on the method described with reference to FIGS. 1 to 15. For example, the first quantizer 5111 may quantize result data generated by the artificial intelligence operation of the first accelerator 5110 to generate first output data, and the first output data is stored in the memory 5200. The second quantizer 5121 may quantize result data generated by the artificial intelligence operation of the second accelerator 5121 to generate second output data, and the second output data is stored in the memory 5200. Each of the first output data and the second output data stored in the memory 5200 may be provided to the first accelerator 5110 or the second accelerator 5120 for a subsequent artificial intelligence operation.

[0116] Figure 17 is a block diagram showing a system according to an embodiment of the present invention. Referring to Figure 17, the system may include a first accelerator 6110, a second accelerator 6120, a memory 6200, a controller 6300, and a quantizer 6400. The memory 5200 and the controller 5300 have been described with reference to Figure 2, so detailed descriptions thereof will be omitted. The first accelerator 6110, the second accelerator 6120, the memory 6200, and the controller 6300 have been described with reference to Figure 16, so detailed descriptions thereof will be omitted.

[0117] In one embodiment, the system 6000 may include a quantizer 6400. The quantizer 6400 may perform quantization on the result data RST generated from the first accelerator 6110 and the second accelerator 6120 to generate output data. The generated output data is stored in the memory 6200. In one embodiment, except that the quantizer 6400 exists outside the accelerators 6110 and 6120, the configuration, structure, and operation method of the quantizer 6400 are similar to those described with reference to FIGS. 1 to 16, and therefore, detailed description thereof will be omitted.

[0118] As described above, according to the present invention, the accelerator may include a quantizer configured to quantize result data. In this case, the quantizer quantizes operation result data or operation intermediate data generated during the accelerator's AI operation, learning, inference, etc., thereby reducing the amount of data accessed from the memory. Therefore, the bandwidth or power consumption required for the memory can be reduced.

[0119] The above-described content is a specific embodiment for implementing the present invention. The present invention includes not only the above-described embodiment but also embodiments that can be simply modified or easily changed. The present invention also includes techniques that can be easily implemented by modifying the embodiment. Therefore, the scope of the present invention should not be limited to the above-described embodiment, but should be defined not only by the following claims but also by equivalents to the claims of the present invention. [Explanation of symbols]

[0120] 101 Memory 102 Controller 1000 Accelerators 1100 Quantizer

Claims

1. 1. An accelerator configured to perform artificial intelligence (AI) operations, comprising: a processing unit configured to perform a first operation on the first activation data and the first weight data loaded from the memory to generate first result data; a quantizer configured to perform quantization on the first result data to generate first output data; the first activation data, the first weight data and the first output data are of low precision type, and the first result data is of high precision type; the first output data is stored in the memory; Accelerator.

2. The size of the first output data is smaller than the size of the first result data. The accelerator of claim 1 .

3. the high-precision type includes at least one of a BF16 (Brain Floating Point Format) type, an FP16 (half-precision IEEE Floating Point Format) type, an FP32 (Single-precision floating-point format) type, and an FP64 (Double-precision floating-point format) type; The low precision type includes at least one of an INT4 type, an INT8 type, and an INT16 type. The accelerator of claim 1 .

4. the first operation is a MAC (multiply and accumulate) operation on the first activation data and the first weight data; The accelerator of claim 1 .

5. The quantizer a round robin switch configured to receive the first result data; a plurality of quantization cores configured to perform quantization on the first result data received from the round robin switch and generate the first output data; a control logic circuit configured to control each of the plurality of quantization cores; the first output data generated by the plurality of quantization cores is transferred to the memory via the round robin switch; The accelerator of claim 1 .

6. each of the plurality of quantization cores performs quantization in parallel; The accelerator of claim 5 .

7. Each of the plurality of quantization cores comprises: an input reformatter configured to change the format of the first result data received from the round robin switch and store intermediate results; a conversion circuit configured to perform an operation on input data received from the input reformatter to generate the first result data; an output reformatter configured to store the first result data generated from the conversion circuit and to output the first result data to the round robin switch. The accelerator of claim 5 .

8. The conversion circuit a code operation module configured to manage the code of the input data; a scalar operation module configured to perform a scalar operation on the input data; a vector-to-scalar operation module configured to perform vector-to-scalar operations on the input data; a vector-vector operation module configured to perform vector-vector operations on the input data; The accelerator of claim 7 .

9. the control logic circuit sequentially controls the input reformatter, the output reformatter, the sign operation module, the scalar operation module, the vector-scalar operation module, and the vector-vector operation module of each of the plurality of quantization cores according to a quantization algorithm performed by each of the plurality of quantization cores; The accelerator of claim 8 .

10. The quantizer performs quantization based on BCQ (Binary Coding based Quantization). The accelerator of claim 1 .

Citation Information

Patent Citations

  • Neural network model quantitative analysis method and system for dedicated processor

    CN116050464A

  • Interlayer circulation data full quantization method for edge hardware equipment calculation

    CN116432712A

  • Deep convolutional neural network mixing precision quantification method applied to FPGA

    CN116502691A

  • US11、461、614

  • US11、556、772