Application processor, neural network device, and method of operating neural network device

By using an integer arithmetic unit and a data converter in a neural network device to convert floating-point data into integer data for computation, the problems of low efficiency and high power consumption of neural network devices in low-power devices are solved, and more efficient computing performance is achieved.

CN111222634BActive Publication Date: 2026-04-17SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2019-11-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Neural network devices are inefficient and power-consuming when processing floating-point data, making it difficult to achieve high-performance computing in low-power devices such as smartphones.

Method used

An integer arithmetic unit and a data converter are used to convert floating-point data into integer data for neural network operations, and direct data transmission and conversion are achieved through a direct memory access controller.

Benefits of technology

It improves the processing speed of neural network operations and reduces power consumption, enhancing computing performance in low-power devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111222634B_ABST
    Figure CN111222634B_ABST
Patent Text Reader

Abstract

The present invention provides an application processor, a neural network device, and a method of operating a neural network device. The neural network device includes a direct memory access (DMA) controller configured to receive floating point data from a memory, a data converter configured to convert the floating point data received via the DMA controller to integer type data, and a processor configured to perform neural network operations based on integer operations using the integer type data provided from the data converter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Cross-reference to related applications]

[0002] This patent application claims priority to Korean Patent Application No. 10-2018-0146610, filed on November 23, 2018, with the Korean Intellectual Property Office, the full disclosure of which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to a neural network. More specifically, this disclosure relates to a method and apparatus for processing floating-point numbers using a neural network device, said neural network device including an integer arithmetic unit. Background Technology

[0004] A neural network is a computational architecture that models the biological networks that make up an animal's brain. Recently, with the development of neural network technology, there has been active research into using neural network-based devices in various types of electronic systems to analyze input data and extract useful information.

[0005] Neural network devices require a large amount of computation on complex input data. A technique is needed to efficiently process neural network computations to analyze inputs (e.g., input data) and extract information from them in real time. Specifically, low-power and high-performance systems (e.g., smartphones) have limited resources and therefore require techniques to maximize the performance of artificial neural networks while reducing the amount of computation required to process complex input data.

[0006] Specifically, floating-point numbers are used as a particular method of encoding numbers. Floating-point numbers can include signed strings of numbers of a given length in a given base, and signed integer exponents that modify the magnitude. Using floating-point numbers provides an approximation of the underlying number by a trade-off between range and precision. That is, floating-point numbers can support a wide range of values ​​and are suitable for representing approximations of real numbers. However, floating-point numbers as a representation can exhibit relative complexity compared to more familiar real numbers. For example, floating-point data can include feature maps, kernels (weight maps), biases, etc., and the nature of handling such floating-point data is not always optimal for neural network devices. Summary of the Invention

[0007] This disclosure describes a method and apparatus for processing floating-point numbers in a neural network device, the neural network device including an integer arithmetic unit.

[0008] According to one aspect of this disclosure, a neural network device for performing neural network operations includes a direct memory access (DMA) controller, a data converter, and a processor. The DMA controller is configured to receive floating-point data from memory. The data converter is configured to convert the floating-point data received via the DMA controller into integer type data. The processor is configured to perform the neural network operations based on integer arithmetic using the integer type data provided from the data converter.

[0009] According to another aspect of this disclosure, a method of operating a neural network device includes: receiving floating-point input data from a memory and converting the floating-point input data into integer type input data. The method further includes using the integer type input data to perform neural network operations based on integer arithmetic.

[0010] According to another aspect of this disclosure, an application processor includes a memory and a neural network device. The memory is configured to store computational parameters and feature values. The computational parameters and feature values ​​are floating-point values. The neural network device is configured to receive the computational parameters and feature values ​​from the memory, convert the computational parameters and feature values ​​into integer values, and perform neural network operations based on the computational parameters and feature values ​​converted to integer values. Attached Figure Description

[0011] The embodiments of this disclosure will be more clearly understood by reading the following detailed description in conjunction with the accompanying drawings, in which:

[0012] Figure 1 This is a block diagram illustrating a neural network system according to an exemplary embodiment.

[0013] Figure 2 This is a diagram illustrating an example of a neural network structure.

[0014] Figure 3A This is a diagram showing an example of a floating-point number.

[0015] Figure 3B This is a diagram showing instances of integers.

[0016] Figure 3C This is a diagram showing an example of a fixed-point number.

[0017] Figure 4 This is a diagram illustrating a method of operating a neural network device according to an exemplary embodiment.

[0018] Figure 5A This is a diagram illustrating the operation of a neural network system according to an exemplary embodiment.

[0019] Figure 5B This is a diagram illustrating the operation of the neural network system based on the comparative example.

[0020] Figure 6 This is a diagram illustrating a neural network device according to an exemplary embodiment.

[0021] Figure 7A This is a circuit diagram illustrating a data converter according to an exemplary embodiment.

[0022] Figure 7B This illustrates an exemplary embodiment. Figure 7A Another circuit diagram of the data converter shown.

[0023] Figure 8A This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment.

[0024] Figure 8B It is shown Figure 8A The diagram shows the input of the data converter.

[0025] Figure 9A This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment.

[0026] Figure 9B It is shown Figure 9A The diagram shows the input of the data converter.

[0027] Figure 10 This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment.

[0028] Figure 11 This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment.

[0029] Figure 12 This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment.

[0030] Figure 13 This is a circuit diagram illustrating a data converter according to an exemplary embodiment.

[0031] Figure 14 This is a block diagram illustrating a neural network device according to an exemplary embodiment.

[0032] Figure 15 This is a block diagram illustrating a data processing system according to an exemplary embodiment.

[0033] Figure 16 This is a block diagram illustrating an application processor according to an exemplary embodiment. Detailed Implementation

[0034] Figure 1 This is a block diagram illustrating a neural network system according to an exemplary embodiment.

[0035] The neural network system 100 can infer information contained in or derived from input data by training (or learning) the neural network or by analyzing input data using the neural network. Based on the inferred information, the neural network system 100 can determine how to resolve or continue a situation, or control components of an electronic device on which the neural network system 100 is mounted. For example, the neural network system 100 can be applied to or used in devices such as smartphones, tablets, smart TVs, augmented reality (AR) devices, Internet of Things (IoT) devices, autonomous vehicles, robots, medical devices, drones, advanced driver assistance systems (ADAS), image display devices, measurement devices, and / or other types of devices. For example, the neural network system 100 can be applied to perform speech recognition, image recognition, image classification, and / or other operations using neural networks. The neural network system 100 can also be mounted on any of various types of electronic devices. In an exemplary embodiment, Figure 1 The neural network system 100 shown can be an application processor.

[0036] Reference Figure 1 The neural network system 100 may include a central processing unit (CPU) 110, a neural network device 120, a memory 130, and a sensor module 140. The neural network system 100 may also include input / output modules, security modules, power control devices, and various types of processors. Some or all of the components of the neural network system 100 (e.g., CPU 110, neural network device 120, memory 130, and sensor module 140) may be formed on a single semiconductor chip. For example, the neural network system 100 may be implemented as a system-on-chip (SoC). The components of the neural network system 100 may communicate with each other via a bus 150.

[0037] CPU 110 controls all operations of neural network system 100. CPU 110 may include one processing core (single core) or multiple processing cores (multi-core). CPU 110 can process or execute programs and / or data stored in a storage area (such as memory 130). CPU 110 can implement or control some or all of the performance of the various methods described herein by executing instructions in such programs and / or data. CPU 110 may also be an integer arithmetic unit or may include an integer arithmetic unit and perform integer operations based on integer type input values ​​converted from floating-point input values ​​by data converter 20. Alternatively, CPU 110 may act as an integer arithmetic unit to control the processor in neural network device 120 to perform integer operations based on integer type input values ​​converted from floating-point input values ​​by data converter 20.

[0038] For example, CPU 110 can control neural network device 120 to execute an application and perform neural network-based tasks required in connection with executing the application. The neural network includes at least one of various types of neural network models, including convolutional neural networks (CNNs), region with convolutional neural networks (R-CNNs), region proposal networks (RPNs), recurrent neural networks (RNNs), stacking-based deep neural networks (S-DNNs), state-space dynamic neural networks (S-SDNNs), deconvolution networks, deep belief networks (DBNs), restricted Boltzmann machines (RBMs), fully convolutional networks, long short-term memory (LSTM) networks, and classification networks.

[0039] The neural network device 120 can perform neural network operations based on the received input data. For example, neural network operations may include convolution, pooling, activation function operations, vector operations, dot product operations, cross product operations, etc. Furthermore, the neural network device 120 can generate information signals based on the results of the neural network operations. The neural network device 120 can be implemented as a neural network operation accelerator, a coprocessor, a digital signal processor (DSP), or an application-specific integrated circuit (ASIC).

[0040] The neural network device 120 according to an embodiment of the present invention may include an integer arithmetic unit and may perform neural network operations based on integer operations. The integer arithmetic unit may perform operations on integers or fixed-point numbers. Compared with a floating-point arithmetic unit that performs floating-point operations, the integer arithmetic unit may exhibit lower power consumption, faster processing speed, and smaller circuit area.

[0041] However, neural networks may include floating-point data represented by real numbers (e.g., floating-point numbers). Neural network device 120 may include such floating-point data by receiving, generating, retrieving, or otherwise acquiring it. For example, floating-point data may include feature maps, kernels (weight maps), biases, etc. As described earlier herein, floating-point numbers can support a wide range of values ​​and are suitable for representing approximations of real numbers. In one or more embodiments, integer arithmetic units may not readily handle floating-point data. Therefore, neural network device 120 according to embodiments of the present invention may include a data converter 20 for processing floating-point data. The data converter 20 may perform data conversion between floating-point data and integer type data. In other words, the data converter 20 may convert floating-point numbers to integers (or fixed-point numbers), or it may convert integers to floating-point numbers. The data converter 20 may include a plurality of floating-point arithmetic units and shifters that are selectively deactivated and activated. For example, deactivation and activation may be performed according to the received operand code so that the activated floating-point arithmetic units and / or shifters perform an operation mode according to the operand code.

[0042] The neural network device 120 can receive floating-point data FPD stored in memory 130 and use data converter 20 to convert the floating-point data FPD into integer type data INTD. The neural network device 120 can use components other than data converter 20 to perform neural network operations based on the integer type data INTD and store the results of the neural network operations in memory 130. In this case, the result of the neural network operation can be generated from the integer type data INTD and output as integer type data INTD (i.e., integer type output data). Data converter 20 can convert the integer type data INTD into floating-point data FPD, where the integer type data INTD is generated and output from other components of the neural network device 120 and delivered to data converter 20. The neural network device 120 can store the floating-point data FPD generated by data converter 20 in memory 130. In an exemplary embodiment, the neural network device 120 can transmit and receive floating-point data FPDs via bus 150 to memory 130 without interference from CPU 110 or at least independently of floating-point data processing performed by CPU 110 or by CPU 110 on floating-point data. In other words, floating-point data FPDs can be directly transmitted and received between the neural network device 120 and memory 130. In embodiments explained below, the direct transmission and reception of floating-point data FPDs can be performed according to DMA (Direct Memory Access), a mechanism that provides access to main memory (e.g., random access memory) independently of the CPU (e.g., CPU 110). In other words, the direct transmission and reception described herein can be an exception to the overall operation of the neural network system 100, and particularly for the overall operation of CPU 110 and memory 130.

[0043] The memory 130 may store programs and / or data used in the neural network system 100. The memory 130 may also store operational parameters, quantization parameters, input data, and output data. Examples of operational parameters include weight values, bias values, etc., of the neural network. Examples of quantization parameters used for quantizing the neural network include scale factors, bias values, etc., and may be referred to as quantization parameters hereinafter. Examples of input data are input feature maps, and examples of output data are output feature maps. Operational parameters, quantization parameters, input data, and output data may be floating-point data (FPDs).

[0044] Memory 130 may be dynamic random access memory (DRAM), but is not limited to this. Memory 130 may include at least one of volatile memory and non-volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable read-only memory (PROM), electrically programmable read-only memory (EPROM), electrically erasable and programmable read-only memory (EEPROM), flash memory, phase-change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FRAM), etc. Examples of volatile memory include DRAM, static RAM (SRAM), synchronous dynamic random access memory (SDRAM), etc. In an exemplary embodiment, the memory 130 may include at least one of a hard disk drive (HDD), a solid state drive (SSD), a compact flash (CF) card, a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, and a memory stick.

[0045] Sensor module 140 can collect information about the electronic device on which neural network system 100 is mounted. Sensor module 140 can sense or receive signals (e.g., video signals, voice signals, magnetic signals, biosignals, touch signals, etc.) from outside the electronic device and convert the sensed or received signals into sensing data. For this purpose, sensor module 140 may include at least one of various types of sensing devices, such as microphones, imaging devices, image sensors, light detection and ranging (LIDAR) sensors, ultrasonic sensors, infrared sensors, biosensors, touch sensors, etc. Examples of signals that can be received by sensor module 140 include addressed signals specifically sent to a unique address of the electronic device on which neural network system 100 is mounted or to the address of a component of the electronic device. Examples of signals that can be sensed by sensor module 140 include electronic information that can be deduced by monitoring the environment around the electronic device, for example, using a camera, microphone, biosensor, touch sensor (e.g., screen or keyboard).

[0046] Sensing data can be considered as data sensed or received by sensor module 140, and can be provided as input data to neural network device 120 or stored in memory 130. Sensing data stored in memory 130 can be provided to neural network device 120. In an exemplary embodiment, neural network device 120 may further include a graphics processing unit (GPU) for processing image data, wherein sensing data can be processed by the GPU and then stored in memory 130 or provided to neural network device 120.

[0047] For example, sensor module 140 may include an image sensor and generate image data by capturing images of the environment outside the electronic device. The image data output from sensor module 140 or the image data processed by the GPU may be floating-point data, and the image data may be provided directly to neural network device 120 or provided to neural network device 120 after being stored in memory 130.

[0048] Figure 2 This is a diagram illustrating an example of a neural network structure. (See reference...) Figure 2A neural network (NN) may include multiple layers L1 to Ln. Such a neural network with a multi-layered structure may be referred to as a deep neural network (DNN) or a deep learning architecture. Each of layers L1 to Ln may be a linear layer or a non-linear layer. In an exemplary embodiment, at least one linear layer and at least one non-linear layer may be combined with each other and referred to as a layer. For example, a linear layer may include convolutional layers and fully connected layers, while a non-linear layer may include pooling layers and activation layers.

[0049] For example, the first layer L1 can be a convolutional layer, the second layer L2 can be a pooling layer, and the nth layer Ln can be an output layer and a fully connected layer. The neural network NN may also include activation layers and layers that perform other types of operations.

[0050] Each of layers L1 to Ln can receive an input image frame or a feature map generated in the previous layer as its input feature map, perform operations on the input feature map, and generate an output feature map or a recognition signal REC. Here, the feature map refers to the data that expresses various features of the input data. Each of the feature maps FM1, FM2, FM3, and FMn can have, for example, a two-dimensional matrix structure or a three-dimensional matrix (or tensor) structure including multiple eigenvalues. Each of the feature maps FM1, FM2, FM3, and FMn can have a width W (or column), a height H (or row), and a depth Dt, where the width W, height H, and depth Dt can correspond to the x-axis, y-axis, and z-axis of the coordinate system, respectively. Here, the depth Dt can be referred to as the number of channels.

[0051] The first layer L1 generates the second feature map FM2 by convolving the first feature map FM1 with the weight map WM. The weight map WM can be a two-dimensional or three-dimensional matrix structure including multiple weight values. The weight map WM can be referred to as a kernel. The weight map WM can filter the first feature map FM1 and can be referred to as a filter or kernel. The depth (i.e., the number of channels) of the weight map WM is equal to the depth (i.e., the number of channels) of the first feature map FM1, and the channels of the weight map WM can be convolved with the same channels of the first feature map FM1. The weight map WM is shifted to traverse the first input feature map FM1 as a sliding window. During each shift, each weight in the weight map WM is multiplied by all feature values ​​in the region overlapping with the first feature map FM1, and each weight in the weight map WM is added to all feature values ​​in the region overlapping with the first feature map FM1. When the first feature map FM1 is convolved with the weight map WM, one channel of the second feature map FM2 is generated. Although Figure 2Only one weight map WM is shown; however, when multiple weight maps are convolved with the first feature map FM1, multiple channels of the second feature map FM2 can be generated. In other words, the number of channels in the second feature map FM2 can correspond to the number of weight maps.

[0052] The second layer L2 generates a third feature map FM3 by modifying the spatial size of the second feature map FM2 using pooling. Pooling can also be referred to as sampling or downsampling. The two-dimensional pooling window PW is shifted over the second feature map FM2 in units of the size of the pooling window PW, and the maximum value (or the average of the eigenvalues) is selected from the eigenvalues ​​in the region overlapping with the pooling window PW. Therefore, a third feature map FM3 is generated, which has a spatial size modified from that of the second feature map FM2. The number of channels in the third feature map FM3 is the same as the number of channels in the second feature map FM2.

[0053] The nth layer Ln classifies the input data into categories CL by combining features from the nth feature map FMn. Additionally, the nth layer Ln generates a recognition signal REC corresponding to each category. For example, when the input data is image data and the neural network NN performs image recognition, the nth layer Ln can identify objects and generate a recognition signal REC corresponding to the identified objects by extracting the categories corresponding to the objects in the image indicated by the image data from the nth feature map FMn provided by the previous layer.

[0054] As referenced above Figure 2 As described above, neural networks (e.g., neural networks NN) can be implemented using complex architectures. Neural network devices that perform neural network operations can perform an extremely large number of operations, ranging from hundreds of millions to tens of billions of operations. Therefore, neural network systems (e.g. Figure 1 The neural network system 100 shown may have a neural network device 120 that performs integer operations. According to an exemplary embodiment, the neural network device 120 performs integer operations, and therefore, compared to a neural network device that performs floating-point operations, power consumption and circuit area can be reduced. Therefore, for a neural network system (e.g., Figure 1 For neural network systems (100) in China, the processing speed can be improved.

[0055] Figure 3A This is a diagram showing an example of a floating-point number. Figure 3B This is a diagram showing instances of integers. Figure 3C This is a graph illustrating an example of a fixed-point number. In Figure 3A , 3B In 3C, FP represents floating-point numbers, INT represents integers, and FX represents fixed-point numbers.

[0056] Reference Figure 3A A floating-point number N(FP) can be expressed as a sign and 1.a × 2^b, where b is the exponent part and a is the fractional part. According to the IEEE 754 standard of the Institute of Electrical and Electronics Engineers (IEEE), a 32-bit floating-point number N(FP) has one bit representing the sign, eight bits representing the exponent part, and 23 bits representing the fractional part. For example... Figure 3A As shown, the most significant bit (MSB) represents the sign, the 8 bits following the MSB represent the exponent, and the remaining 23 bits represent the fractional part (or the fractional or significant digit). However, Figure 3A The floating-point number N(FP) shown is just an example, and the number of bits for the exponent and fractional parts can vary depending on the number of bits in the floating-point number.

[0057] Reference Figure 3B Integer N (INT) can be expressed in various types depending on the presence of a sign and the size of the data (number of bits). For example, integers commonly used in computational operations can be like... Figure 3B The value shown is represented as 8 bits of data. An integer N (INT) can be represented as an unsigned complement of two, where the value of an 8-bit integer N (INT) can be represented by the following Equation 1:

[0058]

[0059] In the case of an integer represented as an unsigned complement of two according to Equation 1, i represents the position of a bit in the integer N (INT). For example, when the integer is represented as 8 bits, i represents the least significant bit (LSB) when i is 0, and the MSB when i is 7. Di represents the value of the i-th bit (i.e., 0 or 1).

[0060] When an integer N(INT) is represented by the two's complement of a signed 1, the MSB represents the sign and the following bits represent the integer part INTn. For example, when the integer N(INT) is an 8-bit data, the eighth bit (i.e., the MSB) represents the sign, and the lower seven bits represent the integer part INTn. The value of an integer N(INT) represented as the two's complement of a signed 1 can be represented by the following Equation 2:

[0061]

[0062] Reference Figure 3CA fixed-point number N(FX) can include an integer part INTn and a fractional part Fn. For example, suppose that in a fixed-point number N(FX) represented by 32 bits, the decimal point is located between the Kth bit [K-1] and the (K+1)th bit [K] (where K is a positive integer), then the upper (32-K) bits can represent the integer part INTn and the lower K bits can represent the fractional part Fn. When the fixed-point number N(FX) is represented by the complement of a signed 1, the 32nd bit (i.e., MSB) can represent the sign.

[0063] The sign and magnitude of a fixed-point number N(FX) can be calculated similarly to those of an integer. However, the decimal point is located below zero in an integer, while in a fixed-point number N(FX), the decimal point is located between the Kth and K+1th digits, that is, after the top 32-K digits. Therefore, the value of a fixed-point number N(FX) including 32 digits can be represented by Equation 3.

[0064]

[0065] When the fixed-point number N(FX) is unsigned, a is 1. When the fixed-point number N(FX) is signed, a is -1.

[0066] In Equation 3, when K is 0, Equation 3 is expressed in the same way as Equation 1 or Equation 2, which indicates the value of an integer. When K is not 0 (i.e., a positive integer) and indicates a fixed-point number, the integer N (INT) and the fixed-point number N (FX) can be handled by the integer arithmetic unit.

[0067] However, in Figure 3A In the floating-point number N(FP) shown, some bits indicate the exponent portion. Therefore, the floating-point number N(FP) has the same characteristics as... Figure 3B The integer N(INT) shown and Figure 3C The fixed-point number N(FX) shown has a different structure and cannot be processed by an integer arithmetic unit. Therefore, a data converter (e.g., Figure 1 The data converter 20 shown converts the floating-point number N(FP) into an integer N(INT) (or into a fixed-point number N(FX)), and the integer arithmetic unit processes the integer N(INT). Therefore, the neural network device 120 can process the floating-point number N(FP) based on integer arithmetic.

[0068] Between floating-point numbers and integers, which include the same number of bits, integers can represent more precise values. For example, in the case of 16-bit floating-point numbers, according to the IEEE 754 standard, the fractional part corresponding to the significant value is represented by 10 bits, while in integers, 16 bits can represent significant data. Therefore, the result of a neural network operation (e.g., the result of a convolution operation) obtained when the data converter 20 converts floating-point data into integer type data and performs operations based on the integer type data can have higher precision than floating-point operations. For practical purposes, using the data converter 20 can make the result of the neural network operation more precise. However, the use of the data converter 20 itself can be improved, for example, when a variable indicating the position of the decimal point is available.

[0069] Figure 4 This is a diagram illustrating a method of operating a neural network device according to an exemplary embodiment. Figure 4 The method shown can be used in neural network devices for neural network operations ( Figure 1 As shown in 120), this will be implemented. Therefore, it will be referred to together. Figure 1 The following explanation is provided.

[0070] Reference Figure 4 The neural network device 120 can receive floating-point input data from the memory 130 (operation S110). The floating-point input data may include input feature values, weight values, and function coefficients required for neural network operations. Therefore, the floating-point input data can be partially or entirely represented as input values. Additionally, when the neural network device 120 processes a quantized neural network, the floating-point input data may include quantization parameters. For example, quantization parameters may include scaling values ​​(or inverse scaling values), bias values, etc. The neural network device 120 (and specifically the data converter 20) can perform quantization and data conversion of the input values ​​from the floating-point input data. As an example, the neural network device 120 can perform quantization and data conversion of the input values ​​based on two quantization parameters.

[0071] The neural network device 120 can convert floating-point input data into integer type input data (operation S120). For example, the data converter 20 can convert floating-point numbers of input data received from the memory 130 into integer type data, such as integers or fixed-point numbers. In an embodiment, the data converter 20 can convert floating-point data into quantized integer type data based on quantization parameters. For example, the neural network device 120 can perform quantization and conversion of input values ​​received as floating-point input data. The quantization and conversion can be performed based on two quantization parameters.

[0072] The neural network device 120 can perform neural network operations based on integer type input data (operation S130). The neural network device 120 includes an integer arithmetic unit and can perform neural network operations by performing integer operations on integer type input data. For example, neural network operations may include convolution, multiply and accumulate (MAC) operations, pooling, etc. When performing neural network operations, integer type output data can be generated.

[0073] The neural network device 120 can convert integer-type output data into floating-point output data (operation S140). For example, the data converter 20 can convert integer or fixed-point numbers of integer-type output data into floating-point numbers. That is, the data converter 20 can convert integer-type output data into floating-point output data or can convert integer-type output data into floating-point output data. In an embodiment, the data converter 20 can convert floating-point output data into dequantized floating-point output data based on quantization parameters. For example, in the process of converting integer-type output data into floating-point output data, the data converter 20 can perform inverse quantization and data conversion of the output value of the floating-point output data. Inverse quantization and data conversion can be performed based on two quantization parameters received as floating-point input data. The neural network device 120 can store the floating-point output data in the memory 130 (operation S150).

[0074] Figure 5A This is a diagram illustrating the operation of a neural network system 100 according to an exemplary embodiment, and Figure 5B This is a diagram illustrating the operation of the neural network system 100' according to the comparative example.

[0075] Reference Figure 5A The neural network device 120 can receive floating-point input data FPID from memory 130 (operation S1). In an embodiment, the floating-point input data FPID can be sent from memory 130 to the neural network device 120 via bus 150 without interference from CPU 110 or at least independently of the processing of the floating-point input data FPID by CPU 110 or by CPU 110. For example, the neural network device 120 may include a direct memory access (DMA) controller that can access and read the floating-point input data FPID from memory 130. Therefore, the DMA controller can be configured to receive the floating-point input data FPID from memory 130. The neural network device 120 can use a data converter 20 to receive the floating-point input data FPID from memory 130 and convert the floating-point input data FPID into integer type input data.

[0076] The neural network device 120 can use the data converter 20 to convert integer-type output data, which is the result of a neural network operation, into floating-point output data FPOD and output the floating-point output data FPOD to the memory 130 (operation S2). As described above, the neural network device 120 may be an integer arithmetic unit or may include an integer arithmetic unit and perform integer operations based on the integer-type input value converted from the floating-point input value by the data converter 20 in S1. For example, the processor included in the neural network device 120 may perform neural network operations on the integer-type data INTD converted from the floating-point input data FPID and output integer-type output data as the result of the neural network operation. The data converter 20 may convert the integer-type output data output from the processor into floating-point output data FPOD or may convert the integer-type output data output from the processor into floating-point output data FPOD. In an embodiment, the floating-point output data FPOD can be sent from the neural network device 120 to the memory 130 via the bus 150 without interference from the CPU 110 or at least independently of the processing of the floating-point output data FPOD by the CPU 110 or the processing of the floating-point output data FPOD by the CPU 110.

[0077] As described above, in the neural network system 100 according to the exemplary embodiment, the neural network device 120 includes a data converter 20. Data conversion can be performed when floating-point output data FPOD is sent to memory 130 and floating-point input data FPID is received from memory 130. Between S1 and S2 described above, the data converter 20 can be configured to convert the floating-point input data FPID received via a DMA controller into integer type input data to be subjected to neural network operations. The neural network operations can be performed by a processor included in the neural network device 120, which is configured to perform neural network operations based on integer operations using the integer type input data provided from the data converter 20.

[0078] According to Figure 5BIn the neural network system 100' of the comparative example shown, the neural network device 120' does not include a data converter. Therefore, since the neural network device 120' includes an integer arithmetic unit, another component (e.g., CPU 110) can perform operations to convert floating-point data into integer type data to process floating-point data. The CPU 110 can receive floating-point input data FPID output from memory 130 via bus 150 (operation S11), convert the floating-point input data FPID into integer type input data INTID, and send the integer type input data INTID to the neural network device 120' via bus 150 (operation S12). In addition, the CPU 110 can receive integer type output data INTOD output from the neural network device 120' via bus 150 (operation S13), convert the integer type output data INTOD into floating-point output data FPOD, and send the floating-point output data FPOD to memory 130 via bus 150 (operation S14).

[0079] As referenced above Figure 5A In the neural network system 100 according to the exemplary embodiment, the neural network device 120 performs data conversion while sending and receiving floating-point data (e.g., floating-point input data FPID and floating-point output data FPOD), and the floating-point data can be directly sent from the memory 130 to the neural network device 120. Therefore, according to the neural network system 100 according to the exemplary embodiment, the frequency of bus 150 usage is reduced, and thus the degradation of processing speed due to the bandwidth BW of bus 150 is reduced. Furthermore, since the neural network device 120 performs data conversion, the computational load on the CPU 110 is reduced, and thus the processing speed of the neural network device 120 can be increased.

[0080] Figure 6 This is a diagram illustrating a neural network device according to an exemplary embodiment. For ease of explanation, memory 130 is also shown.

[0081] Reference Figure 6 The neural network device 120 may include a DMA controller 10, a data converter 20, and a neural network processor 30. Figure 6 In this embodiment, the DMA controller 10 and the data converter 20 are shown as separate components. However, according to an embodiment, the data converter 20 may be implemented as part of the DMA controller 10. According to an embodiment, the neural network device 120 may include multiple data converters 20. The number of data converters 20 may be determined based on the processing speed of the neural network processor 30 and the bandwidth of the memory 130.

[0082] The DMA controller 10 can communicate directly with the memory 130. The DMA controller 10 can receive data (e.g., floating-point data) from and send data to the memory 130 without interference from another processor (e.g., a CPU (e.g., CPU 110) or a GPU, or at least independently of processing of the received data by another processor, or processing of the received data by another processor. The DMA controller 10 can receive floating-point input data (FPID) from the memory 130 and send floating-point output data (FPOD) provided from the data converter 20 to the memory 130. For example, the floating-point input data (FPID) may include an input feature map (IFM), operational parameters (PM), etc. The floating-point input data (FPID) may also include quantization parameters, such as two or more quantization parameters for quantization and / or dequantization.

[0083] Data converter 20 can perform data conversion between floating-point data and integer type data. Data converter 20 can convert fixed-point input data FPID received from DMA controller 10 into integer type input data INTID and send the integer type input data INTID to neural network processor 30. Additionally, data converter 20 can convert integer type output data INTOD output from neural network processor 30 into floating-point output data FPOD. In other words, data converter 20 converts integer type output data from neural network processor 30 into floating-point output data. For example, the floating-point output data FPOD may include an output feature map OFM.

[0084] According to an embodiment, the data converter 20 can generate quantized integer type input data INTID by quantizing fixed-point input data FPID based on quantization parameters (e.g., based on at least two quantization parameters). Additionally, the data converter 20 can generate dequantized floating-point output data FPOD by inverse quantizing integer type output data INTOD based on quantization parameters (e.g., based on at least two quantization parameters).

[0085] According to an embodiment, the data converter 20 can convert only the fixed-point input data FPID into integer type input data INTID. Additionally, the data converter 20 can convert only the integer type output data INTOD into floating-point output data FPOD.

[0086] According to an embodiment, the data converter 20 can perform MAC operations on integer data. For example, the data converter 20 performs MAC operations on integer data received from the memory 130 or integer data output from the neural network processor 30, thereby reducing the computational load on the neural network processor 30. The operations of the data converter 20 according to the above embodiments can be performed by selectively operating some components of the data converter 20. In these embodiments, the data converter 20 may be an integer arithmetic unit or may include an integer arithmetic unit and perform integer operations based on integer input values ​​received by the data converter 20 or converted from floating-point input values ​​by the data converter 20.

[0087] The neural network processor 30 may include a processing element array (PEA) comprising multiple processing elements (PEs). Although not shown in the figures, the neural network processor 30 may include internal memory and a controller for storing neural network parameters, such as bias values, weight values, input features, and output features. Each processing element (PE) may include one or more integer arithmetic units, and the neural network processor 30 may perform neural network operations based on integer arithmetic. For example, the neural network processor 30 may perform convolution operations based on integer weight values ​​and integer input feature values ​​provided by the data converter 20. Therefore, the neural network processor 30 may be an integer arithmetic unit or may include an integer arithmetic unit and perform integer arithmetic based on integer input values ​​converted from floating-point input values ​​by the data converter 20.

[0088] The configuration and operation of the data converter 20 will be described below.

[0089] Figure 7A and Figure 7B This is a circuit diagram illustrating a data converter according to an exemplary embodiment.

[0090] Reference Figure 7A The data converter 20 may include a multiplier 21, an exponent and sign calculator 22, a shifter 23, an adder 24, a leading one detector (LOD) 25, an increment and shift circuit 27, and an exponent updater 28. That is, the data converter 20 may include multiple floating-point units, such as multiplier 21, exponent and sign calculator 22, adder 24, LOD 25, and exponent updater 28. The data converter 20 may also include multiple shifters, for example, including shifter 23 as a first shifter and increment and shift circuit 27 as a second shifter. In one or more embodiments herein, at least some of the floating-point units and shifters may be deactivated while the remaining floating-point units and shifters are activated. In the embodiments described below, activation and deactivation may be performed based on the received arithmetic code.

[0091] The data converter 20 performs data conversion by processing a first input IN_A, a second input IN_B, and a third input IN_C. The first input IN_A, the second input IN_B, and the third input IN_C can be floating-point numbers or integers.

[0092] The following explanation will be given under the assumption that the first input IN_A, the second input IN_B, and the third input IN_C are floating-point numbers. As suggested above, a floating-point number N(FP) can be expressed in the form of a sign, an exponent part, and a fractional part. The signs Sa, Sb, and Sc of the first input IN_A, the second input IN_B, and the third input IN_C, as well as the exponent parts Ea, Eb, and Ec, are applied to the exponent and sign calculator 22, and the calculator 22 can determine the exponent and sign based on the signs Sa, Sb, and Sc and the exponents Ea, Eb, and Ec. The determined exponent and sign are various inputs, including the first input IN_A, the second input IN_B, and the third input IN_C. The exponent and sign calculator 22 can also determine the shift information SH1 provided to the shifter 23. The shift information SH1 may include the shift direction, the shift amount, and the shift number (e.g., the shift direction, the shift amount, and the shift number of the two inputs). The following will refer to... Figure 8A , Figure 11 and Figure 12 This section explains the detailed operation of the exponent and symbol calculator 22.

[0093] Multiplier 21 can multiply two inputs applied to it (e.g., the fractional part Fb of the second input IN_B and the fractional part Fc of the third input IN_C) and output the result of the multiplication. According to an embodiment, multiplier 21 may support MAC operations on integer data and may be implemented as a multiplier with extended bits (e.g., a 32-bit multiplier) rather than a 24-bit multiplier based on a single precision.

[0094] Shifter 23 can receive the output R1 of multiplier 21 and the first input IN_A, and shift the output R1 or the first input IN_A with the smaller exponent value based on the shift information SH1 provided from the exponent and sign calculator 22. Shifter 23 can be configured to output a first output of a first value LV and a second output of a second value SV.

[0095] Adder 24 can sum a first value LV and a second value SV received from shifter 23. The first value LV can be a first output from shifter 23, and the second value SV can be a second output from shifter 23. Adder 24 can be configured to produce a third output by summing the first value LV as the first output and the second value SV as the second output. In this case, adder 24 receives the absolute value of the first value LV and the absolute value of the second value SV, and adder 24 can also be configured to function as a subtractor depending on the signs of the first value LV and the second value SV.

[0096] LOD 25 can receive the output R2 of adder 24 and detect the position of the preceding "1" in the output R2 of adder 24. LOD 25 can generate shift information SH2 so that the preceding "1" becomes the MSB of the fractional part of the output value OUT used for post-normalizing the output R2 of adder 24.

[0097] The increment and shift circuit 27 determines whether to round the output R2 of adder 24 based on the rounding policy reflected by the arithmetic code OPCD. The increment and shift circuit 27 can determine the shift amount based on the shift information SH2 received from LOD 25, information indicating whether to round, and the exponent and sign (or shift information) provided from the exponent and sign calculator 22, and shift the output R2 of adder 24 accordingly. In other words, the increment and shift circuit 27 can be configured to shift the output of adder 24 based on a first shift value and a second shift value, the first shift value being based on the position of leading "1"s contained in the output of adder 24, and the second shift value being based on the exponent and sign of the input. Since at least three inputs have been detailed above, these three inputs include at least two inputs and one or more other inputs.

[0098] The index updater 28 can output the index and sign provided by the index and sign calculator 22 as the output value OUT, or update the index and sign and output the updated index and sign as the output value OUT.

[0099] According to an embodiment, the data converter 20 may have a structure similar to that of a fused multiply-add (FMA) circuit used for floating-point operations, and the components of the data converter 20 may vary. (Compared to the above...) Figure 5A In the illustrated embodiment, the data converter 20 may perform one or more functions between S1 and S2. For example, the data converter 20 may be configured to convert floating-point data received via the DMA controller into integer type data. As described above for... Figure 6 The DMA controller 10 may be implemented as an integrated component of the data converter 20 or as a separate component from the data converter 20. The data converter 20 may be implemented or configured to perform floating-point fused multiplication-addition (FMA) operations and may include floating-point FMA circuitry. The floating-point FMA circuitry may include adders.

[0100] According to the exemplary embodiment, the data converter 20 can employ the following reference based on the operation code OPCD. Figures 8A to 13 The operation is performed in one of the various operation modes. At this time, the components (multiplier 21, exponent and sign calculator 22, shifter 23, adder 24, LOD 25, increment and shift circuit 27, and exponent updater 28) can selectively perform operations. In other words, depending on the operation mode, some components of the data converter 20 can be activated while others can be deactivated, and the input / output relationships between components can be changed.

[0101] Furthermore, although the above description is given under the assumption that the first input IN_A, the second input IN_B, and the third input IN_C are floating-point numbers, in some operation modes, the first input IN_A, the second input IN_B, and the third input IN_C can be integers, and the data converter 20 and even the entire neural network device 120 can perform operations based on the integer values ​​A, B, and C of the first input IN_A, the second input IN_B, and the third input IN_C.

[0102] Figure 7B This is a circuit diagram illustrating a data converter according to an exemplary embodiment. (Refer to...) Figure 7B The data converter 20a may include a multiplier 21, an exponent and sign calculator 22, a shifter 23, an adder 24, a leading zero anticipator (LZA) 26, an increment and shift circuit 27, and an exponent updater 28.

[0103] Configuration and functions of data converter 20a Figure 7A The data converter 20 shown has a similar configuration and function. However, the LZA26 can be used as an alternative. Figure 7A LOD 25 is shown. The first value LV can be a first shift value, which is provided based on the position of the leading 1 contained in the output of adder 24. LZA26 can receive the first value LV and the second value SV output from shifter 23 and output a bit shifting amount based on the first value LV and the second value SV, for example, the shift information SH2 required for rounding and post-regularization of the output R2 of adder 24. Figure 7A The LOD 25 shown can be Figure 7BThe LZA 26 is shown as an alternative. For example, when high processing speed is required, such as... Figure 7B As shown, data converter 20a may include LZA 26. On the other hand, when low power consumption is required, data converter 20a may include LOD 25. However, data converter 20a is not limited to this. In some embodiments, data converter 20a may also include LOD 25 and LZA 26, and one of LOD 25 and LZA 26 may be selectively operated according to the performance required by data converter 20a. For example, LZA 26 may be used when fast processing speed of data converter 20a is required. LOD 25 may be used when low power consumption is required.

[0104] The following section will explain the data converter based on the example. Figure 7A The calculation method is determined by the calculation mode shown in Figure 20). The following explanation also applies to... Figure 7B The data converter 20a is shown.

[0105] Figure 8A This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment, and Figure 8B It is shown Figure 8A The diagram shows the input of the data converter. In detail, Figure 8A The diagram illustrates a method by which the data converter 20 quantizes floating-point data by performing floating-point FMA operations on the first input IN_A, the second input IN_B, and the third input IN_C, and converting the quantized values ​​into integer type data (quantized FP to INT conversion). For example... Figure 8B As shown, the first input IN_A, the second input IN_B, and the third input IN_C can respectively include the symbols Sa, Sb, and Sc, the exponent parts Ea, Eb, and Ec, and the fractional parts Fa, Fb, and Fc.

[0106] According to the prompt, the data converter 20 may include a plurality of floating-point units and shifters that can be selectively deactivated and activated. For example, deactivation and activation can be performed based on the received operand code so that the activated floating-point units and / or shifters perform an operation mode according to the operand code. Based on the received operand code, at least some of the floating-point units and shifters can be deactivated and the remaining floating-point units and shifters are activated.

[0107] The real number R can be expressed as shown in Equation 4, and the quantization equation of the real number R can be expressed as shown in Equation 5.

[0108] R = S × (QZ) = S × QS × Z [Equation 4]

[0109]

[0110] Here, Q represents the quantized integer of the real number R, S represents the scaling factor, and Z represents the zero offset. The real number R, the scaling factor S, and the zero offset Z are floating-point numbers.

[0111] Reference Figure 8A The first input IN_A, the second input IN_B, and the third input IN_C indicate the zero offset Z, the real number R to be converted, and the reciprocal of the scaling factor 1 / S. The arithmetic code OPCD indicates the data type of the output value OUT from the data converter 20 (e.g., whether the output value OUT is a signed or unsigned integer and the number of bits in the integer). The zero offset Z and the reciprocal of the scaling factor 1 / S are instances of two quantization parameters, and the quantization parameters can be provided by a neural network using a quantization model. The first input IN_A, the second input IN_B, and the third input IN_C are floating-point input data and can be retrieved from memory ( Figure 1 (See Figure 130) Receive.

[0112] The data converter 20 can output the quantized integer of the second input IN_B as the output value OUT by performing operations according to Equation 6.

[0113] OUT = INT(IN_A + IN_B × IN_C) [Equation 6]

[0114] Here, INT represents the conversion from floating-point to integer and indicates the integer conversion after the floating-point FMA operation of the first input IN_A, the second input IN_B, and the third input IN_C.

[0115] When the OCD indicator is used to indicate the quantization and data conversion operation mode, multiplier 21, exponent and sign calculator 22, shifter 23, adder 24, LOD 25, and increment and shift circuit 27 can be activated, and exponent updater 28 can be deactivated. Multiplier 21, exponent and sign calculator 22, shifter 23, adder 24, LOD 25, and increment and shift circuit 27 can be activated as described above. Figure 7A The calculations are performed to perform quantization and data conversion.

[0116] Here, the exponent and sign calculator 22 can select the larger value between the sum of the exponent portion Eb of the second input IN_B and the exponent portion Ec of the third input IN_C, or the exponent portion Ea of the first input IN_A, as the exponent portion of the output value OUT. The exponent and sign calculator 22 can perform an XOR operation on the sign Sb of the second input IN_B and the sign Sc of the third input IN_C, and select the larger value between the result of the XOR operation and the sign Sa of the first input IN_A as the sign of the output value OUT. Furthermore, when the data type according to the operation code OPCD is signed, the exponent and sign calculator 22 can perform sign extension. The exponent and sign calculator 22 can provide a flag indicating a shift reflecting the sign extension as shift information SH1 to the shifter 23.

[0117] The value obtained by subtracting the sum of the exponents of the second input IN_B and the third input IN_C, Ec, from the exponent portion Ea of the first input IN_A (i.e., the exponent difference D) may be greater than 0. When this occurs, the exponent and sign calculator 22 can generate shift information SH1, which instructs the output R1 of multiplier 21 to be shifted to the right by an amount equal to the absolute value of the exponent difference D. The output R1 is the value obtained by multiplying the fractional part Fb of the second input IN_B by the fractional part Fc of the third input IN_C. When the exponent difference D is less than 0, the exponent and sign calculator 22 can generate shift information SH1, which instructs the fractional part Fa of the first input IN_A to be shifted to the right by an amount equal to the absolute value of the exponent difference D. The shifter 23 can perform a shift operation based on the shift information SH1.

[0118] In this embodiment, since the output value OUT of the exponent and data converter 20 is an integer, the exponent updater 28 does not perform any calculations. Depending on the data type represented by the operation code OPCD, the output value OUT as an integer can be signed or unsigned, and the number of data bits can be determined.

[0119] Figure 9A This is a diagram illustrating an example of the operation mode of a data converter according to an exemplary embodiment, and Figure 9B It is shown Figure 9A The diagram shows the input of the data converter. In detail, Figure 9A This illustrates a method for converting the result of a neural network operation (i.e., integer data) including the output of a neural network processor 30 into floating-point data in a neural network using a quantized model (dequantized INT to FP conversion).

[0120] The first input IN_A can be the value obtained by multiplying the scaling factor S by the zero offset Z, the second input IN_B can be a quantized integer that is the output of the neural network processor 30, and the third input IN_C can be the scaling factor S.

[0121] like Figure 9A As shown, the first input IN_A and the third input IN_C can be floating-point numbers, and the second input IN_B can be an integer. The exponent and sign calculator 22 can set the sign Sb and the exponent Eb of the second input IN_B so that the second input IN_B can be calculated as a floating-point number. In other words, the second input IN_B, as an integer, can be preprocessed into a floating-point number. For example, when the second input IN_B is a signed 8-bit integer (i.e., a "signed char"), the MSB of the second input IN_B can be set to the sign Sb of the second input IN_B, and 7 can be set to the exponent Eb of the second input IN_B. In addition, the 8-bit data of the second input IN_B (i.e., the integer value INTb) can be provided to the multiplier 21 as the fractional part Fb of the second input IN_B. When the second input IN_B is an unsigned 16-bit integer (i.e., an "unsigned char"), the sign Sb of the second input IN_B can be set to 0, the exponent Eb can be set to 15, and the 16-bit data of the second input IN_B can be set to the fractional part Fb. The data converter 20 can output the dequantized floating-point number of the second input IN_B as the output value OUT by performing a calculation according to Equation 7. OUT = FLP(IN_A + IN_B × IN_C)

[0122] Equation 7

[0123] Here, FLP represents the conversion from integer to floating-point number and indicates the preprocessing of the second input IN_B, which is an integer, into a floating-point number, as well as the floating-point FMA operation performed on the first input IN_A, the second input IN_B, and the third input IN_C.

[0124] In this embodiment, all components of the data converter 20 can be activated and perform operations. Since the operation of the components of the data converter 20 is similar to that described above... Figure 7A and Figure 9A The components described above operate in the same way, so the above descriptions will be omitted.

[0125] Figure 10 This is a diagram illustrating an example of the operational mode of a data converter according to an exemplary embodiment. Specifically, this embodiment instructs the data converter 20 to receive integer type data as input and perform a MAC operation on the integer type data. When the neural network processor ( Figure 6As shown in Figure 30), during general operations, the data converter 20 can process the input data, thereby reducing the workload of the neural network processor ( Figure 6 The load shown in Figure 30). For example, when the input data is YUV type input data and the neural network processor 30 is capable of processing RGB type input data, the data converter 20 can convert the YUV type input data into RGB type input data and provide the RGB type input data to the neural network processor 30 when receiving data from the memory 130.

[0126] Reference Figure 10 The first input IN_A, the second input IN_B, and the third input IN_C are integers. When the OPCD indicator operation mode is used to guide the MAC operation of integers, multiplier 21 and adder 24 can be activated, and other configurations (i.e., exponent and sign calculator 22, shifter 23, LOD 25, increment and shift circuit 27, and exponent updater 28) can be deactivated. Data converter 20 can perform operations based on the data A, B, and C of the first input IN_A, the second input IN_B, and the third input IN_C. Multiplier 21 can multiply the data B of the second input IN_B by the data C of the third input IN_C and output the result of the multiplication. The output R1 of multiplier 21 and the data A of the first input IN_A can bypass shifter 23 and be provided to adder 24 as the first value LV and the second value SV. Adder 24 can output the result of summing the first value LV and the second value SV. The output R2 of adder 24 can bypass LOD 25 and increment and shift circuit 27 and be output as the output value OUT of data converter 20.

[0127] Figure 11 This is a diagram illustrating an example of the operational mode of a data converter according to an exemplary embodiment. Specifically, this embodiment instructs the data converter 20 to convert floating-point numbers into integers.

[0128] Reference Figure 11 It can receive the first input IN_A as a floating-point number, and when the operation code OPCD indicates the operation mode to guide the conversion of floating-point number to integer, it can activate the exponent and sign calculator 22 and the shifter 23, and deactivate other configurations (i.e., multiplier 21, adder 24, LOD 25, increment and shift circuit 27 and exponent updater 28).

[0129] The exponent and sign calculator 22 can receive the sign Sa and exponent Ea of the first input IN_A. When the data type according to the arithmetic code OPCD is signed, the exponent and sign calculator 22 can perform sign extension. The exponent and sign calculator 22 can provide a flag indicating the shift reflecting the sign extension as shift information SH1 to the shifter 23.

[0130] The exponent and sign calculator 22 provides shift information SH1 (which, when the exponent part Ea is greater than 0, instructs the fractional part Fa to be shifted left by the same amount as the exponent part Ea) to the shifter 23. Based on the shift information SH1, the shifter 23 shifts the fractional part Fa of the first input IN_A to the left and outputs the shift result R2 as the output value OUT of the data converter 20. Here, when the value of the exponent part Ea exceeds the number of bits of the output value OUT according to the exponent type (e.g., "char", "short", and "int"), the output value OUT can be the maximum value of the corresponding integer type. For example, when the output value OUT is expressed as a 32-bit integer (i.e., expressed as "int" type) and the value of the exponent part Ea is greater than 32, the maximum value that can be expressed according to the "int" type can be output as the output value OUT.

[0131] The exponent and sign calculator 22 can provide information indicating that the output value OUT is 1 when the exponent part Ea is 1 and information indicating that the output value OUT is 0 when the exponent part Ea is 0, and the shifter 23 can output 0 or 1 as the output value OUT based on the information.

[0132] Therefore, the first input IN_A, which is a floating-point number, can be converted into an integer. Here, depending on the data type represented by the arithmetic code OPCD, the output value OUT, which is an integer, can be signed or unsigned, and the number of data bits can be determined.

[0133] Figure 12 This is a diagram illustrating an example of the operational mode of a data converter according to an exemplary embodiment. Specifically, this embodiment of the invention instructs the data converter 20 to convert integers into floating-point numbers.

[0134] Reference Figure 12 It can receive the first input IN_A as an integer, and when the operation code OPCD indicates the operation mode to guide the conversion of integer to floating point (INT to FP conversion), it can activate the exponent and sign calculator 22, adder 24 and exponent updater 28, and deactivate other configurations (i.e. multiplier 21, shifter 23, LOD 25 and increment and shift circuit 27).

[0135] The exponent and sign calculator 22 can receive the MSB of the first input IN_A and determine the sign of the output value OUT based on the MSB. The data A of the first input IN_A can bypass the shifter 23 and be provided as the first value LV of the adder 24. The data 0 can be provided as the second value SV, and the adder 24 can produce the value obtained by summing the first value LV and the second value SV based on the sign determined by the exponent and sign calculator 22 as the output R2. That is, the adder 24 can be configured to produce a third output by summing the first output and the second output. The LOD 25 can detect a leading "1" in the output R2 of the adder 24 based on the output R2 of the adder 24 and generate shift information SH2 based on the detection result. The shift information SH2 is used to match the output R2 of the adder 24 with the fractional part of the floating-point number. For example, the shift information SH2 may include a shift value for representing the exponent part of the floating-point number.

[0136] The increment and shift circuit 27 can output data bits to represent the valid data in the output R2 of the adder 24. In addition, the increment and shift circuit 27 can generate shift information SH to indicate the value obtained by adding 1 to the shift information SH2 (the shift information SH2 is the shift value calculated by LOD 25), and provide the shift information SH to the exponent and sign calculator 22.

[0137] In this embodiment, since the output value OUT of the data converter 20 is a floating-point number, the exponent updater 28 can perform sign extension, and the exponent portion Ea of the output value OUT can be determined based on the number of bits of the first input IN_A. At this time, the exponent updater 28 can increment the data of the exponent portion Ea by shifting the exponent portion Ea to the left based on the shift information SH provided from the increment and shift circuit 27, thereby updating the exponent portion Ea. The sign Sa, the updated exponent portion Ea output from the exponent updater 28, and the output R3 (i.e., the fractional part Fa) of the increment and shift circuit 27 can constitute the output value OUT as a floating-point number. Therefore, the first input IN_A, which is an integer, can be converted into a floating-point number.

[0138] Figure 13 This is a circuit diagram illustrating a data converter according to an exemplary embodiment.

[0139] Reference Figure 13 The data converter 20b may include a multiplier 21, an exponent and sign calculator 22, a shifter 23, an adder 24, a leading zero predictor (LZA) 26, and an exponent updater 28.

[0140] Figure 13 The configuration and operation of the data converter 20b shown are as follows: Figure 7B The configuration and operation of the data converter 20b shown are similar. However, Figure 7BThe function of the increment and shift circuit 27 shown is the same as or similar to the function of the combination of shifter 23 and adder 24, and therefore shifter 23 and adder 24 can implement the function of increment and shift circuit 27. Figure 13 In the data converter 20b shown, the shifter 23 can perform... Figure 7B The increment and shift circuit 27 shown has a shift function, and the adder 24 can perform... Figure 7B The increment and shift circuit 27 shown above demonstrates its increment function. (Refer to the above...) Figure 8A , Figure 9A and Figure 12 When the data converter 20b performs one of the following conversions: quantized FP to INT, dequantized INT to FP, and INT to FP, the shifter 23 can receive the output R2 of the adder 24 and shift the output R2 of the adder 24 based on the shift information SH2 provided from the LZA 26 and the exponent and sign determined by the exponent and sign calculator 22. Therefore, the output R2 of the adder 24 can be regularized (and quantized), and the regularized output R3 can be output via the adder 24.

[0141] Figure 14 This is a block diagram illustrating a neural network device according to an exemplary embodiment.

[0142] Reference Figure 14 The neural network device 120a may include a DMA controller 10, a first data converter 20_1, a second data converter 20_2, and a neural network processor 30. The functions of the DMA controller 10 and the neural network processor 30 are similar to... Figure 6 The DMA controller 10 and neural network processor 30 shown have the same function. Therefore, the descriptions given above will be omitted.

[0143] exist Figure 14 In this process, the first data converter 20_1 can convert the fixed-point input data FPID received via the DMA controller 10 into integer type input data INTID. The first data converter 20_1 can then provide the integer type input data INTID to the neural network processor 30.

[0144] Additionally, the second data converter 20_2 can convert the integer type output data INTOD from the neural network processor 30 into floating-point output data FPOD. The second data converter 20_2 can then provide the floating-point output data FPOD to the DMA controller 10.

[0145] The structures of the first data converter 20_1 and the second data converter 20_2 can be referenced. Figure 7A , Figure 7B and Figure 13The data converters 20, 20a, and 20b are structurally similar. However, the first data converter 20_1 may not include the exponent updater 28.

[0146] According to an embodiment, the neural network device 120a may include a plurality of data converters. For example, the neural network device 120a may include a plurality of first data converters 20_1 and a plurality of second data converters 20_2. The number of first data converters 20_1 and the number of second data converters 20_2 may be set based on the processing speed of the neural network processor 30 and the bandwidth of the memory 130.

[0147] Figure 15 This is a block diagram illustrating a data processing system according to an exemplary embodiment.

[0148] Reference Figure 15 The data processing system 1000 may include a CPU 1100 and multiple intellectual property (IP) units, including a first IP 1200, a second IP 1300, and a third IP 1400. The term "intellectual property" and the abbreviation "IP" refer to unique circuitry that may be individually protected by intellectual property rights. Therefore, intellectual property or IP may be referred to as IP areas, IP circuits, or other forms, but the terms and abbreviations are used for convenience to refer to circuitry and circuit systems that may be uniquely protected by specific intellectual property rights. Components of the data processing system 1000 (e.g., the CPU 1100 and the first IP 1200, second IP 1300, and third IP 1400) may exchange data with each other via a bus 1500. According to an embodiment, the data processing system 1000 may be implemented as a system-on-a-chip (SoC).

[0149] The CPU 1100 controls the overall operation of the data processing system 1000, and each of the first IP 1200, the second IP 1300 and the third IP 1400 can perform a specific function in the data processing system 1000.

[0150] The first IP 1200, one of the first IPs 1200, 1300, and 1400, can perform operations based on integer arithmetic or integer processing. According to an embodiment, the first IP 1200 may be an accelerator including one or more integer arithmetic units. The first IP 1200 may include a data converter 1210 that performs data conversion between floating-point data (FPD) and integer type data (INTD). The data converter 1210 may be implemented as the data converter 20 shown in FIG. 7a, the data converter 20a shown in FIG. 7b, or according to the various embodiments described above. Figure 13The data converter 20b is shown. Data converter 1210 can convert floating-point data (FPD) received from other components (e.g., CPU 1100, second IP 1300, and / or third IP 1400) into integer type data (INTD). The result of an operation performed by first IP 1200 based on the input data (i.e., integer type data (INTD)) can be output as output data (i.e., integer type data (INTD)), and data converter 1210 can convert integer type data (INTD) back to floating-point data (FPD).

[0151] Although the first IP 1200 in the embodiments of the present invention has been described above as including a data converter 1210, the IPs described herein are not limited thereto. According to embodiments, other IPs (e.g., the second IP 1300 and / or the third IP 1400) may also include a data converter 1210 to convert floating-point data FPD received from an external source into integer type data INTD for processing or performing integer operations on the integer type data INTD. According to embodiments, at least one of the second IP 1300 and the third IP 1400 may be a memory (e.g., DRAM, SRAM, etc.).

[0152] Figure 16 This is a block diagram illustrating an application processor according to an exemplary embodiment.

[0153] Reference Figure 16 The application processor 2000 may include a CPU 2100, random access memory (RAM) 2200, GPU 2300, computing device 2400, sensor interface (I / F) 2500, display interface 2600, and memory interface 2700. Additionally, the application processor 2000 may include a communication module. The components of the application processor 2000 (CPU 2100, RAM 2200, GPU 2300, computing device 2400, sensor interface 2500, display interface 2600, and memory interface 2700) can send and receive data between each other via a bus 2800.

[0154] CPU 2100 controls the overall operation of application processor 2000. CPU 2100 may include one processing core (single core) or multiple processing cores (multi-core). CPU 2100 can process or execute programs and / or data stored in memory 2710. According to an exemplary embodiment, CPU 2100 can control the functions of computing device 2400 by executing programs stored in memory 2710.

[0155] RAM 2200 may temporarily store programs, data, and / or instructions. According to embodiments, RAM 2200 may be implemented as dynamic random access memory (DRAM) or static random access memory (SRAM). RAM 2200 may be input / output via sensor interface 2500 and display interface 2600, or may temporarily store images generated by GPU 2300 or CPU 2100.

[0156] According to an embodiment, the application processor 2000 may include a read-only memory (ROM). The ROM may store continuously used programs and / or data. The ROM may be implemented as an erasable programmable read-only memory (EPROM) or an electrically erasable / programmable read-only memory (EEPROM).

[0157] GPU 2300 can perform image processing on image data. For example, GPU 2300 can perform image processing on image data received via sensor interface 2500. According to an embodiment, GPU 2300 can perform floating-point operations.

[0158] Image data processed by GPU 2300 can be stored in memory 2710 or provided to display device 2610 via display interface 2600. Image data stored in memory 2710 can be provided to computing device 2400.

[0159] The sensor interface 2500 can interface with data (e.g., image data, voice data, etc.) input from the sensor 2510 connected to the application processor 2000.

[0160] Display interface 2600 can interface with data (e.g., images) output to display device 2610. Display device 2610 can output data about images or voice via display device (e.g., liquid crystal display (LCD) or active matrix organic light-emitting diode (AMOLED)).

[0161] The memory interface 2700 interfaces to data input from or output to the memory 2710 external to the application processor 2000. According to embodiments, the memory 2710 may be implemented as volatile memory (such as DRAM or SRAM) or non-volatile memory (such as ReRAM, PRAM, or NAND flash memory). For example, the memory 2710 may be implemented as a memory card (e.g., a multimedia card (MMC), an embedded multimedia card (eMMC), an SD card, a micro SD card, etc.).

[0162] As referenced above Figures 1 to 14 The computing device 2400 may include a data converter 20 for performing data conversion between floating-point data and integer type data, which can convert input floating-point data into integer type data and convert the result of integer operations performed on the integer data (i.e., integer data) into floating-point data. According to an embodiment, the computing device 2400 may be a neural network device that performs neural network operations based on integer operations. For example, the computing device 2400 may be implemented as... Figure 1 The neural network device 120 is shown. When reading floating-point data from and sending floating-point data to memory 2710, the computing device 2400 can perform data conversion internally. Therefore, the load on the CPU 2100 can be reduced, and the processing speed of the application processor 2000 can be increased. In addition, since the computing device 2400 performs integer operations, power consumption can be reduced, and the accuracy of the data processing results can be improved.

[0163] Although the inventive concept of this disclosure has been specifically shown and illustrated with reference to embodiments thereof, it will be understood that various changes in form and detail may be made herein without departing from the spirit and scope of the foregoing claims.

Claims

1. A neural network device for performing neural network operations, the neural network device comprising: The direct memory access controller is configured to receive floating-point data from memory; A data converter is configured to convert floating-point data, including feature values ​​and operational parameters, received via the direct memory access controller into integer data, including the feature values ​​and operational parameters. The data converter includes multiple floating-point units, each including a multiplier, an adder, and multiple shifters. Based on an operand code, the data converter operates according to a first operation mode and a second operation mode. In the first operation mode, quantization and data conversion of the floating-point data are performed using a scaling factor and a zero-point offset based on floating-point fused multiplication-addition operations by activating the multipliers and adders. In the second operation mode, data conversion of the floating-point data is performed by deactivating the multipliers and adders. The integer data includes both signed and unsigned values. as well as The processor is configured to perform the neural network operation based on integer operations using the integer type data, including the feature values ​​and the operation parameters, provided from the data converter.

2. The neural network device of claim 1, wherein the processor comprises a plurality of processing elements, each of the plurality of processing elements performing the integer operation.

3. The neural network device of claim 1, wherein, In the first operation mode, the data converter performs the quantization of the input value and the data conversion based on two quantization parameters, including the scaling factor and the zero-point offset, and the input value, wherein the two quantization parameters and the input value are floating-point numbers.

4. The neural network device of claim 3, wherein the output of the data converter is a quantized integer and the presence of the sign of the quantized integer and the number of data bits of the quantized integer are determined based on the arithmetic code applied to the data converter.

5. The neural network device of claim 1, wherein the data converter converts integer-type output data output from the processor into floating-point output data.

6. The neural network device of claim 5, wherein the direct memory access controller provides the floating-point output data directly to the memory.

7. The neural network device of claim 5, wherein the data converter performs inverse quantization and data conversion on the output value based on two quantization parameters and the output value of the integer type data, wherein the two quantization parameters are floating-point numbers and the output value is an integer.

8. The neural network device of claim 1, wherein the data converter selectively receives integer type input values ​​and performs integer operations based on the integer type input values.

9. The neural network device of claim 1, wherein the data converter comprises the plurality of floating-point units and shifters, and at least some of the floating-point units and shifters are deactivated according to the received operation code, while the remaining floating-point units and shifters are activated to perform the operation mode according to the operation code.

10. The neural network device of claim 1, wherein the data converter comprises: The shifter is configured to output a first output and a second output; as well as A floating-point fused multiplication-adder circuit, including the adder, the adder being configured to produce a third output by summing the first output and the second output. The output of the adder is provided to the shifter.

11. The neural network device of claim 1, wherein the processor performs convolution operations based on integer type weight values ​​and integer type input feature values ​​provided from the data converter.

12. A method for operating a neural network device, the method comprising: Receive floating-point input data, including feature values ​​and calculation parameters, from memory; A data converter transforms the floating-point input data, including the feature values ​​and the operational parameters, into integer input data, also including the feature values ​​and the operational parameters. The data converter includes multiple floating-point units, each comprising a multiplier, an adder, and multiple shifters. The conversion of the floating-point input data into integer input data, based on the operand code, is performed using a first operational mode and a second operational mode. In the first operational mode, the multipliers and adders are activated, and a scaling factor and zero-point offset are used to quantize and convert the floating-point input data using floating-point fused multiplication-addition operations. In the second operational mode, the multipliers and adders are deactivated to perform the data conversion of the floating-point input data. The integer input data includes both signed and unsigned values. The neural network operation is performed using the integer type input data, which includes the feature values ​​and the operation parameters.

13. The method of claim 12, wherein, During the process of receiving the floating-point input data, the floating-point input data is received directly from the memory.

14. The method of claim 12, further comprising: In the first operation mode, during the process of converting the floating-point input data into the integer type input data, the input value received as the floating-point input data is quantized and converted based on two quantization parameters, including the scaling factor and the zero-point offset.

15. The method of claim 12, further comprising: The integer output data based on the operation result of the neural network is converted into floating-point output data. as well as The floating-point output data is sent to the memory.

16. The method of claim 15, further comprising: In the process of converting the integer type output data into the floating-point output data, the output value of the floating-point output data is dequantized and the data is converted based on the two quantization parameters received as the floating-point input data.

17. An application processor, comprising: The memory is configured to store operational parameters and feature values, wherein the operational parameters and feature values ​​are floating-point type values; as well as A neural network device is configured to receive the computational parameters and the feature values ​​from the memory, convert the computational parameters and the feature values ​​into integer type values, and perform neural network operations based on the computational parameters and the feature values ​​converted into integer type values. The neural network device includes: A data converter is configured to convert the operational parameters and the feature values ​​into integer type values. The data converter includes multiple floating-point units, each comprising a multiplier, an adder, and multiple shifters. Based on the operation code, the data converter operates according to a first operational mode and a second operational mode. In the first operational mode, the multipliers and adders are activated, and a scaling factor and zero-point offset are used to perform quantization and data conversion of the floating-point type values ​​based on floating-point fused multiplication-addition operations. In the second operational mode, the multipliers and adders are deactivated to perform data conversion of the floating-point type values. The integer type values ​​include both signed and unsigned values.

18. The application processor of claim 17, wherein the neural network device includes a neural network processor configured to perform the neural network operation by performing integer operations based on the operation parameters and the feature values ​​converted into the integer type values.

19. The application processor of claim 18, wherein the data converter converts the result of the neural network operation, which is an integer type value, into a floating-point number.

20. The application processor of claim 18, wherein the neural network processor comprises a plurality of processing elements, each of the plurality of processing elements comprising an integer arithmetic unit.

21. The application processor of claim 18, wherein the neural network processor further comprises a direct memory access controller configured to communicate directly with the memory and receive the operational parameters and the feature values ​​as floating-point type values ​​from the memory.

22. The application processor of claim 18, wherein the data converter comprises: The multiplier is configured to multiply at least two inputs; The first shifter is configured to shift one of the following based on the exponent and sign of the at least two inputs and the remaining inputs: the output of the multiplier and one of the other inputs; The adder is configured to sum the two values ​​output from the first shifter; as well as A second shifter is configured to shift the output of the adder based on a first shift value and a second shift value, the first shift value being provided based on the position of a leading 1 contained in the output of the adder, and the second shift value being provided based on the exponent and sign of the at least two inputs and the remaining inputs.

23. The application processor of claim 22, wherein the data converter deactivates at least one of the multiplier, the first shifter, the adder, and the second shifter according to the operation mode.

Citation Information

Patent Citations

  • Semiconductor device

    CN107391082A

  • Implementing neural networks in fixed point arithmetic computing systems

    CN108427991A