Method and neural processing unit for artificial neural networks
Patent Information
- Application Number
- CN202111624452.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-06
- Filing Date
- 2021-12-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-12-28
AI Technical Summary
[0011]因此,存在一个问题,即从主存储器中读取必要的参数并执行卷积所消耗的时间和功率非常大
[0048]根据本公开,被配置为处理多个输入通道的批模式可以确定片上存储器和/或内部存储器存储和计算人工神经网络的参数的操作顺序。因此,可以减少主存储器读取操作的次数并且可以降低功耗。
Smart Images

Figure CN114692855B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods and neural processing units for artificial neural networks. Background Technology
[0002] Humans possess intelligence capable of recognition, classification, reasoning, prediction, and control / decision-making. Artificial intelligence (AI) refers to the artificial imitation of human intelligence.
[0003] The human brain is composed of numerous nerve cells called neurons, each connected to hundreds or thousands of other neurons via connectors called synapses. To mimic human intelligence, the modeling of the operating principles of biological neurons and the connections between them is called an artificial neural network (ANN) model. In other words, an artificial neural network is a system of nodes that mimic neurons, connected in a layered structure.
[0004] These artificial neural network models are classified into "single-layer neural networks" and "multi-layer neural networks" based on the number of layers.
[0005] A typical multilayer neural network consists of an input layer, hidden layers, and an output layer. (1) The input layer receives external data; the number of neurons in the input layer is the same as the number of input variables. (2) The hidden layer is located between the input and output layers, receiving signals from the input layer, extracting features, and transmitting them to the output layer. (3) The output layer receives signals from the hidden layers and outputs them to the outside. The input signals between neurons are multiplied by the connection strength of each connection, with values between zero and one, and then summed. If the sum is greater than a neuron threshold, the neuron is activated and implemented as an output value through an activation function.
[0006] Meanwhile, in order to achieve higher levels of artificial intelligence, artificial neural networks with an increased number of hidden layers are called deep neural networks (DNNs).
[0007] There are many types of DNNs, but as is well known, Convolutional Neural Networks (CNNs) are good at extracting features from input data and recognizing feature patterns.
[0008] A convolutional neural network (CNN) is a type of neural network that functions similarly to image processing in the human brain's visual cortex. CNNs are well-known for their suitability for image processing.
[0009] refer to Figure 4Convolutional neural networks (CNNs) are configured with alternating convolutional and pooling channels. In a CNN, most of the computation time is consumed by convolution operations. CNNs identify objects by extracting image features for each channel via matrix-like kernels and providing a dynamic balance, such as dynamics or distortion, through pooling. For each channel, a feature map is obtained by convolving the input data with the kernel, and an activation function, such as a rectified linear unit (ReLU), is applied to generate the corresponding channel's activation map. Pooling can then be applied. The actual neural network classifying the patterns is located at the end of the feature extraction neural network and is called a fully connected layer. In the computational processing of a CNN, most of the computation is performed through convolution or matrix multiplication. At this point, the necessary kernels are fetched from memory quite frequently. A large portion of the operations in a CNN require time to fetch the kernel corresponding to each channel from memory.
[0010] Memory can be divided into main memory, internal memory, and on-chip memory. Each memory consists of multiple memory cells, and each memory cell has a unique memory address. In particular, when a neural processing unit reads weights stored in main memory, there may be a delay of several clock cycles until the memory cell corresponding to that memory address is accessed.
[0011] Therefore, there is a problem that reading the necessary parameters from main memory and performing convolution consumes a lot of time and power. Summary of the Invention
[0012] The inventors of this disclosure have recognized the following matters.
[0013] First, during the inference operation of the artificial neural network model, the neural processing unit (NPU) frequently reads the nodes (i.e., features) and / or weight values (i.e., kernels) of each layer of the artificial neural network model from the main memory.
[0014] The NPU is affected by slow processing speed and high power consumption when reading node and / or kernel weight values of artificial neural network models from main memory.
[0015] As access to on-chip memory or NPU internal memory increases rather than access to main memory, the NPU's processing speed increases and its power consumption decreases.
[0016] When using an NPU and an artificial neural network model to process multiple channels, it is inefficient to repeatedly read the same weights from main memory each time each channel is processed individually.
[0017] Especially when processing batch channels (where data is arranged and processed in queues), the on-chip memory or NPU internal memory can be utilized to the maximum extent according to the characteristics of the processing method and sequence.
[0018] Finally, by maximizing the retention of parameters that are repeatedly used or reused in the convolutional processing of batch channels in on-chip memory or NPU internal memory, processing speed can be maximized and energy consumption reduced. That is, by maintaining the at least one kernel or the at least one weight value until the computation of the corresponding feature maps for multiple batch channels is completed, at least one kernel or at least one weight value corresponding to multiple feature maps for multiple batch channels can be reused.
[0019] Therefore, the problem to be solved by this disclosure is to provide a neural processing unit and its operation method. The neural processing unit can reduce the number of read operations of the main memory and reduce power consumption by determining the storage order of the on-chip memory or the internal memory of the NPU and the calculation order of the artificial neural network parameters.
[0020] Another problem this disclosure aims to address is providing a high-performance and low-power neural processing unit and its operation method in autonomous vehicles, drones, or electronic devices with multiple sensors, where batch channels are frequently processed.
[0021] To address the aforementioned issues, a method for artificial neural networks is provided based on an example of this disclosure.
[0022] According to one aspect of this disclosure, a method for performing multiple operations on an artificial neural network (ANN) is provided. The multiple operations may include: storing in at least one memory a set of weights, at least a portion of a first batch of channels in a plurality of batch channels, and at least a portion of a second batch of channels in a plurality of batch channels; and calculating, using the set of weights, at least a portion of the first batch channels and at least a portion of the second batch channels.
[0023] At least a portion of the first batch of channels and at least a portion of the second batch of channels may be substantially equal in size.
[0024] The weights of the group can correspond to each of the at least portion of the first batch of channels and the at least portion of the second batch of channels.
[0025] The plurality of operations may further include: while calculating at least a portion of the second batch of channels by means of the grouped weights, storing at least another portion of the first batch of channels to be subsequently calculated in at least a portion of the at least one memory in the at least one memory, wherein the at least one memory includes at least one of on-chip memory and internal memory.
[0026] The plurality of operations may further include: storing at least a portion of a third batch channel and at least a portion of a fourth batch channel in the plurality of batch channels in the at least one memory, while maintaining the group weights; and calculating the at least a portion of the third batch channel and the at least a portion of the fourth batch channel using the group weights. The at least one memory may include at least one of on-chip memory and internal memory, wherein the group weights are maintained until at least a portion of each of the plurality of batch channels is calculated.
[0027] The multiple operations may further include: storing subsequent grouped weights, subsequent portions of the first batch of channels, and subsequent portions of the second batch of channels in the at least one memory; and calculating the subsequent portions of the first batch of channels and the subsequent portions of the second batch of channels using the subsequent grouped weights, wherein the at least one memory may include at least one of on-chip memory and internal memory.
[0028] The multiple operations may further include: storing the weights of the groups and a first value of the groups calculated based on at least a portion of the first batch of channels and at least a portion of the second batch of channels in the at least one memory; storing subsequent weights of the groups for subsequent processing steps in the at least one memory; and calculating the first value of the groups and the weights of the subsequent groups. The at least one memory may include internal memory storing the first value of the groups and a second value of the groups obtained by calculating the weights of the subsequent groups.
[0029] The at least portion of the first batch of channels and the at least portion of the second batch of channels may include the complete dataset.
[0030] The multiple operations may further include: tiling the size of the grouped weights, the size of at least a portion of the first batch of channels, and the size of at least a portion of the second batch of channels to fit the at least one memory, and the at least one memory may include internal memory.
[0031] The ANN can be configured to perform at least one of the plurality of operations, the at least one operation including the detection, classification, or segmentation of objects from the plurality of batch channels. The objects may include at least one of vehicles, traffic lights, obstacles, pedestrians, people, animals, roads, traffic signs, and lanes.
[0032] The multiple operations may further include preprocessing the multiple batch channels before storing at least a portion of the first batch channels and at least a portion of the second batch channels in the at least one memory. The ANN may be configured to simultaneously detect objects from the multiple batch channels, while simultaneously preprocessing the multiple batch channels to improve the object detection rate.
[0033] Each of the plurality of batch channels may correspond to a plurality of images. The plurality of batch channels includes at least one batch channel having one of the following formats: IR, RGB, YCBCR, HSV, and HIS. The plurality of batch channels may include at least one batch channel for capturing images of the vehicle interior, and the ANN may be configured to detect at least one of vehicle safety-related objects, functions, driver status, and passenger status. The plurality of images may include at least one of RGB images, IR images, radar images, ultrasonic images, lidar images, thermal images, NIR images, and fused images. The plurality of images may be captured within substantially the same time period.
[0034] Each of the plurality of batch channels corresponds to a plurality of sensor data, and the plurality of sensor data includes data from at least one of a pressure sensor, a piezoelectric sensor, a humidity sensor, a dust sensor, a smoke sensor, a sonar sensor, a vibration sensor, an acceleration sensor, and a motion sensor.
[0035] According to another aspect of this disclosure, a neural processing unit is provided for processing a plurality of batch channels of an artificial neural network including a first batch channel and a second batch channel. The neural processing unit may include: at least one internal memory configured to store at least a portion of the first batch channel, at least a portion of the second batch channel, and a set of weights; and at least one processing element (PE) configured to apply the stored set of weights to the at least a portion of the first batch channel and the at least a portion of the second batch channel.
[0036] At least a portion of the first batch of channels allocated to the at least one internal memory and at least a portion of the second batch of channels allocated to the at least one internal memory may be substantially equal in size.
[0037] The weights of the group can correspond to each of the at least portion of the first batch of channels and the at least portion of the second batch of channels.
[0038] The plurality of batch channels may include a third batch channel and a fourth batch channel. The at least one internal memory may also be configured to store at least a portion of the third batch channel and at least a portion of the fourth batch channel while maintaining the group weights, and the at least one PE may also be configured to calculate the at least a portion of the third batch channel, the at least a portion of the fourth batch channel, and the group weights. The at least one internal memory may also be configured to maintain the group weights until the plurality of batch channels have been calculated.
[0039] The at least one PE can also be configured to calculate the weights of the subsequent portions of the first batch of channels, the subsequent portions of the second batch of channels, and another group, and the at least one internal memory can also be configured to store the weights of the subsequent portions of the first batch of channels, the subsequent portions of the second batch of channels, and the other group.
[0040] The at least one PE can also be configured to calculate weights and values for another group for subsequent stages based on at least a portion of the first batch of channels and at least a portion of the second batch of channels. The at least one internal memory can also be configured to store the calculated values and the weights of the other group, and the weights of the other group can be retained in the internal memory until the plurality of batches of channels are calculated.
[0041] The at least one PE may also be configured to calculate a first value based on at least a portion of the first batch of channels and at least a portion of the second batch of channels, and to calculate subsequent group weights for subsequent processing stages. The at least one internal memory may also be configured to be of a size corresponding to at least a portion of the first batch of channels and at least a portion of the second batch of channels, and to store the first value, the group weights, and the subsequent group weights.
[0042] The neural processing unit further includes a scheduler configured to adjust the size of the grouped weights, the size of at least a portion of the first batch of channels, and the size of at least a portion of the second batch of channels for the internal memory.
[0043] According to another aspect of this disclosure, a neural processing unit is provided for processing an artificial neural network (ANN) comprising a first batch of channels and a second batch of channels. The neural processing unit may include: at least one internal memory configured to store at least a portion of the first batch of channels, at least a portion of the second batch of channels, and groups of weights; and at least one processing element (PE) configured to apply the stored groups of weights to the at least a portion of the first batch of channels and the at least a portion of the second batch of channels, wherein the size of the at least a portion of the first batch of channels may be less than or equal to the size of the at least one internal memory divided by the number of the plurality of batch channels.
[0044] The size of the at least one internal memory may correspond to the size of the maximum feature map of the ANN and the number of the plurality of batch channels.
[0045] The at least one internal memory can also be configured to store the compression parameters of the ANN.
[0046] The neural processing unit may further include a scheduler operatively coupled to the at least one PE and the at least one internal memory and configured to adjust the size of at least a portion of the first batch of channels or the size of at least a portion of the second batch of channels.
[0047] The neural processing unit may further include an activation function processing unit located between the at least one PE and the at least one internal memory and configured to sequentially process feature maps corresponding to the first batch of channels and the second batch of channels to sequentially output activation maps corresponding to the first batch of channels and the second batch of channels.
[0048] According to this disclosure, a batch mode configured to process multiple input channels can determine the order in which on-chip memory and / or internal memory store and compute parameters of the artificial neural network. Therefore, the number of main memory read operations can be reduced, and power consumption can be lowered.
[0049] According to this disclosure, even with an increased number of input channels, processing can be performed using a single neural processing unit, which includes on-chip memory and / or internal memory configured to accommodate multiple input channels.
[0050] Furthermore, according to this disclosure, a high-performance and low-power neural processing unit that frequently processes batch channels can be provided in autonomous vehicles, drones, or electronic devices with multiple sensors.
[0051] Furthermore, according to this disclosure, a neural processing unit specifically designed for batch mode can be provided, wherein the size of on-chip memory or internal memory is determined by considering the number of batch channels and computational performance.
[0052] The effects of this disclosure are not limited to those illustrated above; this specification includes a variety of other effects. Attached Figure Description
[0053] Figure 1 This is a schematic conceptual diagram illustrating an apparatus including a neural processing unit, based on examples of this disclosure.
[0054] Figure 2A This is a schematic conceptual diagram illustrating a neural processing unit (NPU) based on examples of this disclosure.
[0055] Figure 2B This is an example diagram showing the energy consumed during NPU operation.
[0056] Figure 2C This is a schematic conceptual diagram illustrating one of the multiple processing elements that can be applied to this disclosure.
[0057] Figure 3 It is shown Figure 2A Example diagram of a modified NPU shown.
[0058] Figure 4 This is a schematic conceptual diagram illustrating an exemplary artificial neural network model.
[0059] Figure 5 This is an exemplary flowchart illustrating how a neural processing unit (NPU) operates, based on examples of this disclosure.
[0060] Figure 6 It is based on Figure 5 The example illustrates an exemplary diagram of allocating artificial neural network parameters in the memory space of an NPU.
[0061] Figure 7 This is an exemplary flowchart illustrating how a neural processing unit operates, according to another example of this disclosure.
[0062] Figure 8 It is based on Figure 7 The example illustrates an exemplary diagram of allocating artificial neural network parameters in the memory space of an NPU.
[0063] Figure 9 This is an exemplary flowchart illustrating how a neural processing unit operates, according to another example of this disclosure.
[0064] Figure 10 It is based on Figure 9The example illustrates an exemplary diagram of allocating artificial neural network parameters in the memory space of an NPU.
[0065] Figure 11 This is an exemplary flowchart illustrating how a neural processing unit operates, based on various examples of this disclosure.
[0066] Figure 12 It is based on Figure 11 The example illustrates an exemplary diagram of allocating artificial neural network parameters in the memory space of an NPU.
[0067] Figure 13 This is an example diagram illustrating an autonomous driving system equipped with a neural processing unit according to an exemplary embodiment of the present disclosure.
[0068] Figure 14 This is a schematic block diagram of an autonomous driving system equipped with a neural processing unit, according to an example of this disclosure.
[0069] Figure 15 This is a flowchart illustrating the identification of target objects for autonomous driving in an autonomous driving system equipped with a neural processing unit, based on examples of this disclosure. Detailed Implementation
[0070] The specific structure or step-by-step description of the concept of this disclosure as presented in this specification or application is merely illustrative for the purpose of explaining the concept of this disclosure.
[0071] Examples of the concepts based on this disclosure may be embodied in various forms and should not be construed as limited to the examples described in this specification or application.
[0072] Because examples of the concept according to this disclosure can have various modifications and can take various forms, specific examples will be shown in the accompanying drawings and described in detail in this specification or application. However, this is not intended to limit the examples of the concept according to this disclosure in terms of a particular form of disclosure, and should be understood to include all modifications, equivalents, and alternatives included within the spirit and scope of this disclosure.
[0073] Terms such as first and / or second may be used to describe various elements, but these elements should not be limited by these terms.
[0074] The terms used above are only used to distinguish one element from another. For example, without departing from the scope of the concept according to this disclosure, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.
[0075] When one element is said to be "connected" or "in contact" with another element, it should be understood that the other element may be directly connected to or in contact with the other element, but other elements may be positioned between them. On the other hand, when it is said that an element is "directly connected" or "directly in contact" with another element, it should be understood that there are no other elements between them.
[0076] Other expressions describing the relationship between elements, such as “between” and “immediately between” or “adjacent” and “closely adjacent”, should be interpreted similarly.
[0077] As used herein, expressions such as “first,” “second,” and “first or second” may modify various elements regardless of order and / or importance. Furthermore, it is used only to distinguish one element from others and does not limit those elements. For example, a first user device and a second user device may represent different user devices regardless of order and / or importance. For example, a first element may be named a second element without departing from the scope of the claims described in this disclosure, and similarly, a second element may be renamed a first element.
[0078] The terminology used in this disclosure is for the purpose of describing specific examples only and is not intended to limit the scope of other examples.
[0079] Unless the context clearly specifies otherwise, singular expressions may include plural expressions. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by one of ordinary skill in the art as described in this document.
[0080] In this disclosure, terms as defined in general dictionaries may be interpreted as having the same or similar meaning as in the context of the relevant art. Furthermore, unless explicitly defined in this document, they should not be interpreted in an ideal or overly formal sense. In some cases, even terms defined in this disclosure should not be interpreted as excluding examples of this disclosure.
[0081] The terminology used herein is for the purpose of describing specific examples only and is not intended to limit this disclosure.
[0082] Unless the context clearly specifies otherwise, singular expressions may include plural expressions. It should be understood that, as used herein, terms such as “comprising” or “having” are intended to indicate the presence of the stated features, quantities, steps, actions, parts, portions or combinations thereof, but do not preclude the possibility of the addition or presence of at least one other feature or quantity, step, action, part, portion or combination thereof.
[0083] Unless otherwise defined, all terms used herein, including technical or scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Terms such as those defined in common dictionaries shall be construed as having meanings consistent with their meanings in the relevant technical context and shall not be construed as having ideal or overly formal meanings unless expressly defined in this specification.
[0084] Each of the features in the various examples of this disclosure can be combined, either partially or entirely, or with each other. Furthermore, those skilled in the art will fully appreciate that multiple interlocks and drives are possible, and each example can be implemented independently of the others or together in a related manner.
[0085] When describing the implementation scheme, descriptions of technical content that is well-known in the technical field to which this disclosure pertains and is not directly related to this disclosure may be omitted. This is to more clearly convey the main points of this disclosure without obscuring them by omitting unnecessary descriptions.
[0086] In the following text, for the purpose of understanding the disclosure presented in this specification, the terminology used in this specification will be briefly summarized.
[0087] NPU: an abbreviation for Neural Processing Unit, which can refer to a processor specifically designed for computing artificial neural network models, separate from the Central Processing Unit (CPU).
[0088] NPU Scheduler (or Scheduler): An NPU scheduler can refer to a unit that controls the overall tasks of the NPU. To allow the NPU scheduler to run within the NPU, the compiler analyzes the data locality of the ANN model and receives the operation sequence information of the compiled ANN model to determine the NPU's processing order. The NPU scheduler can control the NPU within a static task sequence determined based on the data locality of the static ANN model. Alternatively, the NPU scheduler can dynamically analyze the data locality of the ANN model to control the NPU with a dynamic task sequence. Based on the NPU's memory size and the performance of the processing element array, the tiling information of each layer of the ANN model can be stored in the NPU scheduler. The NPU scheduler can control the overall tasks of the NPU through register mapping. The NPU scheduler can be contained within the NPU or located outside the NPU.
[0089] ANN: An abbreviation for Artificial Neural Network. It can refer to a network that mimics human intelligence by connecting nodes in a layered structure, similar to neurons in the human brain, through synaptic connections.
[0090] Information about the structure of artificial neural networks includes information about the number of layers, the number of nodes in each layer, the value of each node, the operation and processing methods, the weight matrix applied to each node, and so on.
[0091] Information about data locality in artificial neural networks: Information that allows neural processing units to predict the operational sequence of an artificial neural network model processed by the neural processing unit based on the sequence of data access requests to individual memory.
[0092] DNN: an abbreviation for Deep Neural Network, which can refer to an increase in the number of hidden layers in an artificial neural network to achieve a higher level of artificial intelligence.
[0093] CNN: an abbreviation for Convolutional Neural Network, a type of neural network whose function is similar to image processing in the human brain's visual cortex. Convolutional neural networks are well-known for their suitability for image processing, and are also known for their superior ability to extract features from input data and recognize patterns of those features.
[0094] Fused-ANN: An abbreviation for fused artificial neural network, which can refer to an artificial neural network designed to process fused sensor data. Sensor fusion is primarily used in the field of autonomous driving technology. Sensor fusion can be a technique where different types of sensors compensate for poor performance of one sensor under certain conditions. The number of sensors fused can vary, such as camera and thermal camera fusion, camera and radar fusion, camera and LiDAR fusion, camera and radar fusion, and LiDAR fusion. A convergent neural network can be an artificial neural network model where data from multiple sensors is fused by adding additional operators (such as skip connections, squeeze and excitation, and cascading).
[0095] The present disclosure will be described in detail below by referring to examples of the accompanying drawings.
[0096] Figure 1 An apparatus including a neural processing unit is illustrated based on examples of this disclosure.
[0097] refer to Figure 1 Device B, including NPU 1000, includes on-chip region A. Main memory 4000 may be included in device B and located outside the on-chip region. Main memory 4000 may be, for example, system memory such as DRAM. Storage cells including ROM may be included outside on-chip region A.
[0098] In on-chip region A, general-purpose processing units such as a central processing unit (CPU) 2000, on-chip memory 3000, and an NPU 1000 are arranged. The CPU 2000 is operatively connected to the NPU 1000, the on-chip memory 3000, and the main memory 4000.
[0099] However, this disclosure is not limited to the configuration described above. For example, the NPU 1000 may be included in the CPU 2000.
[0100] On-chip memory 3000 is memory mounted on a semiconductor die. On-chip memory 3000 may be a cache memory separate from main memory 4000. On-chip memory 3000 may be memory configured to be accessed by other on-chip semiconductors. On-chip memory 3000 may be a cache memory or buffer memory.
[0101] The NPU 1000 may include internal memory 200, and internal memory 200 may include, for example, SRAM. Internal memory 200 may be memory used only for operations within the NPU 1000. Internal memory 200 may be referred to as NPU internal memory. Here, the term "substantial" may indicate that internal memory 200 is configured to store data related to artificial neural networks processed by the NPU 1000.
[0102] For example, internal memory 200 may be a buffer memory and / or a cache memory configured to store weights, kernels (i.e., weights) and / or feature maps required for NPU 1000 operation. However, this disclosure is not limited thereto.
[0103] For example, internal memory 200 can be configured to read and write SRAM, MRAM, register files, etc., faster than main memory 4000. However, this disclosure is not limited thereto.
[0104] Device B, which includes NPU 1000, may include at least one of internal memory 200, on-chip memory 3000, and main memory 4000.
[0105] The term "at least one memory" as described below means at least one of internal memory 200 and on-chip memory 3000.
[0106] Furthermore, the description of on-chip memory 3000 may refer to memory including internal memory 200 of NPU 1000 or memory external to NPU 1000 but in on-chip region A.
[0107] However, internal memory 200 and / or on-chip memory 3000, representing at least one memory, can be distinguished from main memory 4000 based on memory bandwidth rather than location characteristics.
[0108] Generally speaking, a main memory 4000 refers to a superior memory for storing large amounts of data, with relatively low memory bandwidth and relatively high power consumption.
[0109] Generally speaking, internal memory 200 and on-chip memory 3000 refer to memory with high memory bandwidth and low power consumption, but low efficiency for storing large amounts of data.
[0110] Each component of device B, including NPU 1000, can communicate via bus 5000. Device B may have at least one bus 5000. Bus 5000 may be referred to as a communication bus and / or system bus, or the like.
[0111] The NPU 1000's internal memory 200 and on-chip memory 3000 may also include separate dedicated buses to ensure bandwidth exceeding the specific bandwidth used to process the weights and feature maps of artificial neural network models.
[0112] A separate dedicated bus can also be included between the on-chip memory 3000 and the main memory 4000 to ensure bandwidth exceeding a specific bandwidth. This specific bandwidth can be determined based on the processing performance of the NPU 1000's array of processing elements.
[0113] A separate dedicated bus may be further included between the NPU 1000's internal memory 200 and main memory 4000 to ensure bandwidth exceeds a specific limit. This specific bandwidth can be determined based on the processing performance of the NPU 1000's array of processing elements.
[0114] Device B, which includes NPU 1000, may further include a Direct Memory Access (DMA) module and may be configured to directly control internal memory 200, on-chip memory 3000, and / or main memory 4000. The DMA module may be configured to directly control data transfer between NPU 1000 and on-chip memory 3000 via direct control bus 5000. The DMA module may also be configured to directly control data transfer 4000 between on-chip memory 3000 and main memory via direct control bus 5000. Finally, the DMA module may be configured to directly control data transfer between internal memory 200 and main memory 4000 via direct control bus 5000.
[0115] The Neural Processing Unit (NPU) 1000 is a processor specifically designed to perform operations on artificial neural networks. The NPU 1000 can be referred to as an AI accelerator.
[0116] An artificial neural network (ANN) is a type of artificial neural network that, upon receiving multiple inputs or stimuli, multiplies and adds weights, and then transforms and transmits values obtained by adding additional biases through an activation function. An artificial neural network trained in this way can be used to output inference results based on input data.
[0117] The NPU 1000 can be a semiconductor implemented as an electrical / electronic circuit. The electrical / electronic circuit may include multiple electronic devices, such as transistors or capacitors. The NPU 1000 may include a processing element (PE) array (i.e., multiple processing elements), an NPU internal memory 200, an NPU scheduler, and an NPU interface. Each of the processing element array, the NPU internal memory 200, the NPU scheduler, and the NPU interface can be a semiconductor circuit connected to multiple transistors.
[0118] Therefore, some transistors may be difficult or impossible to identify and distinguish with the human eye, and may only be identifiable through operation. For example, any circuit can operate as an array of processing elements, or as an NPU scheduler.
[0119] NPU 1000 may include a processing element array, NPU internal memory 200 configured to store at least a portion of an artificial neural network model that can be inferred from the processing element array, and an NPU scheduler configured to control the processing element array and NPU internal memory 200 based on data locality information of the artificial neural network model or information about the structure of the artificial neural network model.
[0120] Artificial neural network models can include information about the data locality or structure of the artificial neural network model. An artificial neural network model can refer to an AI recognition model trained to perform a specific reasoning function.
[0121] An array of processing elements can perform operations on an artificial neural network. For example, when input data is input, the array of processing elements can train the artificial neural network. After training, when input data is input again, the array of processing units can perform operations to derive inference results using the trained artificial neural network.
[0122] For example, the NPU 1000 can load data of an artificial neural network model stored in the main memory 4000 into the NPU's internal memory 200 via the NPU interface. The NPU interface can communicate with the main memory 4000 via bus 5000.
[0123] The NPU scheduler is configured to control the operation of the processing unit array for inference operations of the NPU 1000 and the read / write order of the NPU internal memory 200. The NPU scheduler is also configured to adjust the size of at least a portion of the batch channels.
[0124] The NPU scheduler analyzes the structure of the artificial neural network model or the structure of the provided artificial neural network model. Next, the NPU scheduler determines the order of operations for each layer. In other words, once the structure of the artificial neural network model is determined, the sequence of operations for each layer can be determined. The order of operations or data flow based on the structure of the artificial neural network model can be defined as the data locality of the artificial neural network model at the algorithm level.
[0125] The NPU scheduler determines the operation sequence of each layer by reflecting the structure of the artificial neural network model and the number of deployed channels. In other words, once the structure and number of channels of the artificial neural network model are determined, the operation sequence of each layer can be determined. The operation sequence or data flow based on the number of batch channels and the structure of the artificial neural network model can be defined as data locality at the algorithm level or data locality in batch mode artificial neural network models. In the following text, the data locality of batch mode artificial neural network models will be referred to as data locality of the artificial neural network model.
[0126] The data locality of an artificial neural network model can be determined by considering the model's structure, the number of batch channels, and the NPU structure.
[0127] When a compiler compiles an artificial neural network model so that it can be executed in an NPU 1000, the neural network data locality at the neural processing unit-memory level can be reconstructed. For example, the compiler can be executed by a CPU 2000.
[0128] In other words, the weight values and batch channel size loaded into the internal memory can be determined based on the compiler, the algorithm applied to the artificial neural network model, the operating characteristics of the NPU 1000, the size of the weight values, and the size of the feature map or batch channel.
[0129] For example, even with the same artificial neural network model, the computation method for the artificial neural network model to be processed can be configured according to the methods and characteristics of the NPU 1000 in computing the corresponding artificial neural network model. These methods and characteristics include, for example, the tiling method of the feature maps, the stabilization method of the processing elements, the number of processing elements in the NPU 1000, the size of the feature maps and the magnitude of the weights in the NPU 1000, the internal memory capacity, the memory hierarchy of the NPU 1000, and the algorithmic characteristics of the compiler that determine the order of operations performed by the NPU 1000 on the artificial neural network model. This is because even when processing the same artificial neural network model using the above factors, the NPU 1000 may determine the required data sequence for each moment differently, on a clock cycle basis.
[0130] Figure 2A The neural processing unit is illustrated by examples from this disclosure.
[0131] refer to Figure 2A The neural processing unit (NPU) 1000 may include a scheduler 300, a processing element array 100, and an internal memory 200.
[0132] The NPU scheduler 300 can be configured to control the processing element array 100 and the NPU internal memory 200 by taking into account the magnitude of the weight values of the artificial neural network model, the size of the feature map, and the computation sequence of the weight values and feature maps.
[0133] The NPU scheduler 300 can receive the magnitudes of the weight values, the sizes of the feature maps, and the computation sequences of the weight values and feature maps calculated in the array of elements to be reprocessed 100. The artificial neural network data of the artificial neural network model may include node data or feature maps for each layer, as well as weight data for each connection network connecting the nodes of each layer. At least some of the data or parameters of the artificial neural network may be stored in memory provided within the NPU scheduler 300 or the NPU internal memory 200.
[0134] In the parameters of an artificial neural network, the feature map may consist of batch channels. Here, multiple batch channels can be images captured by multiple image sensors within substantially the same time period (e.g., between 10 ms and 100 ms).
[0135] The NPU scheduler 300 can control the processing element array 100 and the internal memory 200 by performing convolution operations, such as those of an artificial neural network. First, the NPU scheduler 300 can load groups of weight values into the weight storage unit 210 of the internal memory 200, and can load a portion of multiple batch channels corresponding to the groups of weight values into the batch channel storage unit 220 of the internal memory 200. After calculating the groups of weight values and some of the batch channels, the NPU scheduler 300 can load the subsequent batch channels while maintaining the groups of weight values in the internal memory 200. Although the internal memory 200 is shown as including the weight storage unit 210 and the batch channel storage unit 220 respectively, this is merely an example. In another example, the internal memory 200 can be logically partitioned or variably allocated via memory addresses, or the internal memory 200 may not be partitioned.
[0136] In various examples, the grouped weight values can be a portion of the total weight value. In this case, a portion of multiple batch channels can be calculated first, such as a portion of the first batch channels and a portion of the second batch channels, and then the next portion of the first batch channels and the next portion of the second batch channels can be calculated. Alternatively, a portion of multiple batch channels can be calculated first, such as a portion of the first batch channels and a portion of the second batch channels, and then a portion of the third batch channels and a portion of the fourth batch channels can be calculated.
[0137] In various examples, although a portion of the second batch of channels is used to calculate the grouped weight values, a portion of the third batch of channels to be calculated next can be loaded into the portion of the first batch of channels that has already been calculated. Processing speed may be faster when the next calculation parameter is loaded into memory simultaneously with the calculation.
[0138] In the example above, the parameters of the artificial neural network are described as being stored in the internal memory 200 of the NPU, but this disclosure is not limited thereto. For example, the parameters may be stored in on-chip memory or main memory.
[0139] The configuration used to improve the processing speed in the NPU 1000 of this disclosure is to minimize the read frequency of the DRAM memory or main memory in the following manner (which will refer to...). Figure 2B (Description): The weight values are stored in memory (i.e., any type of memory), and then the weight values are maintained as much as possible without additional memory access. Since the number of main memory reads of weight values or feature maps is directly proportional to energy consumption and inversely proportional to processing speed, reducing the number of main memory reads of these values can increase processing speed while reducing energy consumption.
[0140] CPU scheduling typically achieves optimal efficiency by considering fairness, efficiency, stability, and response time. In other words, it schedules the most processing jobs to be executed within the same timeframe, taking into account priority and operation time.
[0141] Taking into account data such as the priority order of each process and the processing time of operations, traditional CPUs use algorithms for scheduling tasks.
[0142] Alternatively, the NPU scheduler 300 can determine the processing order based on the calculation method of the parameters of the artificial neural network model, especially based on the calculation characteristics between batch channels and weights.
[0143] Furthermore, the NPU scheduler 300 can determine the processing order based on the requirement that a weight reassembly must be applied to all batch channels until a convolution operation is completed, so that the weight reassembly is no longer accessed from main memory. In other words, a convolution operation in batch mode may mean convolving multiple consecutive batch channels with groups of weights respectively.
[0144] However, this disclosure is not limited to the factors described above for the NPU 1000, and may further be based on data locality information or information about the structure. For example, information about the data locality information or structure of the NPU 1000 may include the memory size of the NPU internal memory 200, the hierarchical structure of the NPU internal memory 200, the amount of data in processing elements PE1 to PE12, and data about at least one of the operator architectures of processing elements PE1 to PE12. The memory size of the NPU internal memory 200 may include information about the memory capacity. The hierarchical structure of the NPU internal memory 200 may include information about the connection relationships between specific levels of each hierarchical structure. The operator structure of processing elements PE1 to PE12 may include information about the elements within the processing elements.
[0145] In other words, the NPU scheduler 300 can determine the processing sequence by utilizing at least one of the following: the memory size of the NPU internal memory 200, the hierarchical structure of the NPU internal memory 200, the number of processing elements PE1 to PE12, and the operator structure of the processing elements PE1 to PE12.
[0146] However, this disclosure is not limited to information regarding the locality or structure of data provided to the NPU 1000.
[0147] According to examples of this disclosure, the NPU scheduler 300 can control at least one processing element and the NPU internal memory 200 based on the calculation method of the parameters of the artificial neural network model, particularly based on the characteristics of the calculation between batch channels and weights.
[0148] On the other hand, the processing element array 100 can be configured to include a plurality of processing elements 110 (i.e., PE1, PE2, ...), which are configured to compute node data and weight data of the connection network of the artificial neural network. Each processing element can be configured to include multiplication and accumulation (MAC) operators and / or arithmetic logic unit (ALU) operators. However, the examples according to this disclosure are not limited thereto.
[0149] exist Figure 2A The diagram illustrates multiple processing elements 110 by way of example. However, by modifying the MAC within a single processing unit, operators can also be configured as a tree of multiple multipliers and adders in parallel. In this case, the processing element array 100 can be referred to as at least one processing element comprising multiple operators.
[0150] also, Figure 2A The multiple processing elements 110 shown may be examples provided for ease of description only, and the number of processing elements is not limited. The size or number of processing element arrays can be determined by the number of processing elements 110. The size of the processing element array can be implemented as an N×M matrix, where N and M are integers greater than zero. Therefore, the processing element array 100 can include N×M processing elements. That is, there can be more than one processing element.
[0151] Furthermore, the processing element array 100 can be composed of multiple sub-modules. Therefore, the processing element array 100 can include processing elements consisting of N×M×L sub-modules. More specifically, L is the number of sub-modules in the processing element array, and can be referred to as cores, engines, or threads.
[0152] The size of the processing element array 100 can be designed taking into account the characteristics of the artificial neural network model operating in the NPU 1000. In other words, the number of processing elements can be determined by considering the data size of the artificial neural network model to be operated, the required operating speed, and the required power consumption. The data size of the artificial neural network model can be determined based on the number of layers in the artificial neural network model and the weight data size of each layer.
[0153] Therefore, the size of the processing element array 100 according to the example of this disclosure is not limited thereto. As the number of processing elements 110 in the processing element array 100 increases, the parallel computing power of the running artificial neural network model increases, but the manufacturing cost and physical size of the NPU 1000 may increase.
[0154] For example, the artificial neural network model operating in the NPU 1000 could be an artificial neural network trained to detect thirty specific keywords, i.e., an AI keyword recognition model. In this case, considering the computational characteristics, the size of the processing element array 100 can be designed to be N×M. In other words, the NPU 1000 can be configured to include twelve processing elements. However, it is not limited to this, and the number of the plurality of processing elements 110 can be selected, for example, in the range of 8 to 16,384. That is to say, the examples of this disclosure are not limited in terms of the number of processing elements.
[0155] The processing element array 100 can be configured to perform functions such as addition, multiplication, and accumulation required for artificial neural network operations. In other words, the processing element array 100 can be configured to perform multiplication and accumulation (MAC) operations.
[0156] Internal memory 200 may be volatile memory. Volatile memory may be a memory that stores data only when powered on and deletes (dumps) the stored data when power is off. Volatile memory may include static random access memory (SRAM), dynamic random access memory (DRAM), etc. Internal memory 200 may preferably be SRAM, but is not limited thereto.
[0157] The following text will primarily describe Convolutional Neural Networks (CNNs), which are a type of deep neural network (DNN) in artificial neural networks.
[0158] A convolutional neural network (CNN) can be a combination of one or more convolutional layers, pooling layers, and fully connected layers. CNNs have a structure suitable for training and inference on two-dimensional data and can be trained using the backpropagation algorithm.
[0159] In the examples disclosed herein, for each channel of the convolutional neural network, there exists a kernel for extracting features from the input image of that channel. The kernel can consist of a two-dimensional matrix. The kernel performs convolution operations while traversing the input data. The size of the kernel (N × M) can be arbitrarily determined, as can the stride of the kernel traversing the input data. The degree of matching between the kernel and all input data for each kernel can be a feature map or an activation map. In the following, the kernel may include a set of weight values or multiple sets of weight values.
[0160] The processing element array 100, i.e., multiple processing elements, can be configured to process convolution operations of artificial neural networks, and activation function operations can be configured to be processed in a separate activation function processing module. In this case, the processing element array 100 can be operated solely for convolution operations. In particular, in this case, the processing element array 100 is configured to process only integer data, thereby maximizing arithmetic efficiency during large-scale convolution operations.
[0161] Since convolution is an operation consisting of a combination of input data and a kernel, activation functions such as ReLU can be applied to add non-linearity. When an activation function is applied to a feature map that is the result of a convolution operation, it can be called an activation map.
[0162] Convolutional neural networks (CNNs) can include AlexNet, SqueezeNet, VGG16, ResNet152, and MobileNet, among others. The number of multiplications required for one inference operation in each CNN model is 727 MFLOPs, 837 MFLOPs, 16 MFLOPs, 11 MFLOPs, 11 MFLOPs, and 579 MFLOPs, respectively, and the data sizes for all weights, including the kernel, are 233 MB, 5 MB, 528 MB, 230 MB, and 16 MB, respectively. Therefore, it is evident that they require a considerable amount of hardware resources and power consumption.
[0163] The activation function processing unit can be further positioned between the processing element array 100 and the internal memory 200 to apply activation functions. The activation function processing unit can be configured to include multiple sub-modules. For example, the activation function processing unit may include at least one of the following units: ReLU unit, Leaky-ReLU unit, ReLU6 unit, Swish unit, Sigmoid unit, Average-Pooling unit, Skip-Connection unit, Squeeze and Excitation unit, Bias unit, Quantization unit, Inverse Quantization unit, Hyperbolic-Tangent unit, Maxout unit, ELU unit, Batch-Normalization unit, and Piecewise-Function-Approximation unit. The activation function processing unit can be used to arrange the various sub-modules into a pipeline structure.
[0164] The activation function processing unit can selectively activate or deactivate each submodule.
[0165] The NPU scheduler 300 can be configured to control the activation function processing unit.
[0166] The NPU scheduler 300 can selectively activate or deactivate each sub-module of the activation function processing unit based on the data locality of the artificial neural network model.
[0167] The activation function processing unit can be configured to sequentially process the feature maps of each batch of channels output from the processing element array 100 to output an activation map for each batch of channels. This will refer to... Figure 2B Detailed description.
[0168] Figure 2B The diagram shows the energy consumed during the operation of the NPU 1000. (Refer to...) Figure 2C The example presented is described by the configuration of the first processing element PE1 110 (e.g., multiplier 641 and adder 642), which will be described later.
[0169] refer to Figure 2B Energy consumption can be divided into memory access operations, addition operations, and multiplication operations.
[0170] "8b Add" refers to the 8-bit integer addition operation of the adder 642. The 8-bit integer addition operation consumes 0.03 pj of energy.
[0171] "16b Add" refers to the 16-bit integer addition operation of the adder 642. The 16-bit integer addition operation consumes 0.05 pj of energy.
[0172] "32b Add" refers to the 32-bit integer addition operation of the adder 642. The 32-bit integer addition operation consumes 0.1 pj of energy.
[0173] "16b FP Add" refers to the 16-bit floating-point addition operation of the adder 642. The 16-bit floating-point addition operation consumes 0.4pJ of energy.
[0174] "32b FP Add" refers to the 32-bit floating-point addition operation of the adder 642. The 32-bit floating-point addition operation consumes 0.9 pj of energy.
[0175] "8b Mult" refers to the 8-bit integer multiplication operation of the multiplier 641. The 8-bit integer multiplication operation consumes 0.2 pj of energy.
[0176] "32b Mult" refers to the 32-bit integer multiplication operation of the multiplier 641. The 32-bit integer multiplication operation consumes 3.1 pj of energy.
[0177] "16b FP Mult" refers to the 16-bit floating-point multiplication operation of the multiplier 641. The 16-bit floating-point multiplication operation consumes 1.1 pj of energy.
[0178] "32b FP Mult" refers to the 32-bit floating-point multiplication operation of the multiplier 641. The 32-bit floating-point multiplication operation consumes 3.7 pj of energy.
[0179] "32b SRAM Read" refers to a 32-bit data read access when the NPU memory system's internal memory is Static Random Access Memory (SRAM). Reading 32 bits of data from the NPU memory system consumes 5 pJ of energy.
[0180] "32-bit DRAM Read" refers to a 32-bit data read access when the vehicle control device's memory cell is DRAM. Reading 32 bits of data from the memory cell into the NPU memory system consumes 640 pJ of energy. The unit of energy is picojoule (pJ).
[0181] Traditional neural processing units store these kernels in memory for each corresponding channel and process input data by fetching it from memory for each convolutional process. For example, in a 32-bit read operation during a convolutional process, the SRAM, which serves as the internal memory of the NPU 1000, consumes 5 pj of power. Figure 2B As shown, the DRAM, serving as main memory, consumes 640 pj of power. This memory consumes 0.03 pj of power for 8-bit addition, 0.05 pj for 16-bit addition, 0.1 pj for 32-bit addition, and 0.2 pj for 8-bit multiplication. Therefore, the power consumption of a traditional neural processing unit is far higher than other operations, leading to a decrease in overall performance. In other words, reading the kernel from the NPU 1000's main memory consumes 128 times the power consumed when reading the kernel from internal memory.
[0182] In other words, the main memory 4000 operates slower than the internal memory 200, but consumes more power per operation. Therefore, minimizing read operations from the main memory 4000 will affect the power consumption reduction of the NPU 1000. In particular, power efficiency may be particularly degraded if multiple channels are processed individually.
[0183] To overcome this inefficiency, this disclosure proposes a neural processing unit with improved computational performance by minimizing data movement between the main memory 4000 and on-chip region A loaded with grouped weight values or multiple batch channels, thereby reducing overall hardware resources and power consumption due to data movement.
[0184] Using batch channels as input to perform object detection using the object detection model in the neural processing unit aims to minimize the number of times the object detection model's weight values are accessed from DRAM. As the batch size increases, the number of accesses to the weight values stored in DRAM increases. In other words, the number of accesses to the weight values stored in DRAM can increase proportionally to the number of batch channels.
[0185] Therefore, this disclosure stores the necessary parts of the object detection model used for object detection in the NPU's internal memory, which consists of SRAM, thereby reducing the power consumption per NPU operation and further improving the NPU's performance.
[0186] Therefore, autonomous vehicles equipped with the NPU disclosed herein can minimize the amount of time required to identify target objects, obstacles, traffic light signals, and pedestrians (which must be identified in real time to ensure safe autonomous driving) in front of, behind, to the left and right of the vehicle, as well as the computational resources consumed by object detection.
[0187] In the following text, Figure 2A The first processing element PE1 of the processing element array will be described as an example to enable the neural processing unit to reduce overall hardware resources and power consumption caused by data movement and improve computing performance.
[0188] NPU 1000 may include a processing element array 100, an NPU internal memory 200 configured to store artificial neural network models that can be inferred from the processing element array 100, and an NPU scheduler 300 configured to control the processing element array 100 and the NPU internal memory 200. The processing element array 100 may be configured to perform MAC operations and quantize and output the MAC operation results. However, the examples in this disclosure are not limited thereto.
[0189] The NPU internal memory 200 can store all or part of the artificial neural network model, depending on the memory size and data size of the artificial neural network model.
[0190] Figure 2C This describes one of the multiple processing elements that can be applied to this disclosure.
[0191] refer to Figure 2C The first processing element PE1 may include a multiplier 641, an adder 642, and an accumulator 643. However, the examples according to this disclosure are not limited thereto, and the processing element array 100 may be modified to take into account the computational characteristics of artificial neural networks.
[0192] Multiplier 641 multiplies the received N-bit data and M-bit data. The output of multiplier 641 is (N+M)-bit data, where N and M are integers greater than zero. The first input unit for receiving N-bit data can be configured to receive values with properties such as variability, and the second input unit for receiving M-bit data can be configured to receive values with properties such as constancy.
[0193] Here, a value or variable with properties similar to a variable means that whenever the input data is updated, the memory address storing the corresponding value is updated. For example, the node data of each layer can be the MAC operation value of the weight data of the artificial neural network model in which it is applied. In the case of object detection using a corresponding artificial neural network model to infer moving image data, the node data of each layer will change because the input image changes every frame.
[0194] Here, values with constant-like characteristics, or constants referring to the memory addresses storing the corresponding values, are preserved regardless of how the input data is updated. For example, the weights of the network are the sole criterion for inference in the artificial neural network model; even if the artificial neural network model infers object detection from moving image data, the weights (i.e., the kernels) of the network can remain unchanged.
[0195] In other words, multiplier 641 can be configured to receive a variable and a constant. More specifically, the variable value input to the first input unit can be node data of a layer in the artificial neural network, such as input data of the input layer, accumulated values of the hidden layers, and accumulated values of the output layer. The constant value input to the second input unit can be weight data of the connection network of the artificial neural network.
[0196] Thus, when the NPU scheduler 300 distinguishes between the characteristics of variable values and constant values, the NPU scheduler 300 can increase the memory reuse rate of the NPU internal memory 200. However, the input data of the multiplier 641 is not limited to constant values and variable values. That is, according to the examples of this disclosure, since the input data of the processing element can be manipulated by understanding the characteristics of constant values and variable values, the operating efficiency of the NPU 1000 can be improved. However, the operation of the NPU 1000 is not limited to the characteristics of constant values and variable values of the input data.
[0197] Based on this, the NPU scheduler 300 can be configured to improve memory reuse by taking into account the characteristics of constant values.
[0198] The variable values are the calculated values of each layer, and the NPU scheduler 300 can control the NPU internal memory 200 to identify reusable variable values and reuse the memory based on the artificial neural network model structure data or the locality information of the artificial neural network data.
[0199] The constant values are the weight data for each network. The NPU scheduler 130 can control the NPU internal memory 200 to identify the constant values of the reusable connected networks and reuse the memory based on the artificial neural network model structure data or the locality information of the artificial neural network data.
[0200] In other words, the NPU scheduler 300 identifies reusable variable values and / or reusable constant values based on the structural data or locality information of the artificial neural network model, and the NPU scheduler 300 can be configured to control the NPU internal memory 200 to reuse data stored in the memory.
[0201] When zero is input to either the first or second input unit of the multiplier 641, the first processing element PE1 knows that the result of the operation is zero even if the operation is not performed. Therefore, the operation of the multiplier 641 can be restricted so that the operation is not performed.
[0202] For example, when zero is input to one of the first and second input units of the multiplier 641, the multiplier 641 can be configured to operate in a zero-jump mode.
[0203] The number of bits of data input to the first and second input units can be determined based on the quantization of the node data and weight data of each layer of the artificial neural network model. For example, the node data of the first layer can be quantized to five bits, and the weight data of the first layer can be quantized to seven bits. In this case, the first input unit can be configured to receive five bits of data, and the second input unit can be configured to receive seven bits of data.
[0204] When quantized data stored in the NPU's internal memory 200 is input to the input of the first processing element PE1, the NPU 1000 can control the number of quantized bits to be converted in real time. That is, the number of quantized bits for each layer can be different. When the number of bits of the input data is converted, the first processing element PE1 can be configured to receive bit information from the NPU 1000 in real time and convert the number of bits in real time to generate the input data.
[0205] Adder 642 adds the calculated value from multiplier 641 and accumulator 643. When loop L is 0, since there is no data to accumulate, the value of adder 642 will be the same as the value of multiplier 641. When loop L is 1, the value obtained by adding the value of multiplier 641 and accumulator 643 can be the value of adder 642.
[0206] Accumulator 643 temporarily stores the data output from the output unit of adder 642, thereby accumulating the operands of adder 642 and multiplier 641 for L cycles. Specifically, the calculated value of adder 642, output from the output unit of adder 642, is input to the input unit of accumulator 643. The operand input to the accumulator is temporarily stored in accumulator 643 and output from the output unit of accumulator 643. The output operand is cyclically input to the input unit of adder 642. At this time, the newly output operand from the output unit of multiplier 641 is input to the input unit of adder 642. That is, the operands of accumulator 643 and the new operand of multiplier 641 are input to the input unit of adder 642, these values are added by adder 642 and output through the output unit of adder 642. The data output from the output unit of adder 642, i.e. the new operation value of adder 642, is input to the input unit of accumulator 643. Subsequent operations are basically the same as the above operations, with the same number of loops.
[0207] In this way, accumulator 643 temporarily stores the data output from the output unit of adder 642, so as to accumulate the operation values of multiplier 641 and adder 642 by the number of loops. Therefore, the data input to the input unit of accumulator 643 and the data output from the output unit can have the same bit width as the data output from the output unit of adder 642, that is, (N+M+log2(L)) bits, where L is an integer greater than 0.
[0208] When the accumulation is complete, the accumulator 643 may receive an initialization reset signal to initialize the data stored in the accumulator 643 to zero. However, the examples according to this disclosure are not limited thereto.
[0209] The output data of the (N+M+log2(L)) bits of the accumulator 643 can be the node data of the next layer or the input data of the convolution.
[0210] In various examples, the first processing element PE1 may also include a bit quantization unit. The bit quantization unit can reduce the number of bits of data output from the accumulator 643. The bit quantization unit can be controlled by the NPU scheduler 300. The number of bits of quantized data can be output as X bits, where X is a positive integer. According to the above configuration, the processing element array 100 is configured to perform a MAC operation, and the processing element array 100 has the effect of quantizing and outputting the result of the MAC operation. In particular, as the number of L loops increases, this quantization has the effect of further reducing power consumption. Furthermore, reducing power consumption also has the effect of reducing heat generation in the edge device. In particular, reducing heat generation has the effect of reducing the possibility of failure due to high temperatures in the NPU 1000.
[0211] The X-bit output data of the bit quantization unit can be node data of the next layer or input data of a convolution. If the artificial neural network model has been quantized, the bit quantization unit can be configured to receive quantization information from the artificial neural network model. However, it is not limited to this; the NPU scheduler 300 can be configured to extract quantization information by analyzing the artificial neural network model. Therefore, the output data X bits can be converted to the number of quantized bits to correspond to the quantized data size and output. The output data X bits of the bit quantization unit can be stored in the NPU internal memory 200 as the number of quantized bits. The bit quantization unit can be included in a processing element or activation function processing unit.
[0212] According to the example of this disclosure, the processing element array 100 of the NPU 1000 can reduce the number of bits of (N+M+log2(L)) bits of data output from the accumulator 643 to X bits via a bit quantization unit. The NPU scheduler 300 can control the bit quantization unit to reduce the number of bits of the output data from the least significant bit (LSB) to the most significant bit (MSB) by a predetermined number of bits. When the number of bits of output data is reduced, the power consumption, computational cost, and memory usage of the NPU 1000 can be reduced. However, when the number of bits is reduced below a certain length, the inference accuracy of the artificial neural network model may rapidly decrease. Accordingly, the reduction in the number of bits of output data (i.e., the degree of quantization) can be determined by comparing the reduction in power consumption, computational cost, and memory usage with the reduction in the inference accuracy of the artificial neural network model. The degree of quantization can also be determined by determining the target inference accuracy of the artificial neural network model and testing it while gradually reducing the number of bits. The degree of quantization can be determined for each operation value of each layer.
[0213] According to the first processing element PE1, by adjusting the number of bits of the N-bit and M-bit data of the multiplier 641 and reducing the number of bits of the operation value X through the bit quantization unit, the processing unit array has the effect of improving the MAC operation speed while reducing power consumption, and has the effect of performing convolution operations of artificial neural networks more efficiently.
[0214] However, the bit quantization unit of this disclosure can be configured to be included in the activation function processing unit instead of the processing element.
[0215] Therefore, the NPU internal memory 200 of the NPU 1000 can be a memory system configured with consideration of the MAC operation characteristics and power consumption characteristics of the processing element array 100.
[0216] For example, considering the MAC operation characteristics and power consumption characteristics of the processing element array 100, the NPU 1000 can be configured to reduce the bit width of the operation values of the processing element array 100.
[0217] The NPU internal memory 200 of the NPU 1000 can be configured to minimize the power consumption of the NPU 1000.
[0218] The NPU internal memory 200 of the NPU 1000 can be a memory system configured to control the memory with low power, taking into account the size of the parameters and the operating step size of the artificial neural network model to be operated.
[0219] The NPU internal memory 200 of the NPU 1000 can be a low-power memory system configured to reuse specific memory addresses where weight data is stored, taking into account the data size and operation step size of the artificial neural network model.
[0220] The NPU 1000 can provide a variety of activation functions to impart non-linearity. For example, activation functions can include the sigmoid function, hyperbolic tangent (tanh) function, ReLU function, Leaky-ReLU function, Maxout function, or ELU function, which derive a non-linear output value with respect to the input value; however, it is not limited to these. Such activation functions can be selectively applied after a MAC operation. The values after applying the activation function can be called the activation map. The values before applying the activation function can be called the feature map.
[0221] Figure 3 It shows Figure 2A The example of NPU modification shown.
[0222] Figure 3 The NPU 1000 shown is... Figure 2A The processing unit 1000 shown in the example is substantially the same as the processing element 110 except for the plurality of processing elements 110. Therefore, for the sake of convenience, redundant descriptions are omitted.
[0223] In addition to multiple processing elements 110', Figure 3 The array of processing elements 100 shown exemplarily may also include corresponding register files RF1, RF2, ... corresponding to each of the processing elements PE1, PE2, ...
[0224] Figure 3 The multiple processing elements PE1, PE2, ... and multiple register files RF1, RF2, ... shown are merely examples for ease of description, and the number of multiple processing elements and multiple register files is not limited thereto.
[0225] The size or number of processing element array 100 can be determined by the number of multiple processing elements PE1, PE2, ... and multiple register files RF1, RF2, ... The size of processing element array 100 and multiple register files RF1, RF2, ... can be implemented in the form of an N × M matrix, where N and M are integers greater than zero.
[0226] The array size of the processing element array 100 can be designed taking into account the characteristics of the artificial neural network model operating in the NPU 1000. In other words, the memory size of the register file can be determined by considering the data size of the artificial neural network model to be operated, the required operating speed, the required power consumption, etc.
[0227] The register files RF1, RF2, ... of the processing element array 100 are static memory units directly connected to the processing elements PE1 to PE12. The register files RF1, RF2, ... can consist of, for example, flip-flops and / or latches. The register files RF1, RF2, ... can be configured to store the MAC operation values of the corresponding processing elements PE1, PE2, ... . The register files RF1, RF2, ... can be configured to provide or receive NPU system memory 200 and weight data and / or node data. The register files RF1, RF2, ... may also be configured to perform accumulator functions.
[0228] An activation function processing unit can also be provided and set between the processing element array 100 and the internal memory 200 to apply the activation function.
[0229] Figure 4 An exemplary artificial neural network model is shown.
[0230] refer to Figure 4 A convolutional neural network may include at least one convolutional layer, at least one pooling layer, and at least one fully connected layer.
[0231] For example, convolution can be defined by two main parameters: the size of the input data (typically a 1×1, 3×3, or 5×5 matrix) and the depth of the output feature map (the number of kernels). These key parameters can be computed through convolution. These convolutions can start at a depth of 32, continue to a depth of 64, and end at a depth of 128 or 256. The convolution operation can be described as sliding a kernel of size 3×3 or 5×5 over the input image matrix, multiplying each element of the kernel by each element of the input image matrix, and then summing them all. Here, the input image matrix can be referred to as a 3D block, and the kernel can be referred to as a trainable weight matrix.
[0232] In other words, convolution refers to the operation of transforming a 3D block into a 1D vector through a tensor product with a trained weight matrix, and then reorganizing this vector space into a 3D output feature map. All spatial locations in the output feature map can correspond to the same locations in the input feature map.
[0233] Convolutional layers can perform convolutions between the input data and the kernel (i.e., the weight matrix) that has been trained through multiple gradient update iterations during the learning process. If (m, n) is the kernel size and W is set as the weight values, the convolutional layer can convolve the input data and the weight matrix by computing the dot product.
[0234] The stride by which the kernel slides across the input data is called the stride, and the kernel region (m × n) can be called the receptive field. Applying the same convolutional kernel at different positions in the input reduces the number of kernels to train. This also enables position-invariant learning, where the convolutional filter (i.e., the kernel) can learn an important pattern if it exists in the input, regardless of the position in the sequence.
[0235] Activation functions can be applied to the output feature maps generated as described above to ultimately output activation maps. Additionally, the feature maps computed in the current layer can be passed to the next layer via convolution. Pooling layers can perform pooling operations to reduce the size of the feature maps by downsampling the output data (i.e., the activation maps). For example, pooling operations can include, but are not limited to, max pooling and / or average pooling. Max pooling uses a kernel and outputs the maximum value in the region where the feature map overlaps with the kernel by sliding the feature map and the kernel. Average pooling outputs the average value in the region where the feature map overlaps with the kernel by sliding the feature map and the kernel. Thus, because the size of the feature map is reduced through pooling operations, the number of weights in the feature map is also reduced.
[0236] Fully connected layers can classify the data output from pooling layers into multiple categories (i.e., inferred values) and output the classified categories and their scores. The data output from pooling layers forms a 3D feature map, which can be converted into a 1D vector and used as input to fully connected layers.
[0237] A convolutional neural network can be tuned or trained so that input data leads to a specific inference output. In other words, a convolutional neural network can be tuned using backpropagation based on a comparison between the inference output and ground reality until the inference output gradually matches or approaches ground reality.
[0238] Convolutional neural networks can be trained by adjusting the weights between neurons based on the difference between ground-based real-world data and the actual output.
[0239] In the following text, reference will be made to Figure 5Sections 12 describe various examples of methods for neural processing units to perform operations on an ANN according to the present disclosure, and the memory space for allocating artificial neural network parameters according to steps.
[0240] In the following text, weights and batch channels A through D are referred to, and their sizes and divisions are merely exemplary, and for ease of description, they are shown as having the same size or relative size to each other. Furthermore, weights and batch channels each have addresses allocated to memory space, and as... Figures 5 to 12 The data flow shown can refer to writing alternative data to memory addresses. Data at the same location within the same category is intended to be maintained without overwriting other data. Furthermore, each computational step should be understood as a computation time of at least one clock cycle, but not limited to this, and can be executed within variable clock cycles; this does not mean that every computational step is executed within the same clock cycle. It should also be noted that each operation step is a state with a very short memory time, rather than a statically latched state.
[0241] Furthermore, the memory space is exemplarily shown as having the same partitions or size, but is not limited thereto, and the memory space can have various partitions (e.g., segmented partitions) and can have different sizes. Additionally, in this example, S can refer to an operation step (i.e., a step).
[0242] In this example, the ANN is configured to perform at least one operation, including object detection, classification, or segmentation from multiple batch channels. Before operating on the ANN, the multiple batch channels can be preprocessed, and each of the multiple batch channels corresponds to each of the multiple images.
[0243] Figure 5 The following illustrates how a neural processing unit operates according to an example of this disclosure. Figure 6 It shows according to Figure 5 The step involves allocating memory space for the artificial neural network parameters within the neural processing unit.
[0244] In the presented example, multiple batch channels can include a first batch channel (Batch A), a second batch channel (Batch B), a third batch channel (Batch C), and a fourth batch channel (Batch D). For example, each batch channel can be divided into four parts (or a part of four parts). Each of these batch channels can contain the complete dataset.
[0245] First, refer to Figure 5The grouped weights, at least a portion of the first batch of channels, and at least a portion of the second batch of channels are stored in at least one memory (S2001). Here, the grouped weights refer to a weight matrix kernel that includes at least one weight value, and the memory may be internal memory 200, on-chip memory, or main memory.
[0246] In this example, the size of the grouped weights, the size of at least a portion of the first batch of channels, and the size of at least a portion of the second batch of channels can be adjusted to fit the at least one memory before being stored in at least one memory.
[0247] In this example, the size of at least a portion of the first batch of channels can be less than or equal to the size of at least one memory divided by the number of batch channels. Furthermore, the size of at least one internal memory can correspond to the maximum feature map size of the ANN and the number of batch channels. In this example, at least one internal memory can store the compression parameters of the ANN.
[0248] More specifically, the size of at least one memory may be determined by the data size of the specific ANN parameters processed by the neural processing unit 1000 and the number of batch channels.
[0249] Next, at least a portion of the first batch of channels and at least a portion of the second batch of channels are calculated, along with the grouped weights (S2003). For example, this calculation can correspond to a convolution operation based on intervals.
[0250] Next, the subsequent portions of the first batch of channels and the subsequent portions of the second batch of channels are stored in at least one memory while maintaining the grouped weights (S2005). Then, the weights of the subsequent portions of the first batch of channels and the subsequent portions of the second batch of channels, as well as the grouped weights, are calculated (S2007).
[0251] Then, while repeating steps S2005 and S2007, the artificial neural network operation is performed (S2009).
[0252] Specifically, refer to Figure 6 In this example, it is assumed that there are five memory spaces. Additionally, in this example, it is assumed that the weights of the groups include weight W, the first batch A includes A1, A2, A3, and A4, the second batch B includes B1, B2, B3, and B4, the third batch C includes C1, C2, C3, and C4, and the fourth batch D includes D1, D2, D3, and D4.
[0253] refer to Figure 6In S1, the weights W, the first portion A1 of the first batch of channels, the first portion B1 of the second batch of channels, the first portion C1 of the third batch of channels, and the first portion D1 of the fourth batch of channels are filled in five memory spaces. The processing element PE performs calculations on the weights W and each first portion of the first, second, third, and fourth batches of channels. Here, the weights W can be a weight matrix that includes at least one weight value.
[0254] In S2, the weight W remains in S1, while the second portions A2 of the first batch of channels, B2 of the second batch of channels, C2 of the third batch of channels, and D2 of the fourth batch of channels are filled into their respective memory spaces. The processing element PE performs calculations on the weight W and each second portion of the first, second, third, and fourth batches of channels.
[0255] In S3, the weight W remains in S1, while the third portion A3 of the first batch of channels, the third portion B3 of the second batch of channels, the third portion C3 of the third batch of channels, and the third portion D3 of the fourth batch of channels are filled into their respective memory spaces. The processing element PE performs calculations on the weight W and each third portion of the first, second, third, and fourth batches of channels.
[0256] In S4, the weight W remains in S1, while the fourth portions A4 of the first batch of channels, B4 of the second batch of channels, C4 of the third batch of channels, and D4 of the fourth batch of channels are filled into their respective memory spaces. The processing element PE performs calculations on the weight W and each fourth portion of the first, second, third, and fourth batches of channels.
[0257] Once the computation of the first, second, third, and fourth batches of channels is complete, feature maps can be generated, and activation functions can be selectively applied to generate activation maps. The feature maps or activation maps generated in this way can be fed into convolutional layers for further convolution operations, into pooling layers for pooling operations, or as input to fully connected layers for classification, but are not limited to these methods. These computations can be performed by the processing element PE as described above.
[0258] In this way, the processing element PE calculates each of the multiple batch channels and the weights of the groups held in memory (i.e., internal memory). That is, by holding at least one core or group of weight values until the calculation of the corresponding feature maps of the multiple batch channels is completed, at least one core or group of weight values corresponding to the multiple feature maps of the multiple batch channels can be reused respectively.
[0259] Figure 5 and 6The batch mode of the proposed operation method can be described as a method that only tiles the feature maps of each layer in each batch channel, and can be called the first batch mode. The first batch mode can be used when the parameter size of the inter-layer feature maps of the artificial neural network model is relatively larger than the parameter size of the kernel.
[0260] Meanwhile, in related technologies, when processing multiple consecutive data or consecutive image data, a new access weight is assigned for each operation. This traditional method is inefficient.
[0261] On the other hand, the neural processing unit according to this disclosure maintains weights in memory, thereby minimizing unnecessary access operations to the weights, thus improving processing speed and reducing power consumption. In this example, in the case of on-chip memory and NPU internal memory, the memory has the same performance improvement and energy saving effect.
[0262] Figure 7 This illustrates how a neural processing unit operates according to another example of this disclosure. Figure 8 It shows according to Figure 7 The step involves allocating memory space for the artificial neural network parameters within the neural processing unit.
[0263] In this example, multiple batch channels can include a first batch channel (Batch A) and a second batch channel (Batch B). For example, each batch channel can be divided into four parts (or a portion of four parts).
[0264] First, refer to Figure 7 The grouped weights, at least a portion of the first batch of channels, and at least a portion of the second batch of channels are stored in at least one memory (S2011). Here, the grouped weights may be referred to as at least one weight value in a weight matrix that includes at least one weight value. The memory may also be internal memory, on-chip memory, or main memory.
[0265] Next, at least a portion of the first batch of channels and at least a portion of the second batch of channels are calculated, along with the grouped weights (S2013). For example, this calculation can correspond to a convolution operation based on the stride value.
[0266] Next, while maintaining the grouped weights, another part of the first batch of channels and another part of the second batch of channels are stored in at least one memory (S2015), and then each of the other parts of the first batch of channels and the second batch of channels is calculated using the grouped weights (S2017).
[0267] Next, another set of weights is stored in at least one memory, and the artificial neural network operation is performed using the other set of weights (S2019). Here, the other set of weights may be referred to as another weight value in a weight matrix that includes at least one weight value.
[0268] Specifically, refer to Figure 8 Assume that the memory in this example has three memory spaces.
[0269] In this example, assume that the weights of the group are at least one of the first weight W1, the second weight W2, the third weight W3, and the fourth weight W4, and the weights of the other group are the other weight of the first weight W1, the second weight W2, the third weight W3, and the fourth weight W4.
[0270] In this example, assume that the first batch of channels, Batch A, includes A1, A2, A3, and A4, and the second batch of channels, Batch B, includes B1, B2, B3, and B4. The weights of another group can be referred to as the weights of the next group (i.e., the weights of subsequent groups).
[0271] refer to Figure 8 The three memory spaces in S1 can be filled with a first weight W1 as a group weight, a first portion A1 of the first batch of channels, and a first portion B1 of the second batch of channels. The processing element PE calculates the first weight W1 using each of the first portions A1 and B1 of the first and second batches of channels.
[0272] In S2, while maintaining the first weight W1 as in S1, the second portion A2 of the first batch of channels and the second portion B2 of the second batch of channels are filled. The processing element PE calculates the first weight W1 using each of the two portions A2 and B2 of the first and second batches of channels.
[0273] When calculating each of the third portions A3 and B3 and the fourth portions A4 and B4 of the first and second batches of channels as needed, the first weight W1 can be further maintained during S3 to S4.
[0274] When the calculation of the first weight W1 is completed using each of the first and second batches of channels, in S5, the three memory spaces are filled with another set of weights, such as the second weight W2, the first portion A1 of the first batch of channels, and the first portion B1 of the second batch of channels. The processing element PE calculates the second weight W2 and each of the first portions of the first and second batches of channels.
[0275] The second weight W2 can be maintained during S6 to S8, during which each of the second portions A2 and B2, the third portions A3 and B3, and the fourth portions A4 and B4 of the first and second batch channels is calculated.
[0276] When the calculation of the second weight W2 is completed using each of the first and second batches of channels, in S9, three memory spaces are filled with another set of weights, such as the third weight W3, the first portion A1 of the first batch of channels, and the first portion B1 of the second batch of channels. The processing element PE calculates the third weight W3 and each of the first portions of the first and second batches of channels.
[0277] The third weight W3 can be maintained during S10 to S12, during which each of the first and second batch channels A1 and B1, the second batch A2 and B2, the third batch A3 and B3, and the fourth batch A4 and B4 is calculated.
[0278] When the calculation of the third weight W3 is completed using each of the first and second batches of channels, in S13, the three memory spaces are filled with another set of weights, such as the fourth weight W4, the first portion A1 of the first batch of channels, and the first portion B1 of the second batch of channels. The processing element PE calculates the fourth weight W4 and each of the first portions of the first and second batches of channels.
[0279] The fourth weight W4 can be maintained during S14 to S16, during which each of the first and second batch channels A1 and B1, the second batch A2 and B2, the third batch A3 and B3, and the fourth batch A4 and B4 is calculated.
[0280] Once the computation of the first and second batches of channels is complete, feature maps are generated, and activation maps can be applied to generate activation maps. The generated activation maps can be fed into a convolutional layer for another convolution operation, into a pooling layer for a pooling operation, or into a fully connected layer for classification, but this disclosure is not limited thereto. These computations can be performed by a processing element (PE).
[0281] In this way, the processing element PE can calculate multiple weight values held in memory for each of the multiple batch channels. That is, by holding at least one core or group of weight values until the calculation of the corresponding feature maps of the multiple batch channels is completed, at least one core or group of weight values corresponding to the multiple feature maps of the multiple batch channels can be reused respectively.
[0282] Figure 7 and 8 The batch mode of the proposed operation method can be described as a method of tiling the weights and feature maps of each layer in each batch channel, and can be called the second batch mode. The second batch mode can be used when the weight parameters and feature map parameters of the artificial neural network model layers are relatively larger than the memory capacity.
[0283] Meanwhile, in related technologies, when processing multiple consecutive data or consecutive image data, multiple weight values are stored in main memory for each new operation request. This traditional method is inefficient.
[0284] On the other hand, the neural processing unit according to this disclosure continuously maintains multiple weight values in memory, thus minimizing the total frequency of new accesses to multiple weight values, thereby improving processing speed and reducing energy consumption. In this example, in the case of on-chip memory and NPU internal memory, the memory achieves the same performance improvement and energy saving effect.
[0285] Figure 9 This illustrates how a neural processing unit operates according to another example of this disclosure. Figure 10 It shows according to Figure 9 The step involves allocating memory space for the artificial neural network parameters within the neural processing unit.
[0286] In the example presented, multiple batch channels may include a first batch (Batch A), a second batch (Batch B), a third batch (Batch C), and a fourth batch (Batch D). The weights of the groups may be, for example, divided into four parts (or a portion of four parts).
[0287] First, referring to Figure 9, the grouped weights, at least a portion of the first batch of channels, and at least a portion of the second batch of channels are stored in at least one memory (S2021), and at least a portion of the first batch of channels and the grouped weights are calculated (S2023). Here, a grouped weight can refer to at least one weight value in a weight matrix, which includes at least one weight value. The memory can also be internal memory, on-chip memory, or main memory.
[0288] Next, while calculating at least a portion of the second batch of channels using grouped weights, at least a portion of the subsequently calculated third batch of channels is stored in the space of at least a portion of the first batch of channels (S2025). That is, the next calculated parameter is loaded into memory simultaneously with the calculation.
[0289] Next, while calculating at least a portion of the third batch of channels with grouped weights, at least a portion of the subsequently calculated fourth batch of channels is stored in the space of at least a portion of the second batch of channels (S2027), and while calculating at least a portion of the fourth batch of channels with grouped weights, at least a portion of the subsequently calculated first batch of channels is stored in the space of at least a portion of the third batch of channels (S2029).
[0290] Specifically, refer to Figure 10In this example, it is assumed that the memory has three memory spaces. Furthermore, in this example, it is assumed that the weights in the group are at least one of the first weight W1, the second weight W2, the third weight W3, and the fourth weight W4.
[0291] refer to Figure 10 In S1, three memory spaces are filled with a first weight W1 (as the weight of the group), a first batch of channels A, and a second batch of channels B. The processing element PE calculates the first weight W1 using the first batch of channels A. In S2, while the processing unit PE is calculating the first weight W1 and the second batch of channels B, it loads the third batch of channels C into the memory space corresponding to the memory address of the first batch of channels A. In this way, the calculation of the batch channels and the loading of subsequent parameters to be calculated are performed simultaneously. Therefore, the computation speed of the neural processing unit can be further improved.
[0292] Next, in S3, while the processing element PE calculates the first weight W1 using the third batch of channels C, the fourth batch of channels D is loaded into the memory space corresponding to the memory address of the second batch of channels B. In S4, while the processing unit PE calculates the first weight W1 using the fourth batch of channels D, the first batch of channels A is loaded into the memory space corresponding to the memory address of the third batch of channels C.
[0293] After the first weight W1 is calculated for each batch of channels, in S5, the second weight W2 is loaded into the memory space corresponding to the memory address of the first weight W1. In S5, while the processing unit PE calculates the weights of the group that constitute the second weight W2 and the first batch of channels A, the second batch of channels B is loaded into the memory space corresponding to the memory address of the fourth batch of channels D. In various examples, the second weight W2 can be loaded into the space corresponding to another memory address while calculating the first weight W1.
[0294] When the calculation of the second weight W2 is completed for each batch channel, the calculation of the third weight W3, which is a group weight, and each batch channel can be performed. This calculation can be performed in a similar manner to the calculation described above.
[0295] After the calculation of the third weight W3 and each batch channel is completed, the fourth weight W4, which is a group weight, and each batch channel can be calculated, and such calculation can be performed in a similar manner to the above calculation. For example, in SN, while the processing unit PE is performing the calculation between the fourth weight W4 and the fourth batch channel D, the calculated value (i.e., the operation value) A' of the first weight W1 and the first batch channel A is loaded into the space corresponding to the memory address of the third batch channel C.
[0296] After the calculation of the fourth weight W4 and each batch channel is completed, in SN+1, the parameter X is loaded into the memory space corresponding to the memory address of the fourth weight W4 for subsequent processing. In SN+1, while the processing unit PE calculates the parameter X and the calculated value A', the calculated value B' of the first weight W1 and the second batch channel B is loaded into the memory space corresponding to the memory address of the fourth batch channel D.
[0297] Once the computation of the first, second, third, and fourth batches of channels is complete, feature maps are generated, and activation maps can be applied to generate activation maps. The activation maps generated as described above can be input to a convolutional layer for another convolution operation, to a pooling layer for a pooling operation, or to a fully connected layer for classification, but this disclosure is not limited thereto. These computations can be performed by the processing element PE as described above.
[0298] In this way, the processing element PE can calculate multiple weights held in memory and in each of the multiple batch channels.
[0299] More specifically, similar to the transition from S1 to S2, at least a portion of the first batch of channels A can be covered by the third batch of channels C. That is, during a specific time period, the memory space storing the first batch of channels A can be gradually filled by the third batch of channels C. In this case, the memory space covered by the third batch of channels C can be the memory space storing data from the first batch of channels A that has been convolved using W1 weights.
[0300] In other words, in the case of a memory space that stores the calculated input feature map, another batch of channels can gradually fill the memory space that stores a specific batch of channels.
[0301] Figure 9 and 10 The batch mode of the proposed operation method can be called a method in which the parameters of at least some channels (e.g., two channels) out of multiple batch channels (e.g., four channels) are stored separately in memory, and then, when calculating the parameters of another batch channel, the parameters of the next batch channel to be calculated are loaded into the memory region of the batch channel in which the parameter calculation has been completed. This method can be called the third batch mode. In the third batch mode, the size of each region allocated to memory can be increased because the number of memory region partitions is not as many as the total number of batch channels, but rather the number of memory region partitions is less than the total number of batch channels.
[0302] Meanwhile, in related technologies, when processing multiple continuous data or continuous image data, a new access weight is applied for each operation. This traditional method is inefficient.
[0303] On the other hand, the neural processing unit according to this disclosure maximizes memory space utilization by loading new batch processing channels or new weights into the space corresponding to the memory addresses used simultaneously in calculating weights and batch processing channels, thereby improving processing speed and reducing power consumption. In this example, in the case of on-chip memory and NPU internal memory, the memory has the same performance improvement and power saving effect.
[0304] Figure 11 The following illustrates how neural processing units operate according to various examples of this disclosure. Figure 12 It shows according to Figure 11 The step involves allocating memory space for the artificial neural network parameters within the neural processing unit.
[0305] In the example presented, multiple batches of channels may include a first batch (Batch A), a second batch (Batch B), a third batch (Batch C), and a fourth batch (Batch D). The weights of the groups may, for example, be divided into two parts.
[0306] First, refer to Figure 11 The grouped weights, at least a portion of the first batch of channels, and at least a portion of the second batch of channels are stored in at least one memory (S2031). The weights of each of the at least a portion of the first batch of channels and the at least a portion of the second batch of channels, along with the grouped weights, are calculated, and the calculated values are stored in at least one memory (S2033). Here, the grouped weights refer to a weight matrix including at least one weight value, and the memory can be internal memory, on-chip memory, or main memory. In this example, the calculated values can be stored in the same memory as the grouped weights, at least a portion of the first batch of channels, and at least a portion of the second batch of channels.
[0307] Next, another set of weights is stored in at least one memory for use in the next processing step S2035, and the calculated value and another set of weights for the next processing step are calculated (S2037). This calculation may correspond to a ReLU operation or a next-level convolution operation, but this disclosure is not limited thereto. The calculated value is stored in at least one memory, and the artificial neural network operation is performed using the calculated value (S2039).
[0308] Since the calculated values are stored in the NPU's internal memory 200, there is no need to access main memory or external memory for calculations.
[0309] Specifically, refer to Figure 12 In this example, it is assumed that the memory can have ten memory spaces. Furthermore, in this example, it is assumed that the weights in one group represent the first weight W5, and the weights in another group represent the second weight W6.
[0310] refer to Figure 12 In S1, five memory spaces are filled with the grouped weights of the first weight W5, the first batch of channels A, the second batch of channels B, the third batch of channels C, and the fourth batch of channels D. The processing element PE calculates the first weight W5 using each of the first, second, third, and fourth batches of channels. The calculated values A', B', C', and D' are stored in four memory spaces, while another grouped weight, the second weight W6, is loaded into one memory space for the next processing step.
[0311] Next, in S2, while maintaining the second weight W6 as in S1, the second weight W6 and the calculated values A', B', C', and D' are calculated respectively. This calculation is performed by the processing element PE and can correspond to, for example, a ReLU operation or a next-level convolution operation.
[0312] The first weight W5 used can be removed from memory. The first calculated values A', B', C', and D', along with the calculated values of the second weight W6, i.e., the second calculated values A", B", C", and D" are filled into the memory space of each of the first batch of channels A, the second batch of channels B, the third batch of channels C, and the fourth batch of channels D. The next parameter, such as parameter X, is loaded into the memory space.
[0313] In various examples, data stored only in internal memory can be used to perform operations in which calculated values are stored in at least one memory, as well as operations in which another set of weights or parameters are calculated for the next processing step.
[0314] This way, the calculated values can be continuously stored in memory.
[0315] Figure 11 and 12 The batch mode of the operation method proposed in the paper can be described as a method of processing each batch channel by utilizing the property that the output feature map is used as the input feature map for the next operation, and can be called the fourth batch mode.
[0316] On the other hand, in related technologies, when processing multiple consecutive data or consecutive image data, the calculated values of the image data are stored in main memory and are accessed again each time an operation is performed for subsequent calculations. This traditional method is inefficient.
[0317] On the other hand, the neural processing unit according to this disclosure continuously stores the computed values in the NPU's internal memory 200, thereby minimizing new accesses to the computed values, thus improving processing speed and reducing power consumption. In this example, the memory is described as the NPU's internal memory 200, but this disclosure is not limited to this, and even greater performance improvements and energy savings are possible with on-chip memory.
[0318] Artificial neural network models consist of multiple layers, and each layer includes weight parameters and feature map parameter information. The NPU scheduler can be provided with parameter information.
[0319] According to this disclosure, the neural processing unit can be configured to process an artificial neural network model by selectively utilizing at least one of the first to fourth batches of patterns described above.
[0320] Neural processing units can apply specific batch patterns to specific layers of an artificial neural network model.
[0321] Neural processing units can process a portion of a layer of an artificial neural network model in a specific batch pattern and process another portion in a different batch pattern.
[0322] Neural processing units can apply specific batch patterns to each layer of an artificial neural network model.
[0323] The neural processing unit can apply multiple batch patterns to a single layer of an artificial neural network model.
[0324] Neural processing units can be configured to provide the optimal batch pattern for each layer of an artificial neural network model.
[0325] Figure 13 An autonomous driving system in which a neural processing unit is installed is shown according to an exemplary embodiment of the present disclosure.
[0326] refer to Figure 13 The autonomous driving system C may include an autonomous vehicle having multiple sensors for autonomous driving and a vehicle controller 10000 that controls the vehicle to perform autonomous driving based on sensing data obtained from the multiple sensors.
[0327] Autonomous vehicles can include multiple sensors and can perform autonomous driving by monitoring the vehicle’s surrounding environment through these sensors.
[0328] The multiple sensors provided in an autonomous vehicle can include a variety of sensors that autonomous driving may require. For example, the various sensors can include image sensors, radar and / or lidar and / or ultrasonic sensors, etc. In addition, the multiple sensors can include multiple identical sensors or multiple different sensors.
[0329] The image sensor may correspond to a front camera 410, a left camera 430, a right camera 420, and a rear camera 440. In various examples, the image sensor may correspond to a 360-degree camera or a surround-view camera.
[0330] Image sensors may include, but are not limited to, image sensors used to capture color images [e.g., RGB (380 nm to 680 nm) images], such as complementary metal-oxide-semiconductor (CMOS) sensors or charge-coupled device (CCD) sensors.
[0331] In various examples, the image sensor may also include, but is not limited to, infrared (IR) sensors and / or near-infrared (NIR) sensors for capturing night vision and daytime environments for autonomous vehicles. These sensors can be used to compensate for the quality of low-light nighttime images captured by a color image sensor in a nighttime environment. Here, the NIR sensor may be implemented in the form of a four-pixel structure combining the RGB and IR sensors of a CMOS sensor, but is not limited to this.
[0332] To capture near-infrared images using NIR sensors, autonomous vehicles may also include NIR light sources (e.g., 850 nm to 940 nm). These NIR light sources are invisible to the human eye and do not interfere with the vision of other drivers, serving as an additional light source for the vehicle's headlights.
[0333] In various examples, the IR sensor is a thermal sensor and can be used to capture thermal images. In various examples, the autonomous vehicle may also include an IR light source corresponding to the IR sensor.
[0334] For example, thermal images can be configured to include RGB images and synchronized thermal sensing information. Furthermore, thermal images can be used to identify road surface temperature, vehicle engines, exhaust vents, wildlife in nighttime environments, and / or icy roads, all of which could be risk factors in autonomous driving processes.
[0335] In various examples, IR sensors can be used to determine whether a driver has a high fever, a cold, a coronavirus infection, and / or the status of the interior air conditioning by detecting the driver's (or user's) temperature through thermal sensing when provided inside an autonomous vehicle.
[0336] The captured thermal images can be used as reference images for training artificial neural network models, which will be described later in the context of object detection.
[0337] In various examples, the IR light source can be synchronized with multiple IR image sensors, and the thermal images captured by the synchronized IR light source and multiple IR image sensors can be used as reference images for training artificial neural network models, which will be described later for object detection.
[0338] In various examples, the IR light source may have an illumination angle at the front, and this illumination angle may be different from the vehicle illumination angle of the headlight.
[0339] In various examples, the NIR light source and / or IR light source are turned on / off for each frame of the image sensor, respectively, and can be used to identify objects with rear reflector characteristics (e.g., vehicle seat belts, traffic signs, and rear reflectors).
[0340] As mentioned above, multiple cameras, including image sensors, can be provided in various numbers at various locations within an autonomous vehicle. Here, the various locations and various numbers can be the locations and numbers required for autonomous driving.
[0341] Multiple cameras 410, 420, 430, and 440 can capture images of the vehicle's surroundings (e.g., the environment around the vehicle) and send the captured images to a vehicle controller. The multiple images may include at least one of infrared images and near-infrared images (or thermal images) and simultaneously captured color images (e.g., RGB images), or images formed by a combination thereof, but this disclosure is not limited thereto.
[0342] In various examples, multiple cameras 410, 420, 430, and 440 can be installed inside the autonomous vehicle. As mentioned above, the multiple cameras installed inside can be arranged in various locations, and the images captured by them can be used for a driver condition monitoring system, but this disclosure is not limited thereto. In various examples, the captured images can be used to determine driver drowsiness, intoxication, neglect of infants, convenience, safety, etc.
[0343] The autonomous vehicle can receive driving instructions from the vehicle controller 10000 and drive the vehicle according to the received driving instructions.
[0344] Next, the vehicle controller 10000 can be an electronic device for controlling an autonomous vehicle based on sensing data obtained from multiple sensors. The vehicle controller 10000 can be implemented as, for example, an electrical system that can be installed in the vehicle, a dashcam that can be connected to the vehicle, or a portable device such as a smartphone, personal digital assistant (PDA), and / or tablet PC (personal computer), but the invention is not limited thereto.
[0345] The vehicle controller 10000 may include a processor. The processor may be configured to include at least one of a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), a digital signal processing unit (DSP), an arithmetic logic unit (ALU), and a neural processing unit (NPU). However, the processor disclosed herein is not limited to the processors described above.
[0346] Specifically, the vehicle controller 10000 can generate autonomous driving map data for autonomous driving by using sensing data obtained from multiple sensors, and can send autonomous driving commands based on the generated autonomous driving map data to the autonomous vehicle.
[0347] Here, autonomous driving map data is map data that accurately or precisely represents the vehicle's surrounding environment based on sensor data measured by at least one sensor (e.g., camera, radar, lidar, and / or ultrasonic sensor), and can be implemented in three-dimensional space.
[0348] To generate such autonomous driving map data, the vehicle controller 10000 can detect road environment data and real-time environment data based on sensor data. Road environment data may include, but is not limited to, lanes, guardrails, road curvature / slope, traffic light / sign positions, and / or traffic signs. Real-time environment data can be environmental data that changes constantly and may include, but is not limited to, approaching vehicles, construction (or accident) areas, road traffic, real-time signal information, road surface conditions, obstacles, and / or pedestrians. In various examples, the road environment data and real-time environment data can be continuously updated.
[0349] Autonomous driving map data can be generated using the sensing data described above, but is not limited to this, and can also use map data previously generated for a specific area.
[0350] The pre-generated map data can include at least some of the road and surrounding environment data previously collected by survey vehicles equipped with various sensors, and can be stored in a cloud-based database. Such map data can be analyzed in real time and continuously updated.
[0351] The vehicle controller 10000 can obtain map data of the area corresponding to the location of the autonomous vehicle from a database, and can generate autonomous driving map data based on sensing data measured by various sensors along with the acquired map data. For example, the vehicle controller 10000 can generate autonomous driving map data by updating the acquired map data about a specific area in real time based on the sensing data.
[0352] The vehicle controller 10000 may need to accurately identify rapidly changing surrounding environments in real time to control the vehicle's autonomous driving. In other words, during autonomous driving, it needs to accurately and continuously identify potential target objects (or objects) around the vehicle in order to proactively address dangerous situations. Here, target objects may include, but are not limited to, at least one of the following: approaching vehicles, traffic lights, real-time traffic light signal information, obstacles, people, animals, roads, signs, road lines, and pedestrians.
[0353] If accurate and continuous identification of target objects cannot be achieved in the environment surrounding the vehicle, dangerous situations such as serious accidents may occur between the vehicle and the target object, which may prevent the vehicle from achieving safe and correct autonomous driving.
[0354] For safe autonomous driving, the vehicle controller 10000 can identify target objects related to autonomous driving based on multiple images received from multiple cameras 410, 420, 430, and 440. Here, the multiple images can be images captured simultaneously by multiple cameras 410, 420, 430, and 440. As described above, such cameras can be processed by an NPU having multiple batch channels and corresponding memory for each batch channel.
[0355] To accurately and continuously identify target objects, an AI-based object detection model (or artificial neural network model, i.e., ANN) can be used, trained to identify target objects based on training data about various surrounding environments of the vehicle. Here, the training data can be, but is not limited to, multiple reference images of various surrounding environments of the vehicle. These multiple reference images can include, but are not limited to, at least two or more of infrared images, near-infrared images, thermal images, and color images. In various examples, the multiple reference images can be formed by a combination of at least two or more of image sensors (e.g., color image sensors, IR sensors, and / or NIR sensors), LiDAR, radar, and ultrasonic sensors.
[0356] The vehicle controller 10000 uses an object detection model to identify target objects from images received from multiple cameras 410, 420, 430, and 440, and can send autonomous driving commands to the autonomous vehicle based on the identified target objects. For example, when a target object such as a pedestrian is identified on the road during autonomous driving, the vehicle controller 10000 can send a command to the autonomous vehicle to stop driving.
[0357] As described above, this disclosure can use an AI-based object detection model to identify potential target objects for autonomous driving of autonomous vehicles, thereby achieving accurate and rapid object detection.
[0358] In the following text, reference will be made to Figure 14 Describe autonomous vehicles in more detail.
[0359] Figure 14 An example of an autonomous driving system in which a neural processing unit is installed, according to the present disclosure, is shown.
[0360] refer to Figure 14 An autonomous vehicle may include a communication unit 600, sensors 400, a storage unit, and a controller. In the presented example, the autonomous vehicle may refer to... Figure 13The autonomous driving system.
[0361] The communication unit 600 connects to the autonomous vehicle to enable communication with external devices. The communication unit 600 can connect to the vehicle controller 10000 via wired / wireless communication to send / receive various data related to autonomous driving. Specifically, the communication unit 600 can send sensing data obtained from multiple sensors to the vehicle controller 10000 and can receive autonomous driving commands from the vehicle controller 10000.
[0362] The location search unit 700 can search for the location of autonomous vehicles. The location search unit 700 can use at least one of satellite navigation and dead reckoning. For example, when using satellite navigation, the location search unit 700 can obtain location information from a location search system that measures vehicle location information, such as Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Galileo, BeiDou, etc.
[0363] When dead reckoning is used in various examples, the position search unit 700 can calculate the vehicle's route and speed from motion sensors such as the vehicle's speedometer, gyroscope sensor and geomagnetic sensor, and estimate the vehicle's position information based on this.
[0364] In various examples, the location search unit 700 can use both satellite navigation and dead reckoning to obtain the vehicle's location information.
[0365] Sensor 400 can acquire sensing data for sensing the environment surrounding the vehicle. Sensor 400 may include an image sensor 410, a lidar sensor 450, a radar sensor 460, and an ultrasonic sensor 470. Multiple identical sensors may be provided, or multiple different sensors may be provided.
[0366] An image sensor 410 is provided to capture the environment surrounding the vehicle, and may be at least one of a CCD sensor, a CMOS sensor, an IR sensor, and / or a NIR sensor. Multiple such image sensors 410 may be provided, and multiple cameras may be positioned at various locations within the autonomous vehicle to correspond to the multiple image sensors. For example, multiple front, left / right, and rear cameras may be provided to capture the environment surrounding the vehicle, or 360-degree cameras or surround-view cameras may be provided, but are not limited to these.
[0367] In various examples, cameras corresponding to CCD and / or CMOS sensors can acquire color images of the vehicle's surroundings.
[0368] In various examples, IR and / or NIR sensors can detect objects by sensing temperature based on infrared and / or near-infrared light. Corresponding to the IR and / or NIR sensors, infrared cameras, near-infrared cameras, and / or thermal imaging cameras are positioned at at least one location on the autonomous vehicle to acquire infrared, near-infrared, and / or thermal images of the environment surrounding the vehicle. The infrared, near-infrared, and / or thermal images acquired in this way can be used for autonomous driving in low-light conditions or in dark areas.
[0369] Radar 460 can detect the position, velocity, and / or orientation of an object by emitting electromagnetic waves and using the echoes reflected and returned from surrounding objects. In other words, lidar 450 can be a sensor used to detect objects in the environment in which a vehicle is located.
[0370] The lidar 450 can be a sensor capable of sensing the surrounding environment (such as the shape of an object and / or the distance to an object) by emitting laser light and using reflected light reflected from and returned from surrounding objects.
[0371] The ultrasonic sensor 470 can detect the distance between a vehicle and an object by emitting ultrasonic waves and using the ultrasonic waves reflected from surrounding objects. The ultrasonic sensor 470 can be used to measure short distances between a vehicle and an object. For example, the ultrasonic sensor 470 can be positioned at the front left, front right, left side, rear left, rear right, and right side of the vehicle, but is not limited to these locations.
[0372] Storage units can store various types of data used for autonomous driving. In various examples, storage units may include at least one type of storage unit, such as flash memory, hard disk, multimedia card micro, card type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disk, optical disk, etc.
[0373] The controller is operatively connected to the communication unit 600, the position search unit 700, the sensor 400, and the storage unit, and can execute various commands for autonomous driving. The controller can be configured to include, in addition to the neural processing unit (NPU), one of the following: a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), a digital signal processing unit (DSP), and an arithmetic logic unit (ALU).
[0374] Specifically, the controller transmits the sensing data acquired by the sensors to the vehicle controller 10000 via the communication unit 600. Here, the sensing data can be used to generate autonomous driving map data or to identify target objects. For example, the sensing data may include, but is not limited to, at least one of the following: image data acquired by the image sensor 410, object position data acquired by the rider 450, data representing the object's speed and / or orientation, etc., data representing the object's shape and / or distance to the object acquired by the radar 460, data representing the distance between the vehicle and the object acquired by the ultrasonic sensor 470, etc. Here, the image data may include multiple images simultaneously captured, including color images, infrared images, near-infrared images, and thermal images. In various examples, the image data may be formed by fusing at least two or more of the image sensor 410, lidar 450, radar 460, and ultrasonic sensor 470.
[0375] The controller can receive autonomous driving instructions from the vehicle controller 10000 and execute autonomous driving of the vehicle according to the received autonomous driving instructions.
[0376] Figure 15 The process for identifying target objects for autonomous driving in an autonomous driving system equipped with a neural processing unit, according to an example of this disclosure, is illustrated.
[0377] refer to Figure 15 The vehicle controller receives multiple images from multiple cameras installed in the autonomous vehicle (S1200). Here, the multiple images are images acquired simultaneously and may include color images, infrared images, and / or thermal images. In other words, multiple images or videos may refer to images captured within substantially the same time period. Thus, when using color images and infrared images, or color images and thermal images, nighttime and daytime autonomous driving of the vehicle can be effectively performed. In other words, the multiple images may include at least one of RGB images, IR images, radar images, ultrasonic images, lidar images, thermal images, and NIR images.
[0378] The vehicle controller can generate batch data in which multiple images are arranged sequentially (S1210). Here, the batch data can correspond to the input nodes of the input layer constituting the object detection model, and can represent multiple batch channels. Each of the multiple batch channels can correspond to each of the multiple images. The vehicle controller can identify target objects from the multiple images by using an object detection model that has been trained to identify target objects related to autonomous driving by inputting the batch data (S1220). Here, the object detection model refers to an artificial neural network model trained to identify target objects from multiple training images by inputting multiple training images related to various surrounding environments of the vehicle. The various surrounding environments of the vehicle can include daytime and / or nighttime environments, and in order to accurately identify target objects in these environments, images capturing the daytime environment and / or images capturing the nighttime environment can be used as training images.
[0379] In this example, the ANN (i.e., Artificial Neural Network Model) can perform at least one operation, including the detection, classification, or segmentation of objects from multiple batch channels. This ANN can preprocess multiple batch channels to improve object detection and can be configured to detect objects from multiple batch channels simultaneously while preprocessing them. Here, the multiple batch channels can correspond to channels of any one of the RGB, YCBCR, HSV, or HIS color spaces and IR channels. Furthermore, each of the multiple batch channels can include an image captured of the vehicle interior, and the ANN can be configured to detect at least one of vehicle safety-related objects, functions, driver states, and passenger states.
[0380] In various examples, each of the multiple batch channels can correspond to each of the multiple sensor data. The multiple sensor data can include data from one or more of a pressure sensor, piezoelectric sensor, humidity sensor, dust sensor, smoke sensor, sonar sensor, vibration sensor, accelerometer, or motion sensor. When a target object is identified in this way, the vehicle control unit can send autonomous driving commands related to the identified target object to the autonomous vehicle, enabling the autonomous vehicle to safely perform autonomous driving.
[0381] The examples shown in the specification and accompanying drawings are for illustrative purposes only and to provide concrete examples to aid in understanding the subject matter of this disclosure, and are not intended to limit the scope of this disclosure. It will be apparent to those skilled in the art to which this disclosure pertains that other modifications may be made based on the technical spirit of this disclosure, in addition to the examples disclosed herein.
Claims
1. A method for performing multiple operations on an artificial neural network, said multiple operations including: In the respective storage spaces formed in a memory, the storage weights and at least a portion of each of the plurality of batch processing channels are stored, the memory being formed by an on-chip memory disposed on a semiconductor die, each of the plurality of batch processing channels corresponding to each of the plurality of images; as well as At least a portion of each of the plurality of batch processing channels is calculated using the stored weights, while at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels is stored in the on-chip memory. The storage in the on-chip memory is a storage step that occurs prior to the execution of computation using at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels. Each of the plurality of batch processing channels consists of multiple parts. Specifically, for each of the plurality of parts constituting each of the plurality of batch processing channels, the storage operation and the computation operation are executed sequentially. The weights are held in the corresponding storage space of the storage space formed in the one memory until the computational operation on all parts of the plurality of parts constituting each of the plurality of batch processing channels is completed.
2. The method according to claim 1, The plurality of batch processing channels include a first batch processing channel and a second batch processing channel, and in, At least a portion of the first batch processing channel and at least a portion of the second batch processing channel are equal in size.
3. The method according to claim 2, wherein, The weights correspond to each of the at least portion of the first batch processing channel and the at least portion of the second batch processing channel.
4. The method according to claim 2, wherein the plurality of operations further comprises: At least a portion of the third batch processing channel and at least a portion of the fourth batch processing channel are stored in the on-chip memory, while maintaining the weights. as well as The weights are used to calculate at least a portion of the third batch processing channel and at least a portion of the fourth batch processing channel.
5. The method according to claim 2, wherein the plurality of operations further comprises: The subsequent grouped weights, the subsequent portions of the first batch processing channel, and the subsequent portions of the second batch processing channel are stored in the on-chip memory. as well as The subsequent portions of the first batch processing channel and the subsequent portions of the second batch processing channel are calculated using the subsequent grouped weights.
6. The method according to claim 2, wherein the plurality of operations further comprises: The weights and a group of first values calculated based on at least a portion of the first batch of processing channels and at least a portion of the second batch of processing channels are stored in the on-chip memory; The on-chip memory stores subsequent grouped weights for subsequent processing steps; as well as The calculation is performed using the first value of the group and the weights of the subsequent groups.
7. The method according to claim 6, further comprising: The first value of the group and the second value of the group obtained by calculating the weight of the subsequent group are stored in the internal memory.
8. The method according to claim 2, wherein, The at least portion of the first batch processing channel and the at least portion of the second batch processing channel comprise the complete dataset.
9. The method according to claim 2, wherein the plurality of operations further comprises: The size of the weight, the size of at least a portion of the first batch processing channel, and the size of at least a portion of the second batch processing channel are tiled to fit the internal memory.
10. The method according to claim 1, wherein, The artificial neural network is configured to perform at least one of the plurality of operations, the at least one operation including the detection, classification, or segmentation of objects from the plurality of batch processing channels.
11. The method according to claim 10, wherein, The objects include at least one of vehicles, traffic lights, obstacles, people, animals, roads, traffic signs, and lanes.
12. The method according to claim 2, wherein the plurality of operations further comprises: The plurality of batch processing channels are preprocessed before storing at least a portion of the first batch processing channel and at least a portion of the second batch processing channel in the on-chip memory.
13. The method of claim 12, wherein the artificial neural network is configured to simultaneously detect objects from the plurality of batch processing channels and simultaneously preprocess the plurality of batch processing channels.
14. The method according to claim 1, wherein, The plurality of batch processing channels include at least one batch processing channel having one of the following formats: infrared (IR), red-green-blue (RGB), luminance-chroma (YCBCR), hue-saturation-brightness (HSV), and hue-saturation-intensity (HIS).
15. The method according to claim 1, in, The plurality of batch processing channels include at least one batch processing channel for capturing images of the vehicle interior, and The artificial neural network is configured to detect at least one of objects, functions, driver status, and passenger status related to vehicle safety.
16. The method of claim 1, wherein the plurality of images comprises at least one of red-green-blue images, infrared images, radar images, ultrasonic images, lidar images, thermal images, near-infrared images, and fused images.
17. The method according to claim 1, wherein, The multiple images were captured within the same time period.
18. The method according to claim 1, Each of the plurality of batch processing channels corresponds to a plurality of sensor data, and in, The multiple sensor data include data from at least one of a pressure sensor, a piezoelectric sensor, a humidity sensor, a dust sensor, a smoke sensor, a sonar sensor, a vibration sensor, an acceleration sensor, and a motion sensor.
19. A neural processing unit for processing multiple batch channels of an artificial neural network, each batch channel corresponding to each of a plurality of images, the neural processing unit comprising: Each is formed in a corresponding storage space in a memory to store weights and at least a portion of each of the multiple batch processing channels, and At least one processing element is configured to: The stored weights are applied to at least a portion of each of the plurality of batch processing channels, each of the at least one processing element including one or both of multiplication and accumulation operators and arithmetic logic unit operators. Calculate at least a portion of each of the plurality of batch processing channels using the stored weights, and simultaneously store at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels in the one memory. The storage in the memory is a storage step that occurs before computation is performed with at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels. Each of the plurality of batch processing channels consists of multiple parts. Specifically, for each of the plurality of parts constituting each of the plurality of batch processing channels, the storage operation and the computation operation are executed sequentially. The weights are held in the corresponding storage space of the storage space formed in the one memory until the computational operation on all parts of the plurality of parts constituting each of the plurality of batch processing channels is completed.
20. The neural processing unit according to claim 19, The plurality of batch processing channels include a first batch processing channel and a second batch processing channel, and A portion of the first batch processing channel allocated to the one memory and a portion of the second batch processing channel allocated to the one memory are equal in size.
21. The neural processing unit according to claim 20, The at least one processing element is further configured to calculate grouped weights and values for subsequent stages based on at least a portion of the first batch of processing channels and at least a portion of the second batch of processing channels. in, The memory is also configured to store the calculated values and the grouped weights, and The grouped weights are maintained in the at least one internal memory until the plurality of batch processing channels are calculated.
22. The neural processing unit according to claim 20, The at least one processing element is further configured to calculate a first value based on at least a portion of the first batch of processing channels and at least a portion of the second batch of processing channels, and to calculate subsequent grouping weights for subsequent processing stages, and in, The memory is further configured to: Corresponding in size to at least a portion of the first batch processing channel and at least a portion of the second batch processing channel, and Store the first value, the weight, and the subsequent grouped weights.
23. The neural processing unit of claim 20, further comprising a scheduler configured to adjust the size of the weights, the size of at least a portion of the first batch of processing channels, and the size of at least a portion of the second batch of processing channels for the one memory.
24. A neural processing unit for processing multiple batch processing channels of an artificial neural network, each batch processing channel corresponding to each of multiple images, the neural processing unit comprising: Each is formed in a corresponding storage space in a memory to store weights and at least a portion of each of the plurality of batch processing channels; and At least one processing element is configured to apply stored weights to at least a portion of each of the plurality of batch processing channels, each of the at least one processing element including one or both of multiplication and accumulation operators and arithmetic logic unit operators. Wherein, the size of at least a portion of each of the plurality of batch processing channels is less than or equal to the size of the at least one internal memory divided by the number of the plurality of batch processing channels. The at least one processing element is further configured to: Calculate at least a portion of each of the plurality of batch processing channels using the stored weights, and simultaneously store at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels in the one memory. The storage in the memory is a storage step that occurs prior to the execution of computation with at least another portion of each of the plurality of batch processing channels or at least a portion of each of the plurality of subsequent batch processing channels. Each of the plurality of batch processing channels consists of multiple parts. Specifically, for each of the plurality of parts constituting each of the plurality of batch processing channels, the storage operation and the computation operation are executed sequentially. The weights are held in the corresponding storage space of the storage space formed in the one memory until the computational operation on all parts of the plurality of parts constituting each of the plurality of batch processing channels is completed.
25. The neural processing unit according to claim 24, wherein, The size of the memory corresponds to the size of the maximum feature map of the artificial neural network and the number of the plurality of batch processing channels.
26. The neural processing unit according to claim 24, wherein, The memory is also configured to store compression parameters of the artificial neural network.
27. The neural processing unit of claim 24, further comprising a processor located between the at least one processing element and the memory and configured to sequentially process feature maps of a first batch processing channel and a second batch processing channel corresponding to the plurality of batch processing channels to sequentially output activation maps corresponding to the first batch processing channel and the second batch processing channel.
Citation Information
Patent Citations
Batch processing in a neural network processor
IN201747034501A