Neural processing unit including post-processing unit and method of performing same

By introducing a post-processing unit and internal processing circuitry into the neural processing unit, the data processing flow is optimized, solving the problems of data transmission delay and high energy consumption in the prior art, and realizing more efficient and accurate neural network model calculation.

CN121009936APending Publication Date: 2025-11-25DEEPX CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510522248.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2025-04-24
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies suffer from data transmission latency and high energy consumption when performing post-processing operations on neural network-like models, especially in convolution operations and subsequent processing, which limits computational efficiency and accuracy.

Method used

A neural processing unit with a post-processing unit is used to perform nonmaximum suppression operation and class confidence score calculation through internal processing circuits. Combined with internal memory and processing element array circuits, the data processing flow is optimized to reduce data transmission latency and power consumption.

Benefits of technology

It improves the computational efficiency and accuracy of neural network-like models, reduces data transmission latency and energy consumption, and enhances overall processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009936A_ABST
    Figure CN121009936A_ABST
Patent Text Reader

Abstract

According to one example of the invention, a neural-like processing unit may include an array of processing elements for performing operations of a neural-like network model, and a post-processing unit configured to process data output from the array of processing elements. The post-processing unit includes a first computing circuit that extracts a subset of categories for each bounding box by comparing category scores of each category, and a second computing circuit that extracts one or more bounding boxes by comparing a category trust score of each bounding box to a threshold trust score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a neural processing unit including a post-processing unit. Background Technology

[0002] Humans possess intelligence such as recognition, classification, inference, prediction, and control / decision-making. Artificial intelligence (AI) is the artificial simulation of human intelligence.

[0003] The human brain is composed of many nerve cells called neurons. Each neuron is connected to hundreds to thousands of other neurons through connections called synapses. To simulate human intelligence, the operation of biological neurons and the connections between them are modeled as neural network (NN) models. In other words, a neural network model is a system of nodes connected within a layered structure that mimics neurons. Summary of the Invention

[0004] This embodiment relates to a neural network-like processing circuit that includes a processing element array circuit, a post-processing circuit, and subsequent circuitry. The processing element array circuit performs multiple convolution operations on a neural network-like model to generate output data. The post-processing circuit is coupled to the processing element array circuit to receive the output data and to extract a subset of the output data. The subsequent circuitry is coupled to the post-processing circuit and is used to selectively store the extracted subset of output data or to perform operations on the extracted subset of output data.

[0005] In one or more embodiments, the output data includes multiple category scores for each bounding box within the image region, the multiple category scores indicating the probability of an object category appearing in each bounding box.

[0006] In one or more embodiments, the post-processing circuitry includes a first computational circuitry for selecting one or more categories of each bounding box as a subset of the output data by comparing category scores of multiple categories for each bounding box.

[0007] In one or more embodiments, the post-processing circuitry further includes second computational circuitry for extracting one or more bounding boxes by comparing the category confidence score of each bounding box with a threshold confidence score. The category confidence score represents the probability that an object of a certain category exists within each bounding box. The category confidence score is derived from the object presence confidence score and the category score.

[0008] In one or more embodiments, the second computation circuit is used to compute the category confidence score into the product of the object existence confidence score and the category scores of a subset extracted by the first computation circuit.

[0009] In one or more embodiments, the post-processing circuitry further includes internal memory coupled to the first arithmetic circuitry and the second arithmetic circuitry. The internal memory stores a subset of the categories of each bounding box extracted by the first arithmetic circuitry, and stores data of one or more bounding boxes extracted by the second arithmetic circuitry.

[0010] In one or more embodiments, the post-processing circuitry further includes internal processing circuitry for performing non-maximum suppression (NMS) operations on one or more bounding boxes extracted by the second arithmetic circuitry.

[0011] In one or more embodiments, the internal processing circuitry performs nonmaximum suppression operations during the convolution operation performed by the processing element array.

[0012] In one or more embodiments, the internal processing circuitry is configured to perform nonmaximum suppression (NMS) operations on subsequent images of the continuation image at the later of the following two time points: (i) the completion time of the NMS operation on the image data, and (ii) the completion time of the convolution operation on the image by the processing element array circuitry.

[0013] In one or more embodiments, the output data further includes coordinate data for each bounding box.

[0014] In one or more embodiments, the post-processing circuitry further includes internal memory for storing a subset of the categories of each bounding box extracted by the first arithmetic circuitry.

[0015] In one or more embodiments, the first computing circuit performs a comparison of category scores during the convolution operation performed by the processing element array circuit.

[0016] In one or more embodiments, the neural network processing circuit further includes one or more processors and memory. The memory stores multiple instructions from the compiler. When the instructions are executed by one or more processors, they cause one or more processors to add a class-argmax layer to generate a neural network model. The subset of classes extracted by the first computation circuit corresponds to multiple operations of the class-argmax layer. Attached Figure Description

[0017] Figure 1 A schematic diagram illustrating a neural-like processing unit including a post-processing unit according to an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram illustrating a processing element applicable to one embodiment of the present invention.

[0019] Figure 3 A schematic diagram illustrating the convolutional neural network related to this invention.

[0020] Figure 4 A schematic diagram illustrating the energy consumption per unit operation of a neural processing unit according to an embodiment of the present invention.

[0021] Figure 5 This is a schematic diagram illustrating a post-processing unit according to an embodiment of the present invention.

[0022] Figure 6 A flowchart illustrating the computational process of a neural processing unit including a post-processing unit according to an embodiment of the present invention.

[0023] Figure 7 A flowchart illustrating a procedural activation function method according to an embodiment of the present invention.

[0024] Figures 8A to 8C A graph is provided to illustrate an approximation of an activation function using a procedural activation function method according to an embodiment of the present invention.

[0025] Figures 9A to 9D A diagram illustrating various scenarios of segmenting an activation function into multiple segments using a procedural activation function method according to an embodiment of the present invention.

[0026] Figures 10A to 10C A diagram illustrating an example of segmenting an activation function into linear and nonlinear segments using slope change data in segment data in a procedural activation function method according to an embodiment of the present invention.

[0027] Figure 11A as well as Figure 11B A graph illustrating an example of segmenting an activation function into substantially linear intervals and nonlinear intervals using slope change data in segment data in a programmed activation function method according to an embodiment of the present invention.

[0028] Figure 12A as well as Figure 12B This is a chart illustrating another example of how the activation function is segmented into substantially linear intervals and nonlinear intervals using slope change data in segment data in a procedural activation function method according to an embodiment of the present invention.

[0029] Figure 13A as well as Figure 13B A graph illustrating another example of segmenting the activation function into nonlinear intervals using gradient change data in segment data in a procedural activation function method according to an embodiment of the present invention.

[0030] Figure 14 A diagram illustrating an example of transforming a segment into a programmable segment using error values ​​in an activation function programming method according to an embodiment of the present invention.

[0031] Figure 15A as well as Figure 15B A diagram illustrating an example of approximating a segment as a programmable segment by finding the maximum error value according to an embodiment of the present invention.

[0032] Figure 16A as well as Figure 16B A diagram illustrating an example of approximating a segment as a programmable segment by integrating an error value in an activation function programming method according to an embodiment of the present invention.

[0033] Figure 17 A diagram illustrating an example of using machine learning to approximate a fragment into an optimized programmable fragment in an activation function programming method according to an embodiment of the present invention.

[0034] Figure 18 A diagram illustrating an example of segmenting an activation function by using an integral threshold of the segment approximation error of the activation function in a procedural activation function method according to an embodiment of the present invention.

[0035] Figure 19 as well as Figure 20 A diagram illustrating the activation functions of the Exponential Linear Unit (ELU) and the Hardswish activation function.

[0036] Figure 21 A flowchart illustrating a procedural activation function method according to an embodiment of the present invention.

[0037] Figure 22 This is a schematic diagram illustrating a neural network for approximating activation functions according to an embodiment of the present invention.

[0038] Figure 23 This is a schematic diagram illustrating the class argmax calculation steps performed by the post-processing unit according to an embodiment of the present invention.

[0039] Figure 24 This is a schematic diagram illustrating the filtering operation steps performed by the post-processing unit according to an embodiment of the present invention.

[0040] Figure 25 This is a schematic diagram illustrating the result of a filtering operation performed by a post-processing unit according to an embodiment of the present invention.

[0041] Figure 26 This is a schematic diagram illustrating the decoding steps performed by the post-processing unit according to an embodiment of the present invention.

[0042] Figure 27This is a schematic diagram illustrating the nonmaximum suppression operation steps performed by the post-processing unit according to an embodiment of the present invention.

[0043] Figure 28 This is a schematic diagram illustrating the data reduction of a neural processing unit including a post-processing unit according to an embodiment of the present invention.

[0044] Figure 29A This is a schematic diagram of a directed acyclic graph (DAG) representing an object detection neural network model input to a neural processing unit including a post-processing unit, according to an embodiment of the present invention.

[0045] Figure 29B This is a schematic diagram of a directed acyclic graph representing a post-processed object detection neural network model in a neural processing unit including a post-processing unit according to an embodiment of the present invention.

[0046] Figure 30 A timing diagram illustrating the operation of a neural processing unit including a post-processing unit on multiple image data according to an embodiment of the present invention is provided.

[0047] [Explanation of Labels in the Attached Image]

[0048] 100: Controller

[0049] 1000: Neural Processing Unit

[0050] 200: Direct Memory Access

[0051] 2000: Central Processing Unit

[0052] 300: Memory

[0053] 320: Embedded Compiler

[0054] 3000: Main Memory

[0055] 3010: External Compiler

[0056] 400: Processing Element Array

[0057] 4000: Image Sensor

[0058] 500: Special Function Unit

[0059] 5000: Decoder

[0060] 600: Post-processing unit

[0061] 610: First arithmetic unit

[0062] 620: Second arithmetic unit

[0063] 630: Internal Memory

[0064] 640: Internal Processing Unit

[0065] 641: Multiplier

[0066] 642: Adder

[0067] 643: Accumulator

[0068] 644: Bit quantization unit

[0069] 6000: Bus

[0070] S: Nonlinear activation function fragment

[0071] S100: Calculation Process

[0072] S110, S120, S130, S140, S150: Steps

[0073] S200, S210, S220: Steps

[0074] S310, S320, S330: Steps

[0075] s1-s6: Fragments

[0076] Sc1: First candidate segment

[0077] Sc2: Second candidate segment

[0078] Sc3: Third candidate fragment

[0079] Sp1: First Programmable Fragment

[0080] Sp2: Second Programmable Fragment

[0081] Sc(x): Candidate fragment

[0082] Sp(x): Programmable fragment

[0083] a1x+b1: Programmable fragment

[0084] a2x+b2: Programmable fragment

[0085] a3x+b3: Programmable fragment

[0086] a4x+b4: Programmable fragment

[0087] d1-d3: Points

[0088] w1: Nonlinear section

[0089] w1-1: Section

[0090] w1-2: Section

[0091] w2: Linear segment

[0092] w3: Linear segment

[0093] w4-w6: Interval

[0094] Th: Threshold of the essentially linear segment

[0095] x1-x5: Fragment boundary values

[0096] Max: Maximum value

[0097] Period 1: The First Period

[0098] Period 2: The Second Period

[0099] Period 3: The Third Period

[0100] Period 4: The Fourth Period

[0101] Bank 1: First memory partition

[0102] Bank 2: Second memory partition

[0103] IMG1: First Image Data

[0104] IMG2: Second Image Data

[0105] IMG3: Third Image Data

[0106] IMG4: Fourth Image Data

[0107] BOX1-N: Bounding Box Detailed Implementation

[0108] The specific structures or step-by-step descriptions disclosed in this specification or application are merely examples intended to illustrate embodiments of the concepts of the present invention.

[0109] Embodiments of the present invention may be embodied in various forms. They should not be construed as being limited to the embodiments described in this specification or application.

[0110] Various modifications can be made to embodiments of the present invention. The invention can take many forms. Accordingly, specific embodiments have been shown in the accompanying drawings and described in detail herein. However, this does not imply that embodiments of the present invention are limited to specific inventive forms. Therefore, it should be understood that all modifications, equivalents, or alternatives falling within the spirit and scope of the present invention are included within the scope of the present invention.

[0111] Terms such as first and / or second may be used to describe various elements. However, the present invention should not be limited to the terms described above. These terms are only used to distinguish different elements. For example, without departing from the scope of the present invention, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element.

[0112] When an element is described as being "connected to" or "in contact with" another element, it should be understood that the element can be directly connected to or in contact with the other element, although other elements may exist between them. On the other hand, when an element is described as being "directly connected to" or "directly in contact with" another element, it should be understood that there are no other elements between them. Other expressions describing the relationship between elements, such as "between" and "immediately between," or "adjacent" and "directly adjacent," should also be interpreted in a similar manner.

[0113] In this invention, expressions such as “A or B”, “at least one A and / or B”, or “one or more A and / or B” can include all possible combinations thereof. For example, “A or B”, “at least one A and B”, or “at least one A or B” can refer to: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.

[0114] In this specification, expressions such as "first," "second," and "first or second" can be used to modify various elements without limiting their order and / or importance. These expressions are only used to distinguish different elements and are not restrictive. For example, a first user device and a second user device can refer to different user devices without being affected by order or importance. As another example, without departing from the scope of protection described in this specification, a first element can be called a second element, and similarly, a second element can be renamed a first element.

[0115] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the scope of other embodiments. Unless the context clearly specifies otherwise, singular expressions may include plural expressions. The terms used in this specification, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art.

[0116] The terms used in this specification, if defined in a general dictionary, should be interpreted as having the same or similar meanings in the relevant technical field. Unless explicitly defined in this document, they should not be interpreted in an idealized or overly formal manner. In some cases, even terms defined in this specification should not be construed as excluding embodiments of the invention.

[0117] The terminology used in this specification is for describing specific embodiments only and is not intended to limit the scope of the invention. Singular expressions may include plural forms unless the context clearly requires it. Terms used in this specification, such as "comprising," "including," etc., indicate an associated feature, component, part, or combination thereof, and should not be construed as excluding the presence or addition of one or more other features, components, parts, or combinations thereof. Accordingly, it should be understood that the presence or addition of other features, components, or combinations thereof within the scope of this invention is not limited.

[0118] Unless otherwise defined, all terms used in this specification, including technical and scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in common dictionaries shall be interpreted in accordance with their meaning in the context of the relevant technical field. Unless expressly defined in this specification, they shall not be interpreted in an idealized or overly formal manner.

[0119] The various features of the embodiments of the present invention can be partially or completely combined with each other, or applied in combination with each other. The various embodiments of the present invention possess different ways of mutual cooperation and driving capabilities, which will be fully understood by those skilled in the art. The various embodiments of the present invention can be implemented independently of each other, or they can be implemented in combination with each other, forming an associated relationship.

[0120] In describing embodiments of the present invention, descriptions of technical content known in the art and not directly related to the present invention may be omitted. Such omissions are intended to more clearly convey the essence of the invention and avoid obscuring core concepts due to unnecessary descriptions.

[0121] Definition of noun

[0122] To facilitate understanding of this invention, the terminology used in this specification is briefly explained below.

[0123] NPU: Short for Neural Processing Unit, it can refer to a processor specifically designed for computing neural network models and independent of the Central Processing Unit (CPU).

[0124] NN: Short for Neural Network, a network in which multiple nodes are connected in a hierarchical structure, mimicking the way neurons in the human brain are connected by synapses in order to simulate human intelligence.

[0125] Neural network information: This information may include the following: network structure information, layer number information, connection relationship information between layers, parameter information of each layer, computation processing method information, activation function information, data type of each layer parameter (e.g., floating-point or integer), and bit width information of each parameter.

[0126] DNN: an abbreviation for Deep Neural Network, which can refer to increasing the number of hidden layers in a neural network to achieve a higher level of artificial intelligence.

[0127] CNN: An abbreviation for Convolutional Neural Network, a type of neural network similar to the human brain's visual cortex in processing images. Convolutional neural networks are particularly well-suited for image processing and are renowned for their ability to extract features from input data and identify feature patterns.

[0128] Transformer: A Transformer-type neural network is a deep neural network based on the attention mechanism. It relies heavily on matrix operations. The Transformer accepts input values ​​and parameters such as a query (Q), a key (K), and a value (V) to obtain an output value, namely the attention weights (Q, K, V). Based on this output value (i.e., the attention weights (Q, K, V)), the Transformer can handle various inference operations.

[0129] Convolutional kernel: refers to the convolutional weights in the form of an NxM matrix. Each layer of a neural network model contains multiple convolutional kernels, and the number of these kernels can also be referred to as the number of channels, the number of filters, etc.

[0130] Neural network models are classified into "single-layer neural networks" and "multi-layer neural networks" based on the number of layers. A typical multi-layer neural network consists of an input layer, hidden layers, and an output layer. (1) The input layer receives external data, and the number of neurons in the input layer is the same as the number of input variables. (2) The hidden layer is located between the input layer and the output layer, receives multiple signals from the input layer, extracts features, and transmits them to the output layer. (3) The output layer receives signals from the hidden layer and outputs them to the outside. The input signals between neurons are multiplied by their corresponding weights, which have values ​​between 0 and 1, and then summed. If the sum is greater than the threshold of the neuron, the neuron is activated and becomes the output value through the activation function.

[0131] On the other hand, increasing the number of hidden layers in a neural network to achieve a higher level of artificial intelligence is called a deep neural network. There are many types of deep neural networks, among which convolutional neural networks are known for extracting features from input data and identifying patterns from those features. A convolutional neural network is a network structure where the operations between neurons in each layer are performed as convolution operations on the input signal matrix and convolution operations on the weight kernel matrix.

[0132] Convolutional neural networks (CNNs) are a type of neural network that functions similarly to the human visual cortex, responsible for image processing. CNNs are renowned for their suitability for image classification, object detection, and other applications. A CNN is constructed by processing convolution operations, activation function operations, and pooling operations in a specific order (e.g., ...). Figure 3 In convolutional neural networks (CNNs), convolution operations account for the majority of computation time. CNNs use matrix-like convolutional kernels to extract image features for each channel, and pooling to provide stability and reduce shifts or distortion. In each channel, the feature map is obtained by convolving the input data with the convolutional kernel, and an activation function is used to generate the activation map for that channel. Pooling can then be applied. The final classification layer is located at the end of the CNN and can be exemplified by a fully connected layer. During the computation of a CNN, most operations are performed through convolution or matrix multiplication.

[0133] However, in order to improve the efficiency and accuracy of neural network models in image classification and object detection, post-processing operations such as additional filtering of output parameters (e.g., feature maps) and removal of duplicate parts can be performed.

[0134] In this case, the post-processing operations described above can be performed on a central processing unit outside the neural processing unit, and the data subsequently processed by the central processing unit can be stored in memory outside the neural processing unit.

[0135] As described above, the bus is used to input multiple output parameters (e.g., multiple feature maps) to a central processing unit outside the neural processing unit, and the inventors of the present invention recognize that transmitting output parameters via the bus may result in data transmission delays.

[0136] The memory outside the neural processing unit (NPU) comprises multiple memory cells, each with a unique memory address. Whenever the NPU calls a feature map, weights, or other parameters stored in main memory, accessing the memory cell corresponding to that address can incur a delay of several clock cycles. These delays may include row address strobe (CAS) latency and column address strobe (RAS) latency. Therefore, the time and power consumption required to read necessary data and parameters (such as weights, feature maps, or convolutional kernels) from external memory to the NPU can be enormous.

[0137] Figure 1 This is a schematic diagram illustrating a neural processing unit 1000 including a post-processing unit 600 according to an embodiment of the present invention. The neural processing unit 1000 may include the post-processing unit 600, and the neural processing unit is coupled to a plurality of peripheral devices. Accordingly, the neural processing unit and the plurality of peripheral devices may be referred to as a system. At least some components of the system may be formed as a system on a chip (SoC).

[0138] Please refer to Figure 1 The neural processing unit 1000 can be used to perform various types of neural network inference functions and communicate with the processor, central processing unit 2000, main memory 3000, image sensor 4000, and decoder 5000. Each of the neural processing unit 1000, central processing unit 2000, main memory 3000, image sensor 4000, or decoder 5000 can be formed with independent circuitry, but is not limited thereto. The neural processing unit 1000 may include circuitry formed on the same semiconductor chip as the central processing unit 2000. Furthermore, the neural processing unit 1000, central processing unit 2000, and main memory 3000 may include circuitry formed on the same semiconductor chip. In addition, the neural processing unit 1000 may include a semiconductor chip connected to the central processing unit 2000 via chiplet technology. When chiplet technology is applied, it may further include an interposer. In other words, the central processing unit 2000 and the main memory 3000 can contain multiple semiconductor chips connected via chiplet technology.

[0139] Each of the aforementioned components can be categorized by the function it performs, and each component can be implemented as a circuit board, silicon substrate, resistor, transistor, etc. Therefore, each component can be a semiconductor circuit with a large number of transistor connections, some of which may be difficult to identify and distinguish with the naked eye, and can only be identified by their operation. Accordingly, Figure 1Each element within can be considered a circuit element.

[0140] Each of the aforementioned central processing unit 2000, main memory 3000, image sensor 4000, and decoder 5000 can communicate via bus 6000 to transmit and receive data to and from the neural processing unit 1000. According to one embodiment of the invention, bus 6000 can be an Advanced Extensible Interface (AXI) bus. However, without limitation, the neural processing unit 1000 can be used to directly couple to at least one of the aforementioned components.

[0141] The neural processing unit 1000 can be defined as a processor specifically designed to perform computations in a neural network-like model. In particular, the neural processing unit 1000 can be specifically designed to handle matrix operations or convolution operations, which account for a large portion of the computation in the neural network-like model.

[0142] The neural processing unit 1000 may include a controller 100, a direct memory access (DMA) unit 200, a memory 300, a processing element array 400, a special function unit (SFU) 500, and a post-processing unit (PPU) 600.

[0143] The components of the neural processing unit 1000 can be distinguished by their functions, and each component can be formed using circuit elements such as resistors and transistors. Therefore, each component can be a semiconductor circuit with a large number of transistor connections.

[0144] Controller 100 can control operations related to the computation of the neural network model via each of direct memory access 200, memory 300, processing element array 400, special function unit 500, and post-processing unit 600. Controller 100 can be directly or indirectly coupled to each of direct memory access 200, memory 300, processing element array 400, special function unit 500, and post-processing unit 600 to communicate with each other. For example, controller 100 can allocate capacity for each parameter in memory 300 based on the capacity of memory 300. Controller 100 can be used to control neural processing unit 1000 based on machine code (e.g., binary code) of a compiled neural network model. For example, compiler 320 can generate machine code based on the hardware characteristics of neural processing unit 1000 (e.g., the number of processing elements, the amount of memory, the functions provided by special function units, whether there are post-processing units, etc.) to determine the following operation sequences: the data read / write order of the neural network model, the processing order of each layer of the neural network, the operation order of convolution multiplication, the operation order of matrix multiplication, and the read and write operation order of direct memory access. Accordingly, controller 100 can control neural processing unit 1000 based on machine code.

[0145] The controller 100 can obtain scheduling information that plans the order in which the neural processing unit 1000 executes operations on the neural network model. This scheduling information is generated based on the directed acyclic graph of the neural network model compiled by the compiler 3010 executed by the central processing unit 2000. The compiler 3010 can determine the operation schedule that can accelerate the operation of the neural network model by judging the number of processing elements (PEs) of the neural processing unit 1000, the size of the memory 300, and the size of the parameters of each layer of the neural network model. According to the operation schedule, the controller 100 can control the number of processing elements required for each calculation step and control the read and write operations of the required parameters in the memory 300 in each calculation step. The compiler 3010 can effectively schedule operations based on the hardware structure and performance information of the neural processing unit 1000. The compiler 3010 can determine the data locality based on the order of the layers of the neural network, the operation order of unit convolution and / or matrix multiplication, and generate compiled machine code based on the order of data required for computing the neural network model.

[0146] In some embodiments, the neural processing unit 1000 may include an embedded compiler 320. In addition to the external compiler 3010, the embedded compiler 320 may perform some computations, or the embedded compiler 320 may replace the external compiler 3010 in performing some computations. According to the above configuration, the compiler 3010 and / or compiler 320 of the neural processing unit 1000 can generate machine code after inputting files in various AI software frame formats. For example, the AI ​​software framework may include TensorFlow, PyTorch, Keras, XGBoost, mxnet, DARKNET, ONNX, etc.

[0147] Direct memory access 200 allows the neural processing unit 1000 to directly access, read, and / or write to its main memory 3000. The neural processing unit 1000 can then read various data related to the neural network model from the main memory 3000 via direct memory access 200. The main memory 3000 can be embedded within a system-on-a-chip or configured as a separate memory device.

[0148] Memory 300 may be located in the on-chip region of the neural processing unit 1000 and may perform caching or store data processed in the on-chip region. Memory 300 may also be referred to as cache memory. Memory 300 may read and store at least some data related to the computation of the neural network model from the main memory 3000. Memory 300 may be used to store all or part of the neural network model, depending on the memory capacity settings for each parameter and the data size of each layer of the neural network model. Among other data, several representative parameters of the data processed in the neural network model may include attention parameters, key-value cache, activation maps, input feature maps, output feature maps, and weights. Specifically, memory 300 may read and store parameters corresponding to input data from the main memory 3000. Furthermore, memory 300 may read and store parameters corresponding to output data from the processing element array 400.

[0149] The memory 300 can be implemented as one or more read-only memory (ROM), static random access memory (SRAM), dynamic random access memory (DRAM), resistive random access memory (RRAM), magneto-resistive RAM (MRAM), phase-change RAM (PRAM), ferroelectric RAM (FRAM), flash memory, and high-bandwidth memory (HBM), etc. According to one embodiment of the present invention, the memory 300 can be implemented as static random access memory, which has advantages in computing speed. Furthermore, the memory 300 can be organized into at least one memory cell (e.g., a memory bank). The memory 300 can include homogeneous memory or heterogeneous memory.

[0150] The data stored in the memory cells of memory 300 is not static but can be dynamically changed. By changing the allocation of memory in memory 300 to different types of parameters and data, the utilization rate of memory 300 can be improved. In addition, the size of the data of each type of parameter stored in memory 300 can change due to each stage of operation.

[0151] The processing element array 400 is a hardware circuit that performs multiplication and accumulation (MAC) operations. The processing element array 400 can be used to receive input feature maps as input data and / or convolutional kernels corresponding to one, some, or multiple layers of a neural network. The processing elements within the processing element array 400 can be used to perform various operations such as addition, multiplication, accumulation, etc., as well as operations defined by the neural network model. Among other elements, the processing elements may include a multiplication and accumulation unit and an arithmetic logic unit (ALU).

[0152] In one embodiment, the processing element may receive an input feature map or a portion thereof, perform a convolution operation using a convolution kernel, and output an output feature map or a portion thereof. The processing element array 400 or the processing element may also be referred to as an Artificial Intelligence (AI) computing unit. In another embodiment, the processing element may perform a General Matrix Multiply (GEMM) operation or matrix multiplication operation on the input feature map using weights to output an output feature map or a portion thereof. More specifically, the processing element may multiply the input feature map, which is in matrix form, with a weight matrix and add a bias value to the matrix to output an output feature map or a portion thereof in matrix form. Within the neural processing unit, matrix multiplication can be performed at high speed using parallel processing, thereby achieving efficient processing of matrix multiplication operations.

[0153] The processing element can contain circuit designs that can only process integer type parameters as input. In this case, the input parameters of the processing element can be converted into integers of a specific bit width and stored in memory 300. This type of processing element can reduce power consumption compared to processing elements that support floating-point numbers and can be more easily implemented as an on-device component.

[0154] Special function unit 500 can process various activation functions to impart a non-linear relationship to the output feature map. The activation functions processed by special function unit 500 may include, but are not limited to, the SiLU function, Softmax function, sigmoid function, hyperbolic tangent (tanh) function, rectified linear unit (ReLU) function, Leaky ReLU function, Maxout function, or exponential linear unit function, causing the output value to exhibit a non-linear relationship with the input value. Supporting all activation functions in neural processing unit 1000 may present technical difficulties. Therefore, neural processing unit 1000 can approximate various activation functions using piecewise linear function approximation algorithms and piecewise linear function processing circuitry. These activation functions can be selectively applied after multiplication-accumulation operations. The result of the operation after applying the activation functions is called an activation map.

[0155] In some embodiments, the special function unit 500 may be configured to include floating-point multiplication circuitry to perform decimal arithmetic. In other embodiments, the special function unit 500 may be used to communicate with a processing element and may include circuitry to receive integer-type parameters from the processing element. In this case, the special function unit 500 may be configured to include dequantization circuitry for converting the integer-type parameters to floating-point-type parameters. The special function unit 500 may use the floating-point-type parameters to process activation function operations. Furthermore, the special function unit 500 may be further configured to include quantization circuitry for converting the floating-point-type parameters to integer-type parameters at the end of the activation function operation. According to the above configuration, when floating-point operations are required, the special function unit 500 may be used to process floating-point operations by dequantizing the integer parameters and requantizing the results. In other words, according to an embodiment of the present invention, a neural processing unit may include processing element circuitry for processing integer-type parameters and a special function circuitry unit connected in series therewith, including quantization and dequantization circuitry, and may be used to process operations of activation functions with floating-point-type parameters. According to the above configuration, the special function unit 500 can communicate effectively with the processing element that only supports integer parameters, and can directly convert and process integer parameters without the need for circuit support outside the neural processing unit.

[0156] In some embodiments, the post-processing unit 600 can be used to process a variety of activation functions to impart nonlinear relationships to the output feature map.

[0157] Figure 2 A schematic diagram illustrating a processing element according to an embodiment of the present invention is provided. Please refer to... Figure 2 In addition to other components, the processing element may include a multiplier 641, an adder 642, an accumulator 643, and a bit quantization unit 644. The processing element can be adjusted according to the computational characteristics of the target neural network model. Figure 2 Various modifications and adjustments were made to the processing components.

[0158] Multiplier 641 is a circuit used to multiply N-bit data and M-bit data from input. The output of multiplier 641 is N+M-bit data, where N and M are integers greater than 0. The first input receives dynamically changing N-bit data, and the second input receives relatively fixed M-bit parameter data. For example, a set of weight parameters trained within a neural network model can be fixed, while the processing element processes the same layer of the neural network while input parameters (such as activation parameters, feature map parameters, attention parameters, and KV cache parameters) calculated using this set of weight parameters can frequently change relative to this set of weight parameters.

[0159] Variable parameters mean that the parameters are updated each time the input data of the neural network is updated. For example, the node data of each layer can be a multiplicative sum of the weight data of the neural network model, where the node data of each layer in the neural network changes as each frame of the input video changes. Static parameters mean that the parameters remain unchanged regardless of whether the input data is updated. For example, if the neural network model is used to infer object detection from video data, the weight data can remain fixed.

[0160] The multiple variable parameters fed to the first input can be node data from a layer of a neural network model. The node data of the neural network model can be one of the following: input data from the input layer, multiple accumulated values ​​from the hidden layer, or multiple accumulated values ​​from the output layer. The multiple fixed parameters fed to the second input can be weight data from the connection network of the neural network model.

[0161] The controller 100 can improve memory reuse by taking into account the nature of fixed parameters. Variable parameters are calculated values ​​for each layer, and the controller 100 can identify reusable variable parameters based on the machine code of the compiled neural network-like model and control the memory 300 to reuse memory.

[0162] Fixed parameters are the weight data for each connection network, and the controller 100 can identify the fixed parameters of the reusable connection networks based on the structural data of the neural network model or the data locality information of the neural network, and can control the memory 300 to reuse the parameters stored in the memory 300. Parameter reuse means that the parameters stored in the memory 300 are not deleted, copied, or moved to the main memory 3000, but are reused in subsequent operations. According to the above configuration, as... Figure 4 As shown, this configuration effectively reduces the power consumption of the main memory 3000. Furthermore, it eliminates the latency caused when the neural processing unit 1000 transmits data to and from the main memory 3000. The controller 100 may have information on reusable variable parameters and fixed parameters based on the compiled neural network model machine code. Accordingly, the controller 100 can be used to control the memory 300 to reuse parameters stored in the memory.

[0163] The processing element can restrict the operation of multiplier 641 so that multiplier 641 does not perform an operation when either the first or second input is zero, because the processing unit knows that the result will be zero even if the operation is not performed. For example, when zero is input to either the first or second input of multiplier 641, multiplier 641 can be used to operate in a zero-skipping mode.

[0164] For the zero-skip mechanism, each processing element included in the processing element array 400 can be individually enabled or disabled. The controller 100 can be used to provide an enable or disable signal to each processing element in clock cycles. When a processing element is disabled, the multiplier 641 can be disabled based on the potential of the first enable signal En1. Accordingly, the power consumption of the multiplier 641 can be reduced. For example, power consumption information for the multiplier can be found in [reference needed]. Figure 4 .

[0165] For the zero-skip mechanism, each processing element in the processing element array 400 can be individually enabled or disabled. The control unit 100 can be used to provide an enable or disable signal to each processing element on a clock cycle basis. When a processing element is disabled, the adder 642 can be disabled based on the potential of the second enable signal En2. Accordingly, the power consumption of the adder 642 can be reduced. For example, information regarding the adder's power consumption can be found at [reference needed]. Figure 4 In some embodiments, each processing element may be designed to receive a corresponding control signal from the control unit 100 to control (i.e., enable or disable) the zero-skip operation.

[0166] In some embodiments, each multiplier 641 of each processing element may receive a corresponding control signal from the controller 100 to control zero-skip operations. According to the above configuration, the power consumption of the multipliers can be reduced by zero-skip operations.

[0167] In some embodiments, each adder 642 of each processing element can be designed to receive a corresponding control signal from the controller 100 to control zero-skip operations. According to the above configuration, the power consumption of the adders can be reduced by zero-skip operations.

[0168] In some embodiments, each multiplier 641 and each adder 642 of each processing element can be designed to simultaneously receive corresponding control signals from the controller 100 to control zero-skip operations. According to the above configuration, the power consumption of the multipliers and adders can be reduced by zero-skip operations.

[0169] In some embodiments, the weights are fixed parameters generated through training, and the machine code of the compiled neural network model containing the weights can be programmed to input corresponding control signals on the processing element that inputs zero weight values ​​to control zero-skip operations.

[0170] The number of bits of data input to the first and second input terminals can be determined based on the quantization results of the node data and weight data of each layer in the neural network model. For example, the node data of the first layer can be quantized to 5 bits, and the weight data of the first layer can be quantized to 7 bits. In this case, the first input terminal can be used to receive 5 bits of data, and the second input terminal can be used to receive 7 bits of data; that is, the number of bits of data input to each input terminal can be different.

[0171] The processing element can be used to receive quantized information of the data input to each input terminal. The data locality information of a neural network can include quantized information of the input and output data of the processing element.

[0172] The neural processing unit 1000 can control the real-time conversion of the quantization bit width when quantized data stored in the memory 300 is input to the processing element. That is, different layers can have different quantization bit widths, and the processing element can be used to generate input data by receiving bit width information from the neural processing unit 1000 in real time and converting the bit width of the input data in real time.

[0173] Accumulator 643 uses adder 642 to perform L loops to accumulate the values ​​of multiplier 641 and accumulator 643. Therefore, the number of input and output data bits for accumulator 643 can be N + M + log2(L) bits, where L is a positive integer. After accumulator 643 completes accumulation, it can receive an initialization reset signal to reset the data stored in accumulator 643 to zero. However, embodiments of the present invention are not limited to this. Accumulator 643 is used to store accumulated values ​​even when zero skipping is enabled in the corresponding processing element. Therefore, even if the zero skipping mechanism is enabled, subsequent values ​​can still be accumulated.

[0174] The bit quantization unit 644 can be used to reduce the bit width of the output data from the accumulator 643. The bit quantization unit 644 can be controlled by the controller 100. The quantized data bit width can be output as X bits, where X is a positive integer. According to the above configuration, the processing element array is used to perform multiply-accumulate operations, and the processing element array can quantize and output the result of the multiply-accumulate operation. This quantization can further reduce power consumption as the number of loops increases by L. Reducing power consumption also reduces heat generation in the edge device. Furthermore, reducing heat generation helps to reduce the possibility of operational failures due to high temperatures in the neural processing unit 1000.

[0175] The X bits of output data from the bit quantization unit 644 can be used as node data for subsequent layers or as input data for convolution operations. If the neural network model has been quantized, the bit quantization unit 644 can be used to receive quantization information from the neural network model. However, the controller 100 can also be used to analyze the neural network model to extract quantization information. Therefore, the X bits of output data can be converted into the number of quantization bits corresponding to the size of the quantized data. The X bits of output data from the bit quantization unit 644 can be used as the quantization bit width stored in the memory 300.

[0176] According to one embodiment of the present invention, the processing element array of the neural processing unit 1000 includes a multiplier 641, an adder 642, an accumulator 643, and a bit quantization unit 644. The bit quantization unit 644 can reduce N+M+log2(L) bit data output by the accumulator 643 in the processing element array to X bit data. The controller 100 can control the bit quantization unit 644 to reduce the number of bits of output data by a predetermined number of bits from the least significant bit (LSB) to the most significant bit (MSB). Reducing the number of bits of output data can effectively reduce power consumption, computational load, and memory usage. However, if the number of bits is reduced below a certain length, the inference accuracy of the neural network model may drop sharply. Therefore, the quantization level (in other words, the extent of reduction in the number of bits of output data) can be determined by comparing the degree of reduction in power consumption, computational load, and memory usage with the degree of decrease in the inference accuracy of the neural network model. Quantization levels can also be determined by setting a target inference accuracy for the neural network model and gradually decreasing the bit width to test the inference accuracy. The quantization level can be determined separately for each layer of the neural network model.

[0177] By adjusting the number of bits of the N-bit and M-bit data of the multiplier 641 and reducing the X-bit operation value through the bit quantization unit 644, the processing element array can increase the multiply-accumulate instruction cycle while reducing power consumption, and also has the advantage of making the convolution operation of neural network-like models more efficient.

[0178] Figure 3 This is a schematic diagram illustrating a convolutional neural network according to the present invention. A convolutional neural network can be a combination of one or more convolutional layers, pooling layers, and fully connected layers. Convolutional neural networks have a structure suitable for learning and inference from two-dimensional data and can be trained using the backpropagation algorithm.

[0179] In one embodiment of the invention, the convolutional neural network has a convolutional kernel for each channel that extracts multiple features of the input image for each channel. The convolutional kernels can be organized as a two-dimensional matrix and perform convolution operations as they traverse the input data. The size of the convolutional kernels can be arbitrary, and the stride of the convolutional kernel traversing the input data can also be arbitrary. The result of each convolutional kernel performing convolution on the entire input data can be called a feature map or activation map.

[0180] Below, a convolutional kernel can contain a single set of weights or multiple sets of weights. The number of convolutional kernels in each layer can be referred to as the number of channels.

[0181] Since convolution is a combination of input data and a convolution kernel, activation functions can be applied to add non-linearity. When an activation function is applied to the feature map of the result of a convolution operation, it can be called an activation map.

[0182] Specifically, please refer to Figure 3 Convolutional neural networks can contain at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. For example, convolution can be defined by two main parameters: the size of the input data (typically a 1×1, 3×3, or 5×5 matrix) and the depth of the output feature map (i.e., the number of convolutional kernels). These key parameters can be computed using convolution. The depth of these convolutions can start at 32, continue to 64, and end at 128 or 256. A convolution operation can represent sliding a 3×3 or 5×5 convolutional kernel across the input image matrix, multiplying each weight of the convolutional kernel by each overlapping element in the input image matrix, and then summing these products.

[0183] Activation functions can be applied to the output feature maps generated in this way to ultimately output activation maps. Furthermore, the weights used in the current layer can be passed to subsequent layers via convolution. Pooling layers can perform pooling operations to reduce the size of the feature maps by downsampling the output data (i.e., the activation maps). For example, pooling operations can include, but are not limited to, max pooling and / or average pooling.

[0184] Max pooling uses a convolution kernel and outputs the maximum value within the range of the feature map overlapping with the kernel by sliding the feature map and kernel. Average pooling outputs the average value within the range of the feature map overlapping with the kernel by sliding the feature map and kernel. Therefore, since the size of the feature map is reduced by the pooling operation, the number of weights in the feature map is also reduced accordingly.

[0185] Fully connected layers can classify the data output from pooling layers into multiple categories (i.e., inference values) and output the classified categories and their corresponding scores. The data output from pooling layers forms a three-dimensional feature map, which can be converted into a one-dimensional vector and input to the fully connected layer.

[0186] Please refer to Figure 1 According to one embodiment of the present invention, the neural network-like model processed by the neural processing unit 1000 can be associated with image classification and object detection. The input data of the processing element array 400 of the neural processing unit 1000 processing the neural network-like model can be image data, and the output data of the processing element array 400 can be bounding box data mapped from the input image. Each plurality of bounding box data can include bounding box coordinate data and category data. The bounding box coordinate data can include height data, width data, x-data, and y-data.

[0187] As mentioned above, assuming the bounding box is rectangular, its coordinate data includes height, width, x-axis, and y-axis data. However, the shape of the bounding box is not limited to rectangles; it can also be pentagonal, more polygonal, or circular. Accordingly, the quantity and type of the bounding box coordinate data can vary depending on the shape of the bounding box.

[0188] In addition, categorical data can include classifications into multiple categories that exist within the bounding box, along with their corresponding scores.

[0189] Figure 4 This is a schematic diagram illustrating the power consumption of each unit of a neural processing unit according to an embodiment of the present invention. Hereinafter, Figure 4 This will be used to explain the power reduction technology of the memory 300 in the neural processing unit 1000. Please refer to [link / reference]. Figure 4 This figure is a table summarizing the energy consumed per unit operation of the neural processing unit 1000. Energy consumption can be divided into memory access, addition operations, and multiplication operations.

[0190] "8b Add" refers to adder 642 performing 8-bit integer addition. 8-bit integer addition consumes 0.03 pj (picojoules) of energy. "16b Add" refers to adder 642 performing 16-bit integer addition. 16-bit integer addition consumes 0.05 pj of energy. "32b Add" refers to adder 642 performing 32-bit integer addition. 32-bit integer addition consumes 0.1 pj of energy. "16b FP Add" refers to adder 642 performing 16-bit floating-point addition. 16-bit floating-point addition consumes 0.4 pj of energy. "32b FP Add" refers to adder 642 performing 32-bit floating-point addition. 32-bit floating-point addition consumes 0.9 pj of energy. "8b Multi" refers to multiplier 641 performing 8-bit integer multiplication. An 8-bit integer multiplication operation consumes 0.2 pJ of energy. "32bMult" refers to multiplier 641 performing 32-bit integer multiplication. 32-bit integer multiplication consumes 3.1 pJ of energy. "16bFPMult" refers to multiplier 641 performing 16-bit floating-point multiplication. 16-bit floating-point multiplication consumes 1.1 pJ of energy. "32bFPMult" refers to multiplier 641 performing 32-bit floating-point multiplication. 32-bit floating-point multiplication consumes 3.7 pJ of energy. "32b SRAM Read" refers to the access operation of reading 32 bits of data when memory 300 is static random access memory. Reading 32 bits of data from memory 300 consumes 5 pJ of energy. "32b DRAM Read" refers to the energy consumption of reading 32 bits of data from main memory 3000 to memory 300 when main memory 3000 is dynamic random access memory. All energy units mentioned above are in picojoules (pJ).

[0191] When the neural processing unit 1000 performs a 32-bit floating-point multiplication operation, the energy consumption per unit operation differs by approximately 18.5 times compared to an 8-bit integer multiplication operation. When reading 32-bit data from main memory 3000 (configured as dynamic random access memory), the energy consumption per unit access operation differs by approximately 128 times compared to reading 32-bit data from memory 300 (configured as static random access memory). In other words, from a power consumption perspective, power consumption increases with the number of bits of data. Furthermore, floating-point operations consume more energy than integer operations. Moreover, reading data from dynamic random access memory significantly increases power consumption.

[0192] Therefore, the memory 300 of the neural processing unit 1000 can be configured to include high-speed static memory, such as static random access memory, and not include dynamic random access memory. However, according to an embodiment of the present invention, the neural network processing unit is not limited to static random access memory. For example, the memory 300 may not include dynamic random access memory, and the memory 300 may be configured to include static memory to have relatively high read and write speeds and consume less power than the main memory 3000. Accordingly, according to an embodiment of the present invention, the memory 300 of the neural processing unit 1000 can be configured to have relatively high read and write speeds in the inference operations of the neural network model and consume less power than the main memory 3000.

[0193] High-speed static memory, such as static random access memory (SRAM), can include SRAM, magnetoresistive random access memory (MRRAM), spin-transfer torque magnetoresistive random access memory (STT-MRAM), embedded magnetic random access memory (eMRAM), and orthogonal spin-transfer magnetic random access memory (OST-MRAM). Furthermore, magnetoresistive RRAM, spin-transfer torque magnetoresistive RRAM, embedded magnetic random access memory, and orthogonal spin-transfer magnetic random access memory are all static memory types and possess non-volatile characteristics. Therefore, high-speed static memory, such as SRAM, can effectively avoid the redundant requirement of providing additional memory in the main memory 3000 for restarting due to power failure. However, embodiments of the present invention are not limited to this.

[0194] According to the above configuration, the neural processing unit 1000 can reduce the power consumption of the dynamic random access memory during the inference operation of the neural network model. Furthermore, the memory cell of the static random access memory of the memory 300 may include, for example, four to six transistors to store one bit of data. However, embodiments of the present invention are not limited thereto. Furthermore, the memory cell of the magnetoresistive random access memory of the memory 300 may include, for example, a magnetic tunnel junction (MTJ) and a transistor to store one bit of data. However, embodiments of the present invention are also not limited thereto.

[0195] The following describes in detail the specific configuration and operation of the post-processing unit included in the neural processing unit, according to an embodiment of the present invention. Figure 5 This is a schematic diagram of a post-processing unit according to an embodiment of the present invention. Please refer to... Figure 5 In addition to other components, the post-processing unit 600 in one embodiment of the present invention may include a first arithmetic unit 610, a second arithmetic unit 620, an internal processing unit 640, and an internal memory 630.

[0196] The first processing unit 610 can extract the highest-scoring category from multiple categories associated with a given bounding box. The first processing unit 610 can perform a class-argmax operation to extract the index of the highest-scoring category within the bounding box and its corresponding category score. For each category corresponding to an object, the category score represents the probability that the object appears within the bounding box.

[0197] The second operation unit 620 can extract bounding boxes from multiple bounding boxes that contain only those whose category confidence scores are higher than a threshold confidence score. The category confidence score represents the probability or confidence that an object of a specific category appears within the bounding box. This category confidence score is the product of the object presence confidence score and the category score. The object presence confidence score indicates the probability that an object exists within the bounding box, regardless of the object's category. The second operation unit 620 performs bounding box filtering operations to extract only bounding boxes whose object presence confidence score and the product of the category score extracted from the first operation unit 610 are higher than a specific threshold confidence score.

[0198] The internal processing unit 640 can post-process the bounding box data extracted from the second arithmetic unit 620; that is, the internal processing unit 640 can decode the extracted bounding box data. Furthermore, the internal processing unit 640 can perform nonmaximum suppression operations on the extracted bounding box data.

[0199] The internal memory 630 can store the data required for the post-processing unit 600 to perform operations. That is, the internal memory 630 can store input or output data from the first arithmetic unit 610, the second arithmetic unit 620, and the internal processing unit 640.

[0200] Please refer to Figure 5The internal memory 630 may contain multiple memory partitions (e.g., DATA, OUTPUT1, OUTPUT2, and Code). A portion of the multiple memory partitions (DATA) may store multiple bounding box data output from the internal processing unit 640. Another portion of the multiple memory partitions (OUTPUT1, OUTPUT2) may store multiple bounding box data received from the first arithmetic unit 610 and the second arithmetic unit 620. Another portion of the multiple memory partitions (Code) may store program code data related to post-processing operations in the internal processing unit 640. However, the data stored in the multiple memory partitions is not limited to the above, and various types of data may be stored as needed.

[0201] Meanwhile, the inputs and outputs of the internal processing unit 640 can be transmitted via the Advanced High-Performance Bus (AHB). The AHB is a high-performance bus protocol primarily used in system-on-a-chip (SoC) designs, offering advantages such as low power consumption and scalability, thereby improving system reliability and operating efficiency.

[0202] Figure 6 This is a schematic diagram illustrating the computational process of a neural processing unit including a post-processing unit according to an embodiment of the present invention. For ease of explanation, reference will be made to... Figure 1 as well as Figure 5 The structure of the neural processing unit 1000, which includes a post-processing unit 600, is shown.

[0203] According to an embodiment of the present invention, the operation process S100 may include an activation function operation step S110, a class-argmax operation step S120, a filtering operation step S130, a decoding operation step S140, and a non-maximum suppression operation step S150.

[0204] In the activation function operation step S110, the special function unit 500 can process multiple activation functions to give the output feature map a nonlinear relationship.

[0205] The activation function processed by the special function unit 500 may include, but is not limited to, the SiLU function, the Softmax function, the sigmoid function, the hyperbolic tangent (tanh) function, the modified linear unit function, the Leaky ReLU function, the Maxout function, or the exponential linear unit function, such that the output value has a non-linear relationship with the input value.

[0206] On the other hand, not all activation functions can be supported by the neural processing unit 1000. Therefore, the neural processing unit 1000 can be programmed to approximate various activation functions using a piecewise linear function approximation algorithm and a piecewise linear function processing circuit. These activation functions can be selectively applied after multiplication-accumulation operations. The operation value after applying the activation function can be called an activation map.

[0207] The following describes in detail a procedural activation function method that enables the neural processing unit 1000 to approximate various activation functions through a piecewise linear function approximation algorithm and a piecewise linear function processing circuit.

[0208] Figure 7 A flowchart illustrating a procedural activation function method according to an embodiment of the present invention is provided. Please refer to... Figure 7 The activation function programming method includes the steps of generating segment data to segment activation functions, S200, dividing the activation function into multiple segments using the generated segment data, and S220 approximating at least one of the multiple segments as a programmable segment.

[0209] In step S200, segment data is generated. Segment data is data used to segment the activation function into multiple segments. In step S210, the generated segment data is used to segment the activation function into multiple segments. In this invention, a "segment" refers to a portion of the activation function into which multiple segments are divided, and can be distinguished from "candidate segments" or "programmable segments," the latter being related to the approximate processing of the activation function.

[0210] In various examples, step S210 may include determining the number and width of multiple segments based on segment data. In step S210, the segment data can be used to determine the number of segments into which the activation function to be transformed will be divided and the width of each segment. The width of at least one of these segments may be the same as or different from the width of the other segments.

[0211] In this invention, a segment of a plurality of segments can be represented as the coordinates of the start and end points along the x-axis. Furthermore, when the number and width of each of the plurality of segments are determined, the coordinates of each segment can be obtained using the number and width of the plurality of segments.

[0212] In step S220, at least one of the plurality of segments is approximated as a programmable segment. The programmable segment can be programmed according to the hardware configuration of the special function unit 500. That is, based on the hardware configuration of the special function unit 500, it can be configured to program activation functions intended to be processed by the neural processing unit 1000 into programmed activation functions (PAFs). For example, the special function unit 500 can be configured to have hardware capable of operating each programmable segment with a specific slope and a specific offset.

[0213] In this case, the special function unit 500 can program the programmable segment in the form of a quadratic or first-order function having at least a slope and an offset. For example, the programmable segment can be approximated as a first-order function according to a specific criterion. In this case, the special function unit 500 can generate a programmable segment expressed in the form of "(slope a)*(input value x)+(offset b)". The specific slope and specific offset mentioned above can be programmable parameters. For a programmable segment that is determined to be approximated as a first-order function, step S220 may include approximating the selected segment using a specific slope and a specific offset.

[0214] Furthermore, in some examples, steps S210 and S220 can be performed simultaneously. Additionally, in some examples, steps S210 and S220 can be modified to include the steps of segmenting the activation function into multiple segments using the generated segment data and approximating at least one of the multiple segments as a programmable segment.

[0215] Figures 8A to 8C This is a schematic diagram illustrating the process of approximating an activation function using a procedural activation function method according to an embodiment of the present invention. (Representative) Figure 8A The activation function's line is like Figure 8B The data shown can be divided into multiple segments s1, s2, s3, and s4. These segments s1, s2, s3, and s4 are approximated as programmable segments a1x+b1, a2x+b2, a3x+b3, and a4x+b4, as shown below. Figure 8C As shown. In this example, special processing unit 500 generates programmable parameters such that all programmable segments correspond to the first function.

[0216] Each programmable fragment can contain corresponding programmable parameters. For example... Figure 8C As shown, all multiple fragments can be approximated as programmable fragments of the form of a linear function. However, in various examples, some fragments among multiple fragments can also be approximated as other types of programmable fragments.

[0217] The special processing unit 500 can program each programmable segment into a quadratic function, cubic function, logarithmic function, etc. For example, only segments s1, s2, s3, and s4 can be approximated as programmable segments, where segment s2 can be approximated using various methods available on the device for processing activation functions. Specifically, if pre-determined and stored lookup tables, nonlinear approximations, etc., are available for segment s2, segment s2 can be approximated using these pre-determined and stored lookup tables, nonlinear approximations, etc. In other words, the special processing unit 500 can independently program each segment s1, s2, s3, and s4.

[0218] The special processing unit 500 can independently determine the approximation method for each segment s1, s2, s3, and s4 based on hardware configuration information. For example, the special function unit 500 can be configured to include circuitry supporting the computation of first-order functions. In this case, the special function unit 500 can program each segment s1, s2, s3, and s4 as a first-order function. For example, the special function unit 500 can also be configured to include circuitry supporting the computation of both first-order and second-order functions. In this case, the special function unit 500 can program each segment s1, s2, s3, and s4 as either a first-order or second-order function.

[0219] Special function unit 500 can be configured to include circuitry supporting first-order, second-order, and logarithmic functions. In this case, special function unit 500 can selectively program each segment s1, s2, s3, and s4 as a first-order, second-order, or logarithmic function. For example, special function unit 500 can also be configured to include circuitry supporting first-order, second-order, logarithmic, and exponential function operations. In this case, special function unit 500 can selectively program each segment s1, s2, s3, and s4 as a first-order, second-order, logarithmic, or exponential function.

[0220] When the special function unit 500 is configured to include circuitry supporting at least one specific function operation, the special function unit 500 can program each segment s1, s2, s3, and s4 in the form of a corresponding specific function. For example, the special function unit 500 can be configured to include at least one hardware design of a first-order function calculation circuit, a second-order function calculation circuit, a third-order function calculation circuit, a logarithmic function calculation circuit, an exponential function calculation circuit, or a similar function calculation circuit.

[0221] Special function unit 500 can be programmed with specific activation functions using different techniques.

[0222] Alternatively, the special function unit 500 may program only specific activation functions as first-order functions. For example, the special function unit 500 may program only specific activation functions as second-order functions.

[0223] In other embodiments, the special function unit 500 may simply program a specific activation function as a third-order function, a logarithmic function, or an exponential function.

[0224] Special function unit 500 can program each of the multiple segments of a specific activation function into a corresponding approximation function. For example, special function unit 500 can program multiple segments of a specific activation function into a set of approximation functions with different formulas.

[0225] Figures 9A to 9D To illustrate a schematic diagram according to an embodiment of the present invention, various cases are shown where the activation function is segmented into multiple fragments using a procedural activation function method. Please refer to... Figure 9A This means that the activation function line can be divided into four segments of equal width. On the other hand, please refer to... Figure 9B This indicates that the activation function line can be segmented into four segments of different widths. Similarly, please refer to... Figure 9C This indicates that the activation function line can be divided into four segments of different widths. Please refer to [reference needed]. Figure 9D This indicates that the activation function line can be divided into six segments of different widths. The number of these segments and the width of each segment can be determined using segment data.

[0226] Special function unit 500 can be used to analyze the nonlinearity of the activation function to segment multiple segments into different widths. Special function unit 500 can also analyze the nonlinearity of the activation function and segment each of the multiple segments into an optimal width. However, the invention is not limited thereto.

[0227] In this invention, the activation function can be implemented in various forms that include feature segments. When the segmented activation function consists of multiple segments, the number and width of the segments can vary depending on the different forms the activation function takes.

[0228] For example, various activation functions, such as SiLU, Softmax, swish, Mish, sigmoid, hyperbolic tangent (tanh), SELU, Gaussian Error Linear Unit (GELU), SOFTPLUS, modified linear unit, LeakyReLU, Maxout, and exponential linear unit, have various shapes that include substantially linear and / or nonlinear intervals and are segmented into multiple characteristic intervals. Therefore, when approximating a nonlinear activation function in a hardware-processable manner, considering these characteristic intervals for segmentation allows for a more efficient or closer approximation of the activation function to correspond to the characteristics of each activation function. For example, the number and width of segments can be determined by considering substantially linear intervals and nonlinear intervals.

[0229] Accordingly, in the approximate activation function method of the present invention, the concept of segment data considers the characteristic interval of the activation function as being used to segment the activation function. The segment data includes discontinuity information of the activation function, derivative data, hardware information for processing the activation function, and data processed therefrom.

[0230] Please refer to Figures 10A to 12B This describes an example of using discontinuity information in segment data to segment an activation function into multiple segments. Figures 10A to 10C This is a schematic diagram illustrating an example of how, according to an embodiment of the present invention, an activation function programming method divides the activation function into linear and nonlinear intervals using slope change data of segment data.

[0231] The point where the slope of the activation function changes can refer to the point where the slope of the activation function changes. For example, special function unit 500 can be used to generate slope change data (e.g., differential data) to analyze the point where the slope of the activation function changes. However, the slope change data in this invention is not limited to differential data and may include other similar data.

[0232] According to embodiments of the present invention, the slope change data may include the nth derivative of the activation function, such as the first, second, and third derivatives. The gradient change data may represent the gradient rate of change and gradient change points associated with the activation function. Furthermore, the slope change points may refer to points where the slope change data is discontinuous (d1, d2, and d3), meaning that the slope of the activation function necessarily changes at these points (d1, d2, and d3). Accordingly, the slope change points in this invention may refer to points where the nth derivative of the activation function is discontinuous, such as the first, second, and third derivatives.

[0233] Figure 10B illustrate Figure 10A The first derivative f'(x) of the differential data of the activation function f(x) is shown. Figure 10C illustrate Figure 10A The second derivative f(x) of the differential data of the activation function f(x) is shown.

[0234] For example, special function unit 500 can be used to extract the first derivative value without changing the start and end points of the interval. Figure 10B As shown, special function unit 500 generates slope change data corresponding to the first derivative value. Furthermore, special function unit 500 determines that the first derivative values ​​in each of the intervals w2 and w3 are different from each other, but the first derivative values ​​themselves do not change. Based on this, special function unit 500 can determine that each of the intervals w2 and w3 is a linear interval, meaning that the slope change data corresponding to the first derivative value does not change within these linear intervals. However, since the first derivative values ​​in each of the intervals w2 and w3 are different from each other, the slope change data corresponding to the first derivative value has discontinuities d1 and d2 at the boundaries of each of the intervals w2 and w3. In other words, the slope change data corresponding to the first derivative value at the boundaries of each of the intervals w2 and w3 are discontinuous points; therefore, the boundaries of each of the intervals w2 and w3 can correspond to slope change points.

[0235] For example, special function unit 500 can be used to extract the start and end points of an interval where the first derivative value remains unchanged. Figure 10B As shown, special function unit 500 generates slope change data corresponding to the first derivative value. Furthermore, special function unit 500 determines that the first derivative values ​​in each of the intervals w2 and w3 are different from each other, but the first derivative values ​​themselves do not change. Accordingly, special function unit 500 can determine that each of the intervals w2 and w3 is a linear interval, meaning that the slope change data corresponding to the first derivative value does not change within these linear intervals. However, since the first derivative values ​​in each of the intervals w2 and w3 are different from each other, the slope change data corresponding to the first derivative value has discontinuities d1 and d2 at the boundaries of each of the intervals w2 and w3. In other words, the slope change data corresponding to the first derivative value at the boundaries of each of the intervals w2 and w3 are discontinuous points; therefore, the boundaries of each of the intervals w2 and w3 can correspond to slope change points.

[0236] In this scenario, the special function unit 500 can convert the linear interval into programmable parameters in the form of a corresponding first-order function. Therefore, the linear interval of the activation function to be programmed can be segmented into first-order functions with specific slopes and offsets. The first derivative of the linear interval can be a constant value. Furthermore, the linear interval can be approximated by a first-order function to achieve zero approximation error. Therefore, the special function unit 500 can determine that there is substantially no approximation error in each of the intervals w2 and w3. In other words, when the special function unit 500 uses a first-order function to approximate each of the intervals w2 and w3, the approximation error can be zero, while minimizing computational complexity and the power consumption of the special function unit 500.

[0237] Special function unit 500 can be used to determine the interval in which the first derivative of the activation function is constant or non-zero, and the interval in which the second derivative is greater than that of a quadratic function or curve (nonlinear function).

[0238] In this invention, the term "linear interval" related to differential data can refer to the interval where the first derivative of the activation function is an integer or zero, or the interval where the activation function is represented by a first-order function, and the term "nonlinear interval" can refer to the interval where the first derivative of the activation function is neither an integer nor zero. However, in embodiments of this invention, the determination of the linear interval is not solely based on the derivative value; that is, the special functional unit 500 can be used to determine or distinguish the linear interval of the activation function in multiple ways.

[0239] Special function unit 500 can be used to preferentially determine whether a linear interval exists. Special function unit 500 can be used to convert the linear interval into programmable parameters expressed in the form of a first-order function, and to convert the remaining nonlinear interval into programmable parameters expressed in the form of a specific function.

[0240] Furthermore, the derivative data described in the embodiments of the present invention is merely one mathematical method for calculating the slope of the activation function. Accordingly, the present invention is not limited to derivatives, and substantially similar methods may be used to calculate the slope.

[0241] The detection of slope change points is not limited to the above methods. The special function unit 500 can be used to determine the point as the slope change point when the change of the first derivative of the activation function exceeds a certain threshold on the x-axis.

[0242] Then, the special function unit 500 can be used to extract the start and end points of the segment where the second derivative value remains unchanged. For example... Figure 10CAs shown, the special function unit 500 generates slope change data corresponding to the second derivative. Next, the special function unit 500 determines that the second derivative values ​​in each of segments w1-1 and w1-2, although different, remain constant. However, since the second derivative values ​​in each of segments w1-1 and w1-2 are different, there is a discontinuity point d3 in the slope change data corresponding to the second derivative at the boundary between segments w1-1 and w1-2. In other words, since the slope change data corresponding to the second derivative at the boundary between segments w1-1 and w1-2 is a discontinuity point d3, the boundary between segments w1-1 and w1-2 can correspond to a gradient change point.

[0243] In this case, the special function unit 500 can convert the nonlinear segment into a programmable parameter in the form of a corresponding quadratic function. Therefore, the nonlinear segment of the activation function to be programmed can be segmented into a quadratic function containing quadratic coefficients, and a linear function containing a specific slope and a specific offset. The second derivative of the nonlinear segment can be a constant value. In other words, even when approximating the nonlinear segment using a quadratic function, the approximation error can still be zero. Accordingly, the special function unit 500 can determine that there is substantially no approximation error in each of segments w1-1 and w1-2. That is, when the special function unit 500 approximates each of segments w1-1 and w1-2 with a quadratic function, the computational cost and power consumption of the special function unit 500 can be minimized, and the approximation error can also be zero.

[0244] However, the embodiments of the present invention are not limited to Figures 10A to 10C In this embodiment, the intervals w1-1 and w1-2 can also be approximated by first-order functions. In this case, the approximation error may increase, but the power consumption of the neural processing unit 1000 can be reduced by decreasing the computational load of the special function unit 500. In other words, the special function unit 500 can determine the programmable parameters based on different priorities between computational load, power consumption, and approximation error.

[0245] The second derivative of the activation function indicates the rate of change of its slope. Since segments with relatively large second derivatives represent segments with large rates of change of slope, the corresponding activation function segments exhibit significant increases or decreases in slope. Conversely, segments with relatively small second derivatives represent segments with small rates of change of slope, and the corresponding activation functions exhibit smaller increases or decreases in slope.

[0246] In particular, the segment where the second derivative of the activation function is less than or equal to a certain threshold is the segment with a very small rate of change of slope.

[0247] Accordingly, the special function unit 500 can be used to determine the activation function of the segment as a substantially linear function segment with an almost constant slope. For example, the special function unit 500 can be used to determine a segment where the second derivative of the activation function is less than or equal to a threshold as a "substantially linear segment". The threshold for the second derivative of the activation function will be described later.

[0248] The derivative order of an activation function, which changes to zero or an integer, represents the degree of change in the slope of the activation function. Specifically, generally, as the degree of the highest-order term of the function increases, the gradient of the function changes rapidly. A segment of activation function with a high degree of highest-order term is a segment with a steep slope change and can be segmented into more segments by distinguishing it from other segments.

[0249] The degree of the highest-order term of the activation function within a specific segment can be determined by the order of the derivative when the differential value becomes zero or an integer within that segment. For example, if the highest-order term of the activation function within a specific segment is third, since the third derivative of the activation function within that segment becomes an integer (i.e., the coefficient of the highest-order term) and the fourth derivative becomes zero, the degree of the highest-order term of the activation function within that segment can be determined to be third if the third derivative is an integer or the fourth derivative is zero.

[0250] In various examples, segments where the highest-order term of the activation function is three or higher can be divided into more segments to distinguish them from other segments. For instance, the number of segments can be determined as the maximum number of segments that the corresponding segment can be divided into in the hardware processing the activation function.

[0251] The gradient change points of the activation function can be identified using slope change data (i.e., the first derivative f'(x)). Using the slope change data (i.e., the first derivative f'(x)), the activation function f(x) can be segmented into three segments (w1, w2, w3), including two linear segments (w2, w3). In other words, the special function unit 500 can use the slope change data of the activation function f(x) to be programmed to determine and segment the linear segments w2 and w3 and the nonlinear segment w1.

[0252] The activation function f(x) can be segmented based on the fact that the first derivative f'(x) is a constant (non-zero), zero, or below a threshold (non-linear function), or points or segments of the curve (non-linear function). In other words, the activation function f(x) can be segmented based on points where the activation function f(x) is not differentiable or points where the first derivative f'(x) is discontinuous.

[0253] although Figure 10BThe result shown is the result of segmenting into three segments, but this is only to briefly illustrate the process of segmenting into linear and nonlinear segments. Therefore, it should be understood that the activation function f(x) can be segmented into four or more segments, i.e., at least four segments, using segmented data.

[0254] For example, a linear segment w1 can be further segmented into multiple segments using the activation function programming method according to embodiments of the present invention. By further segmenting the linear segment w1, the activation function can be segmented into more fragments and approximated to reduce approximation error. In this invention, the term "approximation error" refers to the difference between a specific fragment of the activation function and a programmable fragment that approximates that specific fragment.

[0255] Figure 11A as well as Figure 11B A graph illustrating an example of segmenting an activation function into substantially linear intervals and nonlinear intervals using slope change data in segment data in a programmed activation function method according to an embodiment of the present invention.

[0256] Figure 11B Showing Figure 11A The absolute value of the second derivative f(x) of the activation function f(x). Special function unit 500 can be used to determine the substantially linear segment by setting a specific threshold to the second derivative f(x). Please refer to... Figure 11B When the maximum absolute value of the second derivative f(x) of the activation function f(x) is 0.5, the threshold Th can be set to 10% of the maximum value Max, i.e., 0.05. As the second derivative f(x) decreases, the activation function exhibits linear characteristics. Conversely, as the second derivative f(x) increases, the activation function exhibits nonlinear characteristics.

[0257] The threshold Th can be determined as the relative ratio of the maximum absolute value Max of the second derivative f' ...

[0258] The search for substantially linear segments can be performed after the search for linear segments. However, the present invention is not limited to the order of the search for linear segments and the search for substantially linear segments.

[0259] exist Figure 11B In the example, the relative proportion can be determined to be 10%. However, the invention is not limited to this and can be determined to be 5% of the maximum value Max based on the allowable error of the deep neural network. Using differential data, i.e., the second derivative f(x), the activation function f(x) can be segmented by segments w1 and w3 where the second derivative f(x) is less than the substantive linear segment threshold Th, and segments w2 where the second derivative f(x) is greater than or equal to the substantive linear segment threshold Th. In the activation function f(x), slope change data can be used to determine and segment the substantive linear segments w1 and w3 and the nonlinear segment w2. Once the first to third segments w1, w2, and w3 are determined, the first to third segments s1, s2, and s3 can be programmed as programmable segments using the corresponding programmable parameters.

[0260] exist Figure 11B The diagram shows the segmentation results of three segments s1, s2, and s3 corresponding to the three segments w1, w2, and w3. This is only to briefly illustrate the process of segmenting into substantially linear segments and nonlinear segments. Using segment data, the activation function f(x) can be segmented into four or more segments, i.e., at least four segments. For example, the nonlinear segment w2 can be further segmented into multiple segments using the segment data according to the activation function procedural method in the embodiments of the present invention. By additionally segmenting the nonlinear segment w2, the approximation error can be reduced.

[0261] Figure 12A as well as Figure 12B To illustrate another example of how the activation function is segmented into substantially linear and nonlinear intervals using slope change data in segment data in a procedural activation function method according to an embodiment of the present invention, please refer to the following diagram. Figure 12A as well as Figure 12B In the activation function f(x), nonlinear segments can be determined based on the threshold Th of the substantially linear segment data, i.e., the absolute value of the second derivative f(x). In other words, segments greater than or equal to the substantially linear segment threshold Th can be identified as nonlinear segments. For details, please refer to [reference needed]. Figure 12B Special Function Unit 500 can use differential data, i.e., the second derivative f(x), to segment the activation function f(x) into substantially linear segments and nonlinear segments. Furthermore, as an example, Special Function Unit 500 can segment the nonlinear segment of the activation function f(x) into segments s2 and s3 corresponding to two segments w2 and w3. That is, Special Function Unit 500 can use slope change data of the activation function f(x) to classify substantially linear segments w1 and w4, and nonlinear segments w2 and w3, and then the nonlinear segments w2 and w3 can be segmented.

[0262] Special function unit 500 can be used to search for the optimal programmable parameters corresponding to each segment in various ways. For example, special function unit 500 can search for optimized programmable parameters that achieve specific performance characteristics while maintaining high speed, low power consumption, and minimizing the degradation of inference accuracy.

[0263] exist Figure 12B The image shows segments s1, s2, s3, and s4 divided into four segments w1, w2, w3, and w4. However, this is only to briefly illustrate the process of segmenting into substantially linear segments and nonlinear segments. Accordingly, it should be understood that the activation function f(x) can be segmented into five or more segments, i.e., at least five segments, using segmented data.

[0264] For example, according to one embodiment of the present invention, based on a procedural method of activation function, the nonlinear segments w2 and w3 can be further segmented into multiple segments using segment data. Specifically, the nonlinear segments w2 and w3 can be segmented based on the maximum value Max of the second derivative f(x). That is, the region from the threshold Th of the substantially linear segment to the maximum value Max of the second derivative f(x) is segmented as segment w2. Furthermore, the region from the maximum value Max of the second derivative f(x) to the threshold Th of the substantially linear segment is segmented as segment w3. Further segmentation of the nonlinear segments w2 and w3 can further reduce the approximation error.

[0265] Figure 13A as well as Figure 13B A graph illustrating another example of segmenting the activation function into nonlinear intervals using gradient change data in segmented data in an activation function programming method according to an embodiment of the present invention is provided. Please refer to... Figure 13A as well as Figure 13B In the activation function f(x), nonlinear segments can be determined based on the threshold Th of the substantially linear segments in the data, i.e., the absolute value of the second derivative f(x). In other words, regions greater than or equal to the threshold Th of the substantially linear segments can be identified as nonlinear segments. For details, please refer to [reference needed]. Figure 8B Special function unit 500 can use differential data, i.e., the second derivative f(x), to segment the activation function f(x) into substantially linear segments and nonlinear segments. In addition, special function unit 500 can also segment the nonlinear segments of the activation function f(x), for example, into segments s2, s3, and s4 corresponding to three segments w2, w3, and w4.

[0266] Special function unit 500 can classify the substantially linear segments w1 and w5 and the nonlinear segments w2, w3 and w4, and then use the slope change data of the activation function f(x) to segment the nonlinear segments w2, w3 and w4.

[0267] Embodiments of the present invention are not limited to substantially linear segments, and substantially linear segments can also be segmented into nonlinear segments. That is, in some cases, the step of determining substantially linear segments may be omitted.

[0268] Special function unit 500 can be used to search for the optimal programmable parameters corresponding to each segment in various ways. For example, special function unit 500 can search for the optimal programmable parameters that achieve specific performance, including high-speed operation, low power consumption, and suppression of the degree of degradation in inference accuracy.

[0269] exist Figure 13B The diagram shows segments s1, s2, s3, s4, and s5 divided into five segments w1, w2, w3, w4, and w5. However, this is only for the purpose of briefly illustrating the process of segmenting into substantially linear segments and nonlinear segments. Accordingly, it should be understood that the activation function f(x) can be segmented into six or more segments, i.e., at least six segments, using segmented data. However, the examples of this invention are not limited to substantially linear segments, and substantially linear segments can also be segmented into nonlinear segments.

[0270] For example, according to the activation function programming method in one embodiment of the present invention, nonlinear segments w2, w3, and w4 can be further segmented into multiple segments using segmented data. Specifically, nonlinear segments w2, w3, and w4 can be segmented based on the integral value (∫f(x)dx) of the second derivative f(x). In other words, the special functional unit 500 can segment the nonlinear segments based on the integral value of the slope change data.

[0271] When the integral value (∫f(x)dx) of the second derivative f(x) is large, the approximation error between the programmed activation function and the activation function may increase. In other words, a large integral value (∫f(x)dx) of the second derivative f(x) may introduce errors, leading to a deterioration in the accuracy of the inference. On the other hand, as the integral value (∫f(x)dx) of the second derivative f(x) increases, the width of the fragment can become wider. Conversely, the smaller the integral value (∫f(x)dx) of the second derivative f(x), the narrower the width of the fragment.

[0272] Accordingly, the special function unit 500 can set the integral value (∫f(x)dx) of a specific second derivative f(x) as the integration threshold for the segment approximation error. For example, the special function unit 500 can start integrating the second derivative f(x) from the end position of segment w1. Accordingly, segment w2 can be defined as starting from the end position of segment w1 until the preset integration threshold for the segment approximation error reaches a specific value.

[0273] More specifically, the integral of the second derivative f(x) in the segment w2. It can be segmented into segments s2, corresponding to the integration threshold of the segment approximation error. Furthermore, in segment w3, the integral of the second derivative f(x)... It can be segmented into segments s3, corresponding to the integration threshold of the segment approximation error. Furthermore, in segment w4, the integral of the second derivative f(x)... It can be segmented into fragments s4, corresponding to the integral threshold of the fragment approximation error.

[0274] In other words, all integral values ​​of the second derivative f(x) in segment w2. All integral values ​​of the second derivative f(x) in segment w3 And all integral values ​​of the second derivative f(x) in segment w4. All of these can be the same as the integration threshold of the segment approximation error.

[0275] However, the integration threshold of the segment approximation error may be affected by hardware factors including at least one of the following: the number of comparators in the special function unit 500 of the neural processing unit 1000, the number of gates used to implement the circuit of the special function unit 500, and the type of arithmetic circuit implemented (linear function circuit, quadratic function circuit, cubic function circuit, exponential function circuit, logarithmic function circuit, anti-logarithmic function circuit, etc.). In other words, the special function unit 500 can be used to determine the integration threshold of the segment approximation error while taking into account the hardware factors.

[0276] The smaller the integration threshold of the fragment approximation error, the closer the programmed activation function can be to the activation function. In other words, as the integration threshold of the fragment approximation error decreases, the number of programmable fragments increases, thus further reducing the approximation error value of the programmed activation function.

[0277] However, since the number of programmable segments is limited by hardware specifications, there are certain limitations to reducing the integration threshold of segment approximation error. In other words, the minimum limit of the integration threshold of segment approximation error can be determined based on the hardware specifications.

[0278] When the aforementioned nonlinear segments w2, w3, and w4 are further segmented, the approximation error can be further reduced. However, embodiments of the present invention are not limited to substantially linear segments; substantially linear segments can also be segmented into nonlinear segments. That is, in some cases, the step of determining substantially linear segments may not be performed.

[0279] like Figures 10A to 13BAs shown, the special function unit 500 can determine linear segments from the activation function before segmenting it and approximating it using slope variation data. When the special function unit 500 segments the activation function using slope variation data, it can determine nonlinear segments from the activation function before approximating it. When the special function unit 500 segments the activation function using slope variation data, it can determine substantially linear segments from the activation function before approximating it.

[0280] A segment with a distinct linear segment or a substantially linear segment can be approximated as a programmable segment expressed as "(slope a) × (input value x) + (offset b)". A segment with a linear segment or a substantially linear segment has a substantially constant slope, which is in the form of a linear function or a substantially linear function. Therefore, when comparing an activation function with a programmable segment expressed as a slope and offset, the programmable segment will not produce an approximation error, or will minimize the approximation error.

[0281] By programming the activation function using slope variation data, the computational cost and power consumption of substantially linear segments or linear regions can be significantly reduced. Furthermore, according to embodiments of the present invention, programming using activation functions containing linear segments or substantially linear segments is efficient and minimizes approximation errors, thus improving the instruction cycle for processing deep neural networks in the neural processing unit 1000, minimizing the degradation of inference accuracy, and reducing the power consumption of the neural processing unit 1000.

[0282] In various examples, step S210 may further include determining the linear segment of the activation function based on slope change data of the activation function.

[0283] In various examples, step S210 may further include determining the nonlinear segment of the activation function based on slope change data of the activation function.

[0284] In various examples, step S210 may further include determining the actual linear segment of the activation function based on the slope change data of the activation function.

[0285] In various examples, step S210 may further include determining the linear and nonlinear segments of the activation function based on slope change data of the activation function.

[0286] In various examples, step S210 may further include determining the actual linear and nonlinear segments of the activation function based on the slope change data of the activation function.

[0287] In various examples, step S210 may further include determining the linear segment, substantially linear segment, and nonlinear segment of the activation function based on the differential data of the activation function.

[0288] However, embodiments of the present invention are not limited to differential data of the activation function, and various mathematical analyses can also be performed to analyze the slope change and linearity of the activation function.

[0289] In various examples, the segment data may include hardware information for processing activation functions. According to an embodiment of the activation function programming method of the present invention, hardware information can be used to segment activation functions. This hardware data may include at least one of the following: the number of comparators in the special function units 500 of the neural processing unit 1000, the number of gates used to implement the multiple circuits of the special function units 500, and the type of the implemented arithmetic circuit (linear function circuit, quadratic function circuit, cubic function circuit, exponential function circuit, logarithmic function circuit, antilogarithmic function circuit, etc.).

[0290] For example, the number of segments used to segment the activation function can be limited by the number of comparators in the special function unit 500 of the neural processing unit 1000. Accordingly, the activation function can be segmented into the maximum number of segments that the neural processing unit 1000 can process, or the number of segments corresponding to the resources allocated to the neural processing unit 1000. Accordingly, the special function unit 500 can design activation functions using predetermined hardware resources more efficiently and / or in a more customized manner.

[0291] In various examples, step 220 may further include approximating at least one of the multiple segments as a programmable segment based on gradient change points.

[0292] In various examples, step 220 may further include approximating at least one of the multiple segments as a programmable segment based on the error value.

[0293] In this invention, the term "error value" or "approximation error value" refers to the difference between a specific segment of the activation function and the programmable segment that approximates that specific segment. The approximation error value may further include an average value, a minimum value, a maximum value, and a cumulative value. In other words, the special function unit 500 can be used to calculate the average error value, minimum error value, maximum error value, and cumulative error value between the specific segment and the approximate programmable segment. The cumulative error value can be a value obtained by integrating the error value between the specific segment and the approximate programmable segment.

[0294] Regarding error values, various activation functions can be divided into multiple feature segments, which include (substantially) linear segments and / or nonlinear segments. Furthermore, if these feature segments are divided into segments of equal width, the error value of each segment will vary significantly. Accordingly, in the activation function programming method of embodiments of the present invention, to reduce approximation errors, at least one characteristic of these feature segments can be considered and approximated as programmable segments.

[0295] In various examples, step S220 may further include calculating the error value by comparing the gradient of the programmable segment and the offset with the corresponding segment of the activation function.

[0296] In various examples, step S220 may further include determining programmable parameters to convert at least one segment of the activation function into a programmable segment. In other words, step S220 may further include searching for optimal programmable parameters to convert at least one segment of the activation function into a programmable segment. When the programmable segment is a linear function, the programmable parameters may include the gradient and offset corresponding to the linear function. When the programmable segment is a quadratic function, the programmable parameters may include the coefficients of the quadratic term corresponding to the quadratic function. The coefficients of the quadratic function may include quadratic coefficients, linear coefficients, and a constant term. The approximation function of the programmable parameters may be determined considering performance factors such as high-speed computation, low power consumption, and suppression of inference accuracy degradation. For example, as the formula of the approximation function becomes more complex, the computation speed may decrease and the power consumption may increase. Conversely, as the approximation error decreases, the degree of inference accuracy degradation may also decrease.

[0297] In various examples, step S220 may further include calculating the error value between at least one segment of the activation function and at least one candidate segment with (temporary) gradient and (temporary) offset. As the number of candidate segments increases, the probability of finding better programmable parameter values ​​also increases, and the search time may increase.

[0298] In various examples, step S220 may further include determining the parameters of at least one candidate segment as programmable parameters of a programmable segment based on the calculated error value.

[0299] Accordingly, the special function unit 500 can provide programmed activation function data to the neural processing unit 1000. The programmed activation function data may include at least one programmed activation function. Specifically, the programmed activation function data may include programmable parameters corresponding to each programmable segment of the at least one programmed activation function.

[0300] The following section describes the process of approximating at least one of multiple segments as a programmable segment based on error values. Figures 14 to 16B Detailed explanation.

[0301] During the programming of activation functions, step changes may occur at the boundaries between programmable segments. In the activation function programming method of embodiments of the present invention, approximation errors can be significantly reduced by generating predetermined step changes between programmable segments or at the beginning and / or end of a single programmable segment.

[0302] Accordingly, in this invention, by segmenting the activation function into multiple segments using segment data and approximating at least one of the multiple segments as a programmable segment based on the error value, allowing step variations between programmable segments can significantly reduce the error value.

[0303] Figure 14 A diagram illustrating an example of transforming a segment into a programmable segment using error values ​​in a programmable activation function method according to an embodiment of the present invention. Please refer to... Figure 14 This shows multiple candidate segments S of the nonlinear activation function segment S. c1 S c2 and S c3 .

[0304] In embodiments of the present invention, the term "candidate fragment" refers to a function that can be transformed into a programmable fragment expressed as "programmable parameters" through a programmable activation function method. When the programmable fragment is expressed as a linear function, the programmable fragment can be represented as "(gradient a) × (input value x) + (offset b)". The programmable parameters include gradient a and offset b.

[0305] For example, when a programmable fragment is expressed as a quadratic function, the programmable fragment can be represented as "(quadratic coefficient a) × (input value x)". 2 The programmable parameter is expressed as: a + (first-order coefficient b) × (input value x) + (constant c)". The programmable parameter includes a quadratic coefficient a, a first-order coefficient b, and a constant c. The programmable parameter can be used to express both linear and quadratic functions. However, this invention is not limited to the format of the programmable parameter.

[0306] The following will use linear functions as examples. Candidate segments can be linear functions corresponding to programmable segments after segmentation using segmented data. Candidate segments of a segment can be determined by linear functions passing through the start and end points of that segment.

[0307] For example, a candidate segment of a segment can be a linear function with an adjusted offset and the same gradient as a linear function passing through the start and end points of the segment.

[0308] For example, a candidate segment of a fragment can be a linear function with an adjusted offset and a gradient that is different from that of a linear function passing through the start and end points of a fragment.

[0309] For example, a candidate segment of a segment can be determined as a tangent to that segment.

[0310] exist Figure 14 In order to briefly describe the process of determining a programmable segment from multiple candidate segments, three candidate segments with the same gradient and passing through the start and end points of segment S are shown. The first candidate segment S... c1 The second candidate segment S is obtained by using a linear function of the start and end points of segment S. c2 And the third candidate fragment S c3 This is in the case of having the same characteristics as the first candidate fragment S c1 With the same slope, the offset is adjusted as a linear function, and the third candidate segment S is included. c3 The offset makes the candidate segment S c3 It becomes the tangent to segment S. Figure 14 The candidate segments shown are for simplification purposes and can be transformed into approximate programmable segments. The gradients and / or offsets of the actual candidate segments can be adjusted in various ways to reduce error values.

[0311] In various examples, at least one of a plurality of segments can be approximated as a programmable segment by searching for an error value Δy. Special function unit 500 can determine that the width of each segment of the plurality of segments is a uniform width. Subsequently, special function unit 500 can approximate the at least one segment as a programmable segment by searching for the error value Δy of the at least one segment. However, the invention is not limited thereto.

[0312] Figure 15A as well as Figure 15B A diagram illustrating an example of how, in a programmable activation function method according to an embodiment of the present invention, a segment is approximated as a programmable segment by searching for the maximum error value (max(Δy)) among error values ​​(Δy). Figure 15A The diagram shows the segments s1 and s2 that divide the activation function f(x), and the first candidate segment s corresponding to the first segment s1. c1 (x), and the second candidate segment s corresponding to the second segment s2 c2 (x). In Figure 15A In the middle, each candidate fragment s c1 (x) and s c2 (x) searches for better programmable parameters (i.e., gradient and offset) for a linear function representing the start and end points of each segment s1 and s2.

[0313] like Figure 15A In the embodiment shown, special function unit 500 calculates the second fragment s2 and the second candidate fragment s. c2 The error value Δy between f(x) and s, i.e., "f(x)-s c2 The absolute value of f(x), or |f(x)-s c2 (x)|. Special function unit 500 can calculate the maximum error value max(Δy) among multiple error values ​​Δy. In order to reduce the maximum error value max(Δy) of the second segment s2, such as Figure 15B As shown, by using candidate fragments s c2 (x) Adjust max(Δy) / 2 (i.e., adjust the offset) in the y-axis direction, which is half of the maximum error value max(Δy). The resulting second candidate segment can be judged as the second programmable segment S obtained by approximating the second segment s2. p2 (x).

[0314] When the first programmable segment S is obtained by approximating the first segment s1 p1 (x) such as Figure 15B When shown, the first programmable segment S p1 (x) and the second programmable fragment S p2 There may be steps between (x).

[0315] exist Figure 15B In the context of the programmable segments, the step order at the connection point of adjacent programmable segments on the y-axis can be based on the error value |f(x)-s. c2 The step (x) is intentionally introduced in the process of approximating the second segment s2 of the activation function f(x) as a programmable segment. Steps can be generated at the boundaries between adjacent programmable segments in the process of approximating a particular programmable segment to reduce the maximum error value within that particular programmable segment. In other words, each programmable segment can be approximated independently of each other.

[0316] As the approximation error of the activation function increases, the inference accuracy of the neural processing unit 1000 using the approximate activation function may deteriorate. Conversely, as the approximation error of the activation function decreases, the inference accuracy of the neural processing unit 1000 using the approximate activation function may deteriorate.

[0317] In various examples, at least one of multiple segments can be obtained by using the integral value of the error value ∫[s] c [(x)-f(x)]dx is approximated as a programmable segment. Special function unit 500 can be used to integrate or accumulate multiple approximation error values ​​for each segment.

[0318] More specifically, the first programmable fragment Sp1 (x) and the second programmable fragment S p2 (x) can be programmed in different ways. That is, each programmable segment can be programmed by choosing a linear function, a quadratic function, a logarithmic function, an exponential function, etc. Therefore, each programmable segment can be programmed using the same function or different functions.

[0319] Figure 16A as well as Figure 16B A diagram illustrating an example of how, in a programmable activation function method according to an embodiment of the present invention, a segment is approximated as a programmable segment using the integral of the error value (∫[sc(x)-f(x)]dx). Figure 16A This shows the segments s1 and s2 of the piecewise activation function f(x), and the first candidate segment s corresponding to the first segment s1. c1 (x), and the second candidate segment s corresponding to the second segment s2 c2 (x). In Figure 16A In the middle, for each candidate fragment s c1 (x) and s c2 (x) is used to search for the optimal programmable parameters (i.e., gradient and offset) representing the linear function at the start and end points of each of fragments s1 and s2. The second candidate fragment s... c2 The offset of (x) can be adjusted while having the same gradient as the linear function passing through the start and end points of the second segment s2. Alternatively, the offset can also be adjusted while having a different gradient than the linear function passing through the start and end points of the second segment s2.

[0320] Please refer to Figures 15A to 16B The first segment s1 contains a start point x0 and an end point x1. Here, the start point x0 and the end point x1 can represent the boundary values ​​of the segment. Please refer to [reference needed]. Figures 15A to 16B The second segment s2 contains a start point x1 and an end point x2. The start point x0 and the end point x1 can represent the boundary values ​​of the segment. For example, the first segment s1 can be set from the start point x0 to less than the end point x1. Similarly, the second segment s2 can be set from the start point x1 to less than the end point x2.

[0321] Programmable parameters can be used to include fragment boundary values.

[0322] like Figure 16A As shown, special function unit 500 calculates the second segment s2 and the candidate segment S. c2 The integral value between (x) Using this as an approximate error value, we search for the integral value with the smallest absolute value. Candidate segments. For example... Figure 16BAs shown, in order to reduce the error value, it has the smallest absolute integral value. Candidate segments, i.e. It can be determined as the second programmable segment S p2 (x).

[0323] When the first programmable segment S of the segment approximates the first segment s1 p1 (x) is as follows Figure 16B As shown, discontinuous steps may result in the first programmable segment S appearing on the y-axis. p1 (x) and the second programmable fragment S p2 Between (x). Figure 16B In this context, the step order might be that the second segment s2 of the activation function f(x) is approximated as the second programmable segment S. p2 (x) is generated based on the approximation error value during its execution. However, even with this discontinuous step, the degradation of the inference accuracy of the neural processing unit 1000 using the approximation activation function can be reduced by decreasing the approximation error value for each programmable segment.

[0324] In various examples, step S220 may further include searching for the minimum approximate error value between the programmable segment and the corresponding activation function segment. This approximate error value may be at least one of the following: average error value, minimum error value, maximum error value, and cumulative error value.

[0325] For example, step S220 may further include searching for at least one minimum error value between at least one programmable segment and the corresponding segment of at least one activation function.

[0326] For example, step S220 may further include determining the slope and offset of the programmable segment based on the searched minimum error value.

[0327] For example, step S220 may include approximating at least one segment as a programmable segment based on the determined gradient and offset.

[0328] In several examples, step S220 may further include using machine learning that leverages a loss function to determine programmable segments.

[0329] Figure 17 A diagram illustrating an example of using machine learning to approximate a fragment into an optimizeable programmable region in a programmable activation function method according to an embodiment of the present invention. Please refer to... Figure 17 Special function unit 500 can set the candidate segments s of activation function f(x). c(x) represents the initial value of the loss function. Special function unit 500 can use machine learning to determine the candidate segment with the minimum loss function value as the better programmable segment S. op (x). Based on this, better programmable parameters can be searched.

[0330] To find better parameters, learning can be performed repeatedly. One learning cycle can be represented as one training epoch. As the number of learning cycles increases, the error value can decrease. Too few training cycles may lead to underfitting; conversely, too many training cycles may lead to overfitting.

[0331] As the loss function, mean squared error (MSE), root mean square error (RMSE), etc., can be used, but are not limited to these. In this invention, the candidate segments used as the initial values ​​for the loss function can be, for example, linear functions, quadratic functions, cubic functions, etc., corresponding to segments approximated by segmenting the data. However, embodiments of this invention are not limited to the above functions. The loss function can be used after the activation function f(x) is segmented into multiple segments by segmenting the data.

[0332] Accordingly, machine learning using loss functions can be performed after considering the characteristics of activation functions, such as multiple feature segments including the (substantially) linear and / or nonlinear segments of the activation function, approximation errors, etc. Therefore, the computational cost and search time for optimizing programmable parameter search can be reduced, and the degradation of inference accuracy of neural processing units 1000 due to the use of programmed activation functions can be minimized.

[0333] Furthermore, embodiments of the present invention can reduce the number of unnecessary segments. That is, embodiments of the present invention can also reduce the number of segments. In other words, if the sum of the approximate error values ​​of two adjacent programmable segments is less than a preset threshold, then these two programmable segments can be merged into one programmable segment.

[0334] In various examples, step S210 may further include segmenting the activation function into multiple segments using the integral (cumulative value) of the second derivative of the activation function. The cumulative value of the second derivative can be used as segment data.

[0335] In one embodiment, step S210 may further include calculating the cumulative value of the second derivative of the activation function.

[0336] In one embodiment, step S210 may further include segmenting the activation function into multiple segments based on an integral threshold (i.e., a cumulative threshold of the second derivative) based on the segment approximation error.

[0337] Furthermore, the activation function programming method according to the present invention may include, when the number of segments determined after segmenting the activation function using the accumulated value of the second derivative is greater than or less than a target number, first adjusting the threshold of the accumulated value of the second derivative, and then re-segmenting the activation function into another number of segments based on the adjusted threshold. Specifically, the threshold may be adjusted such that: (1) when the number of obtained segments is greater than the target number, the threshold is adjusted to increase, and (2) when the number of obtained segments is less than the target number, the threshold is adjusted to decrease.

[0338] In various examples, the special function unit 500 can segment the activation function into multiple segments based on a threshold of the accumulated value of the second derivative. In this case, the special function unit 500 can segment all segments of the activation function based on the threshold of the accumulated value of the second derivative, or segment a portion of the activation function based on the threshold of the accumulated value of the second derivative. In particular, the special function unit 500 can determine certain segments of the activation function as nonlinear segments rather than (substantially) linear segments, and can segment only the segments that are nonlinear segments based on the threshold of the accumulated value of the second derivative. The special function unit 500 can segment the remaining nonlinear segments using the procedural activation function methods described in the various examples of the invention.

[0339] Figure 18 This diagram illustrates an example of segmenting an activation function into segments using an integral threshold of the segment approximation error of the activation function in a procedural activation function method according to an embodiment of the present invention. Please refer to... Figure 18 The activation function f(x) can be segmented using the accumulated value of its second derivative, ∫f″(x). The point of minimum (min) of the activation function f(x) on the x-axis can be determined as the starting point, or the point of maximum (max) on the x-axis can be determined as the starting point. However, this invention is not limited to this, and the starting point can also be a specific point.

[0340] Special function unit 500 can be programmed to include multiple segment boundary values ​​x1, x2, x3, x4, and x5 of the activation function. Special function unit 500 can be programmed to further include, for example, a minimum (min) and a maximum (max) of the activation function. According to embodiments of the invention, the minimum (min) and maximum (max) can be utilized during pruning to improve the efficiency of activation function programming. When the x value is less than or equal to the minimum, the activation function can output the minimum value f(min). When the x value is equal to or greater than the maximum, the activation function can output the maximum value f(max).

[0341] The activation function f(x) is segmented starting from the beginning, for each segment where the accumulated value of the second derivative of the activation function f(x) reaches the threshold Eth (i.e., the integral threshold of the segment approximation error). For example, when At that time, special function unit 500 can determine w1, when At that time, special function unit 500 can determine w2, when At that time, special function unit 500 can determine w3, when At that time, special function unit 500 can determine w4, when At that time, special function unit 500 can determine w5, when At that time, special function unit 500 can determine w6. Further explanation: different ETh values ​​can also be set for each segment; and multiple ETh values ​​can be set according to different situations, such as ETh1 and ETh2.

[0342] Furthermore, the programmable activation function used in operations within a neural network can be used to process only input values ​​within a limited range. For example, the minimum (min) value on the x-axis of the input value for the programmable activation function can be -6, and the maximum (max) value can be 6. According to the above configuration, the data size of the programmable activation function can thus be reduced. However, the invention is not limited thereto.

[0343] Please refer to Figure 18 Since the cumulative value of the second derivative of the activation function is the rate of change of the slope of the activation function, it can be determined that: (1) in the activation function f(x), the widths w2, w3 and w4 of the segments corresponding to the relatively large gradient rate of change are determined to be relatively narrow, and (2) in the activation function f(x), the widths w1 and w6 of the segments containing the linear function and without the slope rate of change are determined to be relatively wide.

[0344] Figure 19 as well as Figure 20Charts illustrating the exponential linear unit activation function and the Hardswish activation function respectively. The exponential linear unit activation function f(x) is x when x > 0 and α(e x -1) (where α is a hyperparameter) when x ≤ 0. As Figure 19 shown, the exponential linear unit activation function has a linear segment when the x value is greater than or equal to a certain value, and a non-linear segment when the x value is less than zero. That is to say, the exponential linear unit activation function has the characteristic of being divided into a linear segment and a non-linear segment.

[0345] The Hardswish activation function f(x) is 0 when x ≤ -3, x when x ≥ +3, and x×(x + 3) / 6 when -3 < x < +3. As Figure 19 shown, when the x value is less than negative three or greater than three, the Hardswish activation function has a linear segment, and in other cases it has a non-linear segment. That is to say, the Hardswish activation function has the characteristic of being divided into a linear segment and a non-linear segment.

[0346] However, the present invention is not limited to the exponential linear unit activation function and the Hardswish activation function, and there are various activation functions that have the characteristic of being divided into a linear segment and a non-linear segment. In the field of neural networks, various custom activation functions that combine various linear and non-linear functions to improve the accuracy of neural networks have been proposed. In this case, the activation function programming method according to an embodiment of the present invention can be more effective.

[0347] In the activation function programming method according to the present invention, the special function unit 500 can distinguish the linear segment and the non-linear segment of the activation function, and can further be divided into a substantially linear segment and a non-linear segment, so that the activation function can be selectively segmented into multiple segments. Accordingly, according to the activation function programming method of the present invention, especially when approximating the programming of the activation function for the (substantially) linear segment and the non-linear segment, the efficiency can be improved and the approximation error can be minimized. Therefore, the instruction cycle of the neural network model processed in the neural processing unit 1000 can be improved, the deterioration of the inference accuracy can be minimized, and the power consumption of the neural processing unit 1000 can be reduced. In the activation function programming method according to the present invention, the special function unit 500 can generate multiple programmable parameters for at least one segment. The neural processing unit 1000 can process at least one programmed activation function based on the above information. The neural processing unit 1000 can receive this information and process at least one programmed activation function.

[0348] Figure 21 Flowchart for illustrating the activation function programming method according to an embodiment of the present invention. Figure 22This is a schematic diagram illustrating a neural network for approximating activation functions according to an embodiment of the present invention.

[0349] Please refer to Figure 21 The activation function programming method includes step S310 of setting a target activation function, step S320 of training a neural network with an approximate target activation function as the programmed activation function, and step S330 of converting the programmed activation function into a slope and offset and storing them in a lookup table.

[0350] In step S310, the activation function of the target activation function to be programmed is set. For example, the target activation function can be the swish function, Mish function, sigmoid function, hyperbolic tangent function, SELU function, Gaussian error linear unit function, SOFTPLUS function, square root (SQRT) function, and other nonlinear functions. In step S320, the target activation function is approximated by the programmed activation function through training a neural network.

[0351] Please refer to Figure 22 A neural network-like system used to approximate a target activation function can consist of two layers and multiple modified linear unit functions positioned between these two layers. In other words, a neural network-like system used to perform the approximation operation of the target activation function can be composed of two neural network segments and multiple modified linear unit functions positioned between these two segments.

[0352] The first type of neural network segment refers to the portion between multiple nodes in the input layer and multiple nodes in the hidden layer. In other words, the first type of neural network segment can be called the first layer.

[0353] The second type of neural network segment refers to the portion between multiple nodes in the hidden layer and multiple nodes in the output layer. In other words, the second type of neural network segment can be called the second layer.

[0354] At least one neuron in a first-class neural network segment contains a connection network that includes weights that connect nodes in the input layer to nodes in the hidden layer.

[0355] At least one neuron in the second type of neural network segment contains a connection network, which includes weights that connect nodes in the hidden layer and nodes in the output layer, as well as corresponding activation functions.

[0356] More specifically, the first type of neural network segment contains at least one neuron. Each neuron in the first type of neural network segment has one node as input in the input layer and each of a plurality of nodes as output in the hidden layer.

[0357] For example, the number of neurons in a first-type neural network segment could be fifteen. Accordingly, the number of nodes in the hidden layer could also be fifteen. However, the number of neurons in a first-type neural network segment and the number of nodes in the hidden layer can vary depending on requirements.

[0358] Furthermore, the first type of neural network segment can be a fully connected layer, where a single node in the input layer is fully connected to multiple nodes in the hidden layer that serve as outputs. Accordingly, each neuron in the first type of neural network segment can have weights and biases. That is, the weights of each neuron in the first type of neural network segment can be represented as n1, n2, ... n. 15 Furthermore, the deviation of each neuron can be represented as b1, b2, ... b 15 .

[0359] Therefore, when input x is fed into a first-class neural network segment, each node in the hidden layer can output z. i =n i *x+b i Next, the modified linear unit function can be applied to the output of each neuron in the first type of neural network segment.

[0360] The corrected linear unit (z) can be represented as max(0,z), meaning that when the corrected linear unit function is applied, all negative values ​​are converted to zero. Therefore, the output value of the first type of neural network segment with the corrected linear unit function applied can be represented as ReLU(n i *x+b i ).

[0361] The second type of neural network segment also contains at least one neuron. Each neuron in the second type of neural network segment takes nodes in the hidden layer as input and a single node in the output layer as output. For example, the second type of neural network segment can have fifteen neurons. Accordingly, the number of nodes in the hidden layer can also be fifteen. However, the number of neurons in the second type of neural network segment and the number of nodes in the hidden layer can be changed according to requirements.

[0362] Furthermore, the second type of neural network segment can be a fully connected layer, where multiple nodes in the hidden layer serving as input and a single node in the output layer serving as output are fully connected. Accordingly, each neuron in the second type of neural network segment can have weights. That is, the weights of multiple neurons in the second type of neural network segment can be represented as m1, m2, ..., m15.

[0363] Therefore, the second type of neural network segment can convert the output value of the first type of neural network segment into ReLU(n). i *x+b iTherefore, the output of the second type of neural network segment is the output ReLU(n) of the first type of neural network segment. i *x+b i The sum is calculated by multiplying the weights of the second type of neural network segment by the output value. A node in the output layer, i.e., the output of the second type of neural network segment, can output the calculated value according to Equation 1.

[0364]

[0365] By performing the aforementioned neural network operations, the error between the approximate programmed function and the target activation function can be calculated, and the training of the neural network can be repeated to reduce the error value. Through this training process, the activation function transformation program unit can approximate the target activation function as the programmed activation function.

[0366] Finally, by calculating the inflection points of the programmed activation function, multiple linear segments of the programmed activation function can be defined. Each linear segment can then be further segmented into first-order functions with specific slopes and offsets.

[0367] In step S320, the programmed activation function is converted into multiple slopes and multiple offsets and stored in a lookup table.

[0368] As mentioned above, each programmed activation function can be segmented into a first-order function with a specific slope and a specific offset for each linear segment. Accordingly, the specific slope and specific offset of each linear segment can be stored in a lookup table.

[0369] Figure 23 To illustrate the execution by the post-processing unit according to an embodiment of the present invention Figure 6 A schematic diagram of step 610 in the category maximum value calculation. Figure 6 In the category maximum value calculation step S120, the first calculation unit 610 extracts the category with the highest category score from the multiple categories contained within the bounding box. That is, in the category maximum value calculation step S120, the first calculation unit 610 performs the category maximum value calculation to extract the index of the category with the highest category score within the bounding box and its category score.

[0370] Specifically, within a memory partition of internal memory 630, for each bounding box, the system can store the object existence confidence score, bounding box coordinates, and multiple category indices corresponding to the multiple objects contained within the bounding box, as well as the scores for each category. Please refer to [reference needed]. Figure 23 Memory partition Bank1 can contain data from multiple bounding boxes. This memory partition Bank1 can contain data as described above. Figure 5The memory bank consists of a portion of the DATA memory partition, a portion of the OUTPUT1 memory partition, and a portion of the OUTPUT2 memory partition. For example, memory partition Bank1 can contain data from the first bounding box BOX1 and the second bounding box BOX2. Similarly, please refer to the following... Figure 30 The memory partition Bank2 may contain another part of the DATA memory partition, another part of the OUTPUT1 memory partition, and another part of the OUTPUT2 memory partition.

[0371] exist Figure 23 In this embodiment, it is assumed that the bounding box is rectangular. The data of the first bounding box BOX1 may include a confidence score C for predicting the object presence of the first bounding box BOX1 memory in the object, and bounding box coordinate data of the first bounding box BOX1, such as height data H, width data W, x data X, and y data Y. Here, x data X and y data Y represent the x-coordinate and y-coordinate of the first bounding box BOX1 in the image, respectively. Furthermore, the data of the second bounding box BOX2 may also include a confidence score C for predicting the object presence of the second bounding box BOX2 memory in the object, and bounding box coordinate data of the second bounding box: height data H, width data W, x data X, and y data Y. Memory partition Bank1 may also contain multiple dummy data entries to fill empty or unused bits within the word width.

[0372] The shape of the bounding box is not limited to a rectangle; it can also be transformed into a pentagon, polygon, or circle. The quantity and type of bounding box coordinate data may vary depending on the shape of the bounding box.

[0373] The data for the first bounding box (BOX1) may include class scores (0 to 33) for the multiple objects included in the first bounding box (BOX1). For example, objects included in the first bounding box (BOX1) may be predicted to belong to one of multiple classes, and the data for the first bounding box (BOX1) may include class scores (0 to 33) for these predicted classes. Similarly, the data for the second bounding box (BOX2) may also include class scores (0 to 33) for the multiple objects included in the second bounding box (BOX2). For example, objects within the second bounding box (BOX2) may be predicted to belong to one of multiple classes, and the data for the second bounding box (BOX2) may include class scores (0 to 33) for these predicted classes.

[0374] Then, in the category maximum value calculation step S120, the first calculation unit 610 extracts the category with the highest score from the multiple categories contained in each bounding box. That is, in the category maximum value calculation step S120, the first calculation unit 610 performs a category maximum value calculation to extract the category index and category score of the highest score in the first bounding box BOX1 and the second bounding box BOX2. For example, the first calculation unit 610 extracts from the first bounding box BOX1 the first category index 0' with the highest category score among the category score data 0 to 32 associated with the first bounding box BOX1, and its corresponding category score data 0. The first calculation unit 610 also extracts from the second bounding box BOX2 the last category index 33' with the highest category score among the category score data 0 to 33 associated with the second bounding box BOX2, and its corresponding category score data 33. The category index, category score data, and bounding box coordinate data can be stored in memory partition Bank 1. By extracting only the index data and corresponding score data of one category from each bounding box BOX1, BOX2, and using or transmitting the extracted index data and its score data, the first processing unit 610 can reduce the data size of each bounding box used for subsequent processing. After storing the extracted data in memory partition Bank 1, the remaining data in memory partition Bank 1 is deleted or overwritten by other data, and the data in memory partition Bank 1 will be used for subsequent processing. In other words, after the extracted data is stored in memory partition Bank 1, the remaining data in memory partition Bank 1 will no longer be used. The extracted data in memory partition Bank 1 becomes the target of subsequent processing. In this way, the available data space in internal memory 630 can be utilized more efficiently. Alternatively, memory partition Bank 1 can store the memory locations of bounding box coordinate data, extracted category indexes, and category score data, instead of directly moving these data to memory partition Bank 1, and these stored locations can be referenced during subsequent processing.

[0375] Figure 24 This is a schematic diagram illustrating the filtering operation steps performed on the bounding box BOX1 by the second calculation unit 620 of the post-processing unit according to an embodiment of the present invention. Figure 24 The same processing procedure applies to other bounding boxes. Figure 25 This is a schematic diagram illustrating the filtering operation result performed by the post-processing unit according to an embodiment of the present invention. In the filtering operation step S130, the second operation unit 620 extracts only the bounding boxes whose category confidence scores are higher than a threshold confidence score from a plurality of bounding boxes. This category confidence score can correspond to the product of the object existence confidence score C of the bounding box and the category score data extracted by the first operation unit 610. (Refer to the above...) Figure 23In the described embodiment, the category confidence score of BOX1 is the product of the object presence score C of BOX1 and the category score data 0 of BOX1, while the category confidence score of BOX2 is the product of the object presence score C of BOX2 and the category score data 33 of BOX2. In the filtering operation step S130, the second operation unit 620 only extracts bounding boxes whose product of the object presence confidence score C and the category score data extracted by the first operation unit 610 is higher than a specific threshold confidence score thr. The information of the extracted or filtered bounding boxes is then stored in the memory partition Bank 1 of the internal memory 630. The information of the extracted or filtered bounding boxes may include bounding box coordinate data, category index, and category score. In addition, the memory partition Bank 1 may also store the memory location of the bounding box coordinate data, the extracted category index, and the category score data of the filtered bounding boxes for subsequent processing. In the filtering operation step S130, the second operation unit 620 does not store data of bounding boxes whose product of confidence score C and category score 0 extracted by the first operation unit 610 is less than or equal to the threshold confidence score thr in memory partition Bank 1. Only the data of the filtered bounding boxes will be further processed. In this way, the amount of computation in subsequent processing can be reduced.

[0376] exist Figure 25 In this example, assuming that after the first processing unit 610 completes the category maximum value calculation, there are N bounding boxes. Accordingly, in the filtering operation step S130, the second processing unit 620 can extract from these N bounding boxes only two of them, whose objects have a confidence score C whose product with the category score data extracted by the first processing unit 610 is greater than a certain threshold confidence score thr. Only the data of these two filtered bounding boxes will be further processed by the internal processing unit. Therefore, the amount of data that the internal processing unit 640 needs to process can be reduced, allowing the internal processing unit 640 to use less memory to perform operations faster. Since the performance of the post-processing unit depends on the instruction cycle of the internal processing unit 640, this method can improve the overall performance of the post-processing unit.

[0377] Figure 26 This is a schematic diagram illustrating the decoding steps performed by the post-processing unit according to an embodiment of the present invention. Figure 6 In the subsequent decoding step S140, the internal processing unit 640 can decode the filtered bounding box data. For details, please refer to... Figure 26 The bounding box coordinates will be decoded for subsequent processing. Figure 6 The nonmaximum suppression operation step S150 is performed by multiplying, adding, and subtracting the height data H, width data W, x-coordinate data X, and y-coordinate data Y corresponding to the bounding box.

[0378] Figure 27 This diagram illustrates a non-maximum suppression (NMS) operation performed by a post-processing unit according to an embodiment of the present invention. Subsequently, in the NMS operation step S150, redundant or overlapping bounding boxes generated by the second operation unit 620 can be removed. NMS is a post-processing step in object detection tasks used to remove redundant or overlapping bounding boxes generated by object detection algorithms, typically applied to neural network models such as You Only Look Once (YOLO) or Faster R-CNN (Fast Region-based Convolutional Network). Through the NMS operation step, duplicate bounding boxes can be removed, and only non-duplicate bounding boxes are retained for subsequent processing.

[0379] The non-maximum suppression operation can be divided into a confidence score sorting step and a deduplication step. First, in the confidence score sorting step, the bounding box data is sorted according to its confidence score, where the confidence score is the product of the object existence confidence score and the class score of the bounding box. In one embodiment, the bounding box data with the highest confidence score is sorted first, and the remaining bounding boxes are sorted in descending order of their confidence scores.

[0380] In the duplicate deletion step, the bounding box with the highest confidence score is used as a reference, and its overlap with other bounding boxes is determined. Typically, the overlap between the reference bounding box (REF BOX) and another bounding box is measured by the Intersection over Union (IoU), which is the ratio between the intersection region and the union region of two bounding boxes. If the IoU between the reference bounding box with the highest confidence score and another bounding box exceeds a preset threshold (e.g., 0.5 or higher), it indicates significant overlap between the two boxes, and the other bounding box is deleted. Deleting another bounding box can be done by deleting the associated data of that bounding box from internal memory 630 or by freeing the data from the occupied internal memory 630 space for other data to overwrite. If the IoU between the reference bounding box with the highest confidence score and another bounding box is equal to or less than the preset threshold (e.g., 0.5 or more), that bounding box is retained. By using the nonmaximum suppression operation step, duplicate bounding boxes can be removed from the internal memory 630, while non-duplicate bounding boxes can be retained, thereby improving the accuracy and reliability of the object detection system.

[0381] Figure 28 This is a schematic diagram illustrating the data reduction amount of a neural processing unit including a post-processing unit according to an embodiment of the present invention. Figure 28 The document describes the data fields and sizes of individual bounding boxes in neural network-like models such as YOLO, face recognition, and pose recognition, as well as the overall data size of these neural network-like models and the reduction in data size after performing filtering operations or combining class maximum value operations with filtering operations.

[0382] Taking an application using a YOLO-like neural network model as an example, the post-processing unit can be fed 50KB of data for each of 100 bounding boxes, which includes the object presence confidence score and multiple class indices and object class scores of the objects contained within that bounding box.

[0383] Next, in the category maximum value calculation step, the first operation unit 610 performs the category maximum value calculation to extract the highest-scoring category index and category score from 100 bounding boxes, thereby reducing multiple category data (category, keypoint) to two values. As a result, the data size is reduced from 50KB to 6.25KB. In the filtering operation step, the second operation unit 620 can reduce the number of bounding boxes from 100 to 10 by removing those bounding boxes whose object existence confidence score multiplied by the category score extracted by the first operation unit 610 is lower than a threshold confidence score (thr), further reducing the data size from 6.25KB to 0.625KB.

[0384] In the next example, if the neural network model is a face recognition model (Face), in the filtering step, the second operation unit 620 can reduce the number of bounding boxes from 100 to 10 by filtering bounding boxes whose confidence scores are lower than a certain threshold confidence score (thr), thereby reducing the data size from 6.25KB to 0.625KB. In this example, no additional reduction is performed through the category maximum value operation.

[0385] In the last example, the neural network-like model is a pose detection model. In the filtering step, the second processing unit 620 can reduce the number of bounding boxes from 100 to 10 by removing bounding boxes whose product of the object's confidence score and class score is lower than a specific threshold confidence score (thr), thereby reducing the data size from 25KB to 2.5KB. In this example, no additional reduction is performed through the class maximum value operation.

[0386] Figure 29A This is a schematic diagram of a directed acyclic graph representing an object detection neural network model whose input is fed to a neural processing unit including a post-processing unit, according to an embodiment of the present invention. Figure 29B This is a schematic diagram of a directed acyclic graph representing a post-processed object detection neural network model in a neural processing unit including a post-processing unit according to an embodiment of the present invention.

[0387] Object recognition neural network models using directed acyclic graphs can consist of multiple layers and multiple nodes connected to these layers. For example... Figure 29A As shown, this object detection neural network model can include convolutional layers (Conv), multiplication layers (Mul), and addition layers (Add). Specifically, the output of the convolutional layer (Conv) can be 80×80 bounding box data with 255 channels, the output of the multiplication layer (Mul) can be 80×80 bounding box data with 255 channels, and the output of the addition layer (Add) can be 80×80 bounding box data with 255 channels.

[0388] like Figure 29B As shown, when the post-processing unit 600 in one embodiment of the present invention is applied, the object detection neural network model may include a convolutional layer Conv, a multiplication layer Mul, and an addition layer Add, and may further include a programmed activation function layer DX_PAF, a class maximum value operation layer PP_Argmax, and a filtering layer PP_Filter. That is, in the neural processing unit containing the post-processing unit 600, according to one embodiment of the present invention, the compiler can modify... Figure 29A The object detection neural network model shown is further improved to include... Figure 29B The diagram shows a programmed activation function layer (DX_PAF), a class maximum value operation layer (PP_Argmax), and a filtering layer (PP_Filter). The compiler can modify or optimize the neural network model based on hardware information of the neural processing unit 1000 (e.g., the presence or absence of the post-processing unit 600 or special function unit 500) to accelerate computation by utilizing dedicated circuitry (e.g., post-processing units or special function units) provided in the neural processing unit 1000.

[0389] exist Figure 29B For example, the number of anchor box types in the bounding boxes is three. Therefore, the figure shows three convolutional layers (Conv), three multiplication layers (Mul), three addition layers (Add), three procedural activation function layers (DX_PAF), three class maximum operation layers (PP_Argmax), and three filtering layers (PP_Filter). Anchor boxes are predefined bounding boxes used in object detection to generate candidate regions of different sizes and aspect ratios to identify objects at specific locations. The number of each type of layer can vary depending on the number of anchor box types in the bounding boxes.

[0390] exist Figure 29BIn this process, the output of the convolutional layer Conv can be 128 channels of 80×80 bounding box data, the output of the multiplication layer Mul can be 128 channels of 80×80 bounding box data, and the output of the addition layer Add can be 128 channels of 80×80 bounding box data. However, as mentioned above, after passing through the class maximum operation layer PP_Argmax, the size of the bounding box data is reduced so that only 7 channels of 80×80 bounding box data are retained in the internal memory 630, including the confidence score of the object's presence, the bounding box coordinates, and the predicted class of the object's presence, while other data is removed from the internal memory 630.

[0391] Next, through the PP_Filter layer, only the bounding box data in the 7-channel 80×80 bounding box data with a class confidence score higher than the threshold confidence score will be retained in the internal memory 630, while the rest of the data will be removed from the internal memory 630.

[0392] Figure 30 A timing diagram illustrating the operation of a neural processing unit including a post-processing unit on multiple image data according to an embodiment of the present invention is provided. Figure 30 The process of calculating multiple image data within a neural processing unit including a post-processing unit 600 is shown, divided into a first period (Period 1) in which the neural processing unit receives first image data IMG1, a second period (Period 2) in which the neural processing unit receives second image data IMG2, a third period (Period 3) in which the neural processing unit receives third image data IMG3, and a fourth period (Period 4) in which the neural processing unit receives fourth image data IMG4.

[0393] exist Figure 30 During Period 1, the processing element array performs a convolution operation on the first image data IMG1 to output multiple bounding box data of the first image data IMG1. Simultaneously with the convolution operation, the first processing unit 610 performs a maximum category operation on the output multiple bounding box data, while the second processing unit 620 performs a filtering operation on the bounding box data. Figure 30 In this context, "operation" refers to the maximum value operation of the category and the filtering operation. The bounding box data of the first image data IMG1 output by the first operation unit 610 and the second operation unit 620 can be stored in the first memory partition Bank 1 of the internal memory 630.

[0394] Period 2 begins after the processing element array completes the convolution operation on the first image data IMG1. During Period 2, the processing element array performs a convolution operation on the second image data IMG2 to output multiple bounding box data corresponding to the second image data IMG2. While the processing element array performs the convolution operation, the first processing unit 610 performs a class maximum operation on the output multiple bounding box data, while the second processing unit 620 performs a filtering operation on the bounding box data. The bounding box data of the second image data IMG2 output by the first processing unit 610 and the second processing unit 620 can be stored in the second memory partition Bank 2 of the internal memory 630. At the same time, the internal processing unit 640 performs decoding and non-maximum suppression operation Post on the bounding box data of the first image data IMG1 from the first memory partition Bank 1. During Period 2, since the convolution operation time of the processing element array is longer than the decoding time and non-maximum suppression operation time of the internal processing unit 640, Period 3 will begin after the processing element array completes the convolution operation on the second image data IMG2.

[0395] exist Figure 30 During Period 3, the processing element array performs a convolution operation on the third image data IMG3 to output multiple bounding box data mapped from the image. Simultaneously with the convolution operation, the first processing unit 610 performs a maximum category operation on the output multiple bounding box data, and the second processing unit 620 performs a filtering operation on the bounding box data.

[0396] Next, the bounding box data in the third image data IMG3, namely the output of the first processing unit 610 and the second processing unit 620, can be stored in the first memory partition Bank 1 of the internal memory 630.

[0397] At the same time, Figure 30 During Period 3, while the processing element array performs convolution operations, the internal processing unit 640 performs decoding and non-maximum suppression (NMS) operations on the bounding box data of the second image data IMG2 input from the second memory partition Bank 2. Since the time required for the internal processing unit 640 to perform decoding and NMS operations is longer than the time required for the processing element array to perform convolution operations, Period 4 begins after the internal processing unit 640 has completed the decoding and NMS operations on the bounding box data of the second image data IMG2.

[0398] During Period 4, the processing element array performs a convolution operation on the fourth image data IMG4, outputting data of multiple bounding boxes in the fourth image data IMG4, with the operation order being the same as described in Period 2. When the processing element array performs the convolution operation, the first processing unit 610 performs a class maximum value operation on the output multiple bounding box data (in... Figure 30 The middle label is "ca"), and the second operation unit 620 performs filtering operations (in Figure 30 (The label is "filter"). In addition, when the processing element array performs convolution operations, the internal processing unit 640 performs decoding operations and non-maximum suppression operations (Post) on the bounding box data of the third image data IMG3.

[0399] Again, when the later of the time when the internal processing unit 640 completes the nonmaximum suppression operation on the first image data IMG1 and the time when the processing element array completes the operation on the second image data IMG2 occurs, the processing element array begins to perform the operation on the third image data, or the internal processing unit 640 performs decoding operation and nonmaximum suppression operation (Post) on the bounding box data of the second image data IMG2.

[0400] As described above, the neural processing unit including a post-processing unit according to the present invention can process the multiple bounding box data output by having the first operation unit 610 perform a class maximum operation and the second operation unit 620 perform a filtering operation while the processing element array performs convolution operations. Accordingly, the neural processing unit including a post-processing unit according to the present invention can reduce the processing time of post-processing operations because the class maximum operation and filtering operation, which are part of the post-processing operations, do not require additional processing time. In other words, the neural processing unit including a post-processing unit according to the present invention only requires additional time to perform decoding operations and non-maximum suppression operations, which are different parts of the post-processing operations, without requiring additional time to perform the class maximum operation and filtering operation.

[0401] Specifically, if all post-processing operations, including the maximum value operation, filtering operation, decoding operation, and non-maximum suppression operation, are performed, the data size processed by the post-processing operation can be 8.2MB, and the data processing time can be 24ms. On the other hand, if only the decoding operation and the non-maximum suppression operation are performed during the post-processing operation, the data size processed by the post-processing operation can be 128KB, and the data processing time can be 1.29ms.

[0402] The neural processing unit including a post-processing unit according to the present invention has the advantage that while the processing element array performs convolution operations, the first operation unit 610 performs class maximum value calculations on the output data of multiple bounding box data and the second operation unit 620 performs filtering operations, thereby reducing the additional time required for post-processing operations and the amount of data to be processed. Therefore, the neural processing unit including a post-processing unit according to the present invention has advantages such as improving instruction cycle time.

[0403] Furthermore, when the processing element array performs convolution operations, the internal processing unit 640 can simultaneously perform decoding and non-maximum suppression operations on multiple bounding box data of the previous image data. Therefore, the decoding and non-maximum suppression operations of the previous image data during post-processing can overlap with the convolution operation time of the subsequent image data, which can further improve the instruction cycle of the neural processing unit.

[0404] The neural processing unit according to the present invention may include a post-processing unit, which includes internal memory and an internal processing unit. Accordingly, it can effectively eliminate or reduce the amount of data transferred to external memory and the external processing unit during post-processing operations, such as category maximum value operations, filtering operations, decoding operations, and non-maximum suppression operations. Accordingly, the neural processing unit according to the present invention does not need to transfer data from an external device for post-processing operations, thereby avoiding data delays caused by bus transmission. Therefore, the instruction cycle of the neural processing unit of the present invention can be further improved, and the data transmission power consumption of the external device can be minimized, achieving low-power operation.

[0405] According to one embodiment of the present invention, a neural processing unit can be provided.

[0406] The neural processing unit may include an array of processing elements for performing operations on a neural network-like model, and a post-processing unit for processing output data from the array of processing elements.

[0407] The neural processing unit may include special functional units for performing activation function operations on output data from the processing element array.

[0408] The neural network-like model can be an object detection model, where the output data from the array of processing elements can include data of multiple bounding boxes of image data, and each of the multiple bounding box data can include an object presence confidence score, bounding box coordinate data, and category data.

[0409] The post-processing unit may include a first computational unit for extracting the highest-scoring category from among a plurality of categories contained in each of the plurality of bounding boxes, and a second computational unit for extracting one or more bounding boxes from the plurality of bounding boxes whose category confidence scores are greater than or equal to a threshold confidence score. The category confidence score may be the product of the object presence confidence score of the bounding box and the category score extracted by the first computational unit.

[0410] The post-processing unit may include a first operation unit for performing a operation to extract the highest-scoring category index and the maximum category score of a plurality of bounding boxes, and a second operation unit for performing a bounding box filtering operation, wherein the bounding box filtering operation is used to extract one or more bounding boxes whose object existence confidence score and the category score extracted by the first operation unit are greater than or equal to a threshold confidence score.

[0411] The post-processing unit may include internal memory for storing data output from the first arithmetic unit and the second arithmetic unit.

[0412] The internal memory can contain multiple memory partitions. One part of the memory partitions can be used to store the output data of the first arithmetic unit, and another part of the memory partitions can be used to store the output data of the second arithmetic unit.

[0413] When the processing element array performs operations, the first operation unit can perform the category maximum value operation, and the second operation unit can perform the bounding box filtering operation.

[0414] The post-processing unit may include a function to perform nonmaximum suppression on multiple captured bounding boxes, and through the nonmaximum suppression operation, redundant bounding boxes in the multiple captured bounding boxes may be removed.

[0415] When the processing element array performs operations on subsequent image data, the internal processing unit can perform nonmaximum suppression operations on previous image data.

[0416] The internal processing unit can be used to start performing nonmaximum suppression operation on subsequent image data at the later of the time when the internal processing unit completes the nonmaximum suppression operation on the previous image data and the time when the processing element array completes the operation on the subsequent image data after the previous image data.

[0417] The processing element array can be used to start performing operations on the third image data at the later of the time when the internal processing unit completes the non-maximum suppression operation on the first image data and the time when the processing element array completes the operation on the second image data after the first image data.

[0418] The neural processing unit may contain a compiler for adding class maximization layers and filtering layers to the input neural network-like model.

[0419] The embodiments and accompanying drawings of this invention are only used to illustrate the technical content of this invention and to help understand this invention, and are not intended to limit the scope of this invention. For those skilled in the art, other modifications can be made based on the technical concept of this invention, in addition to the embodiments shown herein.

[0420] [National R&D Programs Supporting This Invention]

[0421] [Task Identifier] 1711193211

[0422] [Task ID] 2022-0-00957-002

[0423] [Supervisory Authority] Ministry of Science and ICT (Korea)

[0424] [Name of the Responsible (Professional) Management Agency] Korea Institute of Information & Communications Technology Planning & Evaluation

[0425] [Research Project Title] Development of PIM Core Technology for Artificial Intelligence Semiconductor (Design)

[0426] [Research Task Title] Development of Distributed On-Chip Memory-Operator Convergence PIMSemiconductor Technology for Edge Computing

[0427] [Contribution Ratio] 1 / 1

[0428] [Name of Implementing Agency] DEEPX CO.,LTD. (Korea)

[0429] [Research Period] 2023 / 01 / 01~2023 / 12 / 31

Claims

1. A neural processing circuit, characterized in that, Include: A processing element array circuit is used to perform multiple convolution operations on a type of neural network model to produce output data; A post-processing circuit, coupled to the processing element array circuit, receives the output data; the post-processing circuit is used to extract a subset of the output data; and A subsequent circuit, coupled to the post-processing circuit, is used to selectively store the subset of the extracted output data or perform multiple operations on the subset of the extracted output data.

2. The neural processing circuit as described in claim 1, characterized in that, The output data includes multiple category scores for each bounding box within a region of an image, the multiple category scores indicating the probability of multiple object categories appearing in each bounding box, and wherein the post-processing circuitry includes a first arithmetic circuitry for selecting one or more categories of each bounding box as the subset of the output data by comparing the multiple category scores of the multiple categories of each bounding box.

3. The neural processing circuit as described in claim 2, characterized in that, The post-processing circuit further includes: A second computational circuit is used to extract one or more bounding boxes by comparing a category confidence score for each bounding box with a threshold confidence score, wherein the category confidence score represents the probability of an object of a category appearing in each bounding box and is derived from an object presence confidence score and multiple category scores, wherein the object presence confidence score is included in the output data and indicates the probability of an object appearing in each bounding box.

4. The neural processing circuit as described in claim 3, characterized in that, The second operational circuit is used to calculate the category confidence score as the product of the object existence confidence score and a category score of the subset extracted by the first operational circuit.

5. The neural processing circuit as described in claim 3, characterized in that, The post-processing circuit further includes an internal memory coupled to the first arithmetic circuit and the second arithmetic circuit, and the internal memory is used for: The subset of each bounding box extracted by the first arithmetic circuit is stored, and The data of one or more bounding boxes extracted by the second operational circuit are stored.

6. The neural processing circuit as described in claim 3, characterized in that, The post-processing circuit further includes an internal processing circuit for performing a nonmaximum suppression operation on the one or more bounding boxes extracted by the second arithmetic circuit.

7. The neural processing circuit as described in claim 6, characterized in that, During the execution of multiple convolution operations by the processing element array, the internal processing circuitry performs the nonmaximum suppression operation.

8. The neural processing circuit as described in claim 7, characterized in that, The internal processing circuitry is used to perform the nonmaximum suppression operation on a subsequent image following the image at the later of the following two time points: (i) the completion time of the nonmaximum suppression operation on the image data, and (ii) the completion time of the processing element array circuitry's multiple convolution operations on the image.

9. The neural processing circuit as described in claim 2, characterized in that, The output data also includes the coordinates of each bounding box.

10. The neural processing circuit as described in claim 2, characterized in that, The post-processing circuit further includes an internal memory for storing the subset of each bounding box extracted by the first arithmetic circuit.

11. The neural processing circuit as described in claim 2, characterized in that, The first computing circuit performs multiple comparisons of the category scores during the period when the processing element array circuit performs multiple convolution operations.

12. The neural processing circuit as described in claim 2, characterized in that, It also includes: One or more processors; and A memory store, when multiple instructions of a compiler are executed by one or more processors, causes the one or more processors to add a class-specific maximum layer to generate a neural network model of that class. The category subset extracted by the first arithmetic circuit corresponds to multiple operations of the maximum value layer of that category.

13. A method for neural-like processing, characterized in that, Include: Output data is generated by performing convolution operations on a neural network model in a processing element array circuit of a type of neural processing circuit; Receive the output data generated by a post-processing circuit of this type of neural processing circuit; Extract a subset of the data output from the post-processing circuit; as well as A subsequent circuit of this type of neural processing circuit selectively stores or performs operations on the extracted subset of output data.

14. The method as described in claim 13, characterized in that, The output data includes, for each bounding box within the image region, a category score indicating the probability of each object category appearing in each bounding box, and wherein, in a first operational circuit of the neural processing circuit, one or more categories of each bounding box are extracted as a subset of the output data by comparing the category scores of each bounding box.

15. The method as described in claim 14, characterized in that, More includes One or more bounding boxes are extracted by comparing a category confidence score of each bounding box with a threshold confidence score in a second operation circuit of the neural processing circuit. The category confidence score represents the probability that an object in a certain category appears in each bounding box and is derived from an object presence confidence score and the category score. The object presence confidence score is included in the output data and indicates the probability that an object appears in each bounding box.

16. The method as described in claim 15, characterized in that, The category confidence score is the product of the object existence confidence score and a category score from the category subset extracted by the first operational circuit.

17. The method as described in claim 15, characterized in that, It also includes: The category subset of each bounding box extracted by the first arithmetic circuit is stored in an internal memory of the post-processing circuit; and The category subset of each bounding box extracted by the second arithmetic circuit is stored in the internal memory of the post-processing circuit.

18. The method as described in claim 15, characterized in that, It also includes: The post-processing circuit performs a nonmaximum suppression operation on one or more bounding boxes extracted by the second arithmetic circuit.

19. The method as described in claim 18, characterized in that, When the processing element array performs a convolution operation, the internal processing circuit performs a nonmaximum suppression operation.

20. A neural processing circuit, characterized in that, Include: An array of processing elements is used to perform convolution operations on a type of neural network model to produce output data for detecting objects within an image. This output data includes, for each bounding box located within the image area: An object has a confidence score, indicating the probability that the object exists within each bounding box, and Multiple category scores indicate the probability that each object category appears within each bounding box; as well as An arithmetic circuit, coupled to the processing element array circuit, is used to extract one or more bounding boxes by comparing a category confidence score of each bounding box with a threshold confidence score, the category confidence score representing the probability that an object in a certain category appears in each bounding box, the category confidence score being derived based on an object presence confidence score and the category score.