Processing element, Neural processing device including same and Method for calculating thereof
Patent Information
- Application Number
- KR1020220030597
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-03-11
Smart Images

Figure 112022026490818-PAT00011_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a processing element, a neural processing device including the same, and a method of computation thereof. Specifically, the present invention relates to a neural processing device that efficiently converts a precision depending on whether there is an overflow or an underflow, and a pruning method thereof. Background Technology
[0003] Over the past few years, Artificial Intelligence (AI) has been cited globally as the most promising technology and a core component of the Fourth Industrial Revolution. The biggest challenge facing this AI technology is computing performance. For AI technology to realize human learning, reasoning, perception, and natural language processing capabilities, rapidly processing large amounts of data is paramount.
[0004] Although the central processing unit (CPU) or graphics processing unit (GPU) of conventional computers was used for deep learning training and inference in early artificial intelligence, there are limitations to deep learning training and inference tasks that have high workloads, so neural processing units (NPUs) that are structurally specialized for deep learning tasks are gaining attention.
[0005] Neural network processing units can generally utilize data of a specific precision. While a larger number of bits allows for more precise data representation, it may require significantly more hardware resources. Prior art literature
[0007] Registered Patent Publication No. 10-2258566 The problem to be solved
[0008] The objective of the present invention is to provide a processing element that improves accuracy through precision conversion during data computation.
[0009] Another objective of the present invention is to provide a neural processing device comprising a processing element that improves accuracy through precision transformation during data computation.
[0010] Another objective of the present invention is to provide a method for computation of a neural processing device that improves accuracy through precision transformation during data computation.
[0011] The objects of the present invention are not limited to those mentioned above, and other unmentioned objects and advantages of the present invention may be understood from the following description and will be more clearly understood by the embodiments of the present invention. Furthermore, it will be readily apparent that the objects and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims. means of solving the problem
[0013] A processing element according to some embodiments of the present invention for solving the above problem includes a weight register that receives and stores a weight, an input activation register that stores an input activation, a flexible multiplier that receives the weight and the input activation and performs a multiplication operation with a first precision or a second precision different from the first precision to generate result data depending on a mode signal, whether there is an overflow and whether there is an underflow, and a saturating adder that receives the result data and generates a partial sum.
[0014] Additionally, the flexible multiplier may include a detection unit that generates a detection result by checking whether an overflow or underflow occurs according to the multiplication operation of the weight and the input activation, a mode select logic that generates a mode select signal considering the detection result and the mode signal, a first multiplier that performs a multiplication operation with the first precision, a second multiplier that performs a multiplication operation with the second precision, and a demultiplexer that receives the mode select signal, selects either the first multiplier or the second multiplier, and transmits the weight and the input activation.
[0015] In addition, the first multiplier may be k, and the second multiplier may be 2k.
[0016] In addition, the first precision may be 2N bits, and the second precision may be N bits.
[0017] In addition, the first precision may be INT4, and the second precision may be INT2.
[0018] Additionally, the flexible multiplier may further include a multiplexer that receives a calculation result from the first multiplier or the second multiplier and generates a sign bit representing a sign and a product bit representing a magnitude.
[0019] In addition, the result data may include the sign bit and the product bit.
[0020] Additionally, the mode signal is either a first mode signal for the first precision or a second mode signal for the second precision, and the detection result includes a first result in which the overflow or underflow occurs and a second result in which the overflow or underflow does not occur, and the mode selection signal is generated identically to the mode signal when the mode selection logic receives the second result, and can be generated as the first mode signal regardless of the mode signal when the mode selection logic receives the first result.
[0021] Additionally, the detection unit may include a bit divider that divides the weight and the input activation into preset bit units, an overflow detector that generates the detection result and outputs the weight and the input activation as the second precision if the detection result is the second result, and a converting module that receives the weight and the input activation and converts them into the first precision and outputs them when the detection result is the first result.
[0022] A neural processing device according to some embodiments of the present invention for solving other problems described above includes at least one neural core, wherein the neural core includes a processing unit that performs operations and an L0 memory that stores input / output data of the processing unit, wherein the processing unit includes a PE array that includes at least one processing element, and wherein the PE array includes a flexible multiplier that receives weights and input activations and performs a multiplication operation with a first precision or a second precision different from the first precision to generate result data according to a mode signal, whether there is an overflow and whether there is an underflow, and a saturating adder that receives the result data and generates a partial sum.
[0023] Additionally, the above weight and the above input activation can be represented by the above second precision.
[0024] Additionally, the flexible multiplier can convert the weight and the input activation into the first precision, respectively, when an overflow or underflow occurs when the result of the multiplication operation of the weight and the input activation is expressed as the second precision.
[0025] In addition, the flexible multiplier can perform a multiplication operation by selecting either the first precision or the second precision according to the mode signal when the result of the multiplication operation does not result in the overflow and underflow.
[0026] Additionally, it may further include an L2 shared memory shared by at least one neural core and a local interconnection for transmitting data between the L2 shared memory and at least one neural core.
[0027] A method of operation of a neural processing device according to some embodiment of the present invention for solving the above-mentioned additional problem comprises determining whether an overflow or underflow occurs during the multiplication of a weight and an input activation, and if a mode signal selects a first precision or if the overflow or underflow occurs, converting the weight and the input activation to the first precision, and if the mode signal selects a second precision and if the overflow or underflow does not occur, maintaining the weight and the input activation as the second precision, multiplying the weight and the input activation to generate result data, and accumulating the result data to generate a partial sum.
[0028] In addition, the first precision can use twice the number of bits as the second precision.
[0029] In addition, the second precision can be expressed as symmetric quantization or asymmetric quantization.
[0030] Additionally, the second precision may include a first bit representing a sign and a second bit representing a magnitude.
[0031] Additionally, before determining whether the overflow or underflow occurs, it may further include dividing the weight and the input data.
[0032] In addition, generating the result data may include selecting either a first multiplier corresponding to the first precision and a second multiplier corresponding to the second precision to generate the result data. Effects of the invention
[0034] The processing element of the present invention, the neural processing device including the same, and the computation method thereof can select the necessary precision according to the mode signal.
[0035] In addition, if an overflow or underflow occurs prior to the mode signal, a precision conversion is performed to increase precision.
[0036] In addition to the above, the specific effects of the present invention are described together with the specific details for implementing the invention below. Brief explanation of the drawing
[0038] FIG. 1 is a block diagram illustrating a neural processing system according to some embodiments of the present invention. Figure 2 is a block diagram for explaining the neural processing unit of Figure 1 in detail. Figure 3 is a block diagram for explaining the neural core SoC of Figure 2 in detail. Figure 4 is a structural diagram for explaining the global interconnection of Figure 3 in detail. Figure 5 is a block diagram for explaining the neural processor of Figure 3 in detail. FIG. 6 is a diagram illustrating the hierarchical structure of a neural processing device according to some embodiments of the present invention. Figure 7 is a block diagram for explaining the neural core of Figure 5 in detail. Figure 8 is a block diagram for explaining the LSU of Figure 7 in detail. FIG. 9 is a block diagram for explaining the processing unit of FIG. 7 in detail. FIG. 10 is a block diagram for explaining the processing element of FIG. 9 in detail. FIG. 11 is a block diagram for explaining the flexible multiplier of FIG. 10 in detail. FIG. 12 is an example diagram for explaining the first and second precisions. FIG. 13 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a first mode signal. FIG. 14 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a second mode signal and a second result. FIG. 15 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a second mode signal and a first result. FIG. 16 is a block diagram for explaining the detection unit of FIG. 11 in detail. Figure 17 is a block diagram for explaining the L0 memory of Figure 7 in detail. FIG. 18 is a block diagram for explaining the local memory bank of FIG. 21 in detail. FIG. 19 is a block diagram for explaining in detail the structure of the neural processing device of FIG. 1. FIG. 20 is a block diagram illustrating the memory reconfiguration of the neural processing system of FIG. 1. FIG. 21 is a block diagram showing an example of memory reorganization of the neural processing system of FIG. 1. FIG. 22 is an enlarged block diagram of part A of FIG. 24. FIG. 23 is a drawing for explaining the first memory bank of FIG. 26 in detail. FIG. 24 is a block diagram illustrating the software layer structure of the neural processing device of FIG. 1. FIG. 25 is a conceptual diagram illustrating a deep learning operation performed by the neural processing unit of FIG. 1. FIG. 26 is a conceptual diagram illustrating the learning and inference operations of the neural network of the neural processing device of FIG. 1. FIG. 27 is a flowchart illustrating a method of computation for a neural processing device according to some embodiments of the present invention. Specific details for implementing the invention
[0039] Terms and words used in this specification and claims shall not be interpreted as being limited to their general or dictionary meanings. In accordance with the principle that an inventor may define the concept of a term or word to best describe their invention, they shall be interpreted in a meaning and concept consistent with the technical spirit of the invention. Furthermore, since the embodiments described in this specification and the configurations illustrated in the drawings are merely one embodiment of the invention and do not represent the entire technical spirit of the invention, it should be understood that various equivalents, modifications, and applicable examples capable of replacing them may exist at the time of filing this application.
[0040] The terms first, second, A, B, etc., as used in this specification and claims may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0041] The terms used in this specification and claims are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" should be understood as not precluding the existence or addition of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification.
[0042] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which this invention pertains.
[0043] Terms such as those defined in commonly used dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.
[0044] In addition, each component, process, procedure, or method included in each embodiment of the present invention may be shared within a scope that is not technically contradictory to one another.
[0046] Hereinafter, a neural processing apparatus according to several embodiments of the present invention will be described with reference to FIGS. 1 to 28.
[0047] FIG. 1 is a block diagram illustrating a neural processing system according to some embodiments of the present invention.
[0048] Referring to FIG. 1, a neural processing system (NPS) according to some embodiments of the present invention may include a first neural processing device (1), a second neural processing device (2), and an external interface (3).
[0049] The first neural processing device (1) may be a device that performs calculations using an artificial neural network. The first neural processing device (1) may be, for example, a device specialized in performing deep learning calculation tasks. However, the present embodiment is not limited thereto.
[0050] The second neural processing device (2) may be a device having the same or similar configuration as the first neural processing device (1). The first neural processing device (1) and the second neural processing device (2) may be connected to each other through an external interface (3) to share data and control signals.
[0051] Although two neural processing devices are illustrated in FIG. 1, the neural processing system (NPS) according to some embodiments of the present invention is not limited thereto. That is, the neural processing system (NPS) according to some embodiments of the present invention may have three or more neural processing devices connected to each other through an external interface (3). Alternatively, the neural processing system (NPS) according to some embodiments of the present invention may include only one neural processing device.
[0052] Figure 2 is a block diagram for explaining the neural processing unit of Figure 1 in detail.
[0053] Referring to FIG. 2, the first neural processing device (1) may include a neural core SoC (10), a CPU (20), an off-chip memory (30), a first non-volatile memory interface (40), a first volatile memory interface (50), a second non-volatile memory interface (60), and a second volatile memory interface (70).
[0054] The neural core SoC (10) may be a System on Chip device. The neural core SoC (10) may be an accelerator as an artificial intelligence computing unit. The neural core SoC (10) may be, for example, any one of a GPU (graphics processing unit), an FPGA (field programmable gate array), and an ASIC (application-specific integrated circuit). However, the present embodiment is not limited thereto.
[0055] The neural core SoC (10) can exchange data with other external computation units through an external interface (3). Additionally, the neural core SoC (10) can be connected to a non-volatile memory (31) and a volatile memory (32), respectively, through a first non-volatile memory interface (40) and a first volatile memory interface (50).
[0056] The CPU (20) may be a control unit that controls the system of the first neural processing unit (1) and executes program operations. As a general-purpose computing unit, the CPU (20) may have low efficiency in performing parallel simple operations that are frequently used in deep learning. Therefore, the neural core SoC (10) can have high efficiency by performing operations on deep learning inference and learning tasks.
[0057] The CPU (20) can exchange data with other external computing units through the external interface (3). Additionally, the CPU (20) can be connected to the non-volatile memory (31) and the volatile memory (32), respectively, through the second non-volatile memory interface (60) and the second volatile memory interface (70).
[0058] The off-chip memory (30) may be memory placed outside the chip of the neural core SoC (10). The off-chip memory (30) may include non-volatile memory (31) and volatile memory (32).
[0059] Non-volatile memory (31) may be a memory that retains stored information even when power is not supplied. Non-volatile memory (31) is, for example, ROM (Read-Only Memory), PROM (Programmable Read-Only Memory), EAROM (Erasable Alterable ROM), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory) (e.g., NAND Flash memory, NOR Flash memory), UVEPROM (Ultra-Violet Erasable Programmable Read-Only Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory), PRAM (Phase-change Random Access Memory), SONOS (silicon-oxide-nitride-oxide-silicon), RRAM (Resistive Random Access Memory), NRAM (Nanotube Random Access Memory), magnetic computer memory devices (e.g., hard disk, floppy disk drive, magnetic tape), optical disc drive and 3D crosspoint It may include at least one of a memory (3D XPoint memory). However, the present embodiment is not limited thereto.
[0060] Volatile memory (32), unlike non-volatile memory (31), may be a memory that continuously requires power to maintain stored information. Volatile memory (32) may include, for example, at least one of DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), SDRAM (Synchronous Dynamic Random Access Memory), and DDR SDRAM (Double Data Rate SDRAM). However, the present embodiment is not limited thereto.
[0061] The first non-volatile memory interface (40) and the second non-volatile memory interface (60) may each include, for example, at least one of PATA (Parallel Advanced Technology Attachment), SCSI (Small Computer System Interface), SAS (Serial Attached SCSI), SATA (Serial Advanced Technology Attachment), and PCIe (PCI Express). However, the present embodiment is not limited thereto.
[0062] The first volatile memory interface (50) and the second volatile memory interface (70) may each be, for example, at least one of SDR (Single Data Rate), DDR (Double Data Rate), QDR (Quad Data Rate), and XDR (eXtreme Data Rate, Octal Data Rate). However, the present embodiment is not limited thereto.
[0063] Figure 3 is a block diagram for explaining the neural core SoC of Figure 2 in detail.
[0064] Referring to FIGS. 2 and 3, the neural core SoC (10) may include at least one neural processor (1000), shared memory (2000), Direct Memory Access (DMA) (3000), a non-volatile memory controller (4000), a volatile memory controller (5000), and a global interconnection (5000).
[0065] A neural processor (1000) may be a computational unit that directly performs computational tasks. If there are multiple neural processors (1000), computational tasks may be assigned to each neural processor (1000). Each neural processor (1000) may be connected to each other through a global interconnection (5000).
[0066] The shared memory (2000) may be a memory shared by multiple neural processors (1000). The shared memory (2000) may store data from each neural processor (1000). Additionally, the shared memory (2000) may receive data from the off-chip memory (30), temporarily store it, and transfer it to each neural processor (1000). Conversely, the shared memory (2000) may receive data from the neural processor (1000), temporarily store it, and transfer it to the off-chip memory (30) of FIG. 2.
[0067] The shared memory (2000) may require a relatively fast memory. Accordingly, the shared memory (2000) may include, for example, SRAM. However, the present embodiment is not limited thereto. That is, the shared memory (2000) may include DRAM.
[0068] The shared memory (2000) may be a memory corresponding to the SoC level, i.e., L3 (level 3). Therefore, the shared memory (2000) may be defined as an L3 shared memory.
[0069] The DMA (3000) can directly control the movement of data without the need for the neural processor (1000) to control the input and output of data. Accordingly, the DMA (3000) can control the movement of data between memories to minimize the number of interrupts of the neural processor (1000).
[0070] The DMA (3000) can control data movement between the shared memory (2000) and the off-chip memory (30). Through the authority of the DMA (3000), the non-volatile memory controller (4000) and the volatile memory controller (5000) can perform data movement.
[0071] The non-volatile memory controller (4000) can control read or write operations on the non-volatile memory (31). The non-volatile memory controller (4000) can control the non-volatile memory (31) through the first non-volatile memory interface (40).
[0072] The volatile memory controller (5000) can control read or write operations on the volatile memory (32). Additionally, the volatile memory controller (5000) can perform refresh operations on the volatile memory (32). The volatile memory controller (5000) can control the non-volatile memory (31) through the first volatile memory interface (50).
[0073] A global interconnection (5000) can connect at least one neural processor (1000), shared memory (2000), DMA (3000), non-volatile memory controller (4000), and volatile memory controller (5000) to each other. Additionally, an external interface (3) can also be connected to the global interconnection (5000). The global interconnection (5000) may be a path through which data travels between at least one neural processor (1000), shared memory (2000), DMA (3000), non-volatile memory controller (4000), volatile memory controller (5000), and the external interface (3).
[0074] The global interconnection (5000) can transmit signals for the transmission and synchronization of control signals as well as data. That is, in some embodiments of the present invention, the neural processing device does not have a separate control processor managing the synchronization signals, but rather each neural processor (1000) can directly transmit and receive the synchronization signals. Accordingly, the latency of the synchronization signals generated by the control processor can be blocked.
[0075] That is, when there are multiple neural processors (1000), there may be dependencies of individual tasks such that the next neural processor (1000) can start a new task only after the task of one neural processor (1000) has been completed. The completion and start of these individual tasks can be verified through a synchronization signal, and in conventional technology, the control processor performed the reception of such synchronization signal and the instruction to start a new task.
[0076] However, as the number of neural processors (1000) increases and the dependencies of the tasks become more complex, the number of requests and instructions for these synchronization tasks increases exponentially. Consequently, the latency associated with each request and instruction can significantly reduce the efficiency of the task.
[0077] Accordingly, in some embodiments of the present invention, the neural processing device may have each neural processor (1000) directly transmit a synchronization signal to another neural processor (1000) according to the dependency of the task, instead of a control processor. In this case, compared to a method managed by a control processor, multiple neural processors (1000) can perform synchronization tasks in parallel, thereby minimizing the latency associated with synchronization.
[0078] In addition, the control processor must perform task scheduling of neural processors (1000) according to task dependencies, and the overhead of such scheduling can increase significantly as the number of neural processors (1000) increases. Therefore, in some embodiments of the present invention, the neural processing device can have the scheduling task performed by individual neural processors (1000), and thus the scheduling burden is eliminated, thereby improving the performance of the device.
[0079] Figure 4 is a structural diagram for explaining the global interconnection of Figure 3 in detail.
[0080] Referring to FIG. 4, the global interconnection (5000) may include a data channel (5100), a control channel (5200), and an L3 sink channel (5300).
[0081] The data channel (5100) may be a dedicated channel for transmitting data. Through the data channel (5100), at least one neural processor (1000), shared memory (2000), DMA (3000), non-volatile memory controller (4000), volatile memory controller (5000), and external interface (3) can exchange data with each other.
[0082] The control channel (5200) may be a dedicated channel for transmitting control signals. Through the control channel (5200), at least one neural processor (1000), shared memory (2000), DMA (3000), non-volatile memory controller (4000), volatile memory controller (5000), and external interface (3) can exchange control signals with each other.
[0083] The L3 sink channel (5300) may be a dedicated channel for transmitting synchronization signals. Through the L3 sink channel (5300), at least one neural processor (1000), shared memory (2000), DMA (3000), non-volatile memory controller (4000), volatile memory controller (5000), and external interface (3) can exchange synchronization signals with each other.
[0084] The L3 sink channel (5300) is set as a dedicated channel within the global interconnection (5000) so that it can rapidly transmit synchronization signals without overlapping with other channels. Accordingly, the neural processing device according to some embodiments of the present invention does not require new wiring work and can smoothly perform synchronization work using the existing global interconnection (5000).
[0085] Figure 5 is a block diagram for explaining the neural processor of Figure 3 in detail.
[0086] Referring to FIGS. 3 to 5, the neural processor (1000) may include at least one neural core (100), an L2 shared memory (400), a local interconnection (200), and an L2 sink path (300).
[0087] At least one neural core (100) can share and perform the work of the neural processor (1000). For example, there may be eight neural cores (100). However, the present embodiment is not limited thereto. Although FIGS. 3 and 5 show that multiple neural cores (100) are included in the neural processor (1000), the present embodiment is not limited thereto. That is, the neural processor (1000) can be configured with only one neural core (100).
[0088] The L2 shared memory (400) may be a memory shared by each neural core (100) within the neural processor (1000). The L2 shared memory (400) may store data of each neural core (100). Additionally, the L2 shared memory (400) may receive data from the shared memory (2000) of FIG. 4, temporarily store it, and transfer it to each neural core (100). Conversely, the L2 shared memory (400) may receive data from the neural core (100), temporarily store it, and transfer it to the shared memory (2000) of FIG. 3.
[0089] L2 shared memory (400) may be memory corresponding to the neural processor level, i.e., L2 (level 2). L3 shared memory, i.e., shared memory (2000), may be shared by the neural processor (1000), and L2 shared memory (400) may be shared by the neural core (100).
[0090] A local interconnection (200) can connect at least one neural core (100) and an L2 shared memory (400) to each other. The local interconnection (200) may be a path for data to travel between at least one neural core (100) and an L2 shared memory (400). The local interconnection (200) can be connected to the global interconnection (5000) of FIG. 3 to transmit data.
[0091] The L2 sink path (300) can connect at least one neural core (100) and an L2 shared memory (400) to each other. The L2 sink path (300) may be a path through which a synchronization signal of at least one neural core (100) and an L2 shared memory (400) travels.
[0092] The L2 sync path (300) can be physically formed separately from the local interconnection (200). Unlike the global interconnection (5000), the local interconnection (200) may not have sufficient internal channels. In such cases, the L2 sync path (300) can be formed separately to perform the transmission of synchronization signals quickly and without delay. The L2 sync path (300) can be used for synchronization performed at a level one step lower than the L3 sync channel (5300) of the global interconnection (5000).
[0093] FIG. 6 is a diagram illustrating the hierarchical structure of a neural processing device according to some embodiments of the present invention.
[0094] Referring to FIG. 6, the neural core SoC (10) may include at least one neural processor (1000). Each neural processor (1000) can transmit data to each other through a global interconnection (5000).
[0095] Each neural processor (1000) may include at least one neural core (100). The neural core (100) may be a processing unit optimized for deep learning computation tasks. The neural core (100) may be a processing unit corresponding to one operation of a deep learning computation task. That is, a deep learning computation task may be expressed as a sequential or parallel combination of multiple operations. Each neural core (100) may be a processing unit capable of processing one operation, and may be the minimum computation unit that can be considered for scheduling from the perspective of a compiler.
[0096] The neural processing device according to the present embodiment can facilitate fast and efficient scheduling and computational operations by configuring the minimum computational unit considered from the perspective of compiler scheduling and the hardware processing unit to have the same scale.
[0097] In other words, if the hardware's divisible processing unit is excessively large compared to the computational task, inefficiency in the operation of the processing unit may occur. Conversely, scheduling a processing unit smaller than an operation, which is the minimum scheduling unit of a compiler, every time is not appropriate because it can lead to scheduling inefficiency and increase hardware design costs.
[0098] Therefore, this embodiment can simultaneously achieve fast scheduling of computational tasks and efficient execution of computational tasks without wasting hardware resources by coordinating the scale of the compiler's scheduling unit and the hardware processing unit to be similar.
[0099] Figure 7 is a block diagram for explaining the neural core of Figure 5 in detail.
[0100] Referring to FIG. 7, the neural core (100) may include an LSU (Load / Store Unit) (110), an L0 memory (120), a weight buffer (130), an activation LSU (140), an activation buffer (150), and a processing unit (160).
[0101] The LSU (110) can receive at least one of data, control signals, and synchronization signals from the outside through the local interconnection (200) and the L2 sink path (300). The LSU (110) can transmit at least one of the received data, control signals, and synchronization signals to the L0 memory (120). Similarly, the LSU (110) can transmit at least one of the data, control signals, and synchronization signals to the outside through the local interconnection (200) and the L2 sink path (300).
[0102] Figure 8 is a block diagram for explaining the LSU of Figure 7 in detail.
[0103] Referring to FIG. 8, the LSU (110) may include a local memory load unit (111a), a local memory store unit (111b), a neural core load unit (112a), a neural core store unit (112b), a load buffer (LB), a store buffer (SB), a load engine (113a), a store engine (113b), and a conversion index buffer (114).
[0104] The local memory load unit (111a) can fetch a load instruction for L0 memory (120) and issue a load instruction. When the local memory load unit (111a) provides the issued load instruction to the load buffer (LB), the load buffer (LB) can sequentially send memory access requests to the load engine (113a) according to the order in which they were input.
[0105] Additionally, the local memory store unit (111b) can fetch store instructions for L0 memory (120) and issue store instructions. When the local memory store unit (111b) provides the issued store instructions to the store buffer (SB), the store buffer (SB) can sequentially send memory access requests to the store engine (113b) according to the order in which they were input.
[0106] The neural core load unit (112a) can fetch load instructions for the neural core (100) and issue load instructions. When the neural core load unit (112a) provides the issued load instructions to the load buffer (LB), the load buffer (LB) can sequentially send memory access requests to the load engine (113a) according to the order in which they were input.
[0107] Additionally, the neural core store unit (112b) can fetch store instructions for the neural core (100) and issue store instructions. When the neural core store unit (112b) provides the issued store instructions to the store buffer (SB), the store buffer (SB) can sequentially send memory access requests to the store engine (113b) according to the order in which they were input.
[0108] The load engine (113a) can receive a memory access request and retrieve data through the local interconnection (200). At this time, the load engine (113a) can quickly find data by using the translation table of recently used virtual and physical addresses in the translation index buffer (114). If the virtual address of the load engine (113a) is not in the translation index buffer (114), address translation information can be found in another memory.
[0109] The store engine (113b) can receive a memory access request and retrieve data through the local interconnection (200). At this time, the store engine (113b) can quickly find the data by using the translation table of recently used virtual and physical addresses in the translation index buffer (114). If the virtual address of the store engine (113b) is not in the translation index buffer (114), the address translation information can be found in another memory.
[0110] The load engine (113a) and the store engine (113b) can send a synchronization signal to the L2 sink path (300). At this time, the synchronization signal may indicate that the operation has been completed.
[0111] Referring again to FIG. 7, L0 memory (120) is a memory located inside the neural core (100) that can receive all input data required for the operation from the outside and temporarily store it. Additionally, L0 memory (120) can temporarily store output data computed by the neural core (100) to transmit it to the outside. L0 memory (120) can perform the role of a cache memory for the neural core (100).
[0112] The L0 memory (120) can transmit an input activation (Act_In) to an activation buffer (150) via an activation LSU (140) and receive an output activation (Act_Out). In addition to the activation LSU (140), the L0 memory (120) can also directly transmit and receive data with a processing unit (160). That is, the L0 memory (120) can exchange data with each of the PE array (163) and the vector unit (164).
[0113] L0 memory (120) may be memory corresponding to the neural core level. In this case, unlike L2 shared memory (400) and shared memory (2000), L0 memory (120) may not be shared and may operate as private memory for the neural core.
[0114] The L0 memory (120) can transmit data such as activation or weight through a data path. The L0 memory (120) can send and receive synchronization signals through a separate dedicated path, the L3 sync path. The L0 memory (120) can send and receive synchronization signals through the L3 sync path with, for example, the LSU (110), the weight buffer (130), the activation LSU (140), and the processing unit (160).
[0115] The weight buffer (130) can receive weights from the L0 memory (120). The weight buffer (130) can transfer weights to the processing unit (160). The weight buffer (130) can temporarily store weights before transferring them.
[0116] Input activation (Act_In) and output activation (Act_Out) can refer to the input and output values of a layer in a neural network. In this case, if there are multiple layers in the neural network, the output value of the previous layer becomes the input value of the next layer, so the output activation (Act_Out) of the previous layer can be used as the input activation (Act_In) of the next layer.
[0117] Weight can refer to a parameter that is multiplied by the input activation (Act_In) in each layer. Weight is adjusted and fixed during the deep learning training phase, and can be used to derive the output activation (Act_Out) through a fixed value during the inference phase.
[0118] The activation LSU (140) can transfer an input activation (Act_In) from the L0 memory (120) to the activation buffer (150) and transfer an output activation (Act_Out) from the activation buffer (150) to the on-chip buffer. That is, the activation LSU (140) can perform both the load and store operations of the activation.
[0119] The activation buffer (150) can provide an input activation (Act_In) to the processing unit (160) and receive an output activation (Act_Out) from the processing unit (160). The activation buffer (150) can temporarily store the input activation (Act_In) and the output activation (Act_Out).
[0120] The activation buffer (150) can quickly provide activation to a processing unit (160) with a large amount of computation, particularly a PE array (163), and can quickly receive activation to increase the computation speed of the neural core (100).
[0121] The processing unit (160) may be a module that performs operations. The processing unit (160) can perform not only one-dimensional operations but also two-dimensional matrix operations, i.e., convolution operations. The processing unit (160) can receive an input activation (Act_In), multiply it by a weight, and then add the result to generate an output activation (Act_Out).
[0122] FIG. 9 is a block diagram for explaining the processing unit of FIG. 7 in detail.
[0123] Referring to FIGS. 7 and 9, the processing unit (160) may include a PE array (163), a vector unit (164), a column register (161), and a row register (162).
[0124] The PE array (163) can receive input activations (Act_In) and weights and perform multiplication. At this time, the input activations (Act_In) and weights can each be computed through convolution in the form of a matrix. Through this, the PE array (163) can generate output activations (Act_Out). However, the present embodiment is not limited thereto. The PE array (163) can generate any other type of output besides output activations (Act_Out).
[0125] The PE array (163) may include at least one processing element (163_1). The processing elements (163_1) may be aligned with each other to perform multiplication for one input activation (Act_In) and one weight (Weight).
[0126] The PE array (163) can generate partial sums by summing the values for each multiplication. These partial sums can be used as output activations (Act_Out). Since the PE array (163) performs two-dimensional matrix multiplication, it may also be referred to as a two-dimensional matrix compute unit (2D matrix compute unit).
[0127] The vector unit (164) can perform one-dimensional operations. The vector unit (164) can perform deep learning operations together with the PE array (163). Through this, the processing unit (160) can be specialized for the necessary operations. That is, the neural core (100) has a computation module that performs a large amount of two-dimensional matrix multiplication and one-dimensional operations, respectively, so that deep learning tasks can be performed efficiently.
[0128] The column register (161) can receive a first input (I1). The column register (161) can receive the first input (I1), divide it, and provide it to each column of the processing element (163_1).
[0129] The row register (162) can receive the second input (I2). The row register (162) can receive the second input (I2), divide it, and provide it to each row of the processing element (163_1).
[0130] The first input (I1) may be an input activation (Act_In) or a weight. The second input (I2) may be a value other than the first input (I1) among the input activation (Act_In) or the weight. Alternatively, the first input (I1) and the second input (I2) may be values other than the input activation (Act_In) and the weight.
[0131] FIG. 10 is a block diagram for explaining the processing element of FIG. 9 in detail.
[0132] Referring to FIG. 10, the processing element (163_1) may include a weight register (WR), an input activation register (ACR), a flexible multiplier (FM), and a saturating adder (SA).
[0133] The weight register (WR) can receive and store the weight input to the processing element (163_1). The weight register (WR) can transfer the weight to the flexible multiplier (FM).
[0134] The Input Activation Register (ACR) can receive and store the Input Activation (Act_In). The Input Activation Register (ACR) can transmit the Input Activation (Act_In) to the Flexible Multiplier (FM).
[0135] The flexible multiplier (FM) can receive weights and input activations (Act_In). The flexible multiplier (FM) can perform multiplication of weights and input activations (Act_In). The flexible multiplier (FM) can receive a mode signal. In this case, the mode signal may be a signal indicating which of the first precision and the second precision will be used to perform the operation.
[0136] A flexible multiplier (FM) can output a multiplication result as result data. The result data may include a sign bit (SB) and a product bit (PB). In this case, the sign bit (SB) may be a bit indicating the sign of the result data. The product bit (PB) may be a bit indicating the magnitude of the result data. The flexible multiplier (FM) can output the result data as a first precision or a second precision.
[0137] The saturating adder (SA) can receive result data. That is, the saturating adder (SA) can receive sign bits (SB) and product bits (PB). The saturating adder (SA) can receive and accumulate result data multiple times. Accordingly, the saturating adder (SA) can generate partial sums (Psum). These partial sums (Psum) can be output from each processing element (163_1) and finally summed. However, the present embodiment is not limited thereto.
[0138] FIG. 11 is a block diagram for explaining the flexible multiplier of FIG. 10 in detail.
[0139] Referring to FIG. 11, the flexible multiplier (FM) may include a detection unit (DU), mode select logic (MSL), a demultiplexer (Dx), a first multiplier (Mul1), a second multiplier (Mul2), and a multiplexer (Mx).
[0140] The detection unit (DU) can receive weights and input activations (Act_In). The detection unit (DU) can detect whether an overflow or underflow occurs as a result of multiplying the weights and input activations (Act_In). In this case, an overflow is an error that occurs when the result is greater than the numerical range according to the data precision, and an underflow is an error that occurs when the result is smaller than the numerical range according to the data precision.
[0141] The detection unit (DU) can transmit the weight and input activation (Act_In) to the demultiplexer (Dx). Additionally, the detection unit (DU) can generate a detection result (DR). The detection result (DR) may be a signal indicating whether an overflow or underflow occurs in the result of multiplying the weight and input activation (Act_In). If an overflow or underflow occurs in the result of multiplying the weight and input activation (Act_In), the detection result (DR) may be a first result. Conversely, if an overflow or underflow does not occur in the result of multiplying the weight and input activation (Act_In), the detection result (DR) may be a second result. The detection unit (DU) can transmit the detection result (DR) to the mode select logic (MSL).
[0142] The Mode Select Logic (MSL) can receive a Mode signal. In this case, the Mode signal may be a signal indicating which of the first and second precisions the multiplication operation will be performed in mode. If the Mode signal is a signal for the first precision, it may be a first mode signal. Conversely, if the Mode signal is a signal for the second precision, it may be a second mode signal.
[0143] The mode select logic (MSL) can also receive a detection result (DR). The mode select logic (MSL) can generate a mode select signal (Ms) based on the mode signal (Mode) and the detection result (DR).
[0144] At this time, the mode selection signal (Ms) may be a signal indicating whether to perform a multiplication operation in a mode for either the first precision or the second precision. Unlike the mode signal (Mode), the mode selection signal (Ms) may be a signal that ultimately selects the mode. That is, the precision of the data for the multiplication operation performed by the flexible multiplier (FM) may be determined according to the mode selection signal (Ms).
[0145] The mode selection signal (Ms) may also be either the first mode signal or the second mode signal for the first precision, similar to the mode signal (Mode). In this case, the mode selection signal (Ms) may be the same signal as the mode signal (Mode) or may be different signals.
[0146] The demultiplexer (Dx) can receive weights and input activations (Act_In) from the detection unit (DU). The demultiplexer (Dx) can also receive a mode selection signal (Ms). The demultiplexer (Dx) can transmit the weights and input activations (Act_In) to either a first multiplier (Mul1) or a second multiplier (Mul2). The demultiplexer (Dx) can determine the path through which the weights and input activations (Act_In) are transmitted based on the mode selection signal (Ms). Additionally, the demultiplexer (Dx) can divide at least one weight and at least one input activation (Act_In) into multiple first multipliers (Mul1) or multiple second multipliers (Mul2) and transmit them.
[0147] The first multiplier (Mul1) can perform calculations with the first precision. That is, the first multiplier (Mul1) can receive input data of the first precision. If the demultiplexer (Dx) transmits weights and input activations (Act_In) to the first multiplier (Mul1), the weights and input activations (Act_In) may be in the form of the first precision.
[0148] The second multiplier (Mul2) can perform operations with the second precision. That is, the second multiplier (Mul2) can receive input data of the second precision. If the demultiplexer (Dx) transmits weights and input activations (Act_In) to the second multiplier (Mul2), the weights and input activations (Act_In) may be in the form of the second precision.
[0149] At this time, the number of first multipliers (Mul1) may be k, and the number of second multipliers (Mul2) may be 2k. At this time, k may be a natural number.
[0150] The multiplexer (Mx) can receive the result of an operation, that is, the result of a multiplication operation, from either the first multiplier (Mul1) or the second multiplier (Mul2). The multiplexer (Mx) can receive the result of a multiplication operation between the input data of the first precision and the input data of the first precision from the first multiplier (Mul1), and can receive the result of a multiplication operation between the input data of the second precision and the input data of the second precision from the second multiplier (Mul2).
[0151] When the mode selection signal (Ms) is a first mode signal, the multiplexer (Mx) can generate result data by receiving k operation results provided by k first multiplexers (Mx). The result data may include a sign bit (SB) and a product bit (PB). That is, the multiplexer (Mx) can combine k operation results to generate one result data.
[0152] When the mode selection signal (Ms) is a second mode signal, the multiplexer (Mx) can generate result data by receiving 2k operation results provided by 2k second multiplexers (Mx). The result data may include a sign bit (SB) and a product bit (PB). That is, the multiplexer (Mx) can combine 2k operation results to generate one result data.
[0153] FIG. 12 is an example diagram for explaining the first and second precisions.
[0154] Referring to FIG. 12, the first precision (Pr1) may be 2N bits. Here, N may be a natural number. The second precision (Pr2) may be N bits. That is, the first precision (Pr1) may have twice the number of bits as the second precision (Pr2). For example, the first precision (Pr1) and the second precision (Pr2) may be INT4 and INT2, respectively. Alternatively, the first precision (Pr1) and the second precision (Pr2) may be at least one of INT8 and INT4, INT16 and INT8, and INT32 and INT16, respectively. The first precision (Pr1) and the second precision (Pr2) may be in the form of INT, i.e., integer precisions. However, the present embodiment is not limited thereto.
[0155] In FIG. 12, INT4 and INT2 are illustrated as examples of the first precision (Pr1) and the second precision (Pr2), respectively. '11' is illustrated as an example of the second precision (Pr2), and when converted to the first precision (Pr1), it can be expressed as '0011'. Of course, this is merely one example and is not limited thereto.
[0156] When the second precision (Pr2) is INT2, the number of possible values to represent a general number may be very small. That is, when using 2 bits, only a total of 4 values can be represented. Therefore, by quantizing the 2 bits, a larger number of values can be represented. For example, the second precision (Pr2) includes 2 bits, and the 2 bits may include a first bit representing a sign and a second bit representing a magnitude.
[0157] Referring to the table below, 2-bit precision can be represented by symmetric quantization and asymmetric quantization. In this case, it can be represented by the following example.
[0159] Quantizer Type # of bits Representation Range Symmetric Quantization 2 -Y, -X, X, Y X=1, Y=2,3,4,5,6X=2, Y=3,5,7,9 Asymmetric quantization 2 -A, -B, C, D Any valueA>B, D>C
[0160] In this case, for the 2-bit second precision (Pr2), overflow or underflow may frequently occur due to the multiplication operation. That is, the result of the multiplication operation between the second precisions (Pr2) may appear as a result in which the number of bits of the second precision (Pr2) is doubled. That is, the result of the multiplication operation between INT2 and INT2 can be expressed as INT4.
[0161] However, for example, if '11' in INT2 represents the decimal number 9, the product of '11' and '11' in INT2 is 81 in decimal, so it cannot be represented by 4-bit INT4, which may cause an overflow. In such a case, '11' can be converted to INT4 as '0011' to clearly represent the decimal number 81 through the result of the multiplication operation in INT8.
[0162] Therefore, the present embodiment can perform a conversion that increases the number of bits of data by changing the precision when such overflow or underflow occurs. Through this, a low number of bits with high efficiency is normally used, but when calculations may become inaccurate, it is possible to convert to a high number of bits to improve the precision of calculations while maintaining optimal efficiency.
[0163] In particular, in the case of INT2, the range is narrow and quantization is frequent, so such overflow or underflow can occur very frequently. Since INT2 has high data efficiency due to its small number of bits, it can be highly useful in cases where hardware resources are limited, such as in mobile devices. Therefore, the present embodiment can prevent a decrease in accuracy caused by overflow or underflow that frequently occurs in areas where low-bit precision such as INT2 is utilized.
[0164] FIG. 13 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a first mode signal.
[0165] Referring to FIG. 13, the mode signal (Mode) may be a first mode signal. In this case, the detection result (DR) may be a first result or a second result. The first result is when an overflow or underflow occurs, and the second result may be when neither an overflow nor an underflow occurs.
[0166] When the mode select logic (MSL) receives a first mode signal, it can adopt the first mode signal as a mode select signal (Ms) regardless of the detection result (DR). This is because, even if the detection result (DR) is the first result, using the first precision as the first mode signal can prevent overflow and underflow. Conversely, even if the detection result (DR) is the second result, there is no problem in using the first precision as the first mode signal. Therefore, when the mode signal (Mode) is the first mode signal, the mode select logic (MSL) can be the first mode signal regardless of the detection result (DR).
[0167] In this case, the detection unit (DU) can convert the weight and input activation (Act_In) into a first precision (Pr1) and transmit it to the demultiplexer (Dx). The demultiplexer (Dx) can transmit the weight and input activation (Act_In) to the first multiplier (Mul1). Since there are k first multipliers (Mul1), the demultiplexer (Dx) can divide the weight and input activation (Act_In) and transmit them to the first multiplier (Mul1) respectively.
[0168] Subsequently, k first multipliers (Mul1) can perform multiplication operations with the first precision (Pr1) and transmit k operation results to the multiplexer (Mx). The multiplexer (Mx) can receive k operation results and generate one result data. The result data may include a sign bit (SB) and a product bit (PB).
[0169] That is, in this case, the weight and input activation (Act_In) can be processed through the first path (Path 1) passing through the first multiplier (Mul1).
[0170] FIG. 14 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a second mode signal and a second result.
[0171] Referring to FIG. 14, the mode signal (Mode) may be a second mode signal. In this case, the detection result (DR) may be a second result. The second result may be a case where no overflow or underflow occurs.
[0172] When the mode select logic (MSL) receives the second mode signal, it can generate a mode select signal (Ms) by considering the detection result (DR). The mode select signal (Ms) can adopt the second mode signal as is if the detection result (DR) is the second result. This is because, since there is no overflow or underflow, precision is not reduced even when performing operations with the second precision (Pr2), thus maximizing efficiency with the second precision (Pr2).
[0173] In this case, the detection unit (DU) can transmit the weight and input activation (Act_In) to the second precision (Pr2) and to the demultiplexer (Dx). The demultiplexer (Dx) can transmit the weight and input activation (Act_In) to the second multiplier (Mul2). Since there are 2k of the second multipliers (Mul2), the demultiplexer (Dx) can divide the weight and input activation (Act_In) and transmit them to the second multiplier (Mul2) respectively.
[0174] Subsequently, 2k second multipliers (Mul2) can perform multiplication operations with second precision (Pr2) and transmit 2k operation results to a multiplexer (Mx). The multiplexer (Mx) can receive 2k operation results and generate 1 result data. The result data may include a sign bit (SB) and a product bit (PB).
[0175] That is, in this case, the weight and input activation (Act_In) can be processed through a second path (Path 2) that passes through a second multiplier (Mul2).
[0176] FIG. 15 is a diagram illustrating the operation of the flexible multiplier of FIG. 10 when it receives a second mode signal and a first result.
[0177] Referring to FIG. 15, the mode signal (Mode) may be a second mode signal. In this case, the detection result (DR) may be a first result. The first result may be a case where overflow and underflow occur.
[0178] When the mode select logic (MSL) receives a second mode signal, it can generate a mode select signal (Ms) by considering the detection result (DR). The mode select signal (Ms) may adopt the first mode signal instead of the second mode signal if the detection result (DR) is the first result. This is because precision is reduced when performing operations with the second precision (Pr2) due to the occurrence of overflow and underflow. Therefore, the second precision (Pr2) can be converted to the first precision (Pr1) to prevent a reduction in precision.
[0179] In this case, the detection unit (DU) can transmit the weight and input activation (Act_In) to the first precision (Pr1) and to the demultiplexer (Dx). The demultiplexer (Dx) can transmit the weight and input activation (Act_In) to the first multiplier (Mul1). Since there are k first multipliers (Mul1), the demultiplexer (Dx) can divide the weight and input activation (Act_In) and transmit them to the first multiplier (Mul1), respectively.
[0180] Subsequently, k first multipliers (Mul1) can perform multiplication operations with the first precision (Pr1) and transmit k operation results to the multiplexer (Mx). The multiplexer (Mx) can receive k operation results and generate one result data. The result data may include a sign bit (SB) and a product bit (PB).
[0181] That is, in this case, the weight and input activation (Act_In) can be processed through the first path (Path 1) passing through the first multiplier (Mul1).
[0182] FIG. 16 is a block diagram for explaining the detection unit of FIG. 11 in detail.
[0183] Referring to FIG. 16, the detection unit (DU) may include a bit divider (Bd), an overflow detector (Od), and an overflow detector (Od).
[0184] The bit divider (Bd) can receive a weight and an input activation (Act_In). The bit divider (Bd) can divide the weight and the input activation (Act_In) by a preset number of bits of a second precision (Pr2). Accordingly, the weight and the input activation (Act_In) may each be data of the second precision (Pr2) in multiple numbers.
[0185] The overflow detector (Od) can detect overflow and underflow. The overflow detector (Od) can determine whether an overflow or underflow occurs as a result of the multiplication of each of the multiple weights of the second precision (Pr2) and the multiple input activations (Act_In) of the second precision (Pr2). Accordingly, the overflow detector (Od) can generate a detection result (DR). The detection result (DR) may be a first result if an overflow or underflow occurs. The detection result (DR) may be a second result if no overflow or underflow occurs.
[0186] If it is the first result, the overflow detector (Od) can transmit the weight and input activation (Act_In) to the overflow detector (Od). If it is the second result, the overflow detector (Od) can transmit the weight and input activation (Act_In) directly to the demultiplexer (Dx) without transmitting them to the overflow detector (Od).
[0187] The overflow detector (Od) can convert the weight of the second precision (Pr2) to the first precision (Pr1). Additionally, the overflow detector (Od) can convert the input activation (Act_In) of the second precision (Pr2) to the first precision (Pr1). The overflow detector (Od) can transmit the weight and input activation (Act_In) directly to the demultiplexer (Dx) without transmitting them to the overflow detector (Od).
[0188] This embodiment allows data to be transmitted and processed at a low number of bits under normal conditions. Additionally, in the event of an overflow or underflow that affects precision, the number of bits can be increased to prevent a decrease in precision.
[0189] Figure 17 is a block diagram for explaining the L0 memory of Figure 7 in detail.
[0190] Referring to FIG. 17, the L0 memory (120) may include an arbiter (121) and at least one local memory bank (122).
[0191] When data is stored in the L0 memory (120), the arbiter (121) can receive data from the load engine (113a). At this time, the data can be allocated to a local memory bank (122) in a round-robin manner. Accordingly, the data can be stored in at least one of the local memory banks (122).
[0192] Conversely, when data is loaded from L0 memory (120), the arbiter (121) can receive data from the local memory bank (122) and transfer it to the store engine (113b). The store engine (113b) can store the data externally through the local interconnection (200).
[0193] Figure 18 is a block diagram for explaining the local memory bank of Figure 18 in detail.
[0194] Referring to FIG. 18, the local memory bank (122) may include a local memory bank controller (122_1) and a local memory bank cell array (122_2).
[0195] The local memory bank controller (122_1) can manage read and write operations through the address of the data stored in the local memory bank (122). That is, the local memory bank controller (122_1) can manage the input and output of data overall.
[0196] The local memory bank cell array (122_2) may have a structure in which cells where data is directly stored are aligned in rows and columns. The local memory bank cell array (122_2) may be controlled by the local memory bank controller (122_1).
[0197] FIG. 19 is a block diagram for explaining in detail the structure of a neural processing device according to some embodiments of the present invention.
[0198] Referring to FIG. 19, the neural core (101) may be a CGRA structure, unlike the neural core (100). The neural core (101) may include an instruction memory (111_1), a CGRA L0 memory (111_2), a PE array (111_3), and an LSU (Load / Store Unit) (111_4).
[0199] The instruction memory (111_1) can receive and store instructions. The instruction memory (111_1) can sequentially store instructions internally and provide the stored instructions to the PE array (111_3). At this time, the instructions can direct the operation of the processing elements (111_3a) included in each PE array (111_3).
[0200] The CGRA L0 memory (111_2) is a memory located inside the neural core (101) that can receive all input data required for the operation from the outside and temporarily store it. Additionally, the CGRA L0 memory (111_2) can temporarily store output data computed by the neural core (101) to transmit it to the outside. The CGRA L0 memory (111_2) can serve as a cache memory for the neural core (101).
[0201] The CGRA L0 memory (111_2) can transmit and receive data with the PE array (111_3). The CGRA L0 memory (111_2) may be a memory corresponding to L0 (level 0), which is lower than L1. In this case, the L0 memory may be a private memory of the neural core (101) that is not shared. The CGRA L0 memory (111_2) can transmit data such as activations or weights, programs, etc., to the PE array (111_3).
[0202] The PE array (111_3) may be a module that performs operations. The PE array (111_3) may perform not only one-dimensional operations but also matrix / tensor operations of two dimensions or more. The PE array (111_3) may include multiple processing elements (111_3a) and specific processing elements (111_3b) internally.
[0203] Processing elements (111_3a) and specific processing elements (111_3b) can be arranged in rows and columns. Processing elements (111_3a) and specific processing elements (111_3b) can be arranged in m columns. Additionally, processing elements (111_3a) can be arranged in n rows, and specific processing elements (111_3b) can be arranged in l rows. Accordingly, processing elements (111_3a) and specific processing elements (111_3b) can be arranged in (n+l) rows and m columns.
[0204] The LSU (111_4) can receive at least one of data, control signals, and synchronization signals from the outside through the local interconnection (200). The LSU (111_4) can transmit at least one of the received data, control signals, and synchronization signals to the CGRA L0 memory (111_2). Similarly, the LSU (111_4) can transmit at least one of the data, control signals, and synchronization signals to the outside through the local interconnection (200).
[0205] The neural core (101) may have a Coarse Grained Reconfigurable Architecture (CGRA) structure. Accordingly, the neural core (101) may have each processing element (111_3a) and a specific processing element (111_3b) of the PE array (111_3) connected to at least one of the CGRA L0 memory (111_2), instruction memory (111_1), and LSU (111_4). That is, the processing element (111_3a) and the specific processing element (111_3b) do not have to be connected to all of the CGRA L0 memory (111_2), instruction memory (111_1), and LSU (111_4), but may be connected to some of them.
[0206] Additionally, the processing element (111_3a) and the specific processing element (111_3b) may be different types of processing devices. Accordingly, among the CGRA L0 memory (111_2), instruction memory (111_1), and LSU (111_4), the device connected to the processing element (111_3a) and the device connected to the specific processing element (111_3b) may be different.
[0207] The neural core (101) of the present invention having a CGRA structure is capable of high-level parallel computation and can have low power consumption because direct data exchange between a processing element (111_3a) and a specific processing element (111_3b) is possible. In addition, optimization according to various computational tasks may be possible by including two or more types of processing elements (111_3a).
[0208] For example, if the processing element (111_3a) is a processing element that performs a two-dimensional operation, a specific processing element (111_3b) may be a processing element that performs a one-dimensional operation. However, the present embodiment is not limited thereto.
[0209] FIG. 20 is a block diagram illustrating memory reconfiguration of a neural processing system according to some embodiments of the present invention.
[0210] Referring to FIG. 20, the neural core SoC (10) may include first to eighth processing units (160a to 160h) and on-chip memory (OCM). FIG. 24 illustrates eight processing units as an example, but this is merely an example and the number of processing units can vary.
[0211] The on-chip memory (OCM) may include first to eighth L0 memories (120a to 120h) and a shared memory (2000).
[0212] The first to eighth L0 memories (120a to 120h) can each be used as dedicated memory for the first to eighth processing units (160a to 160h). That is, the first to eighth processing units (160a to 160h) and the first to eighth L0 memories (120a to 120h) can correspond to each other in a 1:1 ratio.
[0213] The shared memory (2000) may include first to eighth memory units (2100a to 2100h). The first to eighth memory units (2100a to 2100h) may correspond to the first to eighth processing units (160a to 160h) and the first to eighth L0 memories (120a to 120h), respectively. That is, the number of memory units may be eight, which is the same as the number of processing units and L0 memories.
[0214] The shared memory (2000) can operate as either of two types of on-chip memory formats. That is, the shared memory (2000) can operate as either an L0 memory format or a global memory format. That is, the shared memory (2000) can implement two types of logical memory with a single piece of hardware.
[0215] When the shared memory (2000) is implemented in the form of an L0 memory, the shared memory (2000) can operate as a private memory for each of the first to eighth processing units (160a to 160h), such as the first to eighth L0 memories (120a to 120h). The L0 memory can operate at a relatively high clock speed compared to the global memory, and the shared memory (2000) can also use a relatively faster clock speed when operating in the form of an L0 memory.
[0216] When the shared memory (2000) is implemented in the form of a global memory, the shared memory (2000) can operate as a common memory used by the first processing unit (100a) and the second processing unit (100b) together. At this time, the shared memory (2000) can be shared not only by the first to eighth processing units (160a to 160h) but also by the first to eighth L0 memories (120a to 120h).
[0217] Global memory can generally use a lower clock speed compared to L0 memory, but is not limited thereto. When the shared memory (2000) operates in the form of a global memory, the first to eighth processing units (160a to 160h) can share the shared memory (2000). At this time, the shared memory (2000) is connected to the volatile memory (32) of FIG. 2 through a global interconnection (5000) and may operate as a buffer for the volatile memory (32).
[0218] The shared memory (2000) may operate in an L0 memory format for at least a portion and in a global memory format for the remainder. That is, the entire shared memory (2000) may operate in an L0 memory format, or the entire shared memory (2000) may operate in a global memory format. Alternatively, a portion of the shared memory (2000) may operate in an L0 memory format, and the remainder may operate in a global memory format.
[0219] FIG. 21 is a block diagram illustrating an example of memory reconfiguration of a neural processing system according to some embodiments of the present invention.
[0220] Referring to FIGS. 20 and 21, the first, third, fifth, and seventh dedicated areas (AE1, AE3, AE5, AE7) of each of the first, third, fifth, and seventh processing units (100a, 100c, 100e, 100g) may each include only the first, third, fifth, and seventh L0 memories (120a, 120c, 120e, 120g). Additionally, the second, fourth, sixth, and eighth dedicated areas (AE2, AE4, AE6, AE8) of each of the second, fourth, sixth, and eighth processing units (100b, 100d, 100f, 100h) may each include the second, fourth, sixth, and eighth L0 memories (120b, 120d, 120f, 120h). Additionally, the second, fourth, sixth, and eighth dedicated areas (AE2, AE4, AE6, AE8) may include the second, fourth, sixth, and eighth memory units (2100b, 2100d, 2100f, 2100h). The first, third, fifth, and seventh memory units (2100a, 2100c, 2100e, 2100g) of the shared memory (2000) may be utilized as a common area (AC).
[0221] The shared area (AC) may be a memory shared by the first to eighth processing units (160a to 160h). The second private area (AE2) may include the second L0 memory (120b) and the second memory unit (2100b). The second private area (AE2) may be an area where the hardware-separated second L0 memory (120b) and the second memory unit (210b) operate in the same way to logically operate as a single L0 memory. The fourth, sixth, and eighth private areas (AE4, AE6, AE8) may also operate in the same way as the second private area (AE2).
[0222] The shared memory (2000) according to the present embodiment can use the area corresponding to each neural core by converting it into a logical L0 memory and a logical global memory at an optimized ratio. The shared memory (2000) can perform this ratio adjustment at runtime.
[0223] In other words, while each neural core may perform identical tasks, they may also perform different tasks. In this case, the capacity of L0 memory and global memory required for each neural core's task will inevitably vary each time. Consequently, if the configuration ratio of L0 memory to shared memory is fixed, as is the case with conventional on-chip memory, inefficiency may arise due to the computational tasks allocated to each neural core.
[0224] Accordingly, the shared memory (2000) of the neural processing device according to the present embodiment can set the optimal ratio of L0 memory and global memory according to the computational task during runtime, and can improve the efficiency and speed of the computation.
[0225] FIG. 22 is an enlarged block diagram of part A of FIG. 20.
[0226] Referring to FIGS. 20 and 22, the shared memory (2000) may include a first L0 memory controller (122_1a), a second L0 memory controller (122_1b), a fifth L0 memory controller (122_1e), a sixth L0 memory controller (122_1f), first to eighth memory units (2100a to 2100h), and a global controller (2200). Other L0 memory controllers not illustrated may also be included in this embodiment, but are omitted for convenience.
[0227] The first L0 memory controller (122_1a) can control the first L0 memory (120a). Additionally, the first L0 memory controller (122_1a) can control the first memory unit (2100a). Specifically, when the first memory unit (2100a) is implemented in a logical L0 memory format, control by the first L0 memory controller (122_1a) can be performed on the first memory unit (2100a).
[0228] The second L0 memory controller (122_1b) can control the second L0 memory (120b). Additionally, the second L0 memory controller (122_1b) can control the second memory unit (2100b). That is, when the second memory unit (2100b) is implemented in a logical L0 memory format, control by the first L0 memory controller (122_1a) can be performed on the second memory unit (2100b).
[0229] The fifth L0 memory controller (122_1e) can control the fifth L0 memory (120e). Additionally, the fifth L0 memory controller (122_1e) can control the fifth memory unit (2100e). That is, when the fifth memory unit (2100e) is implemented in a logical L0 memory format, control by the fifth L0 memory controller (122_1e) can be performed on the fifth memory unit (2100e).
[0230] The 6th L0 memory controller (122_1f) can control the 6th L0 memory (120f). Additionally, the 6th L0 memory controller (122_1f) can control the 6th memory unit (2100f). That is, when the 6th memory unit (2100f) is implemented in a logical L0 memory format, control by the 6th L0 memory controller (122_1f) can be performed on the 6th memory unit (2100f).
[0231] The global controller (2200) can control all of the first to eighth memory units (2100a to 2100h). Specifically, the global controller (2200) can control the first memory unit (2100a) to the eighth memory unit (2100h) when each of the first to eighth memory units (2100a to 2100h) operates logically in a global memory format (i.e., when they do not operate logically in an L0 memory format).
[0232] That is, depending on what type of memory the first to eighth memory units (2100a to 2100h) are logically implemented, they may be controlled by the first to eighth L0 memory controllers (122_1a to 122_1h) respectively or by the global controller (2200).
[0233] When an L0 memory controller including the first, second, fifth, and sixth L0 memory controllers (122_1a, 122_1b, 122_1e, 122_1f) controls the first to eighth memory units (2100a to 2100h), the first to eighth L0 memory controllers (122_1a to 141h) control the first to eighth memory units (2100a to 2100h) in the same way as the first to eighth L0 memories (120a to 120h), so they can be controlled as dedicated memory for the first to eighth processing units (160a to 160h). Accordingly, the first to eighth memory units (2100a to 2100h) can operate at a clock frequency corresponding to the clock frequency of the first to eighth processing units (160a to 160h).
[0234] The L0 memory controllers, including the first L0 memory controller (122_1a), the second L0 memory controller (122_1b), the fifth L0 memory controller (122_1e), and the sixth L0 memory controller (122_1f), may each include the LSU (110) of FIG. 7.
[0235] When the global controller (2200) controls at least one of the first to eighth memory units (2100a to 2100h), the global controller (2200) can control the first to eighth memory units (2100a to 2100h) as global memory of the first to eighth processing units (160a to 160h). Accordingly, at least one of the first to eighth memory units (2100a to 2100h) can operate at a clock frequency independent of the clock frequency of the first to eighth processing units (160a to 160h). However, the present embodiment is not limited thereto.
[0236] The global controller (2200) can connect the first to eighth memory units (2100a to 2100h) to the global interconnection (5000) of FIG. 3. The first to eighth memory units (2100a to 2100h) can exchange data with the off-chip memory (30) of FIG. 1 or exchange data with the first to eighth L0 memories (120a to 120h) respectively by the global controller (2200).
[0237] The first to eighth memory units (2100a to 2100h) may each include at least one memory bank. The first memory unit (2100a) may include at least one first memory bank (2110a). The first memory bank (2110a) may be an area in which the first memory unit (2100a) is divided into specific sizes. Each first memory bank (2110a) may be a memory element of the same size. However, the present embodiment is not limited thereto. In FIG. 15, four memory banks are shown included in one memory unit.
[0238] Similarly, the second, fifth, and sixth memory units (2100b, 2100e, 2100f) may each include at least one second, fifth, and sixth memory bank (2110b, 2110e, 2110f).
[0239] The following description is based on the first memory bank (2110a) and the fifth memory bank (2110e), which may be identical to other memory banks including the second and sixth memory banks (2110b, 2110f).
[0240] Each of the first memory banks (2110a) may logically operate in an L0 memory format or logically in a global memory format. At this time, the first memory bank (2110a) may operate independently of other memory banks within the first memory unit (2100a). However, the present embodiment is not limited thereto.
[0241] When each memory bank operates independently, the first memory unit (2100a) may include a first region that operates in the same manner as the first L0 memory (120a) and a second region that operates in a different manner from the first L0 memory (120a). In this case, the first region and the second region do not necessarily coexist, and either region may occupy the entire first memory unit (2100a).
[0242] Likewise, the second memory unit (2100b) may include a third region that operates in the same manner as the second L0 memory (120b) and a fourth region that operates in a different manner from the second L0 memory (120b). In this case, the third region and the fourth region do not necessarily coexist, and either region may occupy the entire first memory unit (2100a).
[0243] At this time, the ratio of the first region to the second region may differ from the ratio of the third region to the fourth region. However, the present embodiment is not limited thereto. Accordingly, the ratio of the first region to the second region may be the same as the ratio of the third region to the fourth region. That is, the memory configuration ratio in each memory unit can vary as much as desired.
[0244] Generally, in conventional system-on-chip (SOC) architectures, on-chip memory—excluding high-speed L0 memory—was often configured using high-density, low-power SRAM. This is because SRAM offers high efficiency in terms of chip area and power consumption relative to the required capacity. However, conventional on-chip memory inevitably suffered from significantly slower processing speeds for tasks requiring more data rapidly than the predetermined L0 memory capacity. Furthermore, inefficiencies arose because there was no way to utilize the remaining global memory even when its need was minimal.
[0245] In contrast, the shared memory (2000) according to some embodiments of the present invention may be selectively controlled by either of two controllers depending on the case. At this time, the shared memory (2000) is not controlled entirely by only one of the two controllers, but may be controlled independently at the memory unit level or the memory bank level.
[0246] Through this, the shared memory (2000) according to the present embodiment can obtain an optimal memory configuration ratio according to the computation task during runtime, thereby enabling faster and more efficient computation tasks. In the case of a processing unit specialized for artificial intelligence, the required size of L0 memory and global memory may vary depending on the specific application. Furthermore, even for the same application, if a deep learning network is used, the required size of L0 memory and global memory may vary for each layer. The shared memory (2000) according to the present embodiment can change the memory configuration ratio during runtime even with changes in the computation stage according to each layer, thereby enabling fast and efficient deep learning tasks.
[0247] FIG. 23 is a drawing for explaining the first memory bank of FIG. 22 in detail. FIG. 16 is illustrated with respect to the first memory bank (2110a), but other memory banks may have the same structure as the first memory bank (2110a).
[0248] Referring to FIG. 23, the first memory bank (2110a) may include a cell array (Ca), a bank controller (Bc), a first path unit (P1), and a second path unit (P2).
[0249] The cell array (Ca) may include a plurality of memory elements (Cells) internally. The cell array (Ca) may have a plurality of memory elements arranged in a grid structure. The cell array (Ca) may be, for example, a Static Random Access Memory (SRAM) cell array.
[0250] The bank controller (Bc) can control the cell array (Ca). The bank controller (Bc) can determine whether the cell array (Ca) will operate in L0 memory format or global memory format and control the cell array (Ca) accordingly.
[0251] Specifically, the bank controller (Bc) can determine whether to transmit and receive data in the direction of the first path unit (P1) or the direction of the second path unit (P2) during runtime. The bank controller (Bc) can determine the direction of data transmission and reception according to the path control signal (Spc).
[0252] The path control signal (Spc) can be generated by a pre-designed device driver or compiler. The path control signal (Spc) can be generated based on the characteristics of the operation. Alternatively, the path control signal (Spc) can be generated by input received from the user. That is, the user can directly apply input to the path control signal (Spc) to select the most optimal memory configuration ratio.
[0253] The bank controller (Bc) can determine the path for transmitting and receiving data stored in the cell array (Ca) through the path control signal (Spc). Depending on how the bank controller (Bc) determines the path for transmitting and receiving data, the data exchange interface may vary. That is, when the bank controller (Bc) exchanges data with the first path unit (P1), it may use the first interface, and when it exchanges data with the second path unit (P2), it may use the second interface. At this time, the first interface and the second interface may be different from each other.
[0254] In addition, the address scheme in which data is stored may also change. That is, if a specific interface is selected, read and write operations can be performed using the corresponding address scheme.
[0255] The bank controller (Bc) can operate at a specific clock frequency. For example, if the cell array (Ca) is an SRAM cell array, the bank controller (Bc) can operate at the operating clock frequency of a typical SRAM.
[0256] The first path unit (P1) can be connected to the bank controller (Bc). The first path unit (P1) can directly exchange data of the cell array (Ca) with the first processing unit (100a). Here, “directly” may mean that the data is exchanged without passing through the global interconnection (5000). That is, the first processing unit (100a) can directly exchange data with the first L0 memory (120a), and the first processing unit (100a) can exchange data through the first path unit (P1) when the shared memory (2000) is logically implemented in the L0 memory format. The first path unit (P1) may include an L0 memory controller including the first L0 memory controller (122_1a) and the second L0 memory controller (122_1b) of FIG. 14.
[0257] The first path unit (P1) can configure a multi-cycle sync-path. That is, the operating clock frequency of the first path unit (P1) can be the same as the operating clock frequency of the first processing unit (100a). The first L0 memory (120a) can rapidly exchange data at the same clock frequency as the operating clock frequency of the first processing unit (100a) in order to rapidly exchange data at the same speed as the operation of the first processing unit (100a). The first path unit (P1) can also likewise operate at the same clock frequency as the operating clock frequency of the first processing unit (100a).
[0258] At this time, the operating clock frequency of the first path unit (P1) may be a multiple of the operating clock frequency of the bank controller (Bc). In this case, a separate Clock Domain Crossing (CDC) operation for clock synchronization between the bank controller (Bc) and the first path unit (P1) is not required, and accordingly, no delay in data transmission may occur. Accordingly, faster and more efficient data exchange may be possible.
[0259] In FIG. 23, for example, the operating clock frequency of the first path unit (P1) may be 1.5 GHz. This may be twice the frequency of the bank controller (Bc) at 750 MHz. However, the present embodiment is not limited thereto, and any number of times may be possible as long as the first path unit (P1) operates at an integer multiple of the clock frequency of the bank controller (Bc).
[0260] The second path unit (P2) can be connected to the bank controller (Bc). The second path unit (P2) can exchange data from the cell array (Ca) through the global interconnection (5000) rather than directly with the first processing unit (100a). That is, the first processing unit (100a) can exchange data with the cell array (Ca) through the global interconnection (5000) and the second path unit (P2). At this time, the cell array (Ca) can exchange data not only with the first processing unit (100a) but also with other neural cores.
[0261] That is, the second path unit (P2) may be a data exchange path between the cell array (Ca) and all neural cores when the first memory bank (2110a) is logically implemented in a global memory format. The second path unit (P2) may include the global controller (2200) of FIG. 22.
[0262] The second path unit (P2) can configure an async-path. The operating clock frequency of the second path unit (P2) may be the same as the operating clock frequency of the global interconnection (5000). The second path unit (P2) may also operate at the same clock frequency as the operating clock frequency of the global interconnection (5000).
[0263] At this time, the operating clock frequency of the second path unit (P2) may not be synchronized with the operating clock frequency of the bank controller (Bc). In this case, a Clock Domain Crossing (CDC) operation may be required to synchronize the clock between the bank controller (Bc) and the second path unit (P2). If the operating clock frequency of the bank controller (Bc) and the operating clock frequency of the second path unit (P2) are not synchronized with each other, the degree of freedom in the design of the clock domain may be increased. Therefore, the difficulty of hardware design is lowered, making it easier to derive hardware operation.
[0264] The bank controller (Bc) may use different address schemes when exchanging data through the first path unit (P1) and when exchanging data through the second path unit (P2). That is, the bank controller (Bc) may use the first address scheme through the first path unit (P1) and the second address scheme through the second path unit (P2). At this time, the first address scheme and the second address scheme may be different from each other.
[0265] A bank controller (Bc) does not necessarily need to exist for each memory bank. That is, since the bank controller (Bc) serves the role of transmitting signals rather than being a part for scheduling, it is not an essential part for each memory bank having two ports. Therefore, a single bank controller (Bc) can control multiple memory banks. Multiple memory banks can operate independently even if they are controlled by a bank controller (Bc). However, the present embodiment is not limited thereto.
[0266] Of course, a bank controller (Bc) may exist for each memory bank. In this case, the bank controller (Bc) can control each memory bank individually.
[0267] Referring to FIGS. 22 and 23, the first memory unit (210a) may use a first address scheme when exchanging data through the first path unit (P1) and a second address scheme when exchanging data through the second path unit (P2). Similarly, the second memory unit (210b) may use a third address scheme when exchanging data through the first path unit (P1) and a second address scheme when exchanging data through the second path unit (P2). In this case, the first address scheme and the third address scheme may be identical. However, the present embodiment is not limited thereto.
[0268] The first address scheme and the third address scheme can be used exclusively for the first processing unit (100a) and the second processing unit (100b), respectively. The second address scheme can be applied commonly to the first processing unit (100a) and the second processing unit (100b).
[0269] In FIG. 15, for example, the operating clock frequency of the second path unit (P2) may operate at 1 GHz. This may be a frequency that is not synchronized with the operating clock frequency of the bank controller (Bc) at 750 MHz. That is, the operating clock frequency of the second path unit (P2) may be freely set without being dependent on the operating clock frequency of the bank controller (Bc) at all.
[0270] Conventional global memory inevitably experiences delays due to CDC operations by using slow SRAM (e.g., 750 MHz) and a faster global interconnect (e.g., 1 GHz). In contrast, the shared memory (2000) according to some embodiments of the present invention can avoid delays due to CDC operations by utilizing the first path unit (P1) in addition to the second path unit (P2).
[0271] In addition, since a typical global memory utilizes a single global interconnection (5000) for multiple neural cores, a decrease in overall processing speed can easily occur when data transfer volumes occur simultaneously. In contrast, the shared memory (2000) according to some embodiments of the present invention has the potential to utilize a first path unit (P1) in addition to a second path unit (P2), thereby enabling the effect of appropriately distributing the data processing volume concentrated on the global controller (2200).
[0272] FIG. 24 is a block diagram illustrating the software layer structure of a neural processing device according to some embodiments of the present invention.
[0273] Referring to FIG. 24, the software layer structure of a neural processing device according to some embodiments of the present invention may include a DL framework (10000), a compiler stack (20000), and a backend module (30000).
[0274] The DL framework (10000) may refer to a framework for a deep learning model network used by the user. For example, a neural network that has been trained can be created using a program such as TensorFlow or PyTorch.
[0275] The compiler stack (20000) may include an adaptation layer (21000), a compute library (22000), a frontend compiler (23000), a backend compiler (24000), and a runtime driver (25000).
[0276] The adaptation layer (21000) may be a layer that interacts with the DL framework (10000). The adaptation layer (21000) can quantize the user's neural network model generated by the DL framework (10000) and modify the graph. Additionally, the adaptation layer (21000) can convert the model type to a required type.
[0277] The frontend compiler (23000) can convert various neural network models and graphs received from the adaptation layer (21000) into a certain intermediate representation (IR). The converted IR may be a pre-configured representation that is easy to handle later by the backend compiler (24000).
[0278] The IR of the frontend compiler (23000) can perform optimizations that can be done in advance at the graph level. Additionally, the frontend compiler (23000) can finally generate the IR by converting it into a layout optimized for hardware.
[0279] The backend compiler (24000) optimizes the IR converted by the frontend compiler (23000) and converts it into a binary file so that it can be used by a runtime driver. The backend compiler (24000) can generate optimized code by splitting the job into a scale that fits the hardware details.
[0280] The compute library (22000) can store template operations designed in a form suitable for hardware among various operations. The compute library (22000) provides various template operations required by hardware to the backend compiler (24000) so that optimized code can be generated.
[0281] The runtime driver (25000) can perform continuous monitoring during operation to operate the neural network device according to some embodiments of the present invention. Specifically, it can be responsible for the execution of the interface of the neural network device.
[0282] The backend module (30000) may include an Application Specific Integrated Circuit (ASIC) (31000), a Field Programmable Gate Array (FPGA) (32000), and a C-model (33000). The ASIC (31000) may refer to a hardware chip determined according to a predetermined design method. The FPGA (32000) may be a programmable hardware chip. The C-model (33000) may refer to a model implemented by simulating hardware in software.
[0283] The backend module (30000) can perform various tasks and produce results using binary code generated through the compiler stack (20000).
[0284] FIG. 25 is a conceptual diagram illustrating deep learning operations performed by a neural processing device according to some embodiments of the present invention.
[0285] Referring to FIG. 25, the artificial neural network model (40000) is a statistical learning algorithm implemented based on the structure of a biological neural network in machine learning technology and cognitive science, as an example of a machine learning model, or a structure that executes the algorithm.
[0286] An artificial neural network model (40000) can represent a machine learning model capable of problem-solving by learning to reduce the error between the correct output corresponding to a specific input and the inferred output through the repeated adjustment of synaptic weights by nodes, which are artificial neurons that form a network through the combination of synapses as in biological neural networks. For example, the artificial neural network model (40000) may include arbitrary probability models, neural network models, etc., used in artificial intelligence learning methods such as machine learning and deep learning.
[0287] A neural processing device according to some embodiments of the present invention can perform operations by implementing the form of such an artificial neural network model (40000). For example, the artificial neural network model (40000) can receive an input image and output information about at least a part of an object included in the input image.
[0288] The artificial neural network model (40000) is implemented as a multilayer perceptron (MLP) composed of multiple layers of nodes and connections between them. The artificial neural network model (40000) according to the present embodiment can be implemented using one of various artificial neural network model structures including an MLP. As illustrated in FIG. 25, the artificial neural network model (40000) is composed of an input layer (41000) that receives an input signal or data (40100) from the outside, an output layer (44000) that outputs an output signal or data (40200) corresponding to the input data, and n hidden layers (42000 to 43000) located between the input layer (41000) and the output layer (44000), which receive a signal from the input layer (41000), extract a feature, and transmit it to the output layer (44000) (where n is a positive integer). Here, the output layer (44000) receives a signal from the hidden layer (42000 to 43000) and outputs it to the outside.
[0289] The learning methods of the artificial neural network model (40000) include a supervised learning method that learns to be optimized for solving the problem by inputting a teacher signal (correct answer) and an unsupervised learning method that does not require a teacher signal.
[0290] The neural processing device can directly generate training data for training the artificial neural network model (40000) through simulation. In this way, a plurality of input variables and corresponding plurality of output variables are matched to the input layer (41000) and output layer (44000) of the artificial neural network model (40000), and the synapse values between the nodes included in the input layer (41000), hidden layer (42000 to 43000), and output layer (44000) are adjusted so that the correct output corresponding to a specific input can be extracted. Through this learning process, characteristics hidden in the input variables of the artificial neural network model (40000) can be identified, and the synapse values (or weights) between the nodes of the artificial neural network model (40000) can be adjusted so that the error between the output variable calculated based on the input variables and the target output is reduced.
[0291] FIG. 26 is a conceptual diagram illustrating the learning and inference operations of a neural network of a neural processing device according to some embodiments of the present invention.
[0292] Referring to Fig. 26, during the training phase, multiple training data (TD) can be forwarded to and then backwarded from the artificial neural network model (NN). Through this process, the weights and biases of each node of the artificial neural network model (NN) are adjusted, and training can be performed to derive increasingly accurate results. In this way, through the training phase, the artificial neural network model (NN) can be transformed into a trained neural network model (NN_T).
[0293] In the inference phase, new data (ND) can be input into the retrained neural network model (NN_T). The trained neural network model (NN_T) can derive result data (RD) using the new data (ND) as input and the already learned weights and biases. For this result data (RD), the training materials (TD) used during the training phase and the amount of TD utilized can be important.
[0294] Hereinafter, with reference to FIGS. 16, 19, and 27, a method of computation for a neural processing device according to some embodiments of the present invention will be described. Parts that overlap with the embodiments described above will be omitted or simplified.
[0295] FIG. 27 is a flowchart illustrating a method of computation for a neural processing device according to some embodiments of the present invention.
[0296] Referring to FIG. 27, the weight and input activation are separated (S100).
[0297] Specifically, referring to FIG. 16, the bit divider (Bd) can receive a weight and an input activation (Act_In). The bit divider (Bd) can divide the weight and the input activation (Act_In) by the number of bits of the second precision (Pr2). Accordingly, the weight and the input activation (Act_In) may each be data of the second precision (Pr2).
[0298] Again, referring to FIG. 27, it is determined whether an overflow occurs (S200).
[0299] Specifically, referring to FIG. 16, the overflow detector (Od) can detect overflow and underflow. The overflow detector (Od) can determine whether an overflow or underflow occurs as a result of the multiplication of each of the weights of the second precision (Pr2) and the input activations (Act_In) of the second precision (Pr2). Accordingly, the overflow detector (Od) can generate a detection result (DR). The detection result (DR) may be a first result if an overflow or underflow occurs. The detection result (DR) may be a second result if no overflow or underflow occurs.
[0300] Again, referring to FIG. 27, if an overflow or underflow occurs, the weight and input activation are converted from the second precision to the first precision (S300).
[0301] Specifically, referring to FIG. 16, the overflow detector (Od) can convert the weight of the second precision (Pr2) into the first precision (Pr1). Additionally, the overflow detector (Od) can convert the input activation (Act_In) of the second precision (Pr2) into the first precision (Pr1). The overflow detector (Od) can transmit the weight and the input activation (Act_In) directly to the demultiplexer (Dx) without transmitting them to the overflow detector (Od).
[0302] Again, referring to FIG. 27, if no overflow or underflow occurs, result data is generated by multiplying the weight and the input activation (S400). Also, even if an overflow or underflow occurs, result data is generated by multiplying the weight and the input activation after converting to the first precision (S400).
[0303] Specifically, referring to FIGS. 13 to 15, if the mode signal is a first mode signal for the first precision (Pr1), multiplication can be performed on the first precision (Pr1) regardless of whether an overflow or underflow occurs. If the mode signal is a second mode signal for the second precision (Pr2), multiplication can be performed on the second precision (Pr2) if no overflow or underflow occurs. Additionally, if the mode signal is a second mode signal for the second precision (Pr2), multiplication can be performed on the first precision (Pr1) if an overflow or underflow occurs.
[0304] That is, the demultiplexer (Dx) can receive weights and input activations (Act_In) from the detection unit (DU). The demultiplexer (Dx) can also receive a mode selection signal (Ms). The demultiplexer (Dx) can transmit the weights and input activations (Act_In) to either a first multiplier (Mul1) or a second multiplier (Mul2). The demultiplexer (Dx) can determine the path through which the weights and input activations (Act_In) are transmitted based on the mode selection signal (Ms). Additionally, the demultiplexer (Dx) can divide at least one weight and at least one input activation (Act_In) into multiple first multipliers (Mul1) or multiple second multipliers (Mul2) and transmit them.
[0305] The first multiplier (Mul1) can be operated with the first precision. The second multiplier (Mul2) can be operated with the second precision.
[0306] The multiplexer (Mx) can receive the result of an operation, that is, the result of a multiplication operation, from either the first multiplier (Mul1) or the second multiplier (Mul2). The multiplexer (Mx) can receive the result of a multiplication operation between the input data of the first precision and the input data of the first precision from the first multiplier (Mul1), and can receive the result of a multiplication operation between the input data of the second precision and the input data of the second precision from the second multiplier (Mul2).
[0307] When the mode selection signal (Ms) is a first mode signal, the multiplexer (Mx) can generate result data by receiving k operation results provided by k first multiplexers (Mx). The result data may include a sign bit (SB) and a product bit (PB). That is, the multiplexer (Mx) can combine k operation results to generate one result data.
[0308] When the mode selection signal (Ms) is a second mode signal, the multiplexer (Mx) can generate result data by receiving 2k operation results provided by 2k second multiplexers (Mx). The result data may include a sign bit (SB) and a product bit (PB). That is, the multiplexer (Mx) can combine 2k operation results to generate one result data.
[0309] Again, referring to FIG. 27, result data is added to generate a subtotal (S500).
[0310] Specifically, referring to FIG. 10, the saturating adder (SA) can receive result data. The saturating adder (SA) can receive, that is, sign bits (SB) and product bits (PB). The saturating adder (SA) can receive and accumulate result data multiple times. Accordingly, the saturating adder (SA) can generate partial sums (Psum). These partial sums (Psum) can be output from each processing element (163_1) and finally summed. However, the present embodiment is not limited thereto.
[0311] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment.
Claims
Claim 1 A processing element comprising: a weight register for receiving and storing a weight; an input activation register for storing an input activation; a bit divider for dividing the weight and the input activation into pre-set bit units; a flexible multiplier for receiving the weight and the input activation and generating result data by performing a multiplication operation with a first precision or a second precision different from the first precision according to a mode signal, whether there is an overflow and whether there is an underflow; and a saturating adder for receiving the result data and generating a partial sum, wherein the flexible multiplier performs the multiplication operation with the first precision prior to the mode signal according to whether an overflow or underflow occurs determined based on the output of the bit divider. Claim 2 In claim 1, the flexible multiplier comprises: a detection unit that generates a detection result by checking whether an overflow or underflow occurs according to the multiplication operation of the weight and the input activation; a mode select logic that generates a mode select signal by considering the detection result and the mode signal; a first multiplier that performs a multiplication operation with the first precision; a second multiplier that performs a multiplication operation with the second precision; and a demultiplexer that receives the mode select signal, selects either the first multiplier or the second multiplier, and transmits the weight and the input activation. Claim 3 In claim 2, the processing element, wherein the first multiplier is k and the second multiplier is 2k. Claim 4 In claim 2, the processing element, wherein the first precision is 2N bits and the second precision is N bits. Claim 5 In claim 4, the processing element, wherein the first precision is INT4 and the second precision is INT2. Claim 6 In claim 2, the flexible multiplier further comprises a processing element including a multiplexer that receives a calculation result from the first multiplier or the second multiplier and generates a sign bit indicating a sign and a product bit indicating a magnitude. Claim 7 In claim 6, the result data comprises a processing element including the sign bit and the product bit. Claim 8 In claim 2, the mode signal is either a first mode signal for the first precision and a second mode signal for the second precision, the detection result includes a first result in which the overflow or the underflow occurs and a second result in which the overflow or the underflow does not occur, and the mode selection signal is generated identically to the mode signal when the mode select logic receives the second result, and is generated as the first mode signal regardless of the mode signal when the mode select logic receives the first result, a processing element. Claim 9 In claim 8, the detection unit comprises an overflow detector that generates the detection result and outputs the weight and the input activation as the second precision if the detection result is the second result, and a converting module that receives the weight and the input activation and converts them into the first precision and outputs them when the detection result is the first result. Claim 10 A neural processing device comprising at least one neural core, wherein the neural core comprises a processing unit that performs operations and an L0 memory that stores input / output data of the processing unit, wherein the processing unit comprises a PE array comprising at least one processing element, wherein the PE array comprises a flexible multiplier that receives a weight and an input activation, divides the weight and the input activation into preset bit units, and generates result data by performing a multiplication operation with a first precision or a second precision different from the first precision according to a mode signal, whether there is an overflow and whether there is an underflow, and a saturating adder that receives the result data and generates a partial sum, wherein the flexible multiplier performs the multiplication operation with the first precision prior to the mode signal according to whether an overflow or underflow occurs determined based on the result of dividing the weight and the input activation into preset bit units. Claim 11 In claim 10, the neural processing device wherein the weight and the input activation are represented by the second precision. Claim 12 In claim 11, the flexible multiplier is a neural processing device that converts the weight and the input activation into the first precision, respectively, when an overflow or underflow occurs, when the result of the multiplication operation of the weight and the input activation is expressed as the second precision. Claim 13 In claim 12, the flexible multiplier is a neural processing device that performs a multiplication operation by selecting either the first precision or the second precision according to the mode signal when the result of the multiplication operation does not result in the overflow and the underflow. Claim 14 A neural processing device according to claim 10, further comprising an L2 shared memory shared by at least one neural core and a local interconnection for transmitting data between the L2 shared memory and the at least one neural core. Claim 15 A method of operation of a neural processing device comprising: dividing a weight and an input activation into pre-set bit units; determining whether an overflow or underflow occurs in the multiplication of the weight and the input activation based on the result of dividing the weight and the input activation into the pre-set bit units; if the overflow or the underflow occurs, converting the weight and the input activation into a first precision prior to a mode signal; and for each case where the overflow or the underflow does not occur and the mode signal selects the first precision or the second precision, converting the weight and the input activation into the first precision or maintaining them as a second precision different from the first precision, multiplying the weight and the input activation to generate result data, and accumulating the result data to generate a partial sum. Claim 16 In claim 15, the first precision uses twice the number of bits as the second precision, a method of operation for a neural processing device. Claim 17 In claim 16, the second precision is a method of computation of a neural processing device expressed as symmetric quantization or asymmetric quantization. Claim 18 In claim 17, the second precision comprises a first bit representing a sign and a second bit representing a magnitude, a method of operation for a neural processing device. Claim 19 A method of operation of a neural processing device according to claim 15, further comprising multiplying the weight and the input activation by the first precision based on whether the overflow or underflow occurs when the mode signal selects the second precision. Claim 20 A method of operation of a neural processing device according to claim 15, wherein generating the result data comprises selecting either a first multiplier corresponding to the first precision and a second multiplier corresponding to the second precision to generate the result data.
Citation Information
Patent Citations
Dot product multiplier mechanism
KR1020210058649A
Processing element, method of operation thereof, and accelerator including the same
KR102258566B1