Integrated circuit including direct memory access circuit for distributed operation and electronic device including same

The integration of a DMA controller outside the processing circuit in electronic devices with SOC addresses the increasing complexity of data communication and processing by efficiently managing data operations and reducing bandwidth requirements, thereby enhancing processing speed and reducing power consumption.

WO2025095407A1PCT designated stage expired Publication Date: 2025-05-08SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/015808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-10-17
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

As the number of functions supported by electronic devices containing System On a Chip (SOC) increases, the complexity of the integrated circuit within the SOC also rises, leading to challenges in efficient data communication and processing.

Method used

The integration of a Direct Memory Access (DMA) controller outside the processing circuit, which performs operations on data stored in volatile memory and transmits results to the processing circuit or memory, helping to distribute operations and reduce bandwidth requirements.

Benefits of technology

This solution enhances data communication efficiency between volatile memory and cache memory, allowing for faster processing and reduced power consumption by offloading operations from the processing circuit to the DMA controller.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024015808_08052025_PF_FP_ABST
    Figure KR2024015808_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment may comprise a volatile memory and a processor. The processor may comprise: a first die including a neural processing unit (NPU) circuit; a second die including a cache memory and a direct memory access (DMA) controller for data communication between the cache memory and the volatile memory; and a bus interface including through silicon vias (TSVs) provided between the first die and the second die. The DMA controller may perform an operation on a portion of data related to a neural network. The DMA controller may transmit the remaining portion to the NPU circuit of the first die through the bus interface in order to perform an operation of the NPU circuit for the remaining portion of the data different from the portion.
Need to check novelty before this filing date? Find Prior Art

Description

Integrated circuit including direct memory access circuit for distributed computation and electronic device including same

[0001] The present disclosure relates to an integrated circuit (IC) (e.g., an application processor (AP)) including a direct memory access (DMA) circuit for distributed computing and an electronic device including the same.

[0002] A system on a chip (SOC) is an electronic component that integrates circuits that perform various functions. As the number of functions supported by an electronic device containing an SOC increases, the complexity of the circuits integrated within the SOC can also increase.

[0003] The above information may be provided as background information to aid in understanding the present disclosure. None of the above is claimed to be prior art related to the present disclosure or can be used in making decisions related to prior art.

[0004] According to one embodiment, a portable communication device may include a memory, a processing circuit, and a direct memory access (DMA) controller located external to the processing circuit. The DMA controller may be configured to perform a first operation on first data among first data and second data when controlled by the processing circuit. The DMA controller may be configured to transmit a result of the first operation to the processing circuit or the memory. The processing circuit may be configured to perform a second operation on the result of the first operation or the second data. The processing circuit may be configured to transmit a result of the second operation to the memory.

[0005] In one embodiment, an electronic device may include a volatile memory and a processor. The processor may include a first die including a neural processing unit (NPU), a second die including a cache memory and a direct memory access (DMA) controller for data communication between the cache memory and the volatile memory, and a bus interface including through silicon vias (TSVs) disposed between the first die and the second die. The DMA controller may be configured to store data stored in the volatile memory in the cache memory based on identifying a command related to an operation of data associated with a neural network from the NPU circuit. The DMA controller may be configured to perform an operation on a portion of the data stored in the cache memory. The DMA controller may be configured to transmit the remaining portion to the NPU circuit of the first die through the bus interface so that the NPU circuit performs an operation on the remaining portion of the data, which is different from the portion.

[0006] In one embodiment, a method of a direct memory access (DMA) controller for data communication between a volatile memory and a cache memory within a processor is provided. The method may include storing data stored in the volatile memory in a cache memory based on identifying a command related to an operation of data associated with a neural network from a neural processing unit (NPU) circuit disposed on a first die of the processor. The DMA controller may be disposed on a second die different from the first die. The method may include performing an operation on a portion of the data stored in the cache memory. The method may include transmitting the remaining portion to the NPU circuit of the first die through a bus interface including through silicon vias (TSVs) disposed between the first die and the second die, so as to perform the operation on the remaining portion of the data, which is different from the portion.

[0007] In one embodiment, a processing chip component may include a port for data communication with a volatile memory, a DMA controller connected to the port, and a plurality of data processing circuits configured to control the DMA controller. The DMA controller may be configured to identify a command from any one of the plurality of data processing circuits to perform an operation on a first matrix and a second matrix using the data communication using the DMA controller. The DMA controller may be configured to obtain the first matrix and the second matrix from the volatile memory through the port based on the command. The DMA controller may be configured to perform an operation on the first matrix and the second matrix indicated by the command before transmitting the first matrix and the second matrix to the data processing circuit corresponding to the command in response to obtaining the first matrix from the volatile memory, the first matrix being less than or equal to a size of a buffer included in the DMA controller. The DMA controller may be configured to perform an operation on a first portion of the first matrix and a second portion of the second matrix based on obtaining the first matrix from the volatile memory, the first matrix exceeding the size of the buffer. The DMA controller may be configured to transmit, to a data processing circuit corresponding to the command, a remaining portion of the first matrix different from the first portion and a remaining portion of the second matrix different from the second portion.

[0008] In one embodiment, a method of a processing chip component is provided, the processing chip component including a port for data communication with a volatile memory, a direct memory access (DMA) controller connected to the port, and a plurality of data processing circuits configured to control the DMA controller. The method may include an operation of identifying, from one of the plurality of data processing circuits, a command for performing an operation on a first matrix and a second matrix using the data communication using the DMA controller. The method may include an operation of obtaining, based on the command, the first matrix and the second matrix from the volatile memory through the port. The method may include an operation of performing the operation on the first matrix and the second matrix indicated by the command before transmitting the first matrix and the second matrix to the data processing circuit corresponding to the command, in response to obtaining, from the volatile memory, the first matrix having a size less than or equal to a size of a buffer included in the DMA controller. The method may include performing an operation on a first portion of the first matrix and a second portion of the second matrix based on obtaining the first matrix from the volatile memory, the first matrix exceeding the size of the buffer. The method may include transmitting, to the data processing circuit corresponding to the command, a remaining portion of the first matrix different from the first portion and a remaining portion of the second matrix different from the second portion.

[0009] FIGS. 1A and 1B illustrate exemplary hardware configurations of an electronic device including a direct memory access controller, according to one embodiment.

[0010] FIG. 2 illustrates an exemplary block diagram of a DMA controller, according to one embodiment.

[0011] FIG. 3 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment.

[0012] Figure 4 illustrates an exemplary operation of a DMA controller that transmits data such as a matrix.

[0013] Figures 5a, 5b and 5c illustrate exemplary operations of a DMA controller performing operations on matrices.

[0014] Figure 6 illustrates an exemplary operation of a DMA controller performing element-wise operations.

[0015] Figure 7 illustrates an exemplary operation of an electronic device that controls a DMA controller based on the execution of a software application.

[0016] FIG. 8 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment.

[0017] FIG. 9 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment.

[0018] FIG. 10 is a block diagram of an electronic device within a network environment according to various embodiments.

[0019] Hereinafter, various embodiments of this document are described with reference to the attached drawings.

[0020] FIGS. 1A and 1B illustrate exemplary hardware configurations of an electronic device (101) including a direct memory access controller (132), according to one embodiment. Referring to FIG. 1A, the electronic device (101) may be one of various forms of electronic devices, such as a laptop personal computer (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)), a tablet PC (192), a head-mounted display (HMD) device (193), a watch (194), a cellular phone (not shown), and other similar computing devices (not shown). The electronic device (101) may also be referred to as a mobile device, a user equipment (UE) (or user terminal), a multi-function device, a portable communication device, a portable device, or a server. The form factor of the electronic device (101) is not limited to the exemplary form factors illustrated in FIG. 1A. For example, the electronic device (101) may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). For example, the electronic device (101) may have a form factor that is wearable by a user, such as an earbud (or wireless earphone) and / or a ring, or may have a form factor that is implantable on a body part of a user.

[0021] The components, their relationships, and their functions illustrated in FIG. 1A are merely exemplary and do not limit the implementations described or claimed in this document. Referring to FIG. 1A, an electronic device (101) according to one embodiment may include components such as a processor (e.g., an application processor (AP), a communication processor (CP), or any combination thereof) (110), a memory (e.g., a volatile memory (121) and / or a nonvolatile memory (122)), a display (123), a communication circuit (124), an image sensor (125), and / or a sensor (126). For example, the processor may be implemented as a three-dimensional integrated circuit (3D-IC) including a plurality of circuit layers.

[0022] Referring to FIG. 1A, an embodiment is illustrated in which a processor (110), a volatile memory (121), a nonvolatile memory (122), and / or a communication circuit (124) are disposed over or on a printed circuit board (PCB) (105), but the embodiment is not limited thereto. On the PCB (105), the processor (110) may be electrically and / or operably coupled with the volatile memory (121), the nonvolatile memory (122), and / or the communication circuit (124). The PCB (105) on which the processor (110) is disposed may be referred to as a main board and / or a mother board. Hereinafter, operably coupled components may mean that a direct connection or an indirect connection is established between the components, either wired or wireless, such that a second component is controlled by a first component among the components.

[0023] The components included in the electronic device (101) are not limited to the embodiment of FIG. 1A, and the electronic device (101) may include other components (e.g., a power management integrated circuit (PMIC), an audio processing circuit, a speaker, a microphone, an antenna, a rechargeable battery, or an input / output interface). The location of each of the components electrically and / or operatively connected to the processor (110) is not limited to a location on the PCB (105), and may have a location that is at least partially exposed to the outside through the housing of the electronic device (101). For example, some components may be omitted from the electronic device (101). For example, some components may be integrated into one component.

[0024] Referring to FIG. 1A, a volatile memory (121) and / or a non-volatile memory (122) may be configured to store data and / or instructions based on an address map managed by the processor (110). A display (123) may display images and / or videos provided from the processor (110) on one surface (e.g., flat and / or curved) of a housing of the electronic device (101). A communication circuit (124) may be configured to support wired communication and / or wireless communication between the electronic device (101) including the processor (110) and an external electronic device. An image sensor (125) may be configured to provide an electrical signal representing external light to the processor (110) and / or a memory (e.g., the volatile memory (121) and / or the non-volatile memory (122)). The sensor (126) may be configured to provide an electrical signal based at least on an external environment to the processor (110) and / or memory (e.g., volatile memory (121) and / or non-volatile memory (122)). Each of the volatile memory (121), non-volatile memory (122), display (123), communication circuit (124), image sensor (125) and sensor (126) of FIG. 1A may correspond to the volatile memory (1032), non-volatile memory (1034), display module (1060), communication module (1090), camera module (1080) and sensor module (1076) illustrated in FIG. 10.

[0025] According to one embodiment, an electronic device (101) may include a processor (110) for processing data. The processor (110) may include circuit elements including passive components, transistors, diodes, or a combination thereof. Within the processor (110), logic circuits for processing data may be interconnected. The processor (110) may be referred to as an integrated circuit (IC) and / or a system on chip (SoC). Circuits included in the processor (110) may be classified into units, modules, and / or electronic components depending on their functions. For example, the processor (110) may include electronic components such as a central processing unit (CPU) circuit (136), a graphics processing unit (GPU) circuit, a neural processing unit (NPU) circuit (154), an image signal processor (ISP), a serial communication interface (SCI), a signal interface (I / F), a high speed synchronous serial interface (HSI), a memory management (M / M), a low power dynamic random access memory (LP-DDR) physical interface (PHY) (320), a display controller, a memory controller, a storage controller, an application processor (AP), a communication processor (CP), and / or a sensor interface. Hereinafter, a unit, a module, and / or an electronic component may mean a set of circuits included in the processor (110) and at least a portion of the processor (110) designed to execute a specific function.The processor (110) of FIG. 1A may correspond to the processor (1020) of FIG. 10.

[0026] According to one embodiment, the processor (110) may have a three-dimensional integrated circuit (3D-IC) structure in which a plurality of dies are stacked. In terms of including one or more circuit elements, a die may be referred to as a circuit layer, a semiconductor layer, and / or an integrated circuit layer (IC layer). The yield (e.g., the defect rate of the die) of a die may decrease (e.g., random fail) as the area of ​​the die decreases. A processor (110) designed to include a plurality of dies may be produced with an increased yield than a case designed to include a single die because the processor includes dies having a smaller area. Hereinafter, a “layer” may be used as a term to distinguish dies and / or substrates stacked along a reference direction (the direction of the z-axis in one embodiment of FIG. 1A) within the processor (110). For example, “layer” may be replaced with terms such as “substrate,” “layer,” “semiconductor die,” and / or “die.” The processor (110) may include a plurality of dies based on a silicon wafer. Embodiments are not limited thereto, and dies based on silicon carbide (SiC) wafers, gallium nitride (GaN) wafers, and / or gallium arsenide (GaAs) wafers may be included within the processor (110).

[0027] According to one embodiment, a processor (110) may include a first die (130) and a second die (150). The first die (130) may be disposed (or positioned) above the second die (150) within the processor (110). Referring to FIGS. 1A and / or 1B , an exemplary structure of the second die (150) included in the processor (110) and the first die (130) disposed over the second die (150) along the direction of the z-axis (e.g., vertical direction) is illustrated. The first die (130) and the second die (150) may be arranged parallel to the xy plane within the processor (110). Based on the first die (130) and / or the second die (150) arranged along the z-axis direction, the circuit elements may have distinct three-dimensional locations within the processor (110). In terms of including circuit elements having three-dimensional locations, the processor (110) may be referred to as a three-dimensional integrated circuit (3D-IC), 3D packaging, 3D stacked ICs (SIC), processing circuit, processing chip component, or monolithic 3D IC.

[0028] According to one embodiment, the processor (110) may include at least one interconnect layer (or connection layer) (140) disposed between the first die (130) and the second die (150). The interconnect layer may be referred to as an interposer layer (or interposer). The interconnect layer may be formed based on a material such as silicon, glass, and / or an organic compound.

[0029] In one embodiment, the processor (110) may further include other components not shown in FIG. 1A, in addition to the first die (130), the second die (150), and the interconnect layer (140) exemplarily illustrated in FIG. 1A. For example, the processor (110) may include a substrate, a redistribution layer (RDL), and / or at least one solder ball. The substrate and / or the RDL may be disposed below the second die (150), for example, along the -z axis direction. The processor (110) may include a packaging structure enclosing the components described above.

[0030] In one embodiment, the interconnect layer (140) formed between the first die (130) and the second die (150) may include a component for supporting electrical connection between the first die (130) and the second die (150). The component may include a through silicon via (TSV). The TSV may be referred to as a through chip via and / or a through hole. The component is not limited to a TSV and may include other components (e.g., bonding wires) for electrical connection in the direction of the z-axis. The interconnect layer (140) may be formed in various shapes. For example, the interconnect layer (140) may be formed to have substantially the same width, height, thickness, area, shape, and / or size as the first die (130) and / or the second die (150), but is not limited thereto, and may be formed to have various widths, heights, thicknesses, areas, shapes, and / or sizes. According to one embodiment, the interconnect layer (140) may include at least one connecting member (e.g., a wire, a solder bump, a metal bump, or other conductive adhesive) based on a connection technology different from TSV. The first die (130) and the second die (150) may be electrically connected through the connecting member included in the interconnect layer (140).

[0031] Referring to FIG. 1A, the processor (110) may include a bus interface including a plurality of TSVs (142) formed in an interconnect layer (140). The bus interface, including a plurality of TSVs (142) disposed between a first die (130) and a second die (150), may be configured to support communication (e.g., communication of data signals and / or control signals) between circuits included in the first die (130) and circuits included in the second die (150). In one embodiment, the electronic device (101) and / or the processor (110) may include a DMA controller (132) designed to reduce the bandwidth of the bus interface including the plurality of TSVs (142). To reduce the bandwidth of the bus interface and / or the occupancy of the DMA controller (132), the DMA controller (132) may include a computation circuit.

[0032] In one embodiment, the DMA controller (132) may be disposed within the processor (110) to prevent data processing circuitry, such as the CPU circuit (136), the GPU circuit, and / or the NPU circuit (154), from directly transmitting data (e.g., large images, videos, files, and / or matrices) between the data processing circuitry and memory (e.g., cache memory (134) and / or volatile memory (121)). While the data processing circuitry is directly transmitting the data, the data processing circuitry may not perform other operations related to software applications and / or user input. The DMA controller (132) may transmit the data according to commands of the data processing circuitry, thereby allowing the data processing circuitry to execute the software application more quickly or to react more quickly to the user input.

[0033] In one embodiment, a DMA controller (132) configured to support caching, loading, and / or outputting of data generated by different data processing circuits may be included within a processor (110) including various data processing circuits, such as a CPU circuit (136), a GPU circuit, and / or an NPU circuit (154). In terms of being controllable by a plurality of data processing circuits (e.g., the CPU circuit (136), the GPU circuit, and / or the NPU circuit (154)), the DMA controller (132) may be referred to as a central DMA controller. Referring to FIG. 1A, the DMA controller (132) may be included in a first die (130) of the processor (110). The DMA controller (132) may perform operations related to data communication between a cache memory (134) and a volatile memory (121). Referring to FIG. 1A, one embodiment is shown in which the DMA controller (132) and cache memory (134) are all included in the first die (130), but the embodiment is not limited thereto.

[0034] In one embodiment, the cache memory (134) may be referred to as a last-level cache (LLC). The cache memory (134) included in the processor (110) may store data used by a data processing circuit (e.g., a CPU circuit (136), a GPU circuit, and / or an NPU circuit (154)) included in the processor (110). Hereinafter, caching may refer to an operation in which data is moved from a non-volatile memory (122) and / or a volatile memory (121) to the cache memory (134). The DMA controller (132) may be configured to control data communication occurring between the cache memory (134), the volatile memory (122), and / or the non-volatile memory (122), such as caching.

[0035] Referring to FIG. 1A, the processor (110) may include a CPU circuit (136) as a data processing circuit. The CPU circuit (136) may be configured to perform operations indicated by binary codes referred to as instructions. The instructions may be stored in a cache memory (134), a volatile memory (122), and / or a non-volatile memory (122). The CPU circuit (136) may sequentially perform operations indicated by instructions stored in the cache memory (134) to execute functions supported by a software application including the instructions.

[0036] Referring to FIG. 1A, the processor (110) may include an NPU circuit (154) as a data processing circuit. The NPU circuit (154) may be configured to perform operations related to a neural network. A neural network is a cognitive model implemented in software or hardware that mimics the computational capabilities of a biological system by using a plurality of artificial neurons (or perceptrons and / or nodes). Using a neural network, an electronic device (101) including the processor (110) may simulate human cognitive functions and / or learning processes. Parameters related to a neural network may represent weights assigned to a plurality of nodes and / or connections between the plurality of nodes. Instructions representing the weights and / or one or more operations for simulating a neural network based on the weights may be stored in the electronic device (101) in the form of a software application and / or a system process and / or a library for executing the software application.

[0037] An embodiment of a processor (110) and / or an electronic device (101) including a CPU circuit (136) and an NPU circuit (154) is illustrated, but the embodiment is not limited thereto. For example, the processor (110) may include a GPU circuit configured to perform operations related to graphics rendering. For example, the processor (110) may include an image signal processor (ISP) circuit configured to perform operations related to images and / or videos acquired from a camera. The GPU circuit and the ISP circuit may also, as data processing circuits, generate or transmit signals for controlling the DMA controller (132).

[0038] Referring to FIG. 1A, a DMA controller (132), a cache memory (134), and a CPU circuit (136) may be disposed on a first die (130) of a processor (110). A network on a chip (NOC) circuit (152) and an NPU circuit (154) may be disposed on a second die (150) of the processor (110). A data processing circuit disposed on a different die (e.g., the second die (150)) from the DMA controller (132), such as the NPU circuit (154), may transmit a control signal to the DMA controller (132) through a bus interface including a plurality of TSVs (142).

[0039] The DMA controller (132), which has received the control signal for transmitting data to the NPU circuit (154), can transmit the data stored (or cached) in the cache memory (134) to the NPU circuit (154). If the data corresponding to the control signal is not in the cache memory (134), a cache miss may occur. In response to the cache miss, the DMA controller (132) can copy, move, or store the data stored in a memory different from the cache memory (134) (e.g., the volatile memory (121) and / or the non-volatile memory (122)) to the cache memory (134). While the data corresponding to the control signal is moved from the DMA controller (132) to the NPU circuit (154) via the bus interface, at least one of the plurality of TSVs (142) can be used for transmitting the data.

[0040] Transmitting a signal using a plurality of TSVs (142) may require relatively more power than transmitting a signal within a die, such as the first die (130) and / or the second die (150). In one embodiment, the DMA controller (132) may be configured to at least partially transmit an operation when a data processing circuit disposed on a different die than the DMA controller (132) (e.g., the NPU circuit (154) of the second die (150)) controls the DMA controller (132) to perform the operation. Because the DMA controller (132) at least partially transmits the operation, the size of data transmitted using the plurality of TSVs (142) and to be input to the data processing circuit for the operation may be reduced. Because the size of the data transmitted using the plurality of TSVs (142) is reduced, the power consumed by the processor (110) may be reduced.

[0041] Referring to FIG. 1B, an exemplary chipset (160) is illustrated, which includes a first die (130), an interconnect layer (140), a second die (150), and a volatile memory (121). Within the chipset (160), the first die (130), the interconnect layer (140), and the second die (150) of FIG. 1A may be packaged. The chipset (160) may be referred to as a system on a chip (SOC) and / or an application processor (AP). Within the chipset (160), the second die (150) and the volatile memory (121) may be disposed on an RDL. Through the RDL, data stored in the volatile memory (121) may be transferred to the second die (150). Data moved to the second die (150) can be transmitted to the cache memory (134) through a plurality of TSVs (142) included in the bus interface and another TSV in a different interconnect layer (140). Similarly, data stored in the cache memory (134) can be moved to the volatile memory (121) through the RDL.

[0042] Referring to FIG. 1B, in order to transmit data to the NPU circuit (154), the DMA controller (132) may cache the data. For example, for the caching, data stored in the volatile memory (121) may be transmitted to the cache memory (134) through the RDL of the chipset (160). The DMA controller (132) may transmit the data stored in the cache memory (134) to the NOC circuit (152) on the second die (150) through a bus interface including a plurality of TSVs (142). The NOC circuit (152) may transmit the data received through the bus interface to the NPU circuit (154). The NOC circuit (152) may be configured to control transmission and / or reception of signals between circuits and / or electronic components disposed on the second die (150) on which the NOC circuit (152) is disposed. The NPU circuit (154) can store the result of performing an operation based on the above data in the DMA controller (132) and / or cache memory (134) through the bus interface.

[0043] Referring to FIG. 1B, in one embodiment where the DMA controller (132) at least partially performs an operation of the NPU circuit (154), a first size of data moved to the cache memory (134) through the RDL for the operation and a second size of data moved from the DMA controller (132) to the NPU circuit (154) and / or the NOC circuit (152) through the bus interface including the plurality of TSVs (142) may be different from each other. For example, the second size may be smaller than the first size. Since the second size of data transmitted from the DMA controller (132) to the NPU circuit (154) through the plurality of TSVs (142) is smaller than the first size, power consumed for transmitting data using the plurality of TSVs (142) may be reduced compared to a case where all data of the first size is transmitted.

[0044] In one embodiment, the first die (130) having the DMA controller (132) disposed thereon and the second die (150) having data processing circuitry, such as the NPU circuit (154), disposed thereon may be produced based on different processes. For example, the first die (130) may be produced based on a first process having a minimum line width (or line pitch) of 3 nm. For example, the second die (150) may be produced based on a second process having a minimum line width of greater than 3 nm (e.g., 4 nm and / or 5 nm). In the above example, the minimum size of the circuit elements included in the first die (130) may be larger than the minimum size of the circuit elements included in the second die (150). In one embodiment where the first die (130) including the DMA controller (132) is produced based on a process having a smaller minimum feature width than the second die (150) including the NPU circuit (154), due to the characteristics of the process, the DMA controller (132) of the first die (130) may operate with less power consumption than the NPU circuit (154). In one embodiment described above, when the DMA controller (132) performs at least part of an operation performed by the NPU circuit (154), the power consumed to perform the operation may be reduced compared to when the NPU circuit (154) performs the operation entirely. Due to the characteristics of the process, an operation performed using the first die (130) and the DMA controller (132) produced in a process having a relatively smaller minimum feature width may be performed faster than an operation performed using the second die (150) and the NPU circuit (154) produced in a process having a relatively larger minimum feature width.

[0045] Below, with reference to Fig. 2, an exemplary structure of circuits connected to each other via a bus interface is described.

[0046] FIG. 2 illustrates an exemplary block diagram of a DMA controller (132), according to one embodiment. The electronic device (101) of FIG. 2 may be an example of the electronic device (101) of FIG. 1A and / or FIG. 1B.

[0047] Referring to FIG. 2, a volatile memory (121), a cache memory (134), a DMA controller (132), and a data processing circuit (220) are illustrated, which are electrically connected via a bus interface (210). As described above with reference to FIG. 1A and / or FIG. 1B, the circuits connected via the bus interface (210) illustrated in FIG. 2 may be arranged on different dies and / or ICs. The embodiment is not limited thereto, and the circuits connected via the bus interface (210) illustrated in FIG. 2 may be integrated into a single die and / or IC.

[0048] According to one embodiment, the data processing circuit (220) may include at least one of a CPU circuit (e.g., the CPU circuit (136) of FIGS. 1A and / or 1B), a GPU circuit, and / or an NPU circuit (e.g., the NPU circuit (154) of FIGS. 1A and / or 1B). The data processing circuit (220) may include a DMA controller for directly performing movement of data related to the data processing circuit (220), independently of the DMA controller (132). The DMA controller included in the data processing circuit (220) may be referred to as a DMA controller dedicated to the data processing circuit. The data processing circuit (e.g., CPU-dedicated DMA) included in the CPU circuit may be controlled by the CPU circuit and may not be controlled by other data processing circuits. Meanwhile, the DMA controller (132), as a central DMA controller, may be controlled by a plurality of data processing circuits.

[0049] In one embodiment, the data processing circuit (220) may transmit a control signal indicating a data operation to the DMA controller (132). The control signal may include an address where data to be used for performing the operation is stored. If the address corresponds to a physical address of the volatile memory (121) mapped to the address map, the DMA controller (132) may move or copy (caching) the data stored at the address from the volatile memory (121) to the cache memory (134). The DMA controller (132) may load or fetch the data moved to the cache memory (134) through a signal path (231). The data stored in the cache memory (134) may be moved from the cache memory (134) to the DMA controller (132) through a signal path (231) connecting the cache memory (134) and the DMA controller (132) within the bus interface (210). The signal path (231) can physically connect the cache memory (134) and the DMA controller (132) within the bus interface (210). While data is being transmitted, data stored in the cache memory (134) can be moved from the cache memory (134) to the DMA controller (132) using at least a portion of the signal path (231) according to the address of the data.

[0050] Referring to FIG. 2, the DMA controller (132) may include at least one of an input / output (IO) buffer (240), an operation buffer (250), an operation switch (260), or an operation circuit (270). Data moved by the DMA controller (132) may be stored, at least temporarily, in the input / output buffer (240). As a non-limiting example, the input / output buffer (240) may be divided into an input buffer for storing data transmitted to the DMA controller (132) and an output buffer for storing data to be transmitted from the DMA controller (132). The input / output buffer (240) may be configured to support data communication based on the DMA controller (132).

[0051] In one embodiment, while the DMA controller (132) is at least partially performing an operation of the data processing circuit (220), the operation buffer (250), the operation switch (260), and the operation circuit (270) may be activated or utilized. The operation buffer (250) may be configured to at least temporarily store data to be operated on by the DMA controller (132). In one embodiment, the input / output buffer (240) may be referred to as a first buffer of the DMA controller (132), and the operation buffer (250) may be referred to as a second buffer of the DMA controller (132).

[0052] In one embodiment, the operation switch (260) may be configured to input at least a portion of data stored in the operation buffer (250) to the operation circuit (270). The operation circuit (270) may be connected to the operation buffer (250) via the operation switch (260). The embodiment is not limited thereto, and the operation circuit (270) may be directly connected to the operation buffer (250). The operation circuit (270) may be connected to the input / output buffer (240). The operation circuit (270) may be configured to perform an operation on data input via the operation switch (260) and / or data stored in the input / output buffer (240). The result of the operation performed by the operation circuit (270) may be stored in the input / output buffer (240).

[0053] In one embodiment of FIG. 2, data moved from the volatile memory (121) and / or the cache memory (134) to the DMA controller (132) via the signal path (231) may be stored in the input / output buffer (240). When the control signal indicates an operation of the data based on the DMA controller (132), the DMA controller (132) may perform the operation to input the data stored in the input / output buffer (240) to the operation circuit (270). When the control signal indicates an operation of the data stored in the operation buffer (240) and other data, the DMA controller (132) may move the data stored in the input / output buffer (240) to the operation buffer (250) and then store the other data in the input / output buffer (240). The DMA controller (132) is an operation circuit (270), and can input the other data stored in the input / output buffer (240) and the data moved to the operation buffer (250), and perform the operation indicated by the control signal.

[0054] In one embodiment of FIG. 2, when the control signal indicates a load of data by the data processing circuit (220), the DMA controller (132) can move the data stored in the input / output buffer (240) to the data processing circuit (220). The data can be moved from the DMA controller (132) to the data processing circuit (220) through a signal path (232) connecting the DMA controller (132) and the data processing circuit (220) within the bus interface (210). The data can be moved from the DMA controller (132) to the data processing circuit (220) through the signal path (232) according to the address of the data. When the data processing circuit (220) and the DMA controller (132) are arranged on different dies, the signal path (232) can include one or more TSVs (e.g., the plurality of TSVs (142) of FIG. 1A and / or FIG. 1B) electrically connecting the dies. When data is transmitted to a data processing circuit different from the CPU circuit, such as an NPU circuit, in response to completion of the transmission of the data, the DMA controller (132) may transmit a signal (e.g., a software interrupt (SWI) signal) indicating completion of the transmission of the data to the CPU circuit. The operation of the DMA controller (132) in relation to the control signal of the data processing circuit (220) is described with reference to FIG. 3.

[0055] In one embodiment, the DMA controller (132) may include a plurality of registers for storing information included in the control signal, which may receive a control signal from the data processing circuit (220). The control signal may include an instruction and / or command for assigning a value (or data) to at least one of the plurality of registers. In one embodiment, the DMA controller (132) may include registers having the exemplary names of Table 1.

[0056] Description of data stored in registers with corresponding names SOURCE_ADDR The starting address of the address map for data to be moved to the DMA controller (132) DEST_ADDR The address within the address map indicating the data processing circuit (220) (e.g., GPU circuit, NPU circuit, and / or CPU circuit) to which data stored in the DMA controller (132) will be transmitted DATA_SIZE The size of data to be moved to the DMA controller (132) (e.g., size in units of bytes) MODE The value indicating the mode and / or state of the DMA controller (132) WIDTH The width of the matrix when storing the matrix in the operation buffer (250) HEIGHT The height of the matrix when storing the matrix in the operation buffer (250) DEST_OUT_ADDR The data processing circuit (220) and / or memory (e.g., cache) to which data output from the operation circuit (270) of the DMA controller (132) (e.g., the operation result of the DMA controller (132)) will be stored and / or transmitted An address within an address map representing memory (134) and / or volatile memory (121)

[0057] Referring to Table 1, the physical addresses of the memory cells included in each circuit of the electronic device (101) may be mapped to an address map. For example, within the address map, a specific range of addresses may be mapped to the physical addresses of the cells stored in the volatile memory (121), and an address of another range different from the specific range may be mapped to a register of the data processing circuit (220). In the above example, using the address stored in the register having the name of DEST_OUT_ADDR of Table 1, the DMA controller (132) may transmit data output from the operation circuit (270) to any one of the data processing circuit (220), the cache memory (134), or the volatile memory (121).

[0058] Referring to Table 1, the data processing circuit (220) can store a value in a register named MODE, thereby causing the DMA controller (132) to operate in a mode corresponding to the stored value. The register named MODE may be referred to as a mode register. In the mode register, the mode of the DMA controller (132) corresponding to an operation indicated by a control signal (or command) of the data processing circuit (220) may be stored.

[0059] Referring to Table 1, a register named SOURCE_ADDR may be referred to as a source address register for storing the address of data to be loaded into the DMA controller (132). A register named DEST_ADDR may be referred to as a destination address register for storing an address where data loaded into the DMA controller (132) is to be stored by the DMA controller (132) within the first mode. A register named DATA_SIZE may be referred to as a data size register for storing the size of the data to be loaded into the DMA controller (132). A register named WIDTH may be referred to as a width register for storing the width of at least a portion of a matrix to be stored in the operation buffer (250). A register named HEIGHT may be referred to as a height register for storing the height of at least a portion of a matrix to be stored in the operation buffer (250). The above width register and / or the height register may be referred to as a size register for storing the size of at least a portion of a matrix to be stored in the operation buffer (250). For example, the size register may include the width register and the height register. The register named DEST_OUT_ADDR may be referred to as a start address register for storing the start address of a memory to which the result calculated by the operation circuit (270) is to be transmitted.

[0060] In one embodiment, the mode of the DMA controller (132) may include a first mode in which data transmission is performed without the arithmetic circuit (270). An exemplary operation of the DMA controller (132) within the first mode is described with reference to FIG. 4. The mode of the DMA controller (132) may include a second mode in which multiplication of matrices (e.g., matrix multiplication and / or vector inner product) is performed using the arithmetic circuit (270). An exemplary operation of the DMA controller (132) within the second mode is described with reference to FIGS. 5A to 5C. The mode of the DMA controller (132) may include a third mode in which element-wise operations (e.g., element-wise addition, element-wise subtraction, and / or element-wise multiplication) of matrices are performed using the arithmetic circuit (270). An exemplary operation of the DMA controller (132) within the third mode is described with reference to FIG. 6. In the mode register, any one of the designated numeric values ​​uniquely assigned to each of the first mode to the third mode may be stored. The DMA controller (132) may switch to a mode corresponding to the designated numeric value in response to the numeric value stored in the mode register. The operation of the data processing circuit (220) for generating a control signal for setting a plurality of registers of the DMA controller (132) is described with reference to FIG. 7.

[0061] Below, with reference to FIG. 3, an exemplary operation of a DMA controller (132) including the hardware configuration of FIG. 2 is described.

[0062] FIG. 3 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment. The electronic device (101) and / or the DMA controller (132) of FIGS. 1A, 1B, and 2 may perform the operations described with reference to FIG. 3.

[0063] Referring to FIG. 3, in operation (310), a DMA controller of an electronic device according to an embodiment may receive a command related to an operation of first data and second data from a data processing circuit. The data processing circuit of operation (310) may include the data processing circuit (220) of FIG. 2. The command of operation (310) may include a value to be stored in at least one of the registers of the DMA controller described with reference to Table 1. For example, information related to the command of operation (310) may be stored in the registers of the DMA controller. Based on the registers of the DMA controller adjusted based on the command of operation (310), the DMA controller may confirm or identify the operation indicated by the command. In one embodiment, the command may be provided from a data processing circuit (e.g., an NPU circuit (154) of FIGS. 1A and / or 1B) to perform an operation related to a neural network.

[0064] Referring to FIG. 3, in operation (320), the DMA controller of the electronic device according to one embodiment may determine whether a command for activating the arithmetic circuit of the DMA controller has been received. For example, using a value stored in a mode register (e.g., a register named MODE) described with reference to Table 1, the DMA controller may determine whether a command related to the arithmetic circuit of the DMA controller (e.g., the arithmetic circuit (270) of FIG. 2) has been received. If a command related to the arithmetic circuit has been received, the DMA controller may activate the arithmetic circuit. If a command for activating the arithmetic circuit has been received (320 - YES), the DMA controller may perform operation (340). If a command indicating an operation of the DMA controller independent of the arithmetic circuit has been received (320 - NO), the DMA controller may perform operation (330). For example, a DMA controller supporting the first mode, the second mode, and the third mode described with reference to FIG. 2 may perform operation (330) when a numerical value representing the first mode is identified from a mode register. Within the example, a DMA controller that has identified a numerical value representing the second mode and / or the third mode from a mode register may perform operation (340).

[0065] Referring to FIG. 3, in operation (330), a DMA controller of an electronic device according to an embodiment may transmit first data and second data to a data processing circuit. The DMA controller may transmit all of the first data and the second data to the data processing circuit that transmitted the command of operation (310) without modification of the first data and the second data. While transmitting the first data and the second data, the first data and / or the second data may be stored at least temporarily in an input / output buffer of the DMA controller (e.g., input / output buffer (240) of FIG. 2).

[0066] Referring to FIG. 3, in operation (340), according to one embodiment, a DMA controller of an electronic device may perform an operation on a first portion of first data and a second portion of second data, which are specified by a command, using a computation circuit of the DMA controller. The data processing circuit may transmit a command indicating an operation of the first data and the second data by the DMA controller to the DMA controller in order to use a computation resource of the DMA controller. The operation may include a matrix operation (e.g., matrix multiplication, element-wise addition, and / or element-wise multiplication) performed corresponding to the first data and the second data as matrices. A matrix including a plurality of elements arranged along one or more rows and / or one or more columns may be referred to as a vector (or vector matrix) if either the row or the column is 1. Operations supported by the DMA controller may include operations between different vectors (e.g., dot product, scalar product, or inner product) and / or cross product, tensor product, or outer product).

[0067] Referring to FIG. 3, within operation (350), a DMA controller of an electronic device according to one embodiment may transmit a remaining portion of the first data and a remaining portion of the second data to a data processing circuit. Although an embodiment of a DMA controller performing operation (350) after operation (340) has been described, the order in which the DMA controller performs operations (340, 350) may vary depending on the embodiment. For example, the DMA controller may perform operations (340, 350) substantially simultaneously. For example, the DMA controller may perform operation (340) after performing operation (350). The data processing circuit, which has received the remaining portions of operation (350), may perform an operation on the remaining portion of the first data and the remaining portion of the second data.

[0068] In one embodiment, a bus interface (e.g., bus interface (210) of FIG. 2) may be utilized to transmit the remaining portions of the operation (350). If the DMA controller and the data processing circuitry are disposed on different dies, one or more TSVs (e.g., multiple TSVs (142) of FIGS. 1A and / or 1B) connecting the dies may be utilized. In one embodiment where the DMA controller performs operations on the first portion and the second portion, power consumption of the bus interface may be reduced because the one or more TSVs transmit only the remaining portions.

[0069] Referring to FIG. 3, in operation (360), according to one embodiment, the DMA controller of the electronic device may output the operation results for the first part and the second part to a circuit designated by the command. The designated circuit of operation (360) may be mapped to an address stored in a register (e.g., a start address register) having the name DEST_OUT_ADDR of Table 1. For example, an address where the operation results of the first part and the second part of operation (340) are to be stored may be stored in the start address register. In one embodiment in which the DMA controller performs an operation on first data and second data that are matrices, the operation result of operation (360) may have a matrix form. The DMA controller may transmit the operation result in the matrix form to a circuit corresponding to the address set by the command of operation (310). Although one embodiment of the DMA controller that performs operation (360) after operation (350) has been described, the embodiment is not limited thereto. For example, the DMA controller can perform operation (360) independently of operation (350) after performing operation (340).

[0070] Hereinafter, exemplary operations of a DMA controller controlled by an NPU circuit (e.g., an NPU circuit (154) of FIG. 1A and / or FIG. 1B) and assisting the operation of the NPU circuit are described. The embodiment is not limited thereto, and the DMA controller may assist the operation of a CPU circuit (e.g., a CPU circuit (136) of FIG. 1A and / or FIG. 1B) and / or a GPU circuit.

[0071] In one embodiment, to simulate a neural network, an NPU circuit may perform operations on matrices of n dimensions (n ​​> 1). The operations on the matrices may include convolution, inner products, matrix multiplications, and / or element-wise operations. As the dimensions (e.g., dimension, width, height, and / or depth) of the matrix involved in the operation of the NPU circuit increase, the size of data transmitted to or output from the matrix to the NPU circuit may increase. For example, the size of data transmitted by a DMA controller may increase. In one embodiment that at least partially performs the operations of a data processing circuit, such as an NPU circuit, the DMA controller may at least partially perform the operations of the NPU circuit based on the operation of FIG. 3. A DMA controller that at least partially performs the operations of the NPU circuit may reduce power and / or time consumed for transmitting matrices to the NPU circuit, and may improve the performance of the neural network.

[0072] For example, based on the operation (340) of FIG. 3, the DMA controller may perform a first operation on the first data among the first data and the second data used in the neural network. The result of the first operation may be transmitted to an NPU circuit and / or a memory (e.g., a volatile memory (121) of FIG. 1A and / or FIG. 1B) for driving the neural network. For the first operation, the first data and the third data used in the first operation may be stored in different buffers (e.g., an input / output buffer and / or an operation buffer) of the DMA controller, respectively. The operation circuit of the DMA controller may perform the first operation on the first data and the third data. The NPU circuit may perform a second operation on the result of the first operation or the second data. The NPU circuit may transmit the result of the second operation to a memory. The second operation may include an operation related to the neural network, such as convolution.

[0073] Hereinafter, with reference to FIG. 4, exemplary operations of an electronic device and / or a DMA controller related to operation (330) of FIG. 3 are described.

[0074] Fig. 4 illustrates an exemplary operation of a DMA controller (132) that transmits data such as a matrix (410). The electronic device (101) and / or the DMA controller (132) of Figs. 1a, 1b, and 2 may perform the operation of the electronic device (101) and / or the DMA controller (132) of Fig. 4. The operation of the DMA controller (132) described with reference to Fig. 4 may be related to at least one of the operations of Fig. 3 (e.g., operation (330) of Fig. 3).

[0075] Referring to FIG. 4, a three-dimensional matrix (410) is exemplarily illustrated as an example of data utilized by a data processing circuit (e.g., the data processing circuit (220) of FIG. 2) such as an NPU circuit (154). Although the matrix (410) is illustrated having an exemplary width (w1), height (h1), and depth (d1), the dimensions of the matrix (410) are not limited to the example. The NPU circuit (154) may set the mode of the DMA controller (132) to a first mode for loading data (e.g., the matrix (410)) into the NPU circuit (154) using a control signal for storing data and / or values ​​in registers of the DMA controller (132) (e.g., registers exemplified in Table 1).

[0076] For example, by a control signal (or command) received from the NPU circuit (154), the start address of an area of ​​a circuit (e.g., the cache memory (134) and / or the volatile memory (121) of FIG. 1A and / or FIG. 1B) in which the matrix (410) is stored may be stored in a source register of the DMA controller (132). By the control signal, a designated numerical value indicating a first mode may be stored in a mode register of the DMA controller (132). By the control signal, an address indicating a destination of the matrix (410) to be transmitted by the DMA controller (132) may be stored in a destination address register of the DMA controller (132). For example, in a case where the NPU circuit (154) loads the matrix (410), the start address of an area of ​​a memory included in the NPU circuit (154) (or a DMA controller included in the NPU circuit (154)) may be stored in the destination address register. The size of the matrix (410) can be stored in the data size register of the DMA controller (132) by the above control signal. For example, when the elements of the matrix (410) have a size of 1 byte, a numerical value corresponding to w1×h1×d1 can be stored in the data size register.

[0077] In one embodiment, when the configuration of a plurality of registers by a control signal transmitted from the NPU circuit (154) is completed, the DMA controller (132) may perform an operation related to data stored in the plurality of registers. For example, the DMA controller (132) may access the memory (e.g., the cache memory (134) and / or the volatile memory (121) of FIGS. 1A and 1B) of the electronic device (101) using an address stored in a source register to obtain at least a portion of the matrix (410). The DMA controller (132) may identify the matrix (410) specified by the size stored in the data size register from the address stored in the source register. For example, the DMA controller (132) may load data in the amount of the size stored in the data size register from an address stored in the memory in the source register. The loaded data may be identified as the matrix (410). The DMA controller (132) can copy the matrix (410) at least partially into the input / output buffer (240) based on the size of the input / output buffer (240).

[0078] Referring to FIG. 4, within an exemplary state operating based on the first mode, the operation buffer (250), the operation switch (260), and / or the operation circuit (270) of the DMA controller (132) may be deactivated. The DMA controller (132) may transmit at least a portion of the matrix (410) stored in the input / output buffer (240) to an address stored in the destination address register. By a control signal of the NPU circuit (154), within an exemplary state in which an address corresponding to the NPU circuit (154) is stored in the destination address register, the DMA controller (132) may transmit at least a portion of the matrix (410) stored in the input / output buffer (240) to the NPU circuit (154).

[0079] Hereinafter, exemplary operations of the DMA controller related to operations (340, 350, 360) of FIG. 3 are described with reference to FIG. 5a, FIG. 5b, and / or FIG. 5c.

[0080] FIGS. 5A, 5B, and 5C illustrate exemplary operations of a DMA controller (132) that performs operations on matrices (e.g., a first matrix (510) and / or a second matrix (520)). The electronic device (101) and / or the DMA controller (132) of FIGS. 1A, 1B, and 2 may perform the operations of the electronic device (101) and / or the DMA controller (132) of FIGS. 5A to 5C. The operations of the DMA controller (132) described with reference to FIGS. 5A to 5C may be related to at least one of the operations of FIG. 3 (e.g., operations (340, 350, 360) of FIG. 3).

[0081] Referring to FIGS. 5A to 5C , exemplary states (501, 502, 503) of a DMA controller (132) performing a multiplication operation of matrices (e.g., a first matrix (510) and a second matrix (520)) within a second mode are illustrated. A three-dimensional first matrix (510) having an exemplary width (w1), a height (h1), and a depth (d1), and a three-dimensional second matrix (520) having an exemplary width (w2), a height (h2), and a depth (d2) are exemplarily illustrated, but the sizes of matrices that can be operated on by the DMA controller (132) are not limited to the first matrix (510) and / or the second matrix (520).

[0082] In one embodiment, the NPU circuit (154) may obtain the first matrix (510) and the second matrix (520) from a memory (e.g., the cache memory (134), the volatile memory (121), and / or the non-volatile memory (122) of FIGS. 1A and / or 1B) to perform operations on the first matrix (510) and the second matrix (520). To obtain the first matrix (510) and the second matrix (520), the NPU circuit (154) may transmit commands related to operations on the first matrix (510) and the second matrix (520) to the DMA controller (132). For example, the elements of the first matrix (510) and the second matrix (520) may be referred to as first data and / or second data used in an artificial intelligence model executed by the NPU circuit (154).

[0083] In one embodiment, the command transmitted to the DMA controller (132) may include data and / or at least one value to be stored in a plurality of registers (e.g., registers exemplified in Table 1) of the DMA controller (132). The command may include information indicating a first portion (511) of a first matrix (510) and a second portion (522) of a second matrix (520) related to an operation to be performed by the DMA controller (132). The DMA controller (132) that identifies the command may perform caching of the first matrix (510) and / or the second matrix (520). Based on the caching, the first matrix (510) and the second matrix (520) stored in the volatile memory may be stored in the cache memory.

[0084] Referring to FIG. 5A, an exemplary state (501) of a DMA controller (132) loading a first matrix (510) stored in a memory such as a cache memory is illustrated. Within the state (501), the DMA controller (132) may store a first portion (511) of the first matrix (510) in an operation buffer (250) of the DMA controller (132). The first portion (511) may be distinguished from the remaining portion (513) of the first matrix (510) (e.g., a portion including e51 to e6w as elements) by a size register of Table 1. The first portion (511) may be moved to the operation buffer (250) via the input / output buffer (240).

[0085] Referring to FIG. 5A, elements of the first matrix (510) (e.g., e11 to e3w included in three rows of the first matrix (510)) may be stored in the operation buffer (250). Within the state (501) of FIG. 5A, elements of a specific row of the first matrix (510) (e.g., e41 to e4w) stored in the input / output buffer (240) may be moved to the operation buffer (250). After the elements of a specific row of the first matrix (510) (e.g., e41 to e4w) are moved to the operation buffer (250), elements of four rows of the first matrix (510) (e.g., e11 to e4w) may be stored in the operation buffer (250). The elements stored in the operation buffer (250) may be elements included in the first part (511) of the first matrix indicated by the command provided by the NPU circuit (154).

[0086] Referring to FIG. 5A, a matrix to be stored in the operation buffer (250) among the first matrix (510) or the second matrix (520) may be selected or determined by a data processing circuit, such as the NPU circuit (154) that provides the command. For example, the NPU circuit (154) may select or determine a matrix to be stored in the operation buffer (250) among the first matrix (510) or the second matrix (520), by using the size of the operation buffer (250) and / or the input / output buffer (240).

[0087] Within the state (501) of Fig. 5a, after the first part (511) of the first matrix is ​​stored in the operation buffer (250), the DMA controller (132) can switch to the state (502) of Fig. 5b. Within the state (502), the DMA controller (132) can store at least a part of the second part (522) of the second matrix indicated by the command provided by the NPU circuit (154) in the input / output buffer (240). Within the second mode for performing matrix multiplication, in order to perform the matrix multiplication, the DMA controller (132) can obtain the first part (511) of the first matrix (510) specified along the row direction and the second part (522) of the second matrix (520) specified along the column direction. Within the exemplary state (502) of FIG. 5b, a plurality of rows (e.g., four rows) included in a first portion (511) of a first matrix (510) may be stored in an operation buffer (250), and at least one column (e.g., one column including k elements of f11 to fk1) included in a second portion (522) of a second matrix (520) may be stored in an input / output buffer (240). For example, elements (f11 to fk1) of a specific column of the second matrix (520) may be stored in the input / output buffer (240). In one embodiment, the size of the input / output buffer (240) and / or the operation buffer (250) is not limited to the exemplary sizes of FIGS. 5a to 5c (e.g., the size of the operation buffer (250) of 4 Х w).

[0088] Referring to FIG. 5B, in a state (502) where a first part (511) of a first matrix (510) is stored in a calculation buffer (250) and a second part (522) of a second matrix (520) is stored in an input / output buffer (240), the DMA controller (132) can perform an operation on the first part (511) and the second part (522). Using a calculation switch (260), the DMA controller (132) can sequentially input a plurality of rows stored in the calculation buffer (250) to an operation circuit (270). By controlling the operation circuit (270), the DMA controller (132) can perform matrix multiplication on a specific row of the first matrix (510) input to the calculation circuit (270) through the calculation switch (260) and a specific column of the second matrix (520) stored in the input / output buffer (240). For example, from the operation circuit (270) into which a specific row (a row having elements of e11, e12, ..., e1w) of the first matrix (510) and a specific column (f11, f21, ..., fk1) of the second matrix (520) stored in the input / output buffer (240) are input, the DMA controller (132) can obtain a multiplication of the specific row and the specific column (e.g., g11 = e11 Х f11 + e12 Х f21 + ... + e1w Х fk1). The multiplication result (g11) can be stored in the input / output buffer (240).

[0089] Referring to FIG. 5B, the DMA controller (132) performing matrix multiplication using the operation circuit (270) may obtain or generate at least a portion of a third matrix (530) corresponding to the matrix multiplication of the second matrix (520) by the first matrix (510). In the exemplary state (502) of FIG. 5B, the DMA controller (132) may obtain the multiplication results of each of a plurality of rows of the first matrix (510) by the columns of the second matrix (520).

[0090] In one embodiment, the DMA controller (132) may transmit the result of performing an operation on the first portion (511) of the first matrix (510) and the second portion (522) of the second matrix (520) (e.g., at least a portion of the third matrix (530)) to a circuit corresponding to a destination address set by a command of the NPU circuit (154) (e.g., an address stored in a start address register). The circuit corresponding to the destination address may include any one of a volatile memory, a cache memory, and / or an NPU circuit. The embodiment is not limited thereto, and the circuit corresponding to the destination address may be a circuit including a specific memory cell mapped by an address map among the circuits in the electronic device (101). For example, the DMA controller (132) may transmit at least a portion of the third matrix (530) to a cache memory among the NPU circuit (154) or the cache memory. When the NPU circuit (154) and the cache memory are arranged on different dies of the processor (110), the DMA controller (132) may transmit at least a portion of the third matrix (530) to the cache memory arranged on the first die (e.g., the first die (130) of FIG. 1A and / or FIG. 1B) on which the DMA controller (132) is arranged, or to the NPU circuit (154) arranged on a second die different from the first die (e.g., the second die (150) of FIG. 1A and / or FIG. 1B). For example, the result of performing an operation on the first portion (511) of the first matrix (510) and the second portion (522) of the second matrix (520) may be transmitted to the cache memory arranged on the first die. Before being transmitted to the cache memory, at least a portion of the third matrix (530) may be stored at least temporarily in the input / output buffer (240) (in the exemplary state of FIG. 5b, g11 to g11 stored in the input / output buffer).The states (501, 502) of FIG. 5a and / or FIG. 5b may be states of the DMA controller (132) while performing the operation (340) of FIG. 3.

[0091] Referring to FIG. 5c, an exemplary state (503) of the DMA controller (132) is illustrated that has acquired a remaining portion (513) different from a first portion (511) of the first matrix (510) set by a command of the NPU circuit (154). The remaining portion (513) may correspond to another portion of the first matrix (510) other than the first portion (511) set by the command of the NPU circuit (154). The remaining portion (513) may correspond to at least a portion of the matrix (510) set to be transmitted to the NPU circuit (154) without an operation of the DMA controller (132) by the command of the NPU circuit (154). In a state (503) in which the remaining portion (513) different from the first portion (511) specified by the data stored in the register (e.g., source register and / or size register) of the DMA controller (132) is stored in the input / output buffer (240), the DMA controller (132) can transmit the remaining portion (513) to the NPU circuit (154). Similarly, the DMA controller (132) can transmit the remaining portion of the second matrix (520) different from the second portion (522) of the second matrix (520) to the NPU circuit (154). The remaining portions can be transmitted to the NPU circuit (154) via a bus interface including a plurality of TSVs (e.g., a plurality of TSVs (142) of FIG. 1A and / or FIG. 1B). In one embodiment where the NPU circuit (154) is disposed on a second die different from the first die of the processor (110) where the DMA controller (132) and / or cache memory are disposed, the remaining portions may be transmitted from the first die where the DMA controller (132) is disposed to the second die where the NPU circuit (154) is disposed.

[0092] Within the exemplary state (503) of FIG. 5c, the NPU circuit (154) may perform matrix multiplication on the remaining portion (513) of the first matrix (510) and the remaining portion of the second matrix (520). By the matrix multiplication performed by the NPU circuit (154), the remaining elements different from the elements filled by the operation of the DMA controller (132) within the third matrix (530) may be determined or calculated. The NPU circuit (154) may store the result of the matrix multiplication in an area of ​​memory specified by a command provided to the DMA controller (132).

[0093] For example, a portion of the third matrix (530) operated by the DMA controller (132) and the remaining portion of the third matrix (530) operated by the NPU circuit (154) may be stored in an area of ​​memory including an address stored in a start address register of the DMA controller (132). In the exemplary state (503) of FIG. 5C, the DMA controller (132) may transmit the remaining portion (513) of the first matrix (510) (e.g., the portion including e51 to e6w as elements) to the NPU circuit (154). The NPU circuit (154) may perform matrix multiplication using the remaining portion (513) of the first matrix (510) and the remaining portion of the second matrix (520) additionally transmitted from the DMA controller (132). Based on the matrix multiplication performed by the NPU circuit (154), the elements (e.g., g51 to g5l) of the third matrix (530) can be determined.

[0094] Referring to FIGS. 5A to 5C, exemplary operations of the DMA controller (132) for performing matrix multiplication of a first matrix (510) and / or a second matrix (520) having a size larger than the size of the operation buffer (250) are described. The embodiment is not limited thereto, and the DMA controller (132) may perform matrix multiplication of a first matrix (510) and / or a second matrix (520) having a size smaller than or equal to the operation buffer (250). In one embodiment of performing matrix multiplication of a first matrix (510) and / or a second matrix (520) having a size smaller than or equal to the operation buffer (250), the DMA controller (132) may complete the matrix multiplication without transmitting the first matrix (510) or the second matrix (520) to the NPU circuit (154). For example, the DMA controller (132) may perform the matrix multiplication indicated by the command of the NPU circuit (154) using the operation circuit (270). In one embodiment where the DMA controller (132) completes the matrix multiplication using the control signal provided from the NPU circuit (154) without the NPU circuit (154) (or independently of the NPU circuit (154)), data transmission (e.g., transmission of data through a plurality of TSVs) between the dies on which the DMA controller (132) and the NPU circuit (154) are respectively disposed may be reduced.

[0095] In one embodiment, when the NPU circuit (154) performs operations on all elements of the first matrix (510) and the second matrix (520), the TSVs between the first die on which the DMA controller (132) is disposed and the second die on which the NPU circuit (154) is disposed may be at least temporarily occupied or utilized to transmit the entire first matrix (510) and the entire second matrix (520). In one embodiment where the DMA controller (132) performs operations related to a first portion (511) of the first matrix (510) and a second portion (522) of the second matrix (520), since the NPU circuit (154) performs operations only on the remaining portion (513) of the first matrix (510), the TSVs between the first die and the second die may be at least temporarily occupied or utilized to transmit the remaining portion (513) of the first matrix (510) and the second matrix (520). For example, the time that the TSVs are occupied to transmit data (e.g., the first matrix (510) and / or the second matrix (520)) to the NPU circuit (154) may be reduced.

[0096] Below, with reference to FIG. 6, exemplary operations of the DMA controller related to the operations (340, 350, 360) of FIG. 3 are described.

[0097] Fig. 6 illustrates an exemplary operation of a DMA controller (132) performing element-wise operations. The electronic device (101) and / or the DMA controller (132) of Figs. 1a, 1b, and 2 may perform the operations of the electronic device (101) and / or the DMA controller (132) of Fig. 6. The operation of the DMA controller (132) described with reference to Fig. 6 may be related to at least one of the operations of Fig. 3 (e.g., operations (340, 350, 360) of Fig. 3).

[0098] Referring to FIG. 6, exemplary states of a DMA controller (132) performing element-by-element operations of matrices (e.g., a first matrix (610) and / or a second matrix (620)) within a third mode are illustrated. A three-dimensional first matrix (610) having an exemplary width (w1), a height (h1), and a depth (d1), and a three-dimensional second matrix (620) having an exemplary width (w2), a height (h2), and a depth (d2) are exemplarily illustrated, but the sizes of matrices that can be operated on by the DMA controller (132) are not limited to the first matrix (610) and / or the second matrix (620). In order to perform element-by-element operations (e.g., element-by-element addition, element-by-element subtraction, element-by-element multiplication, and / or element-by-element division), the first matrix (610) and the second matrix (620) can have dimensions that match each other.

[0099] In one embodiment, the NPU circuit (154) may transmit a control signal for an operation to the DMA controller (132) to perform an element-wise operation of the first matrix (610) and the second matrix (620). The DMA controller (132) may identify a command related to the element-wise operation of the first matrix (610) and the second matrix (620). Based on identifying the command, the DMA controller (132) may store a first portion (611) of the first matrix (610) (e.g., a first portion including a plurality of rows) in the operation buffer (250). The DMA controller (132) that has identified the command may store at least a portion of a second portion (622) of the second matrix (620) (e.g., at least one row) in the input / output buffer (240).

[0100] In one embodiment that performs element-by-element operations, a first portion (611) of the first matrix (510) stored in the operation buffer (250) (e.g., a portion of the first matrix (510) based on a data size from a start address of the first matrix (510)) may have a size different from a specific row and / or a specific column of the first matrix (510). For example, unlike matrix multiplication, which requires an entire row of the first matrix (510) and an entire column of the second matrix (520), an element-by-element operation may be performed if only elements in corresponding positions in the first matrix (510) and the second matrix (520) are loaded into the operation buffer (250). In one embodiment, even if the entire operation buffer (250) is not used for storing data, the DMA controller (132) may input at least a portion of the data stored in the operation buffer (250) to the operation circuit (270), or may store data output from the operation circuit (270) (e.g., the result of performing an element-to-element operation) in the operation buffer (250).

[0101] Referring to FIG. 6, a first part (611) of a first matrix (610) specified by an NPU circuit (154) and a second part (622) of a second matrix (620) may be stored in an operation buffer (250) and an input / output buffer (240) of a DMA controller (132), respectively. The DMA controller (132) may obtain a first element (e.g., e11) among a plurality of rows of the first part (611) stored in the operation buffer (250) and a second element (e.g., f11) of at least one row stored in the input / output buffer (240) using an operation circuit (270). The DMA controller (132) may perform an addition operation of the first element and the second element to obtain the addition result (g11 = e11 + f11).

[0102] Within the exemplary state of FIG. 6, the DMA controller (132) may perform element-by-element addition of a specific row of the first matrix (610) (e.g., a row including e11 to e1w as elements) and a specific row of the second matrix (620) (e.g., a row including f11 to f1w as elements). The result of performing the element-by-element addition of the specific rows (e.g., g11 = e11 + f11) may be accumulated in or stored in the input / output buffer (240). In one embodiment, an arithmetic circuit (270) including only one circuit that performs element-by-element operations (e.g., an arithmetic logic unit (ALU) and / or a floating point unit (FPU)) may sequentially perform element-by-element operations on elements included in the specific rows, thereby obtaining an element-by-element operation result for the specific rows. Alternatively, the operation circuit (270) including a plurality of ALUs may perform the element-to-element operation by activating only one of the plurality of ALUs, thereby saving power. The result may be overwritten to a specific element of the input / output buffer (240) input to the operation circuit (270) to perform the element-to-element addition (e.g., f11 stored in the input / output buffer (240) while calculating g11 = e11 + f11). The embodiment is not limited thereto, and the result may be stored in an area of ​​the input / output buffer (240) that is distinct from the specific row.

[0103] In one embodiment, the DMA controller (132) may transmit the result stored in the input / output buffer (240) to a circuit corresponding to an address stored in the start address register. If the size of the result of the inter-element operation stored in the input / output buffer (240) corresponds to a size set by a command transmitted from the NPU circuit (154), the DMA controller (132) may transmit the result. For example, if an address corresponding to a memory of the electronic device (101) (e.g., the volatile memory (121) of FIG. 1A and / or FIG. 1B) is stored in the start address register, the DMA controller (132) may transmit the data (e.g., the result of the inter-element operation) stored in the input / output buffer (240) to the NPU circuit (154), to the memory corresponding to the address.

[0104] As described above, according to one embodiment, the electronic device (101) may include a DMA controller (132) for assisting the operation of a data processing circuit, such as an NPU circuit (154). The DMA controller (132) may be a central DMA controller that may receive commands from various data processing circuits included in the processor. The DMA controller (132) may be disposed on a die that has a relatively high speed and consumes less power within the processor (e.g., a die produced by a process that supports a minimum line width of 3 nm), and may assist the operation of the data processing circuit at a faster speed and with less power than a data processing circuit disposed on a die other than the die.

[0105] Below, with reference to FIG. 7, the operation of a data processing circuit, such as an NPU circuit (154), performed to transmit a control signal to a DMA controller (132) is described.

[0106] FIG. 7 illustrates an exemplary operation of an electronic device (101) controlling a DMA controller (132) based on the execution of a software application (710). The electronic device (101) and / or the DMA controller (132) of FIGS. 1A, 1B, and 2 may perform the operations described with reference to FIG. 7.

[0107] Referring to FIG. 7, the CPU circuit (136) can execute instructions included in a software application (710). The CPU circuit (136) can perform operations corresponding to each of the instructions. Data related to the operations can be moved from a memory such as a volatile memory (121) to the CPU circuit (136) by the CPU circuit (136) (or a DMA controller dedicated to the CPU circuit (136)). The embodiment is not limited thereto, and data related to the operations can be moved from the volatile memory (121) and / or cache memory (134) to the CPU circuit (136) by using the DMA controller (132).

[0108] In one embodiment, the instructions included in the software application (710) may include not only a first type of instructions readable by the CPU circuit (136), but also a second type of instructions readable by data processing circuitry different from the CPU circuit (136), such as the NPU circuit (154) and / or the GPU circuit. For example, the second type of instructions readable by the NPU circuit (154) may be stored within the NPU binary code (720) included in the software application (710). The embodiment is not limited thereto, and the NPU binary code (720) may include text in assembly language that is compilable into the second type of instructions readable by the NPU circuit (154). The CPU circuit (136) that detects the NPU binary code (720) can transmit or upload at least a portion of the NPU binary code (720) to the NPU circuit (154). The NPU circuit (154) can perform an operation indicated by the NPU binary code (720) transmitted from the CPU circuit (136).

[0109] In one embodiment, the NPU circuit (154) performing the operation indicated by the NPU binary code (720) may control the DMA controller (132) to determine whether to perform the operation. For example, when performing a data type and / or operation supported by an operation circuit included in the DMA controller (132) (e.g., operation circuit (270) of FIG. 2), the NPU circuit (154) may determine to perform the operation indicated by the NPU binary code (720) using the DMA controller (132). For example, when the temperature of a processor (e.g., processor (110) of FIG. 1A and / or chipset (160) of FIG. 1B) including the NPU circuit (154) falls within a specified temperature range associated with overheating (or overload), the NPU circuit (154) may determine to perform the operation indicated by the NPU binary code (720) using the DMA controller (132).

[0110] As described above with reference to FIGS. 5A to 5C and / or 6, the operations supported by the DMA controller (132) may include matrix multiplication (or inner product of vectors) and element-by-element operations of matrices. The NPU circuit (154), which detects an operation supported by the DMA controller (132) among the different operations for simulating a neural network, represented by the NPU binary code (720), may control the DMA controller (132) to perform the detected operation. The NPU circuit (154) may store data related to the detected operation in registers of the DMA controller (132), which are described above with reference to Table 1. The NPU circuit (154) may request the DMA controller (132) to perform an operation and / or action represented by the data stored in the registers.

[0111] For example, operations performed to simulate a neural network may include matrix multiplication of different layers, convolution, element-by-element operations, or a combination thereof (e.g., layer fusion). When performing matrix multiplication, the NPU circuit (154) may adjust the mode of the DMA controller (132) to a second mode so that the DMA controller (132) assists the matrix multiplication. When performing convolution, the NPU circuit (154) may adjust the mode of the DMA controller (132) to a first mode so that the DMA controller (132) transmits data related to the convolution to the NPU circuit (154) without operation. In one embodiment, when performing dot products and / or convolutions included in a neural network, a specific layer of the neural network related to dot products may be processed by the DMA controller (132) in the second mode, and another layer of the neural network related to convolutions may be transmitted to the NPU circuit (154) (or CPU (136)) by the DMA controller (132) in the first mode. The NPU circuit (154) may perform convolutions related to the other layers using data transmitted from the DMA controller (132).

[0112] Hereinafter, exemplary operations of the DMA controller (132) described with reference to FIGS. 1A, 1B, 2 to 7 are described with reference to FIGS. 8 and / or 9.

[0113] FIG. 8 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment. The electronic device (101) and / or the DMA controller (132) of FIGS. 1A, 1B, and 2 may perform the operations described with reference to FIG. 8. The operations of the electronic device and / or the DMA controller described with reference to FIG. 8 may be related to at least one of the operations of FIG. 3.

[0114] Referring to FIG. 8, in operation (810), a DMA controller of an electronic device according to an embodiment may store at least a portion of a first matrix corresponding to the size of the operation buffer in an operation buffer (e.g., an operation buffer (250) of FIG. 2) of the DMA controller based on a command provided from a data processing circuit (e.g., a data processing circuit (220) of FIG. 2). The DMA controller may identify a command related to an operation of the first matrix and the second matrix from any one of a plurality of data processing circuits connected to the DMA controller using data communication using the DMA controller. Based on the command, the DMA controller may obtain the first matrix and the second matrix from a memory (e.g., a volatile memory (121) of FIG. 1A, FIG. 1B, and / or FIG. 2) via a port for data communication with the memory. Among the first matrix or the second matrix cached from the memory, the DMA controller may store at least a portion of the first matrix selected by the command of operation (810) in the operation buffer. The state (501) of FIG. 5A may be related to the state of the DMA controller performing operation (810).

[0115] Referring to FIG. 8, in operation (820), the DMA controller of the electronic device according to one embodiment may perform an operation on at least a portion of a first matrix and a portion of a second matrix stored in an operation buffer using an operation circuit of the DMA controller (e.g., the operation circuit (270) of FIG. 2). The operation of operation (820) may be related to a mode of the DMA controller set by the command of operation (810). For example, the operation of operation (820) may be one operation selected by the mode of the DMA controller among the matrix multiplication described with reference to FIGS. 5A to 5C or the element-by-element operation described with reference to FIG. 6. The state (502) of FIG. 5B may be related to a state of the DMA controller performing operation (820).

[0116] Referring to FIG. 8, in operation (830), the DMA controller of the electronic device according to one embodiment may identify or determine whether the size of the first matrix is ​​less than or equal to the size of the computation buffer. If the size of the first matrix is ​​less than or equal to the size of the computation buffer, the entire first matrix may be stored in the computation buffer based on operation (810). If the size of the first matrix exceeds the size of the computation buffer, only a portion of the first matrix may be stored in the computation buffer based on operation (810). If the size of the first matrix is ​​less than or equal to the size of the computation buffer (830—Yes), the DMA controller may perform operation (840). If the size of the first matrix exceeds the size of the computation buffer (830—No), the DMA controller may perform operation (850).

[0117] Referring to FIG. 8, within operation (840), a DMA controller of an electronic device according to one embodiment may perform an operation on the entire first matrix stored in an operation buffer independently of a data processing circuit. For example, in response to obtaining the first matrix from a volatile memory and / or a cache memory, the first matrix having a size less than or equal to the size of the operation buffer, the DMA controller may perform the operation on the first matrix and the second matrix indicated by the command before transmitting the first matrix and the second matrix to the data processing circuit corresponding to the command of operation (810). For example, transmitting the first matrix and the second matrix to the data processing circuit that provided the command of operation (810) may be interrupted, omitted, or bypassed.

[0118] For example, a DMA controller that identifies a command related to a first matrix smaller than or equal to the size of an operation buffer may perform the operation on the entire first matrix and the entire second matrix using the operation circuit of the DMA controller without transmitting the operation to an NPU circuit based on a bus interface (e.g., a bus interface (210) of FIG. 2). The result of performing the operation on the entire first matrix and the entire second matrix may be transmitted to a circuit corresponding to a destination address set by the command of the operation (810). For example, the DMA controller may transmit information and / or data including the result to a circuit corresponding to an address stored in a start address register (e.g., a cache memory (134), a volatile memory (121), and / or a data processing circuit (220) of FIG. 2).

[0119] Referring to FIG. 8, in operation (850), a DMA controller of an electronic device according to an embodiment may transmit at least a portion of a first matrix stored in an operation buffer, a remaining portion different from the first portion, and a remaining portion of a second matrix, to a data processing circuit. The data processing circuit of operation (850) may correspond to the data processing circuit that provided the command of operation (810). For example, the DMA controller, which has obtained a first matrix exceeding the size of the operation buffer from a volatile memory and / or a cache memory, may perform an operation on the first portion of the first matrix and the second portion of the second matrix. The DMA controller may transmit the remaining portion of the first matrix, which is different from the first portion, and the remaining portion of the second matrix, which is different from the second portion, to the data processing circuit corresponding to the command of operation (810). State (503) of FIG. 5c may relate to a state of the DMA controller that transmits the remaining portions based on operation (850).

[0120] In one embodiment, the operation of the DMA controller is described to determine whether to perform the entire operation of the first matrix and the second matrix based on a comparison of the size of the operation buffer and the size of the first matrix, but the embodiment is not limited thereto. Hereinafter, with reference to FIG. 9, various conditions under which the DMA controller assists the operation of a data processing circuit such as an NPU circuit are exemplarily described.

[0121] FIG. 9 illustrates an exemplary flowchart of a DMA controller included in an electronic device according to one embodiment. The electronic device (101) and / or the DMA controller (132) of FIGS. 1A, 1B, and 2 may perform operations described with reference to FIG. 9. Operations of the electronic device and / or the DMA controller described with reference to FIG. 9 may be related to at least one of the operations of FIG. 3 and / or FIG. 8.

[0122] Referring to FIG. 9, in operation (910), a DMA controller of an electronic device according to an embodiment may receive a command related to an operation of first data and second data from a data processing circuit (e.g., data processing circuit (220) of FIG. 2). The first data and the second data of operation (910) may include a matrix related to an operation of an NPU circuit. The embodiment is not limited thereto. The command of operation (910) may include data and / or values ​​to be stored in a plurality of registers included in the DMA controller (e.g., registers having names exemplified with reference to Table 1). The DMA controller may detect the command using registers set based on the command.

[0123] Referring to FIG. 9, in operation (920), the DMA controller of the electronic device according to one embodiment may determine whether a command received based on operation (910) satisfies a specified condition related to an operation using the DMA controller. The specified condition of operation (920) may be set to determine whether to assist the operation of the data processing circuit using the DMA controller. The specified condition of operation (920) may include at least one of the exemplary conditions described below.

[0124] For example, the specified condition of the operation (920) may be related to the data types of the first data and the second data. The data types may have names such as fp (floating point)8 and / or int8 (integer number)8. For example, fp8 may represent a data type for expressing a floating point number using 8-bit binary numbers. For example, int8 may represent a data type for expressing an integer using 8-bit binary numbers. In addition to the above-described fp8 and in8, the numerical values ​​stored in the first data and / or the second data, which are matrices, may be expressed by any one of various data types, such as fp16, fp32, int2, int4, int16, and / or int32. The DMA controller may include an arithmetic circuit (e.g., an arithmetic circuit (270) of FIG. 2) that supports an operation of a specific data type. The specified condition of operation (920) may include whether the data type supported by the operation circuit of the DMA controller and the data types of the first data and the second data match each other. For example, if the data type supported by the operation circuit and the data types of the first data and the second data match each other, the DMA controller may determine that the command of operation (910) satisfies the specified condition of operation (920).

[0125] For example, the specified condition of operation (920) may be related to the operation (or the complexity of the operation) indicated by the command of operation (920). For example, when performing an operation (e.g., convolution and / or softmax) that is different from the matrix multiplication described with reference to FIGS. 5A to 5C and / or the element-by-element operation described with reference to FIG. 6, the DMA controller may determine that the specified condition of operation (920) is not satisfied. When the operation indicated by the command of operation (910) is an operation supported by the operation circuit of the DMA controller, the controller may determine that the specified condition of operation (920) is satisfied.

[0126] For example, the specified condition of operation (920) may relate to the size of the first data and the second data. If the first data and / or the second data is a matrix and / or a vector, the size may include the dimension of the matrix and / or vector. For example, the specified condition of operation (920) may relate to the state of the data processing circuit. For example, if the data processing circuit is operating within a specified state for reducing power consumption of the data processing circuit, such as a sleep state and / or a low power state, the DMA controller may determine that the specified condition of operation (920) is satisfied. For example, if the temperature of the data processing circuit and / or a processor including the data processing circuit exceeds a specified temperature threshold or falls within a specified temperature range indicating overheating, the DMA controller may determine that the specified condition of operation (920) is satisfied. The temperature may be detected or identified by a temperature sensor included in or adjacent to the processor.

[0127] In one embodiment, if the operation indicated by the command of operation (910) satisfies the specified condition of operation (920) (920-Yes), the DMA controller can perform operation (930). If the operation indicated by the command of operation (910) does not satisfy the specified condition of operation (920) (920-No), the DMA controller can perform operation (940).

[0128] Referring to FIG. 9, in operation (940), a DMA controller of an electronic device according to an embodiment may transmit first data and second data to a data processing circuit. In an embodiment, the DMA controller performing operation (940) may transmit the first data and the second data to the data processing circuit without assistance of an operation of the first data and the second data. The embodiment is not limited thereto, and the DMA controller may transmit the first data and the second data to an address indicated by a command of operation (910). For example, if the command of operation (910) does not satisfy a specified condition of operation (920), the DMA controller may not perform, at least partially, the operation indicated by the command. The DMA controller may perform operation (940) within a first mode.

[0129] Referring to FIG. 9, in operation (930), a DMA controller of an electronic device according to an embodiment may perform at least partially an operation of first data and second data indicated by a command using the DMA controller. If the command of operation (910) satisfies a specified condition of operation (920), the DMA controller may at least partially perform the operation indicated by the command. The DMA controller may perform operation (930) in a second mode and / or a third mode. The DMA controller may perform the operation of the first data and the second data with priority over the operation by the data processing circuit.

[0130] In one embodiment, the DMA controller can transmit the result of at least partially performing the operation of the first data and the second data to a circuit (e.g., a data processing circuit and / or a memory) indicated by the command of operation (910). The DMA controller can transmit remaining portions, which are different from at least a portion of the first data and at least a portion of the second data used in the operation of the DMA controller, to the data processing circuit that provided the command of operation (910). By transmitting the remaining portions, the data processing circuit can complete the operation of the first data and the second data.

[0131] As described above, according to one embodiment, a DMA controller of an electronic device may at least partially perform operations of various data processing circuits connected to the DMA controller. The DMA controller may at least partially perform operations on data while transmitting and / or caching data in response to a request of the data processing circuit. The DMA controller may further include an operation circuit and / or an operation buffer controllable by the data processing circuit. The data processing circuit may determine an operation to be performed (or assisted) by the DMA controller and / or data to be used in the operation by adjusting data and / or values ​​stored in registers of the DMA controller (e.g., registers described with reference to Table 1).

[0132] FIG. 10 is a block diagram of an electronic device (1001) within a network environment (1000) according to various embodiments. Referring to FIG. 10, in the network environment (1000), the electronic device (1001) may communicate with the electronic device (1002) via a first network (1098) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1004) or the server (1008) via a second network (1099) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1001) may communicate with the electronic device (1004) via the server (1008). According to one embodiment, the electronic device (1001) may include a processor (1020), a memory (1030), an input module (1050), an audio output module (1055), a display module (1060), an audio module (1070), a sensor module (1076), an interface (1077), a connection terminal (1078), a haptic module (1079), a camera module (1080), a power management module (1088), a battery (1089), a communication module (1090), a subscriber identification module (1096), or an antenna module (1097). In some embodiments, the electronic device (1001) may omit at least one of these components (e.g., the connection terminal (1078)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1076), camera module (1080), or antenna module (1097)) may be integrated into a single component (e.g., display module (1060)).

[0133] The processor (1020) may, for example, execute software (e.g., a program (1040)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1001) connected to the processor (1020) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1020) may store commands or data received from other components (e.g., a sensor module (1076) or a communication module (1090)) in a volatile memory (1032), process the commands or data stored in the volatile memory (1032), and store result data in a non-volatile memory (1034). According to one embodiment, the processor (1020) may include a main processor (1021) (e.g., a central processing unit or an application processor) or an auxiliary processor (1023) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1021). For example, when the electronic device (1001) includes the main processor (1021) and the auxiliary processor (1023), the auxiliary processor (1023) may be configured to use less power than the main processor (1021) or to be specialized for a given function. The auxiliary processor (1023) may be implemented separately from the main processor (1021) or as a part thereof.

[0134] The auxiliary processor (1023) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1060), the sensor module (1076), or the communication module (1090)) of the electronic device (1001), for example, on behalf of the main processor (1021) while the main processor (1021) is in an inactive (e.g., sleep) state, or together with the main processor (1021) while the main processor (1021) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1023) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1080) or a communication module (1090)). In one embodiment, the auxiliary processor (1023) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1001) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1008)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0135] The memory (1030) can store various data used by at least one component (e.g., the processor (1020) or the sensor module (1076)) of the electronic device (1001). The data can include, for example, software (e.g., the program (1040)) and input data or output data for commands related thereto. The memory (1030) can include volatile memory (1032) or non-volatile memory (1034).

[0136] The program (1040) may be stored as software in memory (1030) and may include, for example, an operating system (1042), middleware (1044), or an application (1046).

[0137] The input module (1050) can receive commands or data to be used in a component of the electronic device (1001) (e.g., a processor (1020)) from an external source (e.g., a user) of the electronic device (1001). The input module (1050) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0138] The audio output module (1055) can output audio signals to the outside of the electronic device (1001). The audio output module (1055) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0139] The display module (1060) can visually provide information to an external party (e.g., a user) of the electronic device (1001). The display module (1060) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1060) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0140] The audio module (1070) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1070) can acquire sound through the input module (1050), output sound through the sound output module (1055), or an external electronic device (e.g., electronic device (1002)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1001).

[0141] The sensor module (1076) can detect the operating status (e.g., power or temperature) of the electronic device (1001) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1076) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0142] The interface (1077) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1001) to an external electronic device (e.g., the electronic device (1002)). In one embodiment, the interface (1077) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0143] The connection terminal (1078) may include a connector through which the electronic device (1001) may be physically connected to an external electronic device (e.g., the electronic device (1002)). In one embodiment, the connection terminal (1078) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0144] The haptic module (1079) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1079) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0145] The camera module (1080) can capture still images and videos. In one embodiment, the camera module (1080) may include one or more lenses, image sensors, image signal processors, or flashes.

[0146] The power management module (1088) can manage power supplied to the electronic device (1001). According to one embodiment, the power management module (1088) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0147] A battery (1089) may power at least one component of the electronic device (1001). In one embodiment, the battery (1089) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0148] The communication module (1090) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1001) and an external electronic device (e.g., electronic device (1002), electronic device (1004), or server (1008)), and the performance of communication through the established communication channel. The communication module (1090) may operate independently from the processor (1020) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1090) may include a wireless communication module (1092) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1094) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1004) via a first network (1098) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1099) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1092) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1096) to verify or authenticate the electronic device (1001) within a communication network such as the first network (1098) or the second network (1099).

[0149] The wireless communication module (1092) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1092) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1092) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1092) may support various requirements specified in the electronic device (1001), an external electronic device (e.g., the electronic device (1004)), or a network system (e.g., the second network (1099)). According to one embodiment, the wireless communication module (1092) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0150] The antenna module (1097) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1097) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1097) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1098) or the second network (1099), may be selected from the plurality of antennas, for example, by the communication module (1090). A signal or power may be transmitted or received between the communication module (1090) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1097).

[0151] According to various embodiments, the antenna module (1097) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0152] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0153] According to one embodiment, commands or data may be transmitted or received between the electronic device (1001) and an external electronic device (1004) via a server (1008) connected to a second network (1099). Each of the external electronic devices (1002, or 704) may be the same or a different type of device as the electronic device (1001). According to one embodiment, all or part of the operations executed in the electronic device (1001) may be executed in one or more of the external electronic devices (1002, 704, or 708). For example, when the electronic device (1001) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1001) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1001). The electronic device (1001) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1001) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (1004) may include an Internet of Things (IoT) device. The server (1008) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1004) or server (1008) may be included within the second network (1099). The electronic device (1001) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0154] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0155] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0156] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0157] Various embodiments of the present document may be implemented as software (e.g., a program (1040)) including one or more instructions stored in a storage medium (e.g., an internal memory (1036) or an external memory (1038)) readable by a machine (e.g., an electronic device (1001)). For example, a processor (e.g., a processor (1020)) of the machine (e.g., an electronic device (1001)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0158] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0159] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added. The electronic device (1001) of FIG. 10 may be an example of the electronic device (101) described with reference to FIGS. 1A, 1B, 2 to 9.

[0160] In one embodiment, a method may be required for performing an operation of a data processing circuit, such as an NPU circuit, using a DMA controller that provides data to the data processing circuit for the operation. As described above, according to one embodiment, a portable communication device may include a memory, a processing circuit, and a direct memory access (DMA) controller (e.g., DMA controller (132) of FIGS. 1A, 1B, and 2) located external to the processing circuit. The DMA controller may be configured to perform a first operation on first data among first data and second data when controlled by the processing circuit. The DMA controller may be configured to transmit a result of the first operation to the processing circuit or the memory. The processing circuit may be configured to perform a second operation on the result of the first operation or the second data. The processing circuit may be configured to transmit a result of the second operation to the memory.

[0161] For example, the processing circuit may include a neural network processing unit (NPU) circuit (e.g., the NPU circuit (154) of FIGS. 1A and 1B).

[0162] For example, the first operation or the second operation may include at least one of a dot product, a matrix multiplication, an element-wise addition, an element-wise subtraction, or an element-wise multiplication.

[0163] For example, the types of the first operation and the second operation may be the same.

[0164] For example, the first operation may include at least one of a dot product, a matrix multiplication, an element-wise addition, an element-wise subtraction, or an element-wise multiplication. The second operation may include a convolution.

[0165] For example, the DMA controller may be configured to perform the first operation on the first data with priority over the processing circuit when a specified condition related to the portable communication device, the processing circuit, or the DMA controller is satisfied.

[0166] For example, the specified condition may include when the portable communication device is in a sleep mode or a low power mode, or when heat in an area adjacent to the processing circuit or the DMA controller exceeds a specified temperature.

[0167] For example, the DMA controller may include an input / output buffer (e.g., the input / output buffer (240) of FIG. 2), an operation buffer (e.g., the operation buffer (250) of FIG. 2), and an operation circuit (e.g., the operation circuit (270) of FIG. 2). The input / output buffer may store one of the first data received from the memory or third data on which the first operation is to be performed. The operation buffer may store the other of the first data or the third data. The operation circuit may be configured to obtain the corresponding one of the first data or the third data from the input / output buffer. The operation circuit may be configured to obtain the corresponding other of the first data or the third data from the operation buffer. The operation circuit may be configured to perform the first operation between the first data and the third data.

[0168] For example, the corresponding other one of the first data or the third data stored in the operation buffer may be determined based on the size of the operation buffer.

[0169] For example, the portable communication device may include a central processing unit (CPU) circuit (e.g., the CPU circuit (136) of FIGS. 1A and 1B). The CPU circuit may be configured to determine the corresponding other one of the first data or the third data to be stored in the operation buffer based on the size of the operation buffer.

[0170] For example, the processing circuit and the DMA controller may be included within one chipset (e.g., chipset (160) of FIG. 1B).

[0171] For example, the chipset may include a first die and a second die positioned above or below the first die. The processing circuit may be disposed on the first die. The DMA controller may be disposed on the second die.

[0172] For example, the first die and the second die may be electrically connected through a plurality of through silicon vias (TSVs) (e.g., the plurality of TSVs (142) of FIGS. 1A and 1B).

[0173] For example, the memory and the DMA controller may be electrically connected via a first TSV among the plurality of TSVs. The DMA controller and the processing circuit may be electrically connected via a second TSV among the plurality of TSVs.

[0174] For example, a first semiconductor circuit formed by a first line width process and including the processing circuit may be disposed on the first die. A second semiconductor circuit formed by a second line width process different from the first line width process and including the DMA controller may be disposed on the second die.

[0175] For example, the first line width process may be based on a wider line width than the second line width process.

[0176] For example, the first line width process may be a 4 nm process, and the second line width process may be a 3 nm process.

[0177] For example, the DMA controller may be configured to determine, using a destination address set by the processing circuit, a circuit among the processing circuit or the memory of the portable communication device to which the result of the first operation is to be transmitted, in relation to an address map to which the processing circuit and the memory of the portable communication device are mapped.

[0178] According to one embodiment, as described above, an electronic device (e.g., electronic device (101) of FIG. 1A, FIG. 1B and / or electronic device (1001) of FIG. 10) may include a volatile memory (e.g., volatile memory (121) of FIG. 1A, FIG. 1B, FIG. 2) and a processor (e.g., processor (110) of FIG. 1A). The processor may include a first die including a neural processing unit (NPU) circuit (e.g., an NPU circuit (154) of FIGS. 1A and 1B), a second die including a cache memory (e.g., a cache memory (134) of FIGS. 1A and 1B) and a direct memory access (DMA) controller for data communication between the cache memory and the volatile memory (e.g., a DMA controller (132) of FIGS. 1A, 1B, and 2), and a bus interface (e.g., a bus interface (210) of FIG. 2) including through silicon vias (TSVs) (e.g., a plurality of TSVs (142) of FIGS. 1A and 1B) disposed between the first die and the second die. The DMA controller may be configured to store data stored in the volatile memory in the cache memory based on identifying a command related to an operation of data associated with a neural network from the NPU circuit. The DMA controller may be configured to perform an operation on a portion of the data stored in the cache memory. The DMA controller may be configured to transmit the remaining portion to the NPU circuit of the first die via the bus interface so that the NPU circuit may perform an operation on the remaining portion of the data, which is different from the portion.

[0179] For example, the DMA controller may be configured to be controllable by a central processing unit (CPU) circuit (e.g., CPU circuit (136) of FIG. 1A) included in the first die or the second die.

[0180] For example, the DMA controller may be configured to include a first buffer for the data communication, a second buffer for storing data related to the operation indicated by the command, and an operation circuit for performing an operation on data stored in the first buffer and the second buffer.

[0181] For example, the DMA controller may be configured to store a plurality of rows included in a first portion of the first matrix stored in the second buffer and to store a column included in a second portion of the second matrix stored in the first buffer based on identifying the command related to a multiplication operation of the first matrix and the second matrix. The DMA controller may be configured to obtain multiplication results of each of the plurality of rows by the column using the operation circuit.

[0182] For example, the DMA controller may be configured to store a plurality of rows included in a first portion of the first matrix stored in the second buffer and to store at least one row included in a second portion of the second matrix stored in the first buffer based on identifying the command related to an element-wise operation of the first matrix and the second matrix. The DMA controller may be configured to obtain, using the operation circuit, an operation result of a first element among the plurality of rows stored in the second buffer and a second element of the at least one row stored in the first buffer.

[0183] For example, the DMA controller may be configured to perform the operation on the entire first matrix and the entire second matrix associated with the command using the operation circuit without transmitting the command to the NPU circuit based on the bus interface, based on identifying the command associated with the first matrix less than or equal to the size of the second buffer.

[0184] For example, the DMA controller may include a mode register for storing a mode of the DMA controller corresponding to the operation of the first matrix and the second matrix indicated by the command. The DMA controller may include a size register for storing a size of a first portion of the first matrix to be stored in the second buffer. The DMA controller may include a start address register for storing an address at which a result calculated from the first portion and the second portion based on the operation circuit is to be stored.

[0185] For example, the size register may include a width register and / or a height register for storing the width and height of the first portion, respectively.

[0186] For example, the DMA controller may include a plurality of registers for storing information related to the command. The plurality of registers may include a source address register for storing an address of the data to be loaded into the DMA controller. The plurality of registers may include a destination address register for storing an address at which the data is to be stored by the DMA controller. The plurality of registers may include a data size register for storing a size of the data to be loaded into the DMA controller.

[0187] For example, the DMA controller may be configured to transmit a matrix including a result of performing the operation on the portion to a circuit corresponding to a destination address set by the command among the volatile memory, the cache memory, or the NPU circuit.

[0188] As described above, in one embodiment, a method of a direct memory access (DMA) controller for data communication between a volatile memory and a cache memory within a processor is provided. The method may include storing data stored in the volatile memory in a cache memory based on identifying a command related to an operation of data associated with a neural network from a neural processing unit (NPU) circuit disposed on a first die of the processor. The DMA controller may be disposed on a second die different from the first die. The method may include performing an operation on a portion of the data stored in the cache memory. The method may include transmitting the remaining portion to the NPU circuit of the first die through a bus interface including through silicon vias (TSVs) disposed between the first die and the second die, so as to perform an operation of the NPU circuit on the remaining portion of the data, which is different from the portion.

[0189] For example, the DMA controller may be configured to be controllable by a central processing unit (CPU) circuit included in the first die or the second die.

[0190] For example, the performing operation may include storing a column included in a second portion of the second matrix in a first buffer of the DMA controller for the data communication, and storing a plurality of rows included in a first portion of the first matrix in a second buffer of the DMA controller, based on identifying the command related to a multiplication operation of the first matrix and the second matrix. The performing operation may include obtaining multiplication results of each of the plurality of rows by the column using an operation circuit of the DMA controller for performing an operation on data stored in the first buffer and the second buffer.

[0191] For example, the performing operation may include storing a plurality of rows included in a first portion of the first matrix stored in the second buffer and storing at least one row included in a second portion of the second matrix stored in the first buffer based on identifying the command related to an element-wise operation of the first matrix and the second matrix. The performing operation may include obtaining, by using the operating circuit, an operation result of a first element among the plurality of rows stored in the second buffer and a second element of the at least one row stored in the first buffer.

[0192] For example, the operation to be performed may include an operation of performing the operation on the entire first matrix and the entire second matrix associated with the command using the operation circuit without transmitting the command to the NPU circuit based on the bus interface, based on identifying the command associated with the first matrix less than or equal to the size of the second buffer.

[0193] For example, the storing operation may include an operation of setting a plurality of registers of the DMA controller based on the command. The plurality of registers may include a source address register for storing an address of the data to be loaded into the DMA controller. The plurality of registers may include a destination address register for storing an address at which the data is to be stored by the DMA controller. The plurality of registers may include a data size register for storing a size of the data to be loaded into the DMA controller.

[0194] For example, the transmitting operation may include an operation of transmitting a matrix including a result of performing the operation on the portion to a circuit corresponding to a destination address set by the command among the volatile memory, the cache memory, or the NPU circuit.

[0195] According to one embodiment, a processing chip component as described above may include a port for data communication with a volatile memory, a direct memory access (DMA) controller connected to the port, and a plurality of data processing circuits configured to control the DMA controller. The DMA controller may be configured to identify a command from any one of the plurality of data processing circuits to perform an operation on a first matrix and a second matrix using the data communication using the DMA controller. The DMA controller may be configured to obtain the first matrix and the second matrix from the volatile memory through the port based on the command. The DMA controller may be configured to perform the operation on the first matrix and the second matrix indicated by the command before transmitting the first matrix and the second matrix to the data processing circuit corresponding to the command in response to obtaining the first matrix from the volatile memory, the first matrix being less than or equal to a size of a buffer included in the DMA controller. The DMA controller may be configured to perform an operation on a first portion of the first matrix and a second portion of the second matrix based on obtaining the first matrix from the volatile memory, the first matrix exceeding the size of the buffer. The DMA controller may be configured to transmit, to the data processing circuit corresponding to the command, a remaining portion of the first matrix different from the first portion and a remaining portion of the second matrix different from the second portion.

[0196] For example, the plurality of data processing circuits may include a neural processing unit (NPU) circuit configured to perform operations related to a neural network. The plurality of data processing circuits may include a graphic processing unit (GPU) circuit configured to perform operations related to graphics rendering. The plurality of data processing circuits may include a central processing unit (CPU) circuit configured to perform operations indicated by instructions stored in the volatile memory.

[0197] For example, the processing chip component may include a first die on which the NPU circuit that transmits the command is disposed. The processing chip component may include a second die on which the DMA controller is disposed. The processing chip component may include a bus interface including through silicon vias (TSVs) disposed between the first die and the second die.

[0198] For example, the DMA controller may include a first buffer in which data input to or output from the DMA controller is stored. The DMA controller may include a second buffer corresponding to the buffer and configured to store the first portion of the first matrix. The DMA controller may include an arithmetic circuit connected to the first buffer and the second buffer.

[0199] As described above, in one embodiment, a method of a processing chip component is provided, including a port for data communication with a volatile memory, a direct memory access (DMA) controller connected to the port, and a plurality of data processing circuits configured to control the DMA controller. The method may include an operation of identifying, from any one of the plurality of data processing circuits, a command for performing an operation on a first matrix and a second matrix using the data communication using the DMA controller. The method may include an operation of obtaining, based on the command, the first matrix and the second matrix from the volatile memory through the port. The method may include an operation of performing the operation on the first matrix and the second matrix indicated by the command before transmitting the first matrix and the second matrix to the data processing circuit corresponding to the command, in response to obtaining, from the volatile memory, the first matrix, which is less than or equal to a size of a buffer included in the DMA controller. The method may include performing an operation on a first portion of the first matrix and a second portion of the second matrix based on obtaining the first matrix from the volatile memory, the first matrix exceeding the size of the buffer. The method may include transmitting, to the data processing circuit corresponding to the command, a remaining portion of the first matrix different from the first portion and a remaining portion of the second matrix different from the second portion.

[0200] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0201] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0202] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording media or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0203] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0204] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In a portable communication device, memory; processing circuit; and A direct memory access (DMA) controller (132) located outside the processing circuit, the DMA controller comprising: When controlled by the above processing circuit, a first operation is performed on the first data among the first data and the second data, configured to transmit the result of the first operation to the processing circuit or the memory, The above processing circuit: Performing a second operation on the result of the first operation or the second data, configured to transmit the result of the second operation to the memory, Portable communication device.

2. In claim 1, The above processing circuit includes an NPU (neural network processing unit) circuit (154). Portable communication device.

3. In claims 1 and 2, Wherein the first operation or the second operation comprises at least one of a dot product, a matrix multiplication, an element-wise addition, an element-wise subtraction, or an element-wise multiplication. Portable communication device.

4. In claims 1 to 3, The type of the above first operation and the type of the above second operation are the same. Portable communication device.

5. In claims 1 to 4, The first operation comprises at least one of a dot product, a matrix multiplication, an element-wise addition, an element-wise subtraction, or an element-wise multiplication, The second operation includes convolution, Portable communication device.

6. In claims 1 to 5, the DMA controller: When a specified condition related to the portable communication device, the processing circuit, or the DMA controller is satisfied, the first operation on the first data is configured to be performed with priority over the processing circuit. Portable communication device.

7. In claims 1 to 6, The above specified conditions include when the portable communication device is in a sleep mode or a low power mode, or when heat in an area adjacent to the processing circuit or the DMA controller exceeds a specified temperature. Portable communication device.

8. In claims 1 to 7, The above DMA controller includes an input / output buffer (240), an operation buffer (250), and an operation circuit (270). In the above input / output buffer, one of the first data received from the memory or the third data on which the first operation is to be performed is stored, In the above operation buffer, the other of the first data or the third data is stored, The above operation circuit: Obtaining the corresponding one of the first data or the third data from the input / output buffer, Obtaining the corresponding other one of the first data or the third data from the operation buffer, configured to perform the first operation between the first data and the third data, Portable communication device.

9. In claims 1 to 8, The corresponding other one of the first data or the third data stored in the operation buffer is determined based on the size of the operation buffer. Portable communication device.

10. In claims 1 to 9, Contains a CPU (central processing unit) circuit (136), The CPU circuit is configured to determine the corresponding other one of the first data or the third data stored in the operation buffer based on the size of the operation buffer. Portable communication device.

11. In claims 1 to 10, The above processing circuit and the DMA controller are included in one chipset (160). Portable communication device.

12. In claims 1 to 11, The chipset comprises a first die and a second die positioned above or below the first die, The above processing circuit is arranged on the first die, The above DMA controller is placed on the second die, Portable communication device.

13. In claims 1 to 12, The above first die and the above second die are electrically connected through a plurality of through silicon vias (TSVs) (142). Portable communication device.

14. In claims 1 to 13, The above memory and the DMA controller are electrically connected through a first TSV among the plurality of TSVs, The above DMA controller and the processing circuit are electrically connected through a second TSV among the plurality of TSVs. Portable communication device.

15. A method of a DMA (direct memory access) controller for data communication between volatile memory and cache memory within a processor, An operation of storing data stored in the volatile memory in the cache memory based on identifying a command related to an operation of data related to a neural network from an NPU (neural processing unit) circuit disposed on a first die of the processor, wherein the DMA controller is disposed on a second die different from the first die; An operation for performing an operation on a portion of the data stored in the cache memory; In order to perform an operation of the NPU circuit on the remaining portion of the data different from the above portion, the remaining portion includes an operation of transmitting the remaining portion to the NPU circuit of the first die through a bus interface including through silicon vias (TSVs) arranged between the first die and the second die. method.

Citation Information

Patent Citations

  • dma controller, implementation and computer storage medium

    JP2018511891A

  • Arithmetic circuit, arithmetic device, method, and program

    JP2023009973A

  • Method for synchronizing information between order system and POS system

    KR1020210123159A

  • Lining paper and manufacturing method thereof

    KR102368179B1

  • Stacked die network-on-chip for FPGA

    US8863065B1