Electronic device for accelerating execution of model and method of operating same

By designing in-memory processing (PIM) controllers and PIMs in electronic devices, executing nonlinear functions and generating PIM requests, the operational characteristics limitations and communication overhead problems of PIM accelerated AI model execution in the prior art are solved, and more efficient AI model execution and cost reduction effects are achieved.

CN120104556APending Publication Date: 2025-06-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411779339.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art faces problems of operational characteristics limitations and increased communication overhead when accelerating the execution of AI models using in-memory processing (PIM), which makes it difficult to achieve optimal performance.

Method used

An electronic device is designed, including an in-memory processing (PIM) controller and a PIM controller, which is responsible for performing nonlinear function (NLF) operations in multiple requested operations and generates PIM requests to perform PIM operations, reducing the participation of the host processor.

Benefits of technology

By reducing the communication overhead between the host processor and PIM, the execution efficiency of the AI ​​model is improved, the overall overhead of the system is reduced, and the cost is effectively reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104556A_ABST
    Figure CN120104556A_ABST
Patent Text Reader

Abstract

An electronic device for accelerating execution of a model and a method of operating the same are provided. The electronic device includes: an in-memory processing (PIM) controller; and a PIM configured to perform a PIM operation in response to the PIM request generated by the PIM controller. The PIM controller is configured to perform a non-linear function (NLF) operation among the operations of the plurality of requests, generate a PIM request for the PIM operation, and transmit the PIM request to the PIM.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims the priority benefit of Korean Patent Application No. 10-2023-0175943 filed in the Korean Intellectual Property Office on December 6, 2023, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] The disclosed embodiments are directed to an electronic device and a method of operating the electronic device for accelerating generation of a model. Background Art

[0003] Artificial intelligence (AI) models are computer programs that use a collection of data sets to detect specific patterns. Executing AI models can be time consuming. Processing in memory (PIM) can be used to accelerate the execution of AI models. PIM is a computer architecture in which data operations are available directly on a data memory without having to be executed on an external host processor. PIM is a semiconductor memory device that combines the functionality of a memory with the functionality of a processor for performing arithmetic operations. However, because some operations are difficult to execute in a PIM due to their operational characteristics, they may still need to be executed by a host processor. In addition, due to the increased communication overhead between the memory and the host processor, achieving optimal performance is challenging. Summary of the invention

[0004] According to an embodiment, an electronic device is provided, the electronic device including a process in memory (PIM) controller and a PIM. The PIM is configured to perform a PIM operation in response to a PIM request generated by the PIM controller. The PIM controller is configured to: perform a non-linear function (NLF) operation among a plurality of requested operations, generate a PIM request for the PIM operation, and send the PIM request to the PIM.

[0005] The PIM controller may include a first logic circuit configured to perform an NLF operation among the plurality of requested operations.

[0006] The PIM controller may also include: a command queue configured to store commands for the multiple operations received from the host processor in the electronic device; a second logic circuit configured to generate a PIM request for a PIM operation among the multiple requested operations; and a control register configured to control the operation of the first logic circuit.

[0007] The PIM controller may be configured to classify each of the plurality of requested operations into one of a PIM operation and an NLF operation according to type information of each of the plurality of requested operations.

[0008] The electronic device may further include a host processor configured to send a request for an operation of the plurality of requests to the PIM controller.

[0009] The operations of the plurality of requests may be performed by the PIM controller and the PIM, rather than by the host processor.

[0010] The electronic device may further include a memory controller configured to generate a PIM command and transmit the PIM command to the PIM in response to a PIM request received from the PIM controller.

[0011] The PIM controller may be provided in a memory controller included in the electronic device and configured to manage data input to or output from the PIM, or a direct memory access (DMA) included in the electronic device and configured to access data stored in the PIM.

[0012] The PIM controller may be configured to send results of the plurality of requested operations to the host processor in response to all of the plurality of requested operations being performed.

[0013] The PIM may include a data storage space and an operator configured to perform PIM operations in response to a PIM request.

[0014] The data storage space may be configured to store the results of PIM operations and the results of NLF operations.

[0015] According to an embodiment, a method for operating an electronic device is provided, the method comprising: a process in memory (PIM) controller in the electronic device classifies a target operation to be processed into one of a PIM operation and a non-linear function (NLF) operation based on an order of requests for a plurality of operations; in response to classifying the target operation as an NLF operation, the PIM controller executes the target operation corresponding to the NLF operation; in response to classifying the target as a PIM operation, the PIM controller generates a PIM request for the target operation and sends the PIM request to a PIM in the electronic device; and the PIM executes the target operation corresponding to the PIM operation according to the PIM request received from the PIM controller.

[0016] The step of performing the target operation corresponding to the NLF operation may be performed by a first logic circuit included in the PIM controller and configured to perform the NLF operation among the plurality of operations.

[0017] The method may further include storing commands for the plurality of operations received from a host processor in the electronic device in a command queue included in the PIM controller.

[0018] The classifying of the target operation may include classifying the target operation into one of a PIM operation and an NLF operation according to type information of the target operation.

[0019] The method may also include sending, by a host processor in the electronic device, a request for the plurality of operations to the PIM controller.

[0020] The various operations may be performed by the PIM controller and the PIM, rather than by the host processor.

[0021] The method may further include generating, by a memory controller in the electronic device, a PIM command in response to the PIM request received from the PIM controller, and sending the PIM command to the PIM.

[0022] The method may further include sending, by the PIM controller, results of the plurality of operations to a host processor in the electronic device in response to all of the plurality of operations being performed.

[0023] According to an embodiment, an electronic device is provided, the electronic device including a process in memory (PIM) and a direct memory access (DMA). The PIM performs a PIM operation in response to a PIM request. The PIM controller is configured to: perform a non-linear function (NLF) operation among a plurality of operations requested by a host processor, generate a PIM request for the PIM operation among the plurality of operations, and send the PIM request to the PIM.

[0024] The electronic device may further include a memory controller, wherein the PIM controller sends the PIM request to the memory controller, and the memory controller forwards the PIM request to the PIM.

[0025] The electronic device may further include an interconnect connected to the host processor, the DMA, and the memory controller, wherein the DMA receives commands for the plurality of operations from the host processor through the interconnect, and the memory controller receives the PIM request through the interconnect. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] These and / or other aspects and features disclosed will become clear and more easily understood through the following description of embodiments in conjunction with the accompanying drawings.

[0027] Figure 1 is a diagram illustrating an electronic device according to an embodiment.

[0028] Figure 2 is a diagram illustrating operations of performing a plurality of operations including a process in memory (PIM) operation and a non-linear function (NLF) operation according to an embodiment.

[0029] Figure 3is a diagram illustrating an example of sequentially performing a PIM operation and an NLF operation according to an embodiment.

[0030] Figure 4 is a diagram illustrating an operation of performing a large language model (LLM) based transformer decoder according to an embodiment.

[0031] Figure 5 is a diagram illustrating an example of an LLM acceleration system using a compute express link (CXL)-PIM card according to an embodiment.

[0032] Figure 6 is a diagram illustrating an electronic device according to an embodiment.

[0033] Figure 7 is a diagram illustrating an electronic device according to an embodiment.

[0034] Figure 8 is a diagram illustrating a method of operating an electronic device according to an embodiment. DETAILED DESCRIPTION

[0035] Embodiments will now be described more fully below with reference to the accompanying drawings. However, embodiments may be provided in different forms and should not be construed as limiting. Throughout the disclosure, the same reference numerals may indicate the same components.

[0036] As used herein, "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B and C", "at least one of A, B or C", and "one or a combination of at least two of A, B, and C", each of which may include any one of the items listed together in a corresponding one of the multiple phrases, or all possible combinations thereof.

[0037] It should be noted that if one component is described as being “connected,” “coupled,” or “engaged” to another component, although the first component may be directly connected, coupled, or engaged to the second component, a third component may be “connected,” “coupled,” or “engaged” between the first and second components.

[0038] Unless the context clearly indicates otherwise, the singular forms are intended to include the plural forms as well.

[0039] Hereinafter, the embodiments are described in detail with reference to the accompanying drawings. When describing the embodiments with reference to the accompanying drawings, the same reference numerals denote the same elements, and repeated descriptions thereof are omitted.

[0040] Figure 1 is a diagram illustrating an electronic device according to an embodiment.

[0041] Reference Figure 1, the electronic device 100 may include a host processor 110, a process in memory (PIM) controller 120 (e.g., a first controller circuit), a memory controller 130 (e.g., a second controller circuit), and a PIM (or PIM device) 140. The host processor 110, the PIM controller 120, the memory controller 130, and the PIM 140 may communicate with each other via an interconnect 150 (e.g., an interconnect circuit). For example, the interconnect 150 may include a bus, a compute express link (CXL), and a peripheral component interconnect express (PCIe). However, embodiments are not limited thereto. When the interconnect 150 is an internal bus, the memory controller 130 may be packaged together with the host processor 110 to allow the memory controller 130 to be placed in the host processor 110. Optionally, when the interconnect 150 is a CXL, as described below with reference to Figure 5 As described above, the PIM controller 120 may be placed in the CXL controller. Figure 1 As shown in , an embodiment in which the host processor 110, the PIM controller 120, and the memory controller 130 are separate components is described.

[0042] The electronic device 100 may include various computing devices (such as a mobile phone, a smart phone, a tablet personal computer (PC), an e-book device, a laptop computer, a PC, a desktop computer, a workstation or a server), various wearable devices (such as a smart watch, smart glasses, a head-mounted display (HMD) or smart clothing), various home appliances (such as a smart speaker, a smart television (TV) or a smart refrigerator) and other devices (such as a smart vehicle, a smart self-service terminal, an Internet of Things (IoT) device, a walking assistance device (WAD), a drone or a robot).

[0043] The host processor 110 may be a device configured to control the overall operation of the electronic device 100, and may include other processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), or a digital signal processor (DSP)). The host processor 110 may generate requests to various components (e.g., the PIM controller 120, etc.) in the electronic device 100 through a host program. In one example, the host processor 110 may include a cache.

[0044] The request to the PIM controller 120 generated by the host processor 110 may be related to a plurality of operations of a model to be executed. The model may be a neural network to be executed by the electronic device 100, and may include, for example, a large language model (LLM), a speech recognition model, a translation model, or an advanced virtual assistant model. However, embodiments are not limited thereto.

[0045] The plurality of operations may include a PIM operation and a nonlinear function (NLF) operation requested when the electronic device 100 executes the model. The PIM operation may include at least one of an arithmetic operation (such as addition, multiplication, accumulation, and general matrix vector multiplication (GEMV)) and a logic operation (such as AND, OR, and XOR), and may be performed by the PIM 140. Since the PIM operation is directly performed by the PIM 140 storing the operand data, it is not necessary to read the operand data into the host processor 110 and / or use a separate accelerator for the PIM operation. Therefore, power consumption may be minimized by reducing the data movement distance and minimizing the loss of memory bandwidth. The NLF operation may be an operation where the relationship between variables is not linear, and may include, for example, a hyperbolic tangent function (tanh), a sigmoid, a normalized exponential function (softmax), a dropout, and a Gaussian error linear unit (GELU). The NLF operation may be performed by the PIM controller 120. However, due to the operational characteristics of the NLF operation, it may be difficult for the PIM 140 to perform the NLF operation.

[0046] PIM controller 120 may be a device for managing PIM operations performed by PIM 140. PIM controller 120 may include command queue 121, PIM request generator 122 (eg, logic circuit), NLF hardware block 123 (eg, logic circuit, processor, etc.), and control register 124.

[0047] The command queue 121 may store commands for a plurality of operations received from the host processor 110 according to a first-in-first-out (FIFO) structure. The commands stored in the command queue 121 may be processed sequentially according to the FIFO structure. The commands stored in the command queue 121 may have a command structure or format of [op_type, precision, in_addr_0, in_addr_1, out_addr, in_0_size, in_1_size, out_size].

[0048] In the above command structure, the field "op_type" may represent the type of operation. For example, the field "op_type" may represent the type of PIM operation (such as addition, multiplication, GEMV, AND, and OR), or the type of NLF operation (such as tanh, sigmoid, and GELU). The field "precision" may represent the digital format, and include, for example, integer 4 bits (INT4), integer 8 bits (INT8), integer 16 bits (INT16), floating point 8 bits (FP8), floating point 16 bits (FP16), floating point 32 bits (FP32), binary floating point 8 bits (BF8), binary floating point 16 bits (BF16), etc. However, the embodiment is not limited to this. The field "in_addr_0" may represent the address of storing operand data 0, the field "in_addr_1" may represent the address of storing operand data 1, and the field "out_addr" may represent the address of storing output data. The field 'in_0_size' may represent the size of data stored in in_addr_0, the field 'in_1_size' may represent the size of data stored in in_addr_1, and the field 'out_size' may represent the size of data stored in out_addr.

[0049] When an operation type corresponding to a target command to be processed in a requested order among commands stored in the command queue 121 corresponds to a PIM operation, the PIM request generator 122 may generate a PIM request for performing the PIM operation based on the target command. The PIM request generator 122 may send the generated PIM request to the memory controller 130. The memory controller 130 may forward the generated PIM request to the PIM 140.

[0050] When the operation type corresponding to the target command to be processed in the requested order among the commands stored in the command queue 121 corresponds to an NLF operation, the NLF hardware block 123 may perform the NLF operation based on the target command. The NLF hardware block 123 may be a hardware device that performs a nonlinear operation and may be implemented using, for example, a lookup table (LUT), piecewise linear, and direct calculation techniques. However, the embodiments are not limited thereto, and other implementations may be used without limitation.

[0051] For example, the LUT technique may be a technique for storing the x value and y value of the NLF operation in an internal buffer in a table format in advance and retrieving the output value corresponding to the input value from the table. The NLF hardware block 123 using the LUT technique may perform at least one NLF operation by storing at least one table. When the LUT technique is used, as the table data becomes more accurate, the amount of storage space required increases. The piecewise linear technique may compensate for this shortcoming of the LUT technique, and may perform the operation by linearly interpolating the y value according to the x value based on the reference LUT value. The piecewise linear technique may perform the operation by approximating the NLF with a plurality of linear functions. The direct calculation technique may involve equipping a general-purpose operator or a dedicated NLF operation processor in the NLF hardware block 123 to perform the NLF operation. For example, the NLF hardware block 123 using the direct calculation technique may include an arithmetic logic unit (ALU), a register, a static random access memory (SRAM), and a local memory. An operating system (OS) or a dedicated firmware may exist to operate the NLF hardware block 123.

[0052] The control register 124 may provide an external interface for controlling the PIM controller 120. The control register 124 may control functions of the command queue 121, the PIM request generator 122, and the NLF hardware block 123, and may verify operations of the command queue 121, the PIM request generator 122, and the NLF hardware block 123.

[0053] The memory controller 130 may manage data input to or output from the PIM 140. The memory controller 130 may generate a memory command based on a memory request. The memory controller 130 may receive a memory request from the host processor 110 through the interconnect 150. For example, the memory controller 130 may convert the memory request into a memory command including an activate command, a precharge command, a refresh command, a read command, and a write command. In addition, the memory controller 130 may generate a PIM command based on the PIM request. The memory controller 130 may send the generated memory command and / or the generated PIM command to the PIM 140.

[0054] The PIM 140 may be a device for performing a PIM operation by an internal processor without storing data, and may include, for example, a dynamic random access memory (RAM) (DRAM), a high bandwidth memory (HBM), a graphic double data rate (GDDR), or a low power double data rate (LPDDR). However, the embodiment is not limited thereto. The PIM 140 may be a hardware device for performing a PIM operation other than a general memory operation, and may perform other operations, for example, by being programmed. The PIM 140 may include a data storage space for storing data and an internal processor (e.g., an operator) for performing PIM operations including the above-mentioned logical operations and / or arithmetic operations. For example, the internal processor may perform a PIM operation in response to a PIM request. The PIM operation may use the data storage space and the internal processor. The general memory operation may use the data storage space without using the internal processor.

[0055] For example, the PIM 140 may perform a PIM operation on operands stored in the data storage space according to a PIM command transmitted from the memory controller 130, and store the result of the operation in the data storage space. In addition, the PIM 140 may transmit operands for an NLF operation to be performed by the NLF hardware block 123 to the NLF hardware block 123 according to a memory command transmitted from the memory controller 130, and may store the result of the NLF operation in the data storage space.

[0056] All of the multiple operations including PIM operations and NLF operations requested when executing the model in the electronic device 100 can be executed by the PIM controller 120 and the PIM 140 without the help of the host processor 110. This can effectively reduce the communication overhead between the host processor 110 and the PIM 140, thereby accelerating the execution of the model. By reducing the role of the relatively high-cost host processor (i.e., the host processor 110) during the execution of the model and increasing the role of the relatively low-cost PIM controller 120 and the PIM 140, the cost of the electronic device 100 for executing the model can be effectively reduced. Thus, it is feasible to effectively accelerate the execution of the model without loading the data stored in the PIM 140 to the host processor 110 for NLF operations. By including the NLF hardware block 123 that performs the NLF operation in the PIM controller 120 instead of the PIM 140, it is feasible to prevent the increase in the area of ​​the PIM 140.

[0057] Figure 2 is a diagram illustrating performing a plurality of operations including a PIM operation and an NLF operation according to an embodiment.

[0058] In operation 210, the host processor may send a request to the PIM controller (eg, 120) to perform a plurality of operations including a PIM operation and an NLF operation. A command corresponding to the request from the host processor may be stored in a command queue (eg, 121) in the PIM controller.

[0059] In operation 220, the PIM controller may determine the type of target operation to be processed based on the order in which the commands are stored in the command queue. For example, the PIM controller may determine the type of target operation based on the operation type of the target command to be processed in the command queue. When the type of the target operation is a PIM operation, operation 230 may be subsequently performed. When the type of the target operation is an NLF operation, operation 240 may be subsequently performed.

[0060] In operation 230, the PIM controller may generate a PIM request for performing a PIM operation through a PIM request generator (eg, 122) and transmit the generated PIM request to a PIM (eg, 140). The PIM may perform the PIM operation based on the received PIM request.

[0061] In operation 240 , the PIM controller may perform NLF operations through the NLF hardware block (eg, 123 ).

[0062] In operation 250, the PIM controller may determine whether all of the plurality of operations requested by the host processor have been executed based on whether there are any commands remaining in the command queue. When it is determined that not all of the plurality of operations have been executed because there are commands remaining in the command queue, operation 220 may be subsequently executed. Otherwise, when it is determined that all of the plurality of operations have been executed because there are no commands remaining in the command queue, operation 260 may be subsequently executed.

[0063] In operation 260, the PIM controller may send operation results of the plurality of operations to the host processor.

[0064] Figure 3 is a diagram illustrating an example of sequentially performing a PIM operation and an NLF operation according to an embodiment.

[0065] Reference Figure 3 For ease of description, an example in which a plurality of operations are requested in the order of GEMV operation, GELU operation, and GEMV operation is described. However, the embodiments are not limited thereto.

[0066] In operation 301, the host processor may send a request for a bundle of operations to the PIM controller. Here, for ease of description, a plurality of operations may be referred to as a bundle of operations.

[0067] In operation 302, the PIM controller may determine that the first requested GEMV operation (e.g., 1:GEMV) corresponds to a PIM operation. In operation 303, the PIM controller may generate a GEMV operation request and send the GEMV operation request to the PIM (e.g., the PIM controller may request a 1:GEMV operation). In operation 304, the PIM may perform the GEMV operation based on the GEMV operation request (e.g., the PIM may perform the 1:GEMV operation). The operands may be stored in the PIM, and the results of the GEMV operation may also be stored in the PIM.

[0068] In operation 305, the PIM controller may determine that the GELU operation (e.g., 2:GELU) of the second request corresponds to a NLF operation. In operation 306, the PIM controller may directly perform the NLF operation through the NLF hardware block. The operands for the NLF operation may be loaded into the NLF hardware block by the PIM. The result of the NLF operation may be sent from the NLF hardware block to the PIM and stored in the PIM.

[0069] In operation 307, the PIM controller may determine that the third requested GEMV operation (e.g., 3:GEMV) corresponds to a PIM operation. In operation 308, the PIM controller may generate a GEMV operation request and send the GEMV operation request to the PIM (e.g., the PIM controller may request a 3:GEMV operation). In operation 309, the PIM may perform the GEMV operation based on the GEMV operation request (e.g., the PIM may perform a 3:GEMV operation). The operands may be stored in the PIM, and the results of the GEMV operation may also be stored in the PIM.

[0070] In operation 310, when there are no commands remaining in the command queue, the PIM controller may determine that a bundle of operations is completed. In operation 311, the PIM controller may send operation results of the bundle of operations to the host processor.

[0071] Figure 4 is a diagram illustrating the operation of performing an LLM-based transformer decoder according to an embodiment.

[0072] Figure 4 An example of a LLM-based transformer decoder 400 is shown. Figure 4 In the example of , the PIM may have difficulty in executing some parts 410 corresponding to the NLF operation. Other parts may be executed by the PIM. When the NLF hardware block is not included in the PIM controller, as described above, it may be necessary to load the operands stored in the PIM to the host processor and transfer the operation results from the host processor to the PIM to process some parts 410. This process may increase the communication overhead between the host processor and the PIM, thereby increasing the overall system overhead. Figure 4As shown in , the alternating arrangement of some parts 410 and other parts may lead to a significant increase in system overhead. However, the NLF operations of some parts 410 may be processed by the NLF hardware blocks included in the PIM controller, thereby effectively accelerating the LLM-based converter decoder 400 without the help of the host processor. For example, some parts 410 may include a dropout function as a regularization technique for a neural network model, a Gelu function as an activation function of a neural network; a layer normalization (LayerNorm) function as a technique for normalizing the distribution of the intermediate layer of the neural network, a softmax function that converts a vector of real numbers into a probability distribution, or a mask operation. The mask operation may be in the form of a dropout function, in which the contribution of the node is zero. For example, other parts may include a matrix multiplication (Matmul) function and a linear (Linear) function. For example, a block for performing a linear function may include head 1 to head H. For example, the input of the LLM-based converter decoder 400 may be a converter block input, and the output of the LLM-based converter decoder 400 may be a converter block output.

[0073] Figure 5 is a diagram illustrating an example of an LLM acceleration system using a CXL-PIM card according to an embodiment.

[0074] Figure 5 The LLM acceleration system (or node N 本地_专家 ) 500. The CXL-based in-memory processing (CXL-PIM) card 550 in the LLM acceleration system 500 may include the above-mentioned PIM controller (e.g., PIM controller 520). When operations corresponding to experts in the LLM are performed by the CXL-PIM card instead of an accelerator (e.g., a GPU, etc.), cost savings, reduced power consumption, and improved performance due to reduced overhead can be expected compared to when the operations are performed by the accelerator. Experts include Figure 4 The linear+Gelu+linear feed-forward network (FFN) layer shown in FIG. 1 and can be accelerated by the PIM controller and the PIM without the help of an accelerator.

[0075] like Figure 5 As shown in, by establishing an LLM acceleration system using a CXL-PIM card, low cost, low power consumption, and high performance can be achieved compared to an acceleration system using an accelerator. In addition, these effects can be achieved when an on-device LLM model using a service (e.g., an advanced virtual assistant model that helps with various personal tasks (such as scheduling, shopping, news, health, finance, travel, etc.)), a speech recognition service, or a translation service is executed on a mobile device.

[0076] The LLM acceleration system 500 may include a plurality of CXL-PIM cards, wherein each CXL-PIM card 550 may be interfaced with a host processor 510 and may receive a tokenized sentence as input. The host processor 510 may include a memory unit (MC) 515 and / or an interface connected to a memory device 505. The CXL-PIM card 550 may include a decoder and instruction buffer 525, a CXL controller 518, and a PIM 540. The PIM 540 may be similar to the PIM 140. The PIM controller 520 may include a command queue CMDQ 521 including an NLF path, a PIM request generator 522 similar to the PIM request generator 122, an NLF hardware block 523 similar to the NLF hardware block 123, a control register 524 for controlling NLF operations (e.g., performing NLF operations), and a memory unit 530. However, the embodiment is not limited thereto, and for example, the memory unit 530 may be located outside the PIM controller 520. Host processor 510 may communicate with LLM acceleration system 500 using a CXL protocol such as CSL.io and CXL.mem. PIM controller 520 may be similar to PIM controller 120.

[0077] Figure 6 is a diagram illustrating an electronic device according to an embodiment.

[0078] Figure 6 An embodiment of a system 600 is shown in which the above-described PIM controller (eg, PIM controller 620) is disposed in a direct memory access (DMA) (or DMA device) 610. The PIM controller 620 may be similar to the PIM controller 120 or 520.

[0079] DMA 610 may be a function or module of a computer system that allows a predetermined hardware subsystem to access PIM 640 independently of host processor 630. PIM 640 may be similar to PIM 140. DMA 610 may generate a memory request based on a command from host processor 630. Host processor 630 may include cache 635. System 600 may also include an interconnect 650 similar to interconnect 150 and a memory controller 660 similar to memory controller 130.

[0080] The PIM controller 620 may be provided in the DMA 610 and perform the above-mentioned operation. Therefore, a more detailed description thereof is omitted.

[0081] Figure 7 is a diagram illustrating an electronic device according to an embodiment.

[0082] Figure 7An embodiment of a system 700 in which the above-described PIM controller (e.g., PIM controller 720) is provided in a memory controller 710 is shown. The PIM controller 720 may be provided in the memory controller 710 and perform the above-described operations. The PIM controller 720 may be similar to the PIM controller 620. A host processor 730 of the system 700 may include a cache 735. The system 700 may also include a DMA 760 and an interconnect 750 similar to the interconnect 150. The memory controller 710 may be connected to a PIM 740 interface similar to the PIM 640. Therefore, a more detailed description thereof is omitted.

[0083] Figure 8 is a diagram illustrating a method of operating an electronic device according to an embodiment.

[0084] In the following embodiments, the operations may be performed sequentially, but not necessarily sequentially. For example, the order of the operations may be changed, and at least two of the operations may be performed in parallel. Operations 810 to 840 may be performed by at least one component of the electronic device (e.g., a PIM controller, a PIM, etc.).

[0085] The PIM controller classifies a target operation to be processed based on a requested order among a plurality of operations into one of a PIM operation and an NLF operation in operation 810. The PIM controller may classify the target operation into one of a PIM operation and an NLF operation based on type information of the target operation.

[0086] In response to classifying the target operation as the NLF operation, the PIM controller performs the target operation corresponding to the NLF operation in operation 820. The target operation may be included in the PIM controller and performed by an NLF hardware block that performs the NLF operation among a plurality of operations.

[0087] In response to classifying the target operation as a PIM operation, the PIM controller generates a PIM request (e.g., signal, command, etc.) for the target operation and sends the PIM request to the PIM in the electronic device in operation 830. In response to the PIM request received from the PIM controller, the memory controller in the electronic device generates a PIM command and sends the PIM command to the PIM.

[0088] In operation 840 , the PIM performs a target operation corresponding to the PIM operation according to a PIM request received from the PIM controller (eg, according to a PIM command received from the memory controller).

[0089] In one embodiment, a number of operations are performed by the PIM controller and the PIM, and not by the host processor.

[0090] In response to all of the plurality of operations being performed, the PIM controller may send results of the plurality of operations to a host processor in the electronic device.

[0091] Reference Figures 1 to 7 The description provided can be applied to Figure 8 The operations are shown in , and therefore further detailed description thereof is omitted.

[0092] The embodiments described herein may be implemented using hardware components, software components, and / or combinations thereof. A processing device (e.g., a host processor, a PIM controller, an NLF hardware block, etc.) may be implemented using one or more general or special-purpose computers (such as, for example, a processor, a controller and an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of responding to and executing instructions in a defined manner). The processing device may run an operating system (OS) and one or more software applications running on the OS. The processing device may also access, store, manipulate, process, and create data in response to the execution of the software. For the purpose of simplicity, the description of the processing device is singular; however, it will be understood by those of ordinary skill in the art that the processing device may include multiple processing elements and multiple types of processing elements. For example, the processing device may include multiple processors, or a single processor and a single controller. In addition, different processing configurations (such as parallel processors) are feasible.

[0093] Software may include a computer program, a piece of code, instructions, or some combination thereof, to independently or in unison command or configure a processing device to operate as desired. Software and data may be stored in any type of machine, component, physical or virtual device, or computer storage medium or device that can provide instructions or data to or be interpreted by a processing device. Software may also be distributed on networked computer systems so that the software is stored and executed in a distributed manner. Software and data may be stored by one or more non-transitory computer-readable recording media.

[0094] The method according to the above-described embodiment may be recorded in a non-transitory computer-readable medium including program instructions for implementing various operations of the above-described embodiment. The medium may also include data files, data structures, etc., alone or in combination with program instructions. The program instructions recorded on the medium may be program instructions specially designed and constructed for the purpose of the embodiment, or they may be types known and available to technicians in the field of computer software. Examples of non-transitory computer-readable media include magnetic media (such as hard disks, floppy disks, and tapes); optical media (such as CD-ROM disks and DVDs); magneto-optical media (such as optical magnetic floppy disks); and hardware devices (such as read-only memory (ROM), RAM, flash memory, etc.) specially configured to store and execute program instructions. Examples of program instructions include machine code (such as generated by a compiler) and files containing high-level code that can be executed by a computer using an interpreter.

[0095] The aforementioned hardware devices may be configured to act as one or more software modules to perform the operations of the aforementioned embodiments, and vice versa.

[0096] As described above, although the embodiments have been described with reference to specific drawings, a person skilled in the art may apply various technical modifications and variations based on this. For example, if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or replaced or supplemented by other components or their equivalents, then suitable results may be achieved.

[0097] Accordingly, other implementations, other embodiments, and equivalents of the claims are within the scope of the appended claims.

Claims

1. An electronic device, comprising: an in-memory processing controller; as well as an in-memory processing device configured to: perform an in-memory processing operation in response to an in-memory processing request generated by the in-memory processing controller, The in-memory processing controller is configured to: execute a nonlinear function operation among a plurality of requested operations, generate an in-memory processing request for the in-memory processing operation, and send the in-memory processing request to the in-memory processing device.

2. The electronic device according to claim 1, wherein: The in-memory processing controller includes a first logic circuit configured to perform a non-linear function operation among the plurality of requested operations.

3. The electronic device according to claim 2, wherein: The in-memory processing controller also includes: a command queue configured to store commands received from a host processor in the electronic device for operations of the plurality of requests; A second logic circuit is configured to: generate an in-memory processing request for an in-memory processing operation among the plurality of requested operations; and The control register is configured to control the operation of the first logic circuit.

4. The electronic device according to claim 1, wherein: The in-memory processing controller is configured to classify each of the plurality of requested operations into one of an in-memory processing operation and a non-linear function operation according to type information of each of the plurality of requested operations.

5. The electronic device according to claim 1, further comprising: The host processor is configured to send a request for an operation of the plurality of requests to the in-memory processing controller.

6. The electronic device according to claim 5, wherein: The operations of the plurality of requests are performed by the in-memory processing controller and the in-memory processing device, and not by the host processor.

7. The electronic device according to claim 1, further comprising: The memory controller is configured to generate an in-memory processing command and send the in-memory processing command to the in-memory processing device in response to an in-memory processing request received from the in-memory processing controller.

8. The electronic device according to claim 1, wherein: The in-memory processing controller is set in a memory controller or a direct memory access device, the memory controller is included in the electronic device and is configured to manage data input to or output from the in-memory processing device, and the direct memory access device is included in the electronic device and is configured to access data stored in the in-memory processing device.

9. The electronic device according to claim 1, wherein: The in-memory processing controller is configured to send results of the plurality of requested operations to the host processor in response to all of the plurality of requested operations being performed.

10. The electronic device according to any one of claims 1 to 9, wherein: The in-memory processing apparatus includes an operator configured to perform an in-memory processing operation in response to an in-memory processing request.

11. The electronic device according to claim 10, wherein: The in-memory processing device also includes a data storage space configured to store the results of the in-memory processing operation and the results of the non-linear function operation.

12. A method for operating an electronic device, the method comprising: classifying, by an in-memory processing controller in the electronic device, a target operation to be processed based on an order of requests for a plurality of operations into one of an in-memory processing operation and a non-linear function operation; In response to classifying the target operation as a non-linear function operation, executing, by the in-memory processing controller, the target operation corresponding to the non-linear function operation; In response to classifying the target operation as an in-memory processing operation, generating, by the in-memory processing controller, an in-memory processing request for the target operation and sending the in-memory processing request to an in-memory processing device in the electronic device; as well as The in-memory processing device executes a target operation corresponding to the in-memory processing operation according to the in-memory processing request received from the in-memory processing controller.

13. The method according to claim 12, wherein: The step of performing a target operation corresponding to the non-linear function operation is performed by a first logic circuit included in the in-memory processing controller and configured to perform the non-linear function operation among the plurality of operations.

14. The method according to claim 12, further comprising: Commands for the plurality of operations received from a host processor in the electronic device are stored in a command queue included in the in-memory processing controller.

15. The method according to claim 12, wherein: The step of classifying the target operation includes classifying the target operation into one of an in-memory processing operation and a non-linear function operation according to type information of the target operation.

16. The method according to claim 12, further comprising: Requests for the plurality of operations are sent by a host processor in the electronic device to the in-memory processing controller.

17. A non-transitory computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 12 to 16.

18. An electronic device comprising: an in-memory processing device that performs an in-memory processing operation in response to an in-memory processing request; as well as A direct memory access device comprising an in-memory processing controller, The in-memory processing controller is configured to: execute a nonlinear function operation among multiple operations requested by a host processor, generate an in-memory processing request for the in-memory processing operation among the multiple operations, and send the in-memory processing request to the in-memory processing device.

19. The electronic device according to claim 18, further comprising a memory controller, wherein: The in-memory processing controller sends the in-memory processing request to the memory controller, and the memory controller forwards the in-memory processing request to the in-memory processing device.

20. The electronic device according to claim 19, further comprising: interconnects, connecting to the host processor, direct memory access device, and memory controller, wherein the direct memory access device receives commands for the plurality of operations from the host processor via the interconnect, and The memory controller receives a request for processing within the memory through an interconnect.