Method and electronic device using generative model

By executing input tensor cache in the transformer layer of the generative model, the problem of high computational complexity and resource requirements in real-time inference is solved, and memory usage and performance optimization is achieved.

CN120354894APending Publication Date: 2025-07-22SAMSUNG ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510074157.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-05
Filing Date
2025-01-17
Publication Date
2025-07-22

Smart Images

  • Figure CN120354894A_ABST
    Figure CN120354894A_ABST
Patent Text Reader

Abstract

A method and an electronic device using a generative model are provided. The method comprises: executing, by one or more first processors in a first decoding stage, one or more transformer layers to generate a first output lexical unit, the first output lexical unit based on a first input sequence comprising a first input lexical unit; and executing, by the one or more first processors in a second decoding stage, the one or more transformer layers to generate a second output lexical unit, the second output lexical unit based on the second input sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on and claims priority to Korean Patent Application No. 10-2024-0008854, filed with the Korean Intellectual Property Office on January 19, 2024, and Korean Patent Application No. 10-2024-0046554, filed with the Korean Intellectual Property Office on April 5, 2024, the entire contents of which are hereby incorporated by reference in their entirety. Technical Field

[0002] Embodiments of the present disclosure relate to a method and an electronic device using a generative model. Background Art

[0003] Generative models, especially those leveraging machine learning and artificial intelligence, have revolutionized various fields including natural language processing, image generation, and predictive analytics. These models can learn complex patterns from data and generate new reliable data points that follow the learned distribution.

[0004] In some cases, generative models can be used to automate processing techniques. For example, an artificial intelligence model (e.g., a neural network model) can be implemented to provide a connection between input patterns and output patterns after extensive training. However, despite the capabilities of these models, applying generative models to real-time inference and decision-making remains a challenging task due to computational complexity and resource requirements. Therefore, there is a need in the art for methods that can perform inference of generative models while utilizing reduced computational resources. Summary of the Invention

[0005] The present disclosure describes systems and methods for performing inference of a generative model. Embodiments of the present disclosure include a generative model, the generative model including a transformer layer configured to perform an attention mechanism. In some cases, the transformer layer includes a self-attention sublayer and a multi-layer perceptron (MLP) sublayer. Embodiments include an activation sequence caching technique applicable to the self-attention sublayer. In some cases, the activation sequence caching technique can be used to cache activation sequences generated during the execution of generative inference, thereby enhancing the performance of processing and reducing memory usage.

[0006] According to one aspect, a method performed by one or more first processors using a generative model including one or more transformer layers is provided. The method includes: generating, by the one or more first processors, a first output lexical unit by performing the one or more transformer layers in a first decoding stage, the first output lexical unit being based on a first input sequence including a first input lexical unit; and generating, by the one or more first processors, a second output lexical unit by performing the one or more transformer layers in a second decoding stage, the second output lexical unit being based on a second input sequence, where the second input sequence includes the first input sequence and a second input tensor corresponding to the first output lexical unit.

[0007] According to another aspect, an electronic device is provided. The electronic device includes: a first memory configured to store parameters of a generative model including one or more transformer layers; and one or more first processors configured to generate a first output lexical unit by performing the one or more transformer layers based on a first input sequence including a first input lexical unit in a first decoding stage, and generate a second output lexical unit by performing the one or more transformer layers based on a second input sequence including a second input lexical unit corresponding to the first output lexical unit in a second decoding stage, where the second input sequence includes the first input sequence and a second input tensor corresponding to the first output lexical unit.

[0008] According to one aspect, a method performed by one or more first processors using a generative model including one or more transformer layers is provided. The method includes: obtaining a first input sequence including a first input lexical unit; generating, by the one or more first processors, a first output lexical unit by performing a generative model including one or more transformer layers; caching the first input sequence and the first output lexical unit in a memory corresponding to one or more second processors other than the one or more first processors; obtaining a second input sequence by loading the first input sequence and the first output lexical unit from the memory corresponding to the one or more second processors; and generating, by the one or more first processors, a second output lexical unit by performing the generative model, the second output lexical unit being based on the second input sequence.

[0009] Additional aspects of the embodiments will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] These and / or other aspects, features, and advantages of the invention will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings.

[0011] Figure 1 FIG. 1 is a diagram illustrating an example of an inference operation of a generative model according to an embodiment.

[0012] Figure 2 FIG. 2 is a diagram illustrating an example of a data cache operation according to an embodiment.

[0013] Figure 3 FIG. 3 is a diagram illustrating a data flow between transformer layers using an input tensor cache according to an embodiment.

[0014] Figure 4 FIG. 4 is a diagram illustrating a configuration of a transformer layer according to an embodiment.

[0015] Figure 5 FIG. 5 is a flowchart illustrating an example of an inference method using a generative model according to an embodiment.

[0016] Figure 6 FIG. 6 is a flowchart illustrating an example of generating an output token using a generative model according to an embodiment. DETAILED DESCRIPTION

[0017] The present disclosure describes systems and methods for performing inference of a generative model. Embodiments of the present disclosure include a generative model, the generative model including a transformer layer configured to perform an attention mechanism. In some cases, the transformer layer includes a self-attention sublayer and a multi-layer perceptron (MLP) sublayer. Embodiments include an activation sequence cache technique applicable to the self-attention sublayer. In some cases, the activation sequence cache technique can be used to cache an activation sequence generated during generative inference, thereby enhancing the performance of processing and reducing memory usage.

[0018] Existing techniques for caching (such as key-value cache techniques) can use a large amount of memory. In addition, the existing techniques result in high communication overhead between processors during generative inference processing. For example, the communication overhead is caused by transferring a large-capacity cache between processors. As a result, in the case of communication overhead, performance bottlenecks may occur due to low bandwidth.

[0019] The present disclosure describes systems and methods for a generative model. Embodiments of the present disclosure include an activation cache technique. In some cases, an activation cache technique (such as an input tensor cache technique) can be obtained for the attention mechanism. In some cases, when performing the attention mechanism in each transformer layer of the generative model, the activation cache technique can be used to cache an input tensor (e.g., not caching a key-value tensor).

[0020] In some cases, a trainable generative model is used to map an input tensor and an output tensor. The training ability of the model to generate such a mapping can represent the learning ability of an artificial intelligence model. Additionally, in some cases, a trained artificial intelligence model can have a generalization ability, such as generating a relatively accurate output for an untrained input pattern. According to an embodiment, the generative artificial intelligence model can perform advanced inference using the attention mechanism of a transformer included in the generative model.

[0021] The present disclosure describes an activation sequence caching method for enhancing the performance of generative inference while reducing memory usage. In some cases, memory usage can be performed by caching the activation sequences generated during the execution of generative inference. Embodiments of the present disclosure include activation sequence caching techniques for reducing cache-related memory usage based on caching activation sequences. In some cases, the activation sequence caching techniques reduce cache-related processor communication. Additionally, the activation sequence caching techniques enhance generative inference performance by minimizing the amount of computation using an optimized computational schedule.

[0022] Embodiments of the present disclosure are configured to perform inference of a generative neural network model. In some cases, the generative model includes a transformer layer configured to generate a first output lexical unit by performing one or more transformer layers based on a first input sequence including a first input lexical unit during a first decoding stage. Additionally, one or more processors can be configured to generate a second output lexical unit by performing one or more transformer layers based on a second input sequence including a second input tensor corresponding to the first output lexical unit during a second decoding stage. As described, the second input sequence can include the first input sequence and the second input tensor corresponding to the first output lexical unit.

[0023] Embodiments of the present disclosure are configured to obtain a first input sequence including a first input lexical unit. In some cases, one or more first processors executing the generative model can be used to generate the first output lexical unit. In some cases, the generative model can include one or more transformer layers. According to an embodiment, the first input sequence and the first output lexical unit can be cached in a memory corresponding to one or more second processors other than the one or more first processors. In some cases, the second input sequence is obtained by loading the first input sequence and the first output lexical unit from the memory corresponding to the one or more second processors. One or more first processors executing the generative model can generate the second output lexical unit based on the second input sequence.

[0024] The following detailed structural or functional descriptions are provided only as examples, and various changes and modifications can be made to the examples. Therefore, the embodiments are not to be construed as limited to the disclosure, and should be understood to include all changes, equivalents, and substitutions within the spirit and scope of the disclosure.

[0025] Although terms such as first, second, etc. are used to describe various components, the components are not limited to these terms. These terms are only used to distinguish one component from another. For example, the first component may be referred to as the second component, or similarly, the second component may be referred to as the first component.

[0026] It should be noted that if a component is described as "connected", "coupled", or "joined" to another component, then a third component may be "connected", "coupled", or "joined" between the first component and the second component, although the first component may be directly connected, coupled, or joined to the second component.

[0027] Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. It will also be understood that when used herein, the terms "comprises" and / or "comprising" indicate the presence of stated features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups.

[0028] As used herein, "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in the corresponding one of the phrases, or all possible combinations thereof.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Unless explicitly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with the meaning in the context of the relevant art, and shall not be interpreted in an idealized or overly formal sense.

[0030] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. When describing the embodiments with reference to the accompanying drawings, like reference numerals denote like elements, and the repetitive descriptions related thereto will be omitted.

[0031] Figure 1 An example of an inference operation of a generative model according to an embodiment is shown. In some cases, the generative model includes a machine learning model.

[0032] Machine learning parameters (also known as model parameters or weights) are variables that provide the behavior and characteristics of a machine learning model. Machine learning parameters can be learned or estimated from training data and are used to make predictions or perform tasks based on learned patterns and relationships in the data. Machine learning parameters are typically adjusted during the training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find the optimal values of the parameters that allow the machine learning model to make accurate predictions or perform well on a given task.

[0033] For example, during the training process, an algorithm adjusts the machine learning parameters according to an optimization technique (such as gradient descent, stochastic gradient descent, or other optimization algorithms) to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the machine learning parameters are used to make predictions on new, unseen data.

[0034] An artificial neural network (ANN) has several parameters including weights and biases, which are associated with each neuron in the network that "controls the degree of connection between neurons and affects the ability of the neural network to capture complex patterns in the data". An ANN is a hardware component or software component that includes multiple connected nodes (i.e., artificial neurons) that loosely correspond to neurons in the human brain. Each connection or edge sends a signal from one node to another (such as a physical synapse in the brain). When a node receives a signal, the node processes the signal and then sends the processed signal to other connected nodes.

[0035] In some cases, the signals between nodes include real numbers, and the output of each node is calculated by a function of the sum of the inputs to each node. In some examples, a node can use other mathematical algorithms (such as selecting the maximum value from the inputs as the output or any other suitable algorithm for activating the node) to determine its output. Each node and edge is associated with one or more node weights that determine how the signals are processed and sent.

[0036] In an ANN, the hidden (or intermediate) layer includes hidden nodes and is located between the input layer and the output layer. The hidden layer performs a non-linear transformation of the inputs into the network. Each hidden layer is trained to produce a defined output that contributes to the combined output of the output layer of the ANN. A hidden representation is a machine-readable data representation of the inputs that is "learned from the hidden layers of an ANN and produced by the output layer". As the ANN's understanding of the inputs improves as the ANN is trained, the hidden representation gradually differentiates from earlier iterations.

[0037] During the training process of the ANN, the node weights are adjusted to improve the accuracy of the results (i.e., by minimizing the "loss corresponding in some way to the difference between the current result and the target result"). The weights of the edges increase or decrease the strength of the signals sent between the nodes. In some cases, a node has a threshold below which its signal is not sent at all. In some examples, the nodes are aggregated into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer, and the final layer is called the output layer. In some cases, the signal passes through a particular layer multiple times.

[0038] Referring Figure 1 , the generative model 100 may include multiple stages (i.e., a pre-filling stage 110, a first decoding stage 120, a second decoding stage 130, and a final decoding stage 140). The generative model 100 includes an input embedding layer 101, a transformer layer 102, and an output embedding layer 103 in each of the multiple stages. The generative model 100 may receive an input prompt 111. In some cases, the generative model 100 may generate lexical units 113, 123, 133 to 141, and 143 corresponding to the pre-filling stage 110, the first decoding stage 120, and the second decoding stage 130 to the final decoding stage 140, respectively. An inference result corresponding to the input prompt 111 may be generated based on the lexical unit 143 of the final decoding stage 140.

[0039] The input prompt 111 may be data in various formats (such as text, sound, image, video, etc.). For example, the input prompt 111 may be a user's query or request, and the inference result may be a response to the query or request.

[0040] The input embedding layer 101 may perform embedding and / or positional encoding on the input data. The input data may be the input prompt 111 for the pre-filling stage 110 or the lexical units 113, 123, 133, and 141 for the first decoding stage 120 to the final decoding stage 140. Thus, the input data for the pre-filling stage 110 may be the input prompt 111, and the input data for the first decoding stage 120 to the final decoding stage 140 may be the lexical units 113, 123, 133, and 141. The format of the input data may be converted to a format processable in the transformer layer 102 according to the embedding of the input data (i.e., the embedding of the input data generated using the input embedding layer 101). The temporal relationship and / or spatial relationship between the elements of the input data (e.g., words of text or tiles of an image) may be defined by the positional encoding of the input data. For example, the input embedding layer 101 may add positional information to the input data (e.g., the input prompt 111).

[0041] The transformer layer 102 can execute an attention mechanism. In the field of machine learning, the attention mechanism is a method of assigning different levels of importance to different elements of the input. Some sequence models process the input sequence sequentially, maintaining an internal hidden state for capturing information from previous steps. However, in some cases, this sequential processing makes it difficult to capture long-range dependencies or focus on specific parts of the input sequence.

[0042] The attention mechanism addresses these difficulties by enabling the ANN to selectively focus on different parts of the input sequence, assigning different degrees of importance or attention to each part. The attention mechanism achieves selective focusing by considering the relevance of each input element relative to the current state of the ANN.

[0043] In some cases, an ANN employing an attention mechanism receives an input sequence and maintains the current state of the representation understanding or context of the input sequence. For each element in the input sequence, the attention mechanism calculates an attention score, which indicates the importance or relevance of that element given the current state. The attention scores are transformed into attention weights through a normalization process (e.g., applying the softmax function). The attention weights represent the contribution of each input element to the overall attention. The attention weights are used to calculate the weighted sum of the input elements, resulting in a context vector. The context vector represents the attended information or the part of the input sequence that the ANN considers most relevant for the current step. The context vector is combined with the current state of the ANN to provide additional information and influence the subsequent predictions or decisions of the ANN.

[0044] In some cases, by incorporating the attention mechanism, the ANN dynamically allocates attention to different parts of the input sequence, allowing the ANN to focus on relevant information and capture dependencies across longer distances.

[0045] In some cases, calculating attention involves three basic steps. First, the similarity between a query vector Q and a key vector (or simply key) K obtained from the input is calculated to generate attention weights. In some cases, the similarity functions used for this process include dot product, splice, detector, etc. Next, the softmax function is used to normalize the attention weights. Finally, the attention weights are weighted together with the corresponding values V of the attention weights. In the context of an attention network, the key K and the value V are typically vectors or matrices used to represent the input data. The key K is used to determine which parts of the input the attention mechanism should focus on, while the value V is used to represent the actual data being processed.

[0046] In some cases, the attention mechanism may represent a self-attention mechanism and / or a cross-attention mechanism. The self-attention mechanism enables the network to selectively weight input elements (e.g., based on their relevance to other elements) to emphasize important features during computation. The self-attention mechanism incorporates dynamic attention scores to optimize information processing. Additionally, the cross-attention mechanism facilitates effective interaction between different input sequences in a neural network architecture by dynamically allocating attention scores based on the relevance of different input sequences. The cross-attention mechanism enhances model performance by providing the network with the ability to focus on key features from one sequence while processing another sequence, for more subtle and context-aware information processing.

[0047] Referring again to Figure 1 , specifically, the input tensor can be input into each transformer layer 102. The transformer layer 102 can generate an output tensor by performing an attention mechanism (such as the attention mechanism described herein) on each input tensor. The transformer layer 102 can include a first transformer layer and a second transformer layer. In some cases, when the second transformer layer in the transformer layer 102 follows the first transformer layer, the output tensor of the first transformer layer can be the input tensor of the second transformer layer.

[0048] As described herein, each transformer layer 102 can include query weights, key weights, and value weights. Each transformer layer 102 can perform the attention mechanism by applying the query weights, key weights, and value weights to the input tensor. An attention result can be generated based on the attention mechanism, and an output tensor can be generated based on the attention result. The number of layers of the transformer layer 102 is not limited to any number. The index j can be used to identify the transformer layer from multiple transformer layers 102 (i.e., N transformer layers 102).

[0049] The output embedding layer 103 can convert the output tensor obtained from the transformer layer 102 into lexical units 113, 123, 133 to 143. The output embedding layer 103 can be considered the inverse of the input embedding layer 101 (e.g., can operate in the reverse of the input embedding layer 101). For example, the output embedding layer 103 can perform the inverse conversion of the format conversion of the input embedding layer 101 and / or remove the position information.

[0050] As referred to Figure 1As shown, the inference operation can be performed through the prefill stage 110, the first decoding stage 120, and the second decoding stage 130 to the final decoding stage 140 of the generative model 100. The input prompt 111 can include a plurality of lexical units. In the case of the prefill stage 110, the generative model 100 can analyze the correlation between the lexical units of the input prompt 111 based on the attention mechanism (e.g., using the transformer layer 102). In the first decoding stage 120 and the second decoding stage 130 to the final decoding stage 140, the generative model 100 can generate lexical units 123, 133 to 143 based on the attention mechanism (e.g., using the transformer layer 102 corresponding to each of the first decoding stage 120 and the second decoding stage 130 to the final decoding stage 140). The number of decoding stages can be unrestricted. The index i can be used to identify any one of the decoding stages among the plurality of decoding stages 120 to 140 (i.e., among the M decoding stages). For example, referring to Figure 1 , the index i of the final decoding stage 140 can be M.

[0051] According to an embodiment, when the transformer layer 102 is executed, the input tensor cache (i.e., instead of the key-value (KV) cache) can be executed. As the decoding stage progresses from the first decoding stage 120 to the final decoding stage 140, a repetitive operation can occur when the key tensor and the value tensor are obtained. The KV cache can be a technique for caching the key tensor sequence and the value tensor sequence to prevent repetitive operations. The sequence can represent a data format in which tensors are connected to each other (e.g., concatenation). For example, a sequence can be formed by adding a second tensor to a first tensor. The key tensor can form a key tensor sequence, and the value tensor can form a value tensor sequence.

[0052] Caching can represent the operation of storing a sequence. When the size of the generative model 100 is large, the memory of the processor that executes the generative model 100 can be used to store the generative model 100, and the sequence can be cached in an additional storage space (i.e., instead of the memory of the processor that executes the generative model 100). For example, the inference operation using the generative model 100 can be performed by an auxiliary processor (e.g., a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), an accelerator, etc.), and the sequence can be cached in the memory (e.g., CPU memory) or storage device (e.g., disk) of the main processor (e.g., a central processing unit (CPU)), rather than in the memory of the auxiliary processor (e.g., GPU memory, NPU memory, TPU memory, accelerator memory, etc.).

[0053] In some cases, the KV cache is effective in reducing the amount of computation, but additional storage space may be required to store the sequences. Additionally, when sequences are moved between the main processor area and the auxiliary processor area, the KV cache may use additional communication. As the sizes of the key tensor sequence and the value tensor sequence increase, the resource consumption due to the additional storage space and additional communication may increase.

[0054] Table 1 may show the attention operation of the transformer layer using the KV cache.

[0055] [Table 1]

[0056]

[0057] As shown in Table 1, represents the key tensor sequence, represents the value tensor sequence, represents the input sequence, represents the key weight, and represents the value weight. (Since the key tensor, value tensor, and input tensor are initially generated in the pre-fill stage) The key tensor, value tensor, and input tensor may be the same as the key tensor sequence, value tensor sequence, and input sequence in the pre-fill stage. Referring to Table 1, the "..." in the pre-fill stage may indicate the corresponding operation in the decoding stage of Table 1. For example, may be determined by the corresponding operation in the decoding stage. represents the output tensor sequence. The transformer layer may include an attention layer. The input tensor may be the input of the attention layer, and the output tensor may be the output of the attention layer. As described herein, in the pre-fill stage, the output tensor may be the same as the output tensor sequence.

[0058] Referring to Table 1, represents the query tensor, represents the key tensor, represents the value tensor, represents the input tensor of the i-th decoding stage, represents the query weight, represents the key weight, and represents the value weight. Each of the key tensor and the value tensor is generated in a single decoding stage. In some cases, the key tensor sequence and the value tensor sequence may be combined (e.g., cascaded) through multiple decoding stages. The query weight, key weight, value weight, etc. can be distinguished by the transformer layer. However, for ease of description and understanding, the layer index j is omitted in Table 1. Thus, Table 1 may show the description for one layer. corresponds to the product of and , corresponds to corresponds to the product of and, and is associated with and 's product correspondence.

[0059] Referring back to Table 1 again, can represent a combined function (e.g., a cascading function). For example, using to and are combined to generate . In some examples, can be used to combine and , and can be generated. can represent a sequence of key tensors in the (i - 1)-th decoding stage (e.g., the previous decoding stage), and can represent a sequence of value tensors in the (i - 1)-th decoding stage.

[0060] During the inference process, the KV cache for and can be executed. For example, when and are determined in the prefill stage, the KV cache for and can be executed. When i = 0, the prefill stage can correspond to the decoding stage. The and cached in the prefill stage can be loaded as and in the first decoding stage. In the first decoding stage, and can be used to update and , and the KV cache for the updated and can be executed.

[0061] In Table 1, can correspond to the result obtained by multiplying by the transpose result of . can be called the query-key tensor. A non-linear operation on can be executed. can be divided by the square root of , and then a non-linearization operation (e.g., softmax) is performed. can represent the size of the hidden dimension of the attention layer. can represent Nonlinear operation Can correspond to the result obtained by multiplying by Obtained result corresponds to Can be referred to as the attention result

[0062] In Table 1, it can be determined based on To determine For example, by adding the product of And And To determine Can be the output adjustment weight Can represent the output tensor. The query weight, key weight, value weight, etc. can be distinguished by the transformer layer. However, for the sake of simplicity in description and easy understanding, the layer index j is omitted in Table 1. Therefore, Table 1 can show the description for one layer

[0063] According to the embodiment, the input tensor cache can represent a technique for caching the input sequence instead of the key tensor sequence and the value tensor sequence. Referring back to Figure 1 , the caching operation of the input sequences 112, 122, 132 to 142 (or referred to as the input tensor sequences) can be performed. For example, the input sequence 112 generated by the transformer layer 102 in the pre-fill stage 110 can be cached in the memory or cache level of the main processor, and then loaded into the auxiliary processor or the auxiliary processor memory when needed by the transformer layer 102 in the first decoding stage 120. The transformer layer 102 can determine the input sequence 122 by adding the input tensor based on the lexical unit 113 to the input sequence 112, use the input sequence 122 to generate the output tensor, and cache the input sequence 122. Similarly, the transformer layer 102 can determine the input sequence 132 by adding the input tensor based on the lexical unit 123 to the input sequence 122, use the input sequence 132 to generate the output tensor, and cache the input sequence 132

[0064] Table 2 shows the attention operation of the transformer layer using the input tensor cache

[0065] [Table 2]

[0066] In the case of the input tensor cache, the input sequence can be cached instead of the key tensor sequence and the value tensor sequence. Referring to Table 2, Can represent the input sequence in the i-th decoding stage. Can be used to combine To combine And To determine ​。 may represent the input sequence in the (i-1)-th decoding stage (e.g., the previous decoding stage), and may represent the input tensor in the i-th decoding stage (e.g., the current decoding stage). Due to the difference, based on and (i.e., instead of and ), and can be determined.

[0067] In some cases, the input tensor cache may use a larger amount of computation than the KV cache. In some cases, from the viewpoints of memory usage and communication volume, it may be beneficial to use the input tensor cache. For example, the memory usage and communication volume may be reduced by 50%. However, the embodiments are not limited thereto, and the reduction in memory usage and communication volume may be greater than or less than 50%. In some cases, when the operation order of the input tensor cache is adjusted, compared with the KV cache, the increase in the amount of computation can be significantly reduced. For example, the amount of computation can be reduced by 98.8%. However, the embodiments are not limited thereto, and the reduction in computation can vary.

[0068] For example, the operation order of the input tensor cache can be adjusted. In some examples, the operation order can be performed as in Equation 2 and Equation 3 instead of Equation 1 and Equation 3.

[0069] [Equation 1] [Equation 2]

[0070] Referring to Table 2, can be obtained by multiplying by the transposed result of (as shown in Equation 1). When the operation order of Equation 1 is adjusted to Equation 2, the amount of computation can be significantly reduced. Referring to Equation 2, the first intermediate computation result can be determined by multiplying by , the second intermediate computation result can be determined by multiplying the first intermediate computation result by the transposed result of , and can be determined by multiplying the second intermediate computation result by the transposed result of .

[0071] [Equation 3]

[0072] [Equation 4]

[0073] Referring again to Table 2, it can be obtained by multiplying by as shown in Equation 3. When the operation order of Equation 3 is adjusted as in Equation 4, the computational amount can be significantly reduced. According to Equation 4, the third intermediate calculation result can be determined by multiplying by and the can be determined by multiplying the third intermediate calculation result by .

[0074] Figure 2 An example of data cache operation according to an embodiment is shown. Referring to Figure 2 , the electronic device 200 may include a first processing area 210 and a second processing area 220. The first processing area 210 may include one or more first processors 211 and a first memory 212 of the one or more first processors 211. The second processing area 220 may include one or more second processors 221, a second memory 222 of the one or more second processors 221, and a storage device 223.

[0075] A processor is an intelligent hardware device (such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof).

[0076] In some cases, the processor is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in the memory to perform various functions. In some aspects, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0077] The memory includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and a hard disk drive. In some examples, the memory is used to store computer-readable and computer-executable software including instructions that "when executed cause at least one processor of the processor to perform the various functions described herein".

[0078] In some cases, the memory includes a basic input / output system (BIOS) that controls basic hardware or software operations, such as interactions with peripheral components or devices. In some cases, the memory includes a memory controller for operating the memory cells of the memory. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within the memory store information in the form of logical states.

[0079] According to an example embodiment, one or more first processors 211 may include a GPU, an NPU, a TPU, an accelerator, or a combination thereof, however, the embodiments are not limited thereto. One or more first processors 211 may execute a generative model. For example, one or more first processors 211 may execute operations related to the inference of the generative model.

[0080] The second processing region 220 may include one or more second processors 221, a second memory 222, and a storage device 223 for the one or more second processors 221. The one or more second processors 221 may be distinguishable from the one or more first processors 211. The second memory 222 and the storage device 223 may be distinguishable from the first memory 212. For example, one or more second processors 221 may include a CPU, however, the embodiments are not limited thereto.

[0081] The first processing region 210 and the second processing region 220 may be connected via an interface 230. The interface 230 may include, for example, a bus, however, the embodiments are not limited thereto. Compared with the internal data transfer within the first processing region 210 and the second processing region 220, the external data transfer between the first processing region 210 and the second processing region 220 may incur a greater cost in terms of time and energy. When the size of the generative model is large, the first memory 212 may be used to store the generative model, and the second memory 222 or the storage device 223 may be used for data caching.

[0082] For example, the generative model can be executed by one or more first processors 211. In some examples, the input sequence generated during the processing of the generative model can be cached in the second memory 222 or the storage device 223 through the interface 230. For example, the first input sequence generated in the first decoding stage and the second input sequence generated in the second decoding stage can be cached in the second memory 222 or the storage device 223. In some cases, the input sequence can be stored in the second memory 222 through the first memory 212. In some cases, the input sequence can be stored in the second memory 222 without passing through the first memory 212. In some cases, the input sequence can be stored in the storage device 223 through the first memory 212 and the second memory 222. In some cases, the input sequence can be stored in the storage device 223 without passing through the first memory 212 and / or the second memory 222.

[0083] In some cases, the input sequence can be loaded into the first processing area 210 during the decoding stage. For example, the first input sequence and the second input sequence can be loaded from the second memory 222 or the storage device 223 into one or more first processors 211 or the first memory 212. For example, the second input sequence can be generated by adding (e.g., appending) the second input tensor to the first input sequence loaded in the first memory during the second decoding stage.

[0084] The electronic device 200 can be implemented as at least a part of a mobile device (e.g., a mobile phone, a smart phone, a personal digital assistant (PDA), a netbook, a tablet computer, a laptop computer, etc.), a wearable device (e.g., a smart watch, a smart bracelet, smart glasses, etc.), a computing device (e.g., a desktop computer, a server (e.g., a cloud server or a data center server), etc.), a household appliance (e.g., a television (TV), a smart TV, a refrigerator, etc.), a security device (e.g., a door lock, etc.), or a vehicle (e.g., an autonomous vehicle, a smart vehicle, etc.).

[0085] The first memory 212 can store a generative model including a transformer layer. One or more first processors 211 can generate a first output lexical unit by executing the transformer layer using the first input sequence based on the first input lexical unit in the first decoding stage of the generative model. In addition, one or more first processors 211 can generate a second output lexical unit by using the second input sequence to execute the transformer layer. In some cases, in the second decoding stage of the generative model, the second input sequence can be based on a second input tensor, and the second input tensor is based on a second input lexical unit corresponding to the first output lexical unit. As referred to Figure 1As described above, the second input sequence can be determined by adding the second input tensor of the second decoding stage to the first input sequence. For example, the first input sequence can be retrieved during the second decoding stage to obtain the second input sequence.

[0086] The first memory 212 can be a high bandwidth memory (HBM) and / or a Compute Express Link (CXL) memory. The first memory 212 can have a hierarchical memory structure. The first memory 212 can use the hierarchical memory structure to accelerate operations. For example, the first memory 212 can provide computing functions (such as, Processing-in-Memory (PIM) or Processing Near Memory (PNM)). For example, the first memory 212 can include HBM-PIM, HBM-PNM, CXL-PIM memory, CXL-PNM, or a combination thereof. PIM or PNM can accelerate operations through internal logic.

[0087] Processing-in-Memory (PIM) is an advanced computing architecture that directly integrates processing capabilities within a memory chip. PIM alleviates the memory bottleneck problem in conventional computing systems where "data transfer between the processor and the memory can be a significant performance limitation". By embedding processing units within the memory, PIM enables data to be processed directly where it is stored, thereby reducing the latency and energy consumption associated with data movement. PIM architectures are used in data-intensive applications such as artificial intelligence, machine learning, big data analytics, and scientific simulations. PIM can perform operations such as data filtering, aggregation, and even complex computations, which can significantly accelerate overall system performance. Additionally, PIM improves the efficiency of memory usage and the scalability of high-performance computing systems by alleviating the von Neumann bottleneck.

[0088] Processing Near Memory (PNM) is a computing paradigm that involves placing processing units near the memory module rather than integrating them directly within the memory chip. Similar to PIM, PNM reduces the latency and energy costs associated with data transfer between the processor and the memory. By placing the processing units near the memory, the PNM architecture can utilize shorter data paths and higher bandwidth connections, thus enhancing computing efficiency. The PNM architecture is advantageous for applications that require high-speed data access and processing such as real-time data analytics, edge computing, and complex simulations. Since PNM provides a modular approach to enhancing computing power by adding more processing units near the memory as needed, PNM can provide a flexible and scalable solution. PNM offers a substantial improvement in performance and energy efficiency over traditional memory architectures.

[0089] According to an embodiment, the first memory 212 may use PIM or PNM to process operations of the input tensor cache. In the case of the KV cache, the output tensors of the previous transformer layer (i.e., the input tensors of the current transformer layer) may be multiplied by the key weights and value weights respectively to determine the key tensor and the value tensor. In some cases, the key tensor and the value tensor may be concatenated (e.g., cascaded) with the key tensor sequence and the value tensor sequence of the previous decoding stage (or pre-fill stage) to determine the key tensor sequence of the current decoding stage and the value tensor sequence of the current decoding stage. Further, the key tensor sequence of the current decoding stage and the value tensor sequence of the current decoding stage may be cached.

[0090] In the case of the input tensor cache, the input sequence of the previous decoding stage may be concatenated with the output tensors of the previous transformer layer to determine the input sequence of the current decoding stage. Further, the input sequence of the current decoding stage may be multiplied by the key weights and value weights respectively to determine the key tensor sequence and the value tensor sequence. By implementing the features of the input tensor cache, the concatenation operation (e.g., the cascade operation) may be offloaded in PIM or PNM, which increases the operation efficiency.

[0091] For example, the concatenation operation of the input sequence of the previous decoding stage and the output tensors of the previous transformer layer may be effectively processed based on PIM or PNM. In some cases, one or more first processors 211 may not have to process the concatenation operation by loading the input sequence of the previous decoding stage and the output tensors of the previous transformer layer into the operation space of the first processor 211. In some cases, the input sequence of the current decoding stage corresponding to the result of the concatenation operation in the first memory 212 may be loaded into one or more first processors 211. Accordingly, the bandwidth and the computational overhead may be effectively offloaded.

[0092] Embodiments of the present disclosure include instructions or operations for performing concatenation. In some cases, concatenation may be performed while storing the output tensors in the memory area of the first memory 212 that stores the input sequence of the previous decoding stage. In some cases, instructions or operations may be defined for the concatenation of the input sequence and the input tensor when loading the input sequence into the memory area of the first memory 212. For example, the instructions may be configured to include information about the input / output tensors and the input sequence.

[0093] Figure 3 Shows the data flow between transformer layers using the input tensor cache according to an embodiment. Refer to Figure 3, the (i - 1)-th decoding stage of the generative model 300 is shown. In the case of the (i - 1)-th decoding stage, the (i - 1)-th output lexical unit can be generated by executing the transformer layers 301, 302, and 303 (e.g., the (j - 1)-th transformer layer, the j-th transformer layer, and the (j + 1)-th transformer layer) using the (i - 1)-th input sequence based on the (i - 1)-th input lexical unit. Similarly, in the case of the i-th decoding stage of the generative model 300, the i-th output lexical unit can be generated by executing the transformer layers 301 to 303 using the i-th input sequence based on the i-th input lexical unit corresponding to the (i - 1)-th output lexical unit.

[0094] An input sequence can be determined for each of the transformer layers 301 to 303. The input sequence can be a set of input tensors, and the input tensor can be the input data of the transformer layers 301 to 303. The number of input sequences in each decoding stage can be the same as the number of transformer layers of the generative model 300. For example, when the number of transformer layers is N, the number of input sequences can be N in each decoding stage. In some cases, since the input tensor of the first transformer layer can be the input lexical unit, the input sequence can be determined based on the input lexical unit, and the input tensors of the remaining transformer layers can be determined based on the output tensor of the first transformer layer.

[0095] The i-th input sequence can be determined by adding the i-th input tensor of the i-th decoding stage to the (i - 1)-th input sequence. For example, the i-th input sequence of the j-th transformer layer can be determined by adding the i-th input tensor of the j-th transformer layer to the (i - 1)-th input sequence of the j-th transformer layer.

[0096] The data flow of the transformer layer 402 in the i-th decoding stage in Figure 3 will be described in more detail. The transformer layer 402 can correspond to the j-th transformer layer and can output the i-th output tensor based on the i-th input tensor. Referring to Figure 3 , can represent the (i - 1)-th input sequence of the j-th transformer layer, can represent the i-th input tensor of the j-th transformer layer, and can represent the i-th output tensor of the j-th transformer layer. can represent the query weight of the j-th transformer layer, can represent the key weight of the j-th transformer layer, and can represent the value weight of the j-th transformer layer. The (i - 1)-th input sequence of the j-th transformer layer can be represented as . In some cases, an input tensor cache can be applied to and .

[0097] More specifically, in the case of the i-th decoding stage, the generative model 300 may determine the i-th input sequence among the i-th input sequences for the j-th transformer layer. In some cases, the generative model 300 may determine the i-th input sequence by adding the i-th input tensor for the j-th transformer layer among a plurality of i-th input tensors to the (i - 1)-th input sequence for the j-th transformer layer in the (i - 1)-th input sequence.

[0098] The generative model 300 may determine a query-key tensor corresponding to a result obtained by multiplying a query tensor by a transposed result of the i-th key tensor sequence. In addition, the generative model 300 may determine an attention result corresponding to a product of the query-key tensor and the i-th value tensor sequence. In some cases, the attention result may be obtained based on multiplying a non-linear calculation result of the query-key tensor by the i-th value tensor sequence. In some cases, the query tensor may correspond to a result obtained by multiplying the j-th query weight of the j-th transformer layer by the i-th input tensor. The i-th key tensor sequence may be obtained by multiplying the j-th key weight of the j-th transformer layer by the i-th input sequence of the j-th transformer layer. The i-th value tensor sequence may be obtained by multiplying the j-th value weight of the j-th transformer layer by the i-th input sequence of the j-th transformer layer.

[0099] The generative model 300 may determine a first intermediate calculation result by multiplying the i-th input tensor of the j-th transformer layer by the j-th query weight of the j-th transformer layer. In addition, the generative model 300 may determine a second intermediate calculation result by multiplying the first intermediate calculation result by a transposed result of the j-th key weight of the j-th transformer layer. In some cases, the generative model 300 may determine the query-key tensor by multiplying the second intermediate calculation result by a transposed result of the i-th input sequence of the j-th transformer layer. The generative model 300 may determine a third intermediate calculation result by multiplying a non-linear calculation result of the query-key tensor by the i-th input sequence. In addition, the generative model 300 may determine the attention result by multiplying the third intermediate calculation result by the j-th value weight.

[0100] Figure 4 Shows a configuration of a transformer layer according to an embodiment. The transformer layer of the generative model (e.g., Figure 1 's transformer layer 102 and / or Figure 3 's transformer layers 301 to 303) may be configured as Figure 4 's transformer layer 400 in Figure 4 . Referring to Figures 1 to 3 , Figure 5 and Figure 6The operation of the described transformer layer 400 can be the operation of the attention sub-layer 420. However, embodiments are not limited thereto.

[0101] Figure 5 is a flowchart illustrating an example of an inference method using a generative model according to an embodiment of the present disclosure. One or more first processors (such as the first processor 211 described with reference to Figure 2 can perform inference on an input prompt using a generative model including a transformer layer.

[0102] With reference to Figure 5 , in operation 510, one or more first processors can generate a first output token by executing a transformer layer using a first input sequence based on a first input token in a first decoding stage of the generative model.

[0103] In operation 520, one or more first processors can generate a second output token by executing a transformer layer using a second input sequence based on a second input token corresponding to the first output token in a second decoding stage of the generative model. As described with reference to Figure 1 and Figure 3 , the second input sequence can be determined by adding a second input tensor to the first input sequence in the second decoding stage.

[0104] The first input sequence and the second input sequence can be cached in a second memory or storage device of one or more second processors. The one or more second processors can be distinct from the one or more first processors, and the second memory and storage device of the one or more second processors can be distinct from the first memory of the one or more first processors.

[0105] For example, the one or more first processors can include a GPU, an NPU, a TPU, an accelerator, or a combination thereof, and the one or more second processors can include a CPU.

[0106] The first memory can be configured to provide computing functions (such as PIM and / or PNM), and the second input sequence can be determined by adding the second input tensor to the first input sequence in the second decoding stage using PIM and / or PNM.

[0107] One or more first processors may be configured to cache a first input sequence in a second memory or storage device of one or more second processors that are different from the one or more first processors. In some cases, one or more first processors may be configured to load the first input sequence from the second memory or storage device into the one or more first processors or a first memory of the one or more first processors. One or more first processors may be configured to add a second input tensor to the first input sequence loaded in the one or more first processors or the first memory of the one or more first processors to determine a second input sequence.

[0108] In addition, to generate a second output token, one or more first processors may be configured to determine a second input sequence for a first transformer layer in the second input sequence by adding a second input tensor of the first transformer layer among the second input tensors to a first input sequence of the first transformer layer among the first input sequences.

[0109] In some cases, to generate a second output token, one or more first processors may be configured to determine a query-key tensor corresponding to a product of a query tensor and a transposed result of a first key tensor sequence. In addition, one or more first processors may be configured to determine an attention result corresponding to a product of a non-linear calculation result of the query-key tensor and a first value tensor sequence. The query tensor may correspond to a product of a first query weight of the first transformer layer and the second input tensor. The first key tensor sequence may correspond to a product of a first key weight of the first transformer layer and the second input sequence. The first value tensor sequence may correspond to a product of a first value weight of the first transformer layer and the second input sequence.

[0110] One or more first processors may be configured to determine a first intermediate calculation result by multiplying the second input tensor by the first query weight. In some cases, one or more first processors may be configured to determine a second intermediate calculation result by multiplying the first intermediate calculation result by a transposed result of the first key weight. In addition, one or more first processors may be configured to determine the query-key tensor by multiplying the second intermediate calculation result by a transposed result of the second input sequence.

[0111] One or more first processors may be configured to determine a third intermediate calculation result by multiplying a non-linear calculation result of the query-key tensor by the second input sequence. In some cases, one or more first processors may be configured to determine the attention result by multiplying the third intermediate calculation result by the first value weight.

[0112] One or more first processors may be configured to generate a first input token based on an input prompt during a prefill phase of a generative model. In some cases, one or more first processors may be configured to generate an inference result corresponding to the input prompt based on a second output token when a final decoding phase of the generative model terminates.

[0113] Figure 6 is a flowchart illustrating an example of an inference method using a generative model according to an embodiment of the present disclosure. One or more first processors (such as the first processor 211 described with reference to Figure 2 may generate a second output token based on an input sequence. Further details regarding each of operations 610 to 650 have been provided with reference to Figures 1 to 4

[0114] In operation 610, the system obtains a first input sequence. In some cases, the first input sequence includes a first input token. In some cases, the operation of this step may be performed by one or more processors executing a generative model including one or more transformer layers. In some cases, the input sequence may be determined based on an input tensor of a transformer layer and a previous input sequence of the transformer layer. In some examples, the first input sequence may be determined by adding a first input tensor of a transformer layer in a first decoding phase to an input sequence of the transformer layer generated during the prefill phase. For example, the input sequence during the prefill phase may be an input prompt.

[0115] In operation 620, the system is configured to generate a first output token. In some cases, the operation of this step is performed by one or more first processors executing the generative model. In some cases, one or more first processors may generate a first output token by executing a transformer layer using a first input sequence based on a first input token in a first decoding phase of the generative model.

[0116] In operation 630, the system is configured to cache the first input sequence and the first output token. In some cases, the caching is performed in a memory corresponding to one or more second processors other than the one or more first processors. In some cases, the input sequence and the output token may be cached in additional storage space (i.e., other than the memory of the processors executing the generative model).

[0117] In operation 640, the system is configured to obtain a second input sequence. In some cases, the second input sequence is obtained based on loading the first input sequence and the first output token from a memory corresponding to one or more second processors.

[0118] ​In some cases, the input sequence can be determined based on the input tensor of the transformer layer and the input sequence of the transformer layer from a previous decoding stage (i.e., the previous input sequence). In some examples, the second input sequence can be determined by adding the second input tensor of the transformer layer in the second decoding stage to the input sequence of the transformer layer generated in the first decoding stage (such as the input sequence in the first decoding stage obtained in operation 610).

[0119] In operation 650, the system is configured to generate a second output lexical unit. In some cases, the operations of this step are performed by one or more first processors that execute a generative model. In some cases, the second output lexical unit is generated based on the second input sequence. In some cases, one or more first processors can generate the second output lexical unit by executing a transformer layer using the second input sequence based on the second input lexical unit in the second decoding stage of the generative model.

[0120] The embodiments described herein can be implemented using hardware components, software components, and / or combinations thereof. The processing device can be implemented using one or more general-purpose or special-purpose computers (such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field-programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of responding and executing instructions in a defined manner). The processing device can run an operating system (OS) and one or more software applications running on the OS. The processing device can also access, store, manipulate, process, and generate data in response to the execution of the software. For simplicity, the description of the processing device is used in the singular; however, those skilled in the art will understand that the processing device can include multiple processing elements and / or multiple types of processing elements. For example, the processing device can include multiple processors, or a single processor and a single controller. Additionally, different processing configurations (such as parallel processors) are feasible.

[0121] The software can include a computer program, fragments of code, instructions, or some combination thereof to independently or uniformly direct or configure the processing device to operate as needed. The software and / or data can be stored in any type of machine, component, physical or virtual device, computer storage medium, or device that can provide instructions or data to be interpreted or being interpreted by the processing device. The software can also be distributed over a networked computer system such that the software is stored and executed in a distributed manner. The software and data can be stored by one or more non-transitory computer-readable recording media.

[0122] The method according to the above examples can be recorded in a non-transitory computer-readable medium including program instructions to implement the various operations of the above examples. The medium may also separately include program instructions, data files, data structures, etc., or in combination with program instructions, data files, data structures, etc. The program instructions recorded on the medium can be program instructions specifically designed and constructed for the purpose of the examples, or they can be of the kinds well-known and available to those skilled in the computer software art. Examples of non-transitory computer-readable media include magnetic media (such as hard disks, floppy disks, and magnetic tapes); optical media (such as CD-ROM disks, DVDs, and / or Blu-ray disks); magneto-optical media (such as optical disks); and hardware devices specifically configured to store and execute program instructions (such as read-only memory (ROM), random access memory (RAM), flash memory (e.g., USB flash drives, memory cards, memory sticks, etc.)). Examples of program instructions include both machine code (such as generated by a compiler) and files containing higher-level code that can be executed by a computer using an interpreter).

[0123] The above hardware devices can be configured to act as one or more software modules to perform the operations of the above embodiments, and vice versa.

[0124] As described above, although the embodiments have been described with reference to a limited number of drawings, those skilled in the art can apply various technical modifications and variations based on the embodiments. For example, if the described techniques are performed in a different order, and / or if the components in the described system, architecture, device, or circuit are combined in a different manner, or replaced or supplemented by other components or their equivalents, appropriate results can be achieved.

[0125] Therefore, other embodiments, other examples, and equivalents of the claims are within the scope of the appended claims.

Claims

1. A method performed by one or more first processors using a generative model including one or more transformer layers, the method comprising: Executing the one or more transformer layers by the one or more first processors in a first decoding stage based on a first input sequence including first input lexical units to generate first output lexical units; And Executing the one or more transformer layers by the one or more first processors in a second decoding stage based on a second input sequence to generate second output lexical units, Wherein the second input sequence includes the first input sequence and a second input tensor corresponding to the first output lexical unit.

2. The method according to claim 1, further comprising: Caching the first input sequence during the first decoding stage; And Retrieving the cached first input sequence during the second decoding stage to obtain the second input sequence.

3. The method according to claim 2, wherein The one or more first processors include a graphics processor, a neural processor, a tensor processor, an accelerator, or a combination thereof, and The first input sequence is cached in a memory corresponding to one or more second processors including a central processing unit.

4. The method according to claim 2, wherein In-memory processing, near-memory processing, or a combination thereof is used to generate the second input sequence.

5. The method according to claim 1, further comprising: Caching the first input sequence in a memory of a second processor different from the one or more first processors; Loading the first input sequence into a first memory of the one or more first processors; And Generating the second input sequence based on the first input sequence and the second input tensor.

6. The method according to claim 1, wherein, The step of generating the second output lexical units includes appending the second input tensor to the first input sequence.

7. The method according to claim 1, wherein The step of generating the second output lexical units includes: Determining a query tensor corresponding to the product of the second input tensor and a first query weight; Determining a sequence of key tensors corresponding to the product of a first key weight and the second input sequence; Determining a sequence of value tensors corresponding to the product of a first value weight and the second input sequence; Determining a query-key tensor corresponding to the product of the query tensor and the transpose of the first sequence of key tensors; and Determining an attention result corresponding to the product of the non-linear calculation result of the query-key tensor and the sequence of value tensors.

8. The method according to claim 7, wherein, The step of determining the query-key tensor includes: Determining a first intermediate calculation result by multiplying the second input tensor by the first query weight; Determining a second intermediate calculation result by multiplying the first intermediate calculation result by the transpose of the first key weight; and Determining the query-key tensor by multiplying the second intermediate calculation result by the transpose of the second input sequence.

9. The method according to claim 7, wherein, The step of determining the attention result includes: Determining a third intermediate calculation result by multiplying the non-linear calculation result of the query-key tensor by the second input sequence; and Determining the attention result by multiplying the third intermediate calculation result by the first value weight.

10. The method according to any one of claims 1 to 9, further comprising: Generating the first input lexical units based on an input prompt in a prefill stage; And Generate an inference result corresponding to the input prompt based on the second output lexical unit in the final decoding stage.

11. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 10.

12. An electronic device, comprising: A first memory configured to store parameters of a generative model including one or more transformer layers; And One or more first processors configured to generate a first output lexical unit by executing the one or more transformer layers based on a first input sequence including a first input lexical unit in a first decoding stage, and generate a second output lexical unit by executing the one or more transformer layers using a second input sequence based on a second input lexical unit corresponding to the first output lexical unit in a second decoding stage, wherein the second input sequence includes the first input sequence and a second input tensor corresponding to the first output lexical unit.

13. The electronic device according to claim 12, wherein, The one or more first processors are further configured to: Cache the first input sequence during the first decoding stage; and Retrieve the cached first input sequence during the second decoding stage to obtain the second input sequence.

14. The electronic device according to claim 13, wherein The one or more first processors include a graphics processor, a neural processor, a tensor processor, an accelerator, or a combination thereof, and The first input sequence is cached in a memory corresponding to one or more second processors including a central processing unit.

15. The electronic device according to claim 13, wherein The second input sequence is generated using in-memory processing, near-memory processing, or a combination thereof.

16. The electronic device according to claim 12, wherein, The one or more first processors are further configured to: Cache the first input sequence in a memory of a second processor different from the one or more first processors; Load the first input sequence into the first memory of the one or more first processors; And Generate the second input sequence based on the first input sequence and the second input tensor.

17. The electronic device according to any one of claims 12 to 16, wherein The one or more first processors are configured to generate a second output lexical unit by appending the second input tensor to the first input sequence.

18. The electronic device according to any one of claims 12 to 16, wherein, The one or more first processors are configured to generate a second output lexical unit based on the following steps: Determine a query tensor corresponding to the product of the second input tensor and the first query weight; Determine a sequence of key tensors corresponding to the product of the first key weight and the second input sequence; Determine a sequence of value tensors corresponding to the product of the first value weight and the second input sequence; Determine a query-key tensor corresponding to the product of the query tensor and the transpose of the first sequence of key tensors; And Determine an attention result corresponding to the product of the non-linear calculation result of the query-key tensor and the sequence of value tensors.

19. The electronic device according to claim 18, wherein, The one or more first processors are configured to: Determine a first intermediate calculation result by multiplying the second input tensor by the first query weight; Determine a second intermediate calculation result by multiplying the first intermediate calculation result by the transpose of the first key weight; And The query-key tensor is determined by multiplying a second intermediate calculation result by the transpose of a second input sequence.

20. The electronic device according to claim 18, wherein, The one or more first processors are configured to: determine a third intermediate calculation result by multiplying a non-linear calculation result of the query-key tensor by the second input sequence; and determine an attention result by multiplying the third intermediate calculation result by a first value weight.

21. A method using a generative model, comprising: obtaining a first input sequence including a first input lexical unit; generating a first output lexical unit by executing, by one or more first processors, a generative model including one or more transformer layers; caching the first input sequence and the first output lexical unit in a memory corresponding to one or more second processors other than the one or more first processors; obtaining a second input sequence by loading the first input sequence and the first output lexical unit from the memory corresponding to the one or more second processors; and generating a second output lexical unit by executing, by the one or more first processors, the generative model based on the second input sequence.

22. The method of claim 21, further comprising: receiving an input prompt, wherein the first input sequence corresponds to the input prompt; and generating a synthetic output corresponding to the input prompt based on the second output lexical unit.

Citation Information

Patent Citations

  • Method for producing an anion exchanger

    KR1020240008854A

  • display device

    KR1020240046554A