Method for improving processing efficiency of generative model and electronic device for performing same
Prompt simplification and utilization of intermediate results in generative models address inefficiencies, enabling efficient on-device processing by reducing computation and memory demands.
Patent Information
- Application Number
- PCT/KR2025/000136
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-01
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-10
AI Technical Summary
Generative models face inefficiencies due to increased computation and memory requirements as prompt length increases, particularly in on-device implementations, leading to processing delays and cache limitations.
A method that simplifies prompts by reducing token numbers and utilizes intermediate computation results (keys and values) from similar prompts, reducing computation and memory transfer during processing.
Enhances processing efficiency by allowing generative models to operate effectively on devices with lower specifications, improving computational speed and memory utilization.
Smart Images

Figure KR2025000136_10072025_PF_FP_ABST
Abstract
Description
Method for improving the processing efficiency of a generative model and an electronic device for performing the same
[0001] The present disclosure relates to a method for improving the processing efficiency of a generative model and an electronic device for performing the same, and more particularly, to a method for improving the processing efficiency by reducing the amount of computation that a generative model must perform and the amount of data transferred between memories during the processing.
[0002] Generative artificial intelligence (AI) technology is widely used in various fields such as summarizing text, answering questions, translation, or image generation.
[0003] Here's a brief explanation of how generative AI works: When a user inputs a prompt, a command containing a request or question, into a generative model, the generative model can generate a response corresponding to the prompt by performing operations between the matrix corresponding to the input prompt and the matrices contained in the generative model's layers.
[0004] Based on the principle that generative models generate keys and values for new tokens, a method can be used to personalize or fine-tune the generative model by setting the prompt area to suit its purpose. Users can personalize or fine-tune the generative model to perform specific tasks by entering prompts that include descriptions and examples of the tasks it will perform.
[0005] Generative models contain a large number of matrices, which inherently require significant computational effort when processing prompts. As the length of the prompt increases, the corresponding matrix size also increases, dramatically increasing the computational effort. Furthermore, as the computational effort increases, the amount of data transferred between memories during the generative model's computations also increases, potentially leading to longer processing times.
[0006] A method for improving the processing efficiency of a generative model disclosed as a technical means for achieving a technical task may include a step of obtaining a prompt, a step of searching for a key and a value corresponding to the prompt, a step of executing the generative model by performing attention using input data that is the target of the prompt and the discovered key and value when the key and value corresponding to the prompt are found, and a step of outputting a first execution result of the generative model.
[0007] An electronic device disclosed as a technical means for achieving a technical task may include a memory storing a program or at least one instruction and at least one processor. When the at least one processor executes the program or at least one instruction stored in the memory, the electronic device may obtain a prompt, search for a key and a value corresponding to the prompt, and, when the key and value corresponding to the prompt are found, perform attention using input data that is a target of the prompt and the found key and value to execute a generation model, and output a first execution result of the generation model.
[0008] A computer-readable recording medium disclosed as a technical means for achieving a technical task may have stored thereon a program for executing at least one of the embodiments of the disclosed method on a computer.
[0009] A computer program disclosed as a technical means for achieving a technical task may be stored on a recording medium for performing at least one of the embodiments of the disclosed method on a computer.
[0010] FIG. 1 is a drawing for explaining modules included in an electronic device according to one embodiment of the present disclosure.
[0011] FIG. 2 is a drawing for explaining a hardware configuration included in an electronic device according to one embodiment of the present disclosure.
[0012] FIG. 3 is a drawing for explaining a generation model according to one embodiment of the present disclosure.
[0013] FIG. 4 is a diagram illustrating a method for executing a generation model by performing attention by an electronic device according to an embodiment of the present disclosure.
[0014] FIG. 5 is a diagram illustrating an electronic device according to one embodiment of the present disclosure executing a generation model using keys and values.
[0015] FIG. 6 and FIG. 7 are diagrams for explaining keys and values according to one embodiment of the present disclosure.
[0016] FIG. 8 is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to search for a key and value corresponding to a prompt.
[0017] FIG. 9 is a diagram illustrating a method for an electronic device to simplify a prompt according to one embodiment of the present disclosure.
[0018] FIG. 10 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0019] FIG. 11 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0020] FIG. 12 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0021] FIGS. 13 to 16 are flowcharts for explaining a method for improving the processing efficiency of a generation model according to an embodiment of the present disclosure.
[0022] In this disclosure, the expression "at least one of a, b, or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof. In this disclosure, key and value may refer to "key," "value," both "key" and "value," or variations thereof.
[0023] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.
[0024] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.
[0025] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.
[0026] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.
[0027] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be performed substantially simultaneously or in reverse order depending on their functionality. It is also possible to perform the flowchart by omitting at least some of the depicted blocks.
[0028] The term '~ module' or '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the '~ unit' may perform a specific role. Meanwhile, the '~ module' or '~ unit' is not limited to software or hardware. The '~ module' or '~ unit' may be configured to be in an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ module' or '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided by a particular component or a particular '~ module' or '~ unit' may be combined or separated into additional components to reduce their number. Furthermore, in one embodiment, a '~ module' or '~ unit' may include one or more processors.
[0029] Below, the meanings of terms used in this disclosure are explained.
[0030] "Generative AI" can refer to AI technology capable of generating new text, images, etc. in response to prompts and input data (e.g., text, images, etc.). Representative examples of generative AI are described in the "Generative Models" section below.
[0031] A "generative model" may refer to a neural network model that implements generative AI technology. The generative model can generate text or images based on the intent contained in a prompt. Furthermore, the generative model can learn the patterns and structures of training data to generate new data with similar characteristics to the input data or new data corresponding to the input data. For example, if the prompt is text containing a question, the generative model can generate and output an answer to the question. Or, for example, if the prompt is text containing a request, the generative model can output text or images generated according to the request. A transformer executed in an electronic device according to an embodiment of the present disclosure corresponds to a generative model. Instead of "generative model," terms such as "generative artificial intelligence model," "language model," "neural network model," or "model" may also be used.
[0032] A 'prompt' is a sentence or keyword for interaction between a user and a model, and can be text for the user to ask a question, request, or command to the model. In other words, a prompt can mean text or other forms of input that guide the model on what type of output it should generate. In the present disclosure, 'execute a prompt' or 'execute a generative model according to a prompt' can mean an action in which the generative model performs a task according to the request of the prompt, that is, an action in which the generative model performs a calculation and generates a result corresponding to the prompt when a prompt is input to the generative model. A prompt can include 'intent' and 'details', which will be described in detail below. Terms such as 'instruction' can also be used instead of 'prompt'.
[0033] The "intent" of a prompt can refer to the purpose or goal that the user wants to achieve through the model, the intention inherent in the context of the prompt, or the part that instructs the model on the task that it should perform. Furthermore, the remaining parts of a prompt, excluding the intent, can be referred to as the "details" of the prompt. In other words, the details of a prompt can refer to additional information or conditions that specify the intent of the prompt. "Details" can further include example information about the tasks that the generative model will perform, such as examples of the purpose or goal to be achieved, examples of the intent, examples showing the patterns or methods that the model should perform, or examples of pairs of input data and corresponding outputs.
[0034] For example, if the prompt is "Draw a picture of a bird flying in the sky," the intent might be "Draw a picture" or "Draw a picture of a bird," the details might be "bird flying in the sky" or "flying in the sky," and the details might include at least one pair of text descriptions and images as examples. Instead of "intention," terms like "goal," "purpose," "request," "task," or "inquiry" might be used. Furthermore, instead of "details," terms like "specifics" or "conditions" might be used.
[0035] If the prompt asks for a stylistic change task, such as changing a casual sentence to a formal one, the prompt details might include example information such as (Informal: I messed up. / Formal: I made a mistake.), (Informal: That's a lot of stuff. / Formal: That is a considerable amount.), (Informal: You know what I mean? / Formal: Do you understand what I am trying to say?), (Informal: I'm so happy! / Formal: I am overjoyed.), (Informal: I'm kinda tired. / Formal: I am feeling somewhat fatigued.). Or, if the prompt asks you to summarize a sentence: (Original: "The quick brown fox jumps over the lazy dog." / Summary: A fox jumps over a dog.), (Original sentence: "Artificial intelligence is the simulation of human intelligence processes by machines, especially computer systems." / Summary: AI simulates human intelligence in machines.), (Original sentence: "Climate change is a long-term shift in temperature and weather patterns." / Summary: Climate patterns are changing over. time.), (Original sentence: "The COVID-19 pandemic has had a significant impact on global economies and societies." / Summary: COVID-19 has affected the world economy and society.), (Original sentence: "Scientists are working to develop a vaccine for the new coronavirus." / Summary: Researchers are creating a COVID-19 vaccine.) etc. The prompts may include example information such as, "The COVID-19 pandemic has had a significant impact on global economies and societies." / Summary: COVID-19 has affected the world economy and society.," (Original sentence: "Scientists are working to develop a vaccine for the new coronavirus." / Summary: Researchers are creating a COVID-19 vaccine.). However, these are just examples, and the actions and example information that the generative model performs are not limited to these.
[0036] Prompts can include system prompts and user prompts. System prompts can be prompts that provide guidance on what role the generative model should play and how it should respond when interacting with a user. System prompts can be used to establish specific behavioral guidelines, such as default behavior, initial settings for the model, or response styles and ranges. User prompts can be questions, requests, or commands that the user directly inputs to the generative model.
[0037] "Input data" can refer to the actual data that a model must process or analyze. Input data can take various forms, such as text, images, or audio. For example, if a user requests translation by prompting the model with "Translate the following sentence into Korean," the text to be translated can be the input data. Or, if a user requests editing of an image by prompting the model with "Erase the clouds in the sky," the image to be edited can be the input data. Terms such as "source data" or "input values" can also be used instead of "input data."
[0038] An "input sequence" refers to the actual input supplied to a model, and can refer to the entire input that the model must process. In other words, an input sequence can refer to the entire data passed to the model's input layer, and can include not only text prompts but also other forms of input data, such as images and audio. In other words, an input sequence can be a combination of prompts and input data. For example, an input sequence for a text-to-image model can include image data to be edited and a prompt (text) instructing the editing.
[0039] For example, if a user inputs the sentence "He always inspires me" as input data along with the prompt "Translate the following sentence into Korean," the input sequence could be "Translate the following sentence into Korean. He always inspires me." Alternatively, if a user inputs the image to be edited along with the prompt "Erase the clouds in the sky," the input sequence could be a combination of "Erase the clouds in the sky in the photo" and the image. Alternatively, the input sequence could be a combination of the text and the prompt, which is the image input data. Alternatively, if a user inputs the audio to be summarized along with the prompt "Summarize the conversation," the input sequence could be a combination of the phrase "Summarize the conversation" and the audio. Alternatively, the input sequence could be a combination of the text and the prompt, which is the audio input data. Instead of 'input sequence', the terms 'complete input', 'input stream', or 'input series' may also be used.
[0040] A prompt simplification module may refer to a configuration that analyzes an input prompt and performs necessary transformations on the prompt. In the present disclosure, the prompt simplification module can enable a generation model to operate more efficiently by simplifying the prompt. The prompt simplification module can understand the intent and context of the prompt and simplify the prompt based on the results. For example, the prompt simplification module can restructure the prompt by removing, modifying, or extracting some of the tokens contained in the prompt.
[0041] In the present disclosure, an electronic device can use a prompt simplification module to clarify or simplify prompts, thereby removing unnecessary information from the prompt, simplifying sentence structures, replacing them with simpler words, or emphasizing important information, thereby improving the accuracy of a model's responses. The prompt simplification module can generate various prompts and select appropriate prompts from among them. The rules for converting prompts by the prompt simplification module can be implemented in various ways, and various techniques can be used when converting prompts.
[0042] The prompt simplification module can be implemented as part of the generative model or separately, external to the generative model. Furthermore, the prompt simplification module can be implemented as a rule-based system or can include a neural network. Instead of the "prompt simplification module," terms such as "token reasoner" or "prompt compression module" may be used.
[0043] 'Prompt simplification' can refer to a configuration that performs token pruning, i.e., the operation of removing unnecessary tokens from a prompt or input sequence to improve model efficiency. The 'prompt simplification module' can evaluate the importance of each token and remove tokens with low importance. In other words, the prompt simplification module can leave key tokens or main tokens among the tokens included in the prompt or input sequence and remove dummy tokens or auxiliary tokens. At this time, the importance of each token can be determined by an attention mechanism or various other evaluation criteria.
[0044] A 'task database' may refer to a space where data related to tasks performed using a model are stored. The task database may store information on a user's previous use of a model. According to one embodiment of the present disclosure, the task database may store information on prompts corresponding to previously performed tasks (e.g., an embedding matrix corresponding to a prompt, a key (key, K) and a value (value, V) corresponding to a prompt, information on mapping between prompts and keys and values, etc.). In addition, the task database may store intermediate computation results generated in the process of performing a previous task. For example, a hidden state matrix output from each layer of a model (e.g., an attention value matrix, an activation matrix, etc.) may be stored. A key (key, K) and a value (value, V) calculated for a given prompt may be stored. At least one of a given prompt, the calculated key and value, or information on mapping between a prompt and a key and value may be stored in the task database.
[0045] The term 'Hidden State Matrix (HSM)' may refer to the intermediate operation results output from each layer (e.g., self-attention layer, etc.) included in the model when the model executes a prompt. In other words, the hidden state matrix may refer to the result value of performing an operation on tokens included in the input sequence using the weights included in the hidden layer of the model. There may be a corresponding HSM for each of the various layers included in the model (e.g., attention layer, activation layer, etc.). Accordingly, the HSM may include an attention score matrix, an attention value matrix, an activation matrix, a key matrix, a value matrix, etc. Instead of 'hidden state matrix', terms such as 'latent matrix', 'latent variable matrix', or 'intermediate matrix' may also be used.
[0046] For each attention head, a query (Query, Q), a key (key, K), and a value (value, V) can be generated using an embedding matrix X corresponding to at least a portion of an input sequence (e.g., a prompt, input data), or a simplified input sequence (e.g., a simplified prompt, a simplified input data). For example, by performing an embedding transformation on an input sequence including a prompt, the electronic device (1000) can generate an embedding matrix X=[X1, X2, ..., X n ] can be obtained. For each attention head, Q i =W q X X i Through the key matrix Q=[ Q1, Q2,.., Q n ] can be calculated. Here X iis the embedding vector of the i token, Q i is X i The query vector corresponding to W q can mean a query weight matrix for computing a query. Instead of 'query', terms such as 'query data', 'query matrix', 'query cache', 'query cache', or 'one or more query vectors' can also be used. For each attention head, K i =W k X X i Through the key matrix K=[K1, K2,.., K n ] can be calculated. Here X i is the embedding vector of the i token, K i is X i The corresponding key vector, W k can mean a key weight matrix for calculating a key. Instead of 'key', terms such as 'key data', 'key matrix', 'key cache', 'key cache', 'trainable key', 'trainable key cache', 'LK (learnable key)', 'LK cache' or 'one or more key vectors' can also be used. For each attention head, V i =W v X i Through the value matrix V=[V1, V2, ..,,V n ] can be calculated. Here X i is the embedding vector of the i token, V i is X i The corresponding value vector, W v can refer to a value weight matrix for calculating values. Instead of 'value', terms such as 'value data', 'value matrix', 'value cache', 'trainable value', 'trainable value cache', 'LV (learnable key)', 'LV cache', or 'one or more value vectors' can also be used.
[0047] According to one embodiment of the present disclosure, the task database may be a personalized database implemented in the cloud or on-device. That is, the task database may be configured to correspond to each user account. Instead of "task database," terms such as "personal database," "personal knowledge graph," or "database" may also be used.
[0048] When users fine-tune a generative model for specific downstream tasks, they can further improve its performance, create domain-specific models, and guide them toward producing desired outputs. Consequently, the need for fine-tuning generative models for specific downstream tasks, such as through prompts, is increasing.
[0049] Fine-tuning with prompts can refer to providing specific prompts to a previously trained generative model to guide the model toward producing the desired output. Fine-tuning with prompts allows the model to adapt to a specific task without collecting a large amount of new data, increasing model flexibility and accelerating its adaptation to new tasks. For example, users can fine-tune a generative model using few-shot, zero-shot, or predefined prompts that provide a few example data sets for a task, or by instructing the generative model to produce a specific output.
[0050] However, as downstream tasks become more complex, the length of the prompt increases, which can increase the computational load the generative model must perform. Larger prompts also increase the size of the queries, keys, and values computed for that prompt, requiring more memory allocation and computation. This can also lead to limitations in the cache size required for processing input data.
[0051] In an on-device environment, limited cache size can make it difficult to process complex prompts in a generative model. Long prompts require significant memory, potentially exceeding the cache size and making processing impossible. Complex operations also require significant memory, which can slow down processing or even lead to errors if the cache size is small.
[0052] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0053] The features of the embodiments of the present disclosure are briefly summarized as follows.
[0054] - If there is a key and value corresponding to the input prompt, the electronic device executes the generated model using the corresponding key and value.
[0055] - If there is no key and value corresponding to the entered prompt, the electronic device simplifies the entered prompt and generates a key and value using the simplified prompt.
[0056] - Update keys and values using the input prompts
[0057] The present disclosure relates to a method for improving the processing efficiency of a generative model, and embodiments of the present disclosure have the characteristic of reducing the amount of computation that the generative model must perform and the amount of data transferred between memories during the processing by simplifying prompts input to the generative model and utilizing intermediate computational results (e.g., keys and values) obtained in the process of executing the same or similar prompts previously. Therefore, by using a method according to embodiments of the present disclosure, the generative model can be executed even on electronic devices with relatively low specifications, thereby enabling efficient on-device implementation of the generative model.
[0058] According to embodiments of the present disclosure, the computational complexity of a generative model can be reduced through prompt simplification. While generative models are used in embodiments of the present disclosure, the methods according to embodiments of the present disclosure can also be applied to various other neural network models.
[0059] FIG. 1 is a diagram illustrating modules included in an electronic device according to one embodiment of the present disclosure. Referring to FIG. 1, an electronic device (1000) according to one embodiment of the present disclosure may include a task management module (200), a generation model (300), an output module (400), and a task database (500).
[0060] The modules (200, 300, 400, 500) included in the electronic device (1000) of FIG. 1 are components classified based on their functions or roles. The modules (200, 300, 400, 500) of the electronic device (1000) of FIG. 1 may be software components implemented by the processor (1300) of the electronic device (1000), which will be described later with reference to FIG. 2, executing a program stored in the memory (1400), or may be virtual components in which no matching hardware device actually exists. In other words, the operations performed by the processor (1300) of the electronic device (1000) executing a program or instruction stored in the memory (1400) may be classified into a plurality of groups based on their functions or purposes, and the entities that perform the operations included in each of the classified groups may be expressed as the modules (200, 300, 400, 500) of FIG. 1. Accordingly, the operations described as being performed by the modules (200, 300, 400, 500) of the electronic device (1000) illustrated in FIG. 1 can be seen as actually being performed by the processor (1300) of the electronic device (1000) executing a program or instruction stored in the memory (1400).
[0061] In FIG. 1, one electronic device (1000) is illustrated as including all modules (200, 300, 400, 500), but this is not limited to the embodiment, and at least some of the modules (200, 300, 400, 500) may be implemented to be included in a separate device, or one module may be implemented to be included in another module.
[0062] Additionally, although the task management module (200) in FIG. 1 is illustrated as including a KV search module (210), a prompt simplification module (220), and a KV update module (230), the present invention is not limited thereto, and at least some of the modules (210, 220, 230) may be implemented to be included in a separate device, or one of the modules may be implemented to be included in another module described above.
[0063] In this way, the modules (200, 300, 400, 500) included in the electronic device (1000) according to one embodiment of the present disclosure may be a hardware configuration or a software configuration, and may be implemented in the form of various electronic devices (e.g., one electronic device or a combination of two or more electronic devices).
[0064] An electronic device (1000) according to one embodiment of the present disclosure may be a user's terminal (e.g., a smartphone, a laptop, a desktop, etc.), but is not limited thereto, and may also be a server that performs communication with the user's terminal.
[0065] For example, the electronic device can be implemented as various electronic devices such as a laptop computer, a desktop, an e-book reader, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a navigation device, an MP3 player, a camcorder, an IPTV (Internet Protocol Television), a DTV (Digital Television), a TV, a set-top box, a smart monitor, a tablet PC, a laptop, a digital signage, a large display, a 360-degree projector, a MS (Mobile Station), a vehicle, a satellite, an airborn, a cellular phone, a smart phone, a wearable device, etc. Alternatively, the electronic device can be an augmented reality device. An 'augmented reality device' is a device that can express augmented reality, and can be implemented as, for example, augmented reality glasses in the shape of glasses that a user wears on the face. However, it is not limited thereto, and the augmented reality device may be implemented as a head-mounted display apparatus (HMD) worn on the user's head, an augmented reality helmet, etc.
[0066] The task management module (200) is configured to manage the execution of tasks using the generation model (300). The task management module can receive prompts from a user. The prompts may be received together with input data, or may be received before or after the input data. For example, when an input sequence including a prompt and input data is received, the task management module can identify the received prompt.
[0067] The KV search module (210) of the task management module (200) can search for keys and values corresponding to the received prompt. If a key and value corresponding to the prompt are found, the task management module (200) can reduce the amount of computation, computation time, etc. of the generation model (300) and reduce memory allocation for the intermediate computation results (e.g., keys and values) by caching the intermediate computation results (e.g., keys and values) stored in the task database (500). If the task management module (200) does not have a key and value corresponding to the received prompt, the electronic device can simplify the prompt and generate a key and value using the simplified prompt. The task management module (200) can update the key and value using the input prompt.
[0068] The KV search module (210) can find the same or similar prompt that was previously performed, read the key and value corresponding to the prompt from the flash memory, store them in the cache memory (e.g., the cache memory of the GPU, the cache memory of the CPU, the cache memory of the NPU, etc.), and allow the generation model (300) to use the key and value stored in the cache memory when performing operations.
[0069] The task database (500) to be described later may store information related to tasks performed in the past (e.g., prompts corresponding to previous tasks, intermediate operation results generated when performing previous tasks, mapping information between prompts and intermediate operation results, etc.), and the KV search module (210) may determine whether there are keys and values corresponding to the prompt by searching the task database (500) based on the prompt. In one embodiment of the present disclosure, the KV search module (210) may also determine whether there are keys and values corresponding to the prompt by searching the task database (500) based on a prompt simplified by the prompt simplification module (220) to be described later.
[0070] When a key and value corresponding to a prompt are found, the work management module (200) (e.g., KV search module (210)) can cache intermediate operation results (e.g., keys and values) stored in the work database (500) so that the generation model (300) can use them during operation. The work management module (200) (e.g., KV update module (230)) can update the found keys and values using the prompt before simplification.
[0071] If a key and value corresponding to a prompt are not found, the task management module (200) (e.g., prompt simplification module (220)) can simplify the prompt received from the user. The task management module (200) (e.g., KV update module (230)) can generate a key and a value using the simplified prompt. The task management module (200) (e.g., KV update module (230)) can update the key and the value generated using the prompt before simplification. The task management module (200) can control the generation model (300) to perform an operation using the key and the value.
[0072] The prompt simplification module (220) is configured to simplify a prompt received from a user. The prompt simplification module (220) can simplify a prompt by reducing the number of tokens included in the prompt. When converting a prompt into an embedding matrix, the embedding dimension (embedding size) is often set to a large value because the generation model (300) can understand the complex meaning and contextual information of the prompt as the embedding dimension (embedding size) increases. However, if the embedding dimension has a large value, even a small increase in the number of tokens included in the prompt can significantly increase the computational load of the generation model (300). Therefore, the token reasoner (100) can primarily reduce the computational load of the generation model (300) by reducing the number of tokens included in the prompt.
[0073] According to an embodiment of the present disclosure, the prompt simplification module (220) can determine whether the prompt includes example information about a task to be performed by the generation model (300). If the prompt includes example information, the prompt simplification module (220) can simplify the prompt by removing or extracting the example information from the prompt. The prompt simplification module (220) can simplify the prompt by removing or extracting tokens corresponding to the example information or reducing the number of tokens.
[0074] According to one embodiment of the present disclosure, the prompt simplification module (220) can simplify a prompt based on a user's history of using the electronic device (1000). In other words, the prompt simplification module (220) can simplify a prompt based on previously executed prompts on the electronic device (1000). If a user frequently and repeatedly inputs the same or similar prompts, the prompt simplification module (220) can simplify a newly input prompt based on the prompts previously input by the user. Furthermore, according to one embodiment of the present disclosure, the prompt simplification module (220) can simplify a prompt by removing some tokens based on the importance of the tokens included in the prompt or by reducing the number of tokens while maintaining the same intent and context.
[0075] When the generative model (300) supports multimodal input, the prompt simplification module (220) can also simplify an input sequence that combines a prompt (e.g., text requesting editing of an image, text requesting a summary of a text or image) and input data (e.g., image data to be edited, text or image to be summarized). For example, input data in the form of an image (e.g., video) or audio is converted to text and then combined with a prompt to generate an input sequence in the form of a text, and the token reasoner (100) can simplify the input sequence generated in this way. The method by which the prompt simplification module (220) simplifies the input sequence can be the same as the method of simplifying the prompt.
[0076] According to one embodiment of the present disclosure, the prompt simplification module (220) can handle a natural language processing process with the purpose of 'simplifying prompts'. The prompt simplification module (220) may include a neural network, but may also operate without a neural network using other algorithms or rule-based methods. According to embodiments of the present disclosure, the prompt simplification module (220) can improve the processing efficiency of the generation model (300) by reducing the number of tokens in advance based on usage history, etc. before executing the generation model (300).
[0077] In this disclosure, the prompt simplification module (220) is expressed as 'simplifying' the prompt, but other expressions such as prompt optimization, prompt compression, prompt extraction, prompt summary, etc. may also be used.
[0078] The KV update module (230) can update keys and values using a prompt. The KV update module (230) can generate a data set including at least one data sample using a prompt. The KV update module (230) can execute a generation model using the keys and values generated using a simplified prompt and the input data of the data sample. The generated keys and values can be updated using the execution results of the generation model and the output data of the data sample.
[0079] In one embodiment, if the prompt includes example information, the KV update module (230) may extract the example information and store it as a data sample of the data set (550) in the working database (500). For example, if the example information (44) includes a pair of input data examples and output examples, the electronic device (1000) may store the input data examples and the output examples as data samples of the data set (550) in the working database (500).
[0080] In one embodiment, the electronic device (1000) may obtain output corresponding to the execution result of the generation model by executing the generation model using prompts and input data. The electronic device (1000) may store pairs of outputs obtained using the input data and prompts as data samples of a data set (550) in the work database (500).
[0081] The generative model (300) may be a generative AI model for generating text, images, audio, etc. according to a prompt input by a user. According to one embodiment of the present disclosure, the generative model (300) may be implemented in the form of a transformer and may include an encoder (310) and a decoder (320). The encoder (310) and the decoder (320) of the generative model (300) may each include one or more attention layers and feedforward layers.
[0082] In the process in which the generation model (300) performs operations according to the prompt, the layers included in the encoder (310) and decoder (320) output matrices and pass them to the next layer. The matrices generated in the middle of the operation process are called hidden state matrices (HSMs). For example, in the process in which the generation model (300) performs operations, the attention layers may output an attention value matrix, and the feedforward layers may output an activation matrix. Both the attention value matrix and the activation matrix are included in the HSM. The HSMs output from the layers of the generation model (300) may be stored in the work database (500) as intermediate operation results.
[0083] According to one embodiment of the present disclosure, the generation model (300) may be an on-device model installed in an electronic device (1000), which is a user's terminal. For example, the generation model (300) may be executed by the processor (1300) of the electronic device (1000), which will be described later, executing a program stored in a memory (1400).
[0084] The output module (400) may be configured to generate output to be provided to a user according to the output of the generation model (300) or a request from the task management module (200).
[0085] When the generation model (300) performs an operation, the output module (400) can appropriately transform and output the output of the generation model (300) according to the input / output interface (1100) of the electronic device (1000). For example, when the output of the generation model (300) is in the form of text, the output module (400) can output the text on the screen in a predetermined font and size, or convert it into voice and output it through a speaker.
[0086] The task database (500) may be configured to store information related to tasks performed by the generation model (300). According to one embodiment of the present disclosure, the task database (500) may store information related to previously performed tasks. For example, keys and values calculated for a given prompt may be stored in the task database (500). In addition, mapping information between prompts and keys and values may be stored in the task database (500). In addition, key weight matrices and value weight matrices calculated for a given prompt may be stored in the task database (500). When a prompt similar to a previously executed prompt is executed, the task management module (200) may cache the keys and values stored in the task database (500) so that the generation model (300) may use them during calculation.
[0087] Keys and values can be stored in flash memory, and the electronic device (1000) can store some of the data stored in the work database (500) in cache memory during the process of performing a task to quickly utilize it for calculation.
[0088] The working database (500) may include a data set. Data samples of the data set may be used to update keys and values. In one embodiment, if the prompt includes example information, the KV update module (230) may extract the example information and store it as a data sample of the data set (550) in the working database (500). In one embodiment, the electronic device (1000) may execute a generation model using the prompt and input data to obtain an output corresponding to the execution result of the generation model. The electronic device (1000) may store pairs of outputs obtained using the input data and the prompt as data samples of the data set (550) in the working database (500).
[0089] According to one embodiment of the present disclosure, the task database (500) may store token sequences representing the intent of frequently used prompts or the intent and details of frequently used prompts. Here, the term "token sequence" may refer to a set of one or more tokens or a unit listing one or more tokens. The stored token sequences may be used in the process of simplifying prompts.
[0090] According to one embodiment of the present disclosure, the task database (500) may include one or more personalized knowledge graphs, and relationships between the intent and details of previously executed prompts, example information, and intermediate operation results (e.g., keys, values, key weight matrix, value weight matrix) may be established through nodes of the knowledge graph. Accordingly, it is possible to check whether there are keys and values corresponding to the prompts through the knowledge graph, and what intermediate operation results (e.g., keys, values, key weight matrix, value weight matrix) are obtained when executing the corresponding prompts.
[0091] According to one embodiment of the present disclosure, the work database (500) may include databases having various structures and forms, such as a relational database. Data stored in the work database (500) will be further described below, if necessary.
[0092] FIG. 2 is a diagram illustrating a hardware configuration included in an electronic device according to one embodiment of the present disclosure. Referring to FIG. 2, an electronic device (1000) according to one embodiment may include a communication interface (1100), an input / output interface (1200), a processor (1300), and a memory (1400). However, the components of the electronic device (1000) are not limited to the above-described examples, and the electronic device (1000) may include more or fewer components than the above-described components. Some or all of the communication interface (1100), the input / output interface (1200), the processor (1300), and the memory (1400) may be implemented in the form of a single chip.
[0093] The communication interface (1100) is a configuration for transmitting and receiving signals (such as control commands and data) with an external device via wire or wirelessly, and may be implemented to include a communication chipset that supports various communication protocols. The communication interface (1100) may receive signals from the outside and output them to the processor (1300), or transmit signals output from the processor (1300) to the outside. The electronic device (1000) may communicate with external devices via the communication interface (1100).
[0094] The input / output interface (1200) may include an input interface (e.g., a touch screen, a keyboard, a microphone, etc.) for receiving commands or information from a user, and an output interface (e.g., a display panel, a speaker, etc.) for displaying the results of execution of an operation according to a user's command or the status of the electronic device (1000). According to one embodiment of the present disclosure, the electronic device (1000) may receive a prompt and input data from a user through the input / output interface (1200), and when a task is completed, may output the results of performing the task (e.g., an answer to a request or question of a prompt, an image or audio edited according to a request of a prompt, etc.) through the input / output interface (1200).
[0095] The processor (1300) controls a series of processes to operate the electronic device (1000) according to the embodiments described below, and may be composed of one or more processors. The one or more processors included in the processor (1300) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. The one or more processors included in the processor (1300) may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP). When the one or more processors included in the processor (1300) are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0096] The processor (1300) can write data to the memory (1400) or read data stored in the memory (1400), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program or at least one instruction stored in the memory (1400). Accordingly, the processor (1300) can perform operations described in the following embodiments, and operations described as being performed by the electronic device (1000) or modules (200, 300, 400, 500) included in the electronic device (1000) in the following embodiments can be regarded as being performed by the processor (1300) unless otherwise specifically described.
[0097] The memory (1400) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (1400) may not exist separately and may be configured to be included in the processor (1300). The memory (1400) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (1400). The memory (1400) may also provide stored data to the processor (1300) upon request of the processor (1300).
[0098] FIG. 3 is a drawing for explaining a generation model according to one embodiment of the present disclosure.
[0099] Referring to FIG. 3, the generative model (300) may be a generative AI model for generating text, images, audio, etc. according to a prompt input by a user. According to one embodiment of the present disclosure, the generative model (300) may be implemented in the form of a transformer and may include an encoder (310) and a decoder (320). The encoder (310) and the decoder (320) of the generative model (300) may each include one or more attention layers and feedforward layers. Each attention layer may include one or more attention heads. For example, in the case of multi-head attention, multiple attention heads may operate in parallel, and the outputs obtained from each attention head may be combined to obtain a final output.
[0100] In the process in which the generation model (300) performs operations according to the prompt, the layers included in the encoder (310) and decoder (320) output matrices and pass them to the next layer. The matrices generated in the middle of the operation process are called hidden state matrices (HSMs). For example, in the process in which the generation model (300) performs operations, the attention layers may output an attention value matrix, and the feedforward layers may output an activation matrix. Both the attention value matrix and the activation matrix are included in the HSM. The HSMs output from the layers of the generation model (300) may be stored in the work database (500) as intermediate operation results.
[0101] According to one embodiment of the present disclosure, the generation model (300) may be an on-device model installed in an electronic device (1000) which is a user's terminal.
[0102] FIG. 4 is a diagram illustrating a method for executing a generation model by performing attention by an electronic device according to an embodiment of the present disclosure.
[0103] The electronic device (1000) can generate a query (Query, Q), a key (key, K), and a value (value, V) for each attention head using an embedding matrix X corresponding to a prompt or a simplified prompt. For example, by performing an embedding transformation on an input sequence including a prompt, the electronic device (1000) can generate an embedding matrix X=[X1, X2, ..., X n ] can be obtained. For each attention head, Q i =W q X X i Through the key matrix Q=[ Q1, Q2,.., Q n ] can be calculated. Here X i is the embedding vector of the i token, Q i is X i The query vector corresponding to W q can mean a query weight matrix for computing a query. Instead of 'query', terms such as 'query data', 'query matrix', 'query cache', 'query cache', or 'one or more query vectors' can also be used. For each attention head, K i =W k X X i Through the key matrix K=[K1, K2,.., K n ] can be calculated. Here X i is the embedding vector of the i token, K i is X i The corresponding key vector, W kcan mean a key weight matrix for calculating a key. Instead of 'key', terms such as 'key data', 'key matrix', 'key cache', 'key cache', 'trainable key', 'trainable key cache', 'LK (learnable key)', 'LK cache' or 'one or more key vectors' can also be used. For each attention head, V i =W v X X i Through the value matrix V=[V1, V2, ..,,V n ] can be calculated. Here X i is the embedding vector of the i token, V i is X i The corresponding value vector, Wv, can mean a value weight matrix for calculating the value. Instead of 'value', terms such as 'value data', 'value matrix', 'value cache', 'trainable value', 'trainable value cache', 'LV (learnable key)', 'LV cache', or 'one or more value vectors' can also be used.
[0104] The generative model (300) can perform scaled dot product attention for each attention head. For example, the generative model (300) can obtain an attention score matrix by calculating the dot product of a query and a key. The generative model (300) can perform scaling on the attention score matrix. For example, the generative model (300) can scale by dividing by the square root of the dimension of the key. The generative model (300) can apply softmax to the scaled attention score matrix and calculate a weight for each key. The generative model (300) can multiply the weight for each key by the attention value matrix to calculate a weighted sum and obtain an attention output.
[0105] The generative model (300) can concatenate the attention outputs calculated from all heads. The generative model (300) can perform a linear transformation on the combined attention outputs and pass them to the output of the attention layer.
[0106] According to an embodiment of the present disclosure, the calculated keys and values may be stored in a work database (500). The electronic device (1000) may cache the stored keys and values when performing calculations, thereby allowing the generation model (300) to utilize the keys and values during calculations. For example, when performing scaled dot product attention, the generation model (300) may perform calculations by utilizing the cached keys and values.
[0107] FIG. 5 is a diagram illustrating an electronic device according to one embodiment of the present disclosure executing a generation model using keys and values.
[0108] The generative model (300) can use a query (Q), a key (K), and a value (V) to calculate relationships between tokens. The generative model (300) can generate an output of a specific token by calculating relationships with all previous tokens, including a specific token. However, as the input sequence becomes longer, calculating relationships with all previous tokens each time can increase the amount of computation. Accordingly, the keys and values calculated for existing tokens can be stored in the working database (500) and reused in calculations for subsequent tokens. When a new token is input, the electronic device (1000) can cache the stored keys and values and calculate the queries, keys, and values only for the new token. The electronic device (1000) can obtain keys and values corresponding to the new token and existing tokens by adding the keys and values calculated for the new token to the cached keys and values. Terms such as KV cache, memory cache, and attention cache may be used for this process.
[0109] For example, assuming that a sequence called "It was a" is processed, a key and a value may be calculated for "It was a", and the calculated key and value for "It was a" may be stored in the working database (500). Now, if the word "dark" is additionally entered into the sequence, the electronic device (1000) can obtain the key and value corresponding to "It was a dark" by caching the calculated key and value for "It was a" and combining the calculated key and value for "dark" with the calculated key and value for "It was a". Thereafter, when the generative model (300) performs attention, the generative model (300) can calculate an attention score matrix using the query corresponding to "dark" and the key corresponding to "It was a dark". The generative model (300) can generate an output by calculating a weighted sum of values corresponding to “It was a dark” based on the attention score matrix.
[0110] The electronic device (1000) can calculate keys and values for predetermined prompts and store them in the work database (500). When a prompt is acquired later, the electronic device (1000) can search for keys and values corresponding to the prompt. When keys and values corresponding to the prompt are found in the work database (500), the electronic device (1000) can cache intermediate calculation results (e.g., keys and values) stored in the work database (500) so that the generation model (300) can use them during calculation. A method for searching for keys corresponding to prompts will be described in detail with reference to FIG. 8 below.
[0111] Referring to FIG. 5, when a prompt (41) is acquired, the electronic device (1000) can search for a key and value corresponding to the prompt (41). When input data (45) is received, the electronic device can obtain a key and value corresponding to the prompt (41) and the input data (45) by caching the key and value (60k, 60v) corresponding to the prompt (41) and combining the key and value (65k, 65v) calculated for the input data (45) with the key and value (60k, 60v) corresponding to the prompt (41). Thereafter, when the generation model (300) performs attention, the generation model (300) can calculate an attention score matrix using the query (65q) corresponding to the input data (45) and the key (60k, 65k) corresponding to the prompt (41) and the input data (45). The generative model (300) can generate an output (50) corresponding to the execution result of the generative model by calculating a weighted sum of values (60v, 65v) corresponding to the prompt (41) and input data (45) based on an attention score matrix. The generated output (50) can be reused when obtaining the next input sequence, and in the process, the key and value corresponding to the output (50) can be added to the cache again.
[0112] FIG. 6 and FIG. 7 are diagrams for explaining keys and values according to one embodiment of the present disclosure.
[0113] As the length of the prompt increases, the cache size required for the keys and values corresponding to the prompt area may also increase, which may increase both memory usage and computational load during inference. Furthermore, when executing a generative model, the maximum usable cache size may be limited on the electronic device. In this case, if the cache size required for the keys and values corresponding to the prompt area is large, the user may input data or output the execution results of the generative model to the user, thereby limiting the maximum sentence length that can actually be utilized by the user, such as reusing the output. For example, referring to Figure 6, if the cache size of a generative model that can be run on an electronic device is limited to 1024 tokens and the sequence length of the prompt is 768 tokens, the sequence length that the generative model can actually process may be limited to 256 tokens.
[0114] According to an embodiment of the present disclosure, the electronic device (1000) can reduce the cache size of keys and values corresponding to a prompt area by having the generative model use keys and values corresponding to simplified prompts when performing calculations. This can reduce both memory usage and computational effort during inference. This can also increase the maximum sentence length that a user can actually utilize. For example, a user can input more or larger input data (45) into the electronic device (1000). Furthermore, the generative model can output more or larger outputs and reuse the outputs. The model's performance in predicting the next token can be improved, and by maintaining context, the model can generate more natural outputs and process more complex sequences.
[0115] Referring to Figure 7, a user can input a prompt (41) that indicates ["Identify and extract the key terms or phrases that represent the most important concepts or ideas from the sentences provided." Input Example: "The migration of birds is influenced by changes in weather and daylight." Output Example: migration, birds, influenced, weather, daylight].
[0116] The electronic device (1000) can search for a key and value corresponding to a prompt (41). The key and value corresponding to the prompt (41) may include a key and value calculated for a simplified prompt (42) indicating [Extract the main keywords from the following sentences.]. A method for searching for a key and value corresponding to a prompt is described in detail with reference to FIG. 8 below.
[0117] When a key and value corresponding to the prompt (41) are stored in the work database (500), the electronic device (1000) can cache the key and value corresponding to the prompt (41). When input data (45) indicating "Sunset means the sun is setting." is received, the electronic device can obtain the key and value corresponding to the prompt (41) and the input data (45) by combining the key and value calculated for the input data (45) with the key and value corresponding to the prompt (41) by caching the key and value corresponding to the prompt (41).
[0118] Thereafter, when the generation model (300) performs attention, the generation model (300) can generate an output (50) corresponding to the execution result of the generation model using a query and a prompt (41) corresponding to the input data (45) and a key corresponding to the input data (45). The generated output (50) can be reused when acquiring the next input sequence, and in the process, the key and value corresponding to the output (50) can be added to the cache again.
[0119] The key and value corresponding to the prompt (41) may correspond to the key and value calculated for the simplified prompt (42). Compared to using the key and value calculated for the prompt (41), the cache size of the key and value corresponding to the prompt area may be reduced when using the key and value calculated for the simplified prompt (42).
[0120] Additionally, the electronic device (1000) can update the computed keys and values for the simplified prompt (42) using the prompt (41). By updating the computed keys and values for the simplified prompt using the prompt (41), the cache size of the keys and values corresponding to the prompt area can be reduced, while the model's inference results can be similar to or better than those before the prompt was simplified. This can improve the performance of the electronic device and help the electronic device provide a better user experience. A method for simplifying the prompt will be described in detail with reference to FIG. 9 below. A method for updating keys and values will be described in detail with reference to FIGS. 10 to 12 below.
[0121] FIG. 8 is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to search for a key and value corresponding to a prompt.
[0122] When a prompt (41) is transmitted to the task management module (200) of the electronic device (1000), the KV search module (210) of the task management module (200) can search for a key and value corresponding to the prompt (41). In one embodiment, searching for a key and value corresponding to the prompt may include searching for a key and value corresponding to a simplified prompt.
[0123] If a key and value corresponding to the prompt (41) are found, the job management module (200) can control the generation model (300) to perform an operation using the intermediate operation results (e.g., key and value) stored in the job database (500). If a key and value corresponding to the prompt (41) are not found, the job management module (200) can create or update a key and value and control the generation model (300) to perform an operation using the key and value.
[0124] In one embodiment, the job management module (200) can control the generation model (300) to perform an operation using the keys and values corresponding to the prompt if there is a prompt identical to the prompt (41) (or a simplified prompt (42)) in the job database (500). If there is a prompt in the job database (500) that is different in details but has the same intent as the prompt (41) (or the simplified prompt (42)), the key and value corresponding to the prompt can be updated using the prompt (41), and the generation model (300) can be controlled to perform an operation using the intermediate operation result (e.g., key and value). If a prompt identical to or similar to the prompt (41) (or the simplified prompt (42)) has never been executed, the key and value can be generated for the simplified prompt (42), and the generation model (300) can be controlled to perform an operation using the generated key and value.
[0125] The KV search module (210) can search for a key and value corresponding to a prompt (41) in the work database (500). The work database (500) can store prompts, keys and values corresponding to the prompts, or mapping information between prompts and keys and values. For example, data 1 representing (prompt 1, key 1, value 1), data 2 representing (prompt 2, key 2, value 2), etc. can be stored in the database.
[0126] Referring to FIG. 8, the task database (500) may store prompts, keys, and values for the English study intent, prompts, keys, and values for the summary intent, and prompts, keys, and values for the style transformation intent. However, these are merely examples for illustrative purposes, and the purposes of the prompts are not limited to the examples mentioned, and prompts for various intents such as translation, code generation, dialogue, image generation, audio generation, video generation, and drug design may be stored.
[0127] In one embodiment, each intent's prompt may include one or more prompts identified by details and key-value pairs. For example, even within the same intent, one or more prompts identified by details such as additional information, conditions, or rules may be stored in the task database (500). For example, prompts such as those in [Table 1] and [Table 2] may be included in the English study intent's prompt or the conversation intent's prompt.
[0128] From now on, you are my English teacher, and we will have conversations in English only. Please follow these rules to continue our dialogue:1. Response Length: Keep your answers to about three to four sentences.2. Response Format and Topic: Always end your response with a question related to the conversation, and try to stay on topic.3. Vocabulary Learning: Include one new word or phrasal verb in each conversation for me to infer its meaning. If I need to remember a particular word or expression, repeat it three times at the end of your response.4. Answer Correction: If there is a grammatical mistake in my response, provide the corrected version first and then answer my original question.5. Content Summary: Occasionally, I will ask for a review of my answers, and when I do, summarize them in the form of a report.
[0129] You are now my interview coach, and we will practice for an upcoming job interview in English. Follow these guidelines to help me improve:1. Interview Simulation: Conduct mock interview sessions by asking me common interview questions for the role I'm applying for.2. Answer Feedback: After I respond, give me feedback on my answers, focusing on clarity, grammar, and the professionalism of my response. Correct any mistakes and suggest improvements.3. Vocabulary and Phrasing: Introduce key phrases or vocabulary that are commonly used in interviews, such as "initiative," "problem-solving," or "team-oriented," and ask me to incorporate them into my responses.4. Follow-up Questions: Ask follow-up questions based on my answers to mimic a real interview setting and help me think on my feet.5. Confidence and Tone: Occasionally comment on my tone and confidence level, providing tips on how to sound more assertive or calm under pressure.
[0130] According to one embodiment of the present disclosure, the task database (500) may store embedding matrices corresponding to prompts. The task management module (200) may compare the embedding matrix corresponding to the prompt (41) (or the simplified prompt (42)) with the embedding matrices stored in the task database (500). If the task database (500) contains a prompt with the same intent and details as the prompt (41) (or the simplified prompt (42)), the task management module (200) may control the generation model (300) to perform an operation using the key and value corresponding to the prompt.
[0131] If there is a prompt in the work database (500) that has different details but the same intent as the prompt (41) (or simplified prompt (42)), the work management module (200) can update the key and value corresponding to the prompt that has different details but the same intent using the prompt (41), and control the generation model (300) to perform the operation using the intermediate operation result (e.g. key and value).
[0132] If there is no prompt identical to or similar to the prompt (41) (or the simplified prompt (42)) in the work database (500), the work management module (200) can generate a key and a value for the simplified prompt (42), update the generated key and value using the prompt (41), and control the generation model (300) to perform an operation using the generated key and value.
[0133] FIG. 9 is a diagram illustrating a method for an electronic device to simplify a prompt according to one embodiment of the present disclosure.
[0134] The prompt simplification module (220) can reduce the amount of data transferred between memories during the task execution process by reducing the number of tokens included in the prompt, and can also increase computational efficiency. The prompt simplification module (220) can convert the prompt entered by the user into a more concise and clearer form, thereby enabling the generation model (300) to generate accurate and relevant output.
[0135] When a user inputs a prompt (41) into an electronic device (1000), a prompt simplification module (220) of the electronic device (1000) can obtain a simplified prompt (42) by removing some tokens from the prompt and changing some tokens. For example, when a user inputs a prompt (41) such as "Could you tell me about the weather in Seoul tomorrow?" into the electronic device (1000), the simplified prompt (42) can include a token sequence indicating an intent ("weather forecast") and a token sequence indicating details ("Seoul tomorrow").
[0136] A simplified prompt may not include at least some of the details (e.g., example information about what the generative model should do). Referring to FIG. 9, for the prompt (41) ["Identify and extract the key terms or phrases that represent the most important concepts or ideas from the sentences provided. Input Example 1: "The migration of birds is influenced by changes in weather and daylight." Output Example 1: migration, birds, influenced, weather, daylight Example 2 (...) Example 3 (...)"], the simplified prompt of FIG. 9 may include [Extract the main keywords from the following sentences.] and not include Example 1, Example 2, etc.
[0137] The prompt simplification module (220) can remove tokens that have a small impact on the intent or context of the prompt based on the importance of each token included in the prompt. The prompt simplification module (220) can simplify the prompt by extracting key tokens from among the tokens included in the prompt and reducing the number of tokens that indicate the intent. For example, the importance of tokens can be determined based on frequency of use or using an attention mechanism, and tokens with an importance above a certain standard can be classified as key tokens and tokens with low importance can be removed, thereby leaving only the core request portion.
[0138] The prompt simplification module (220) can simplify the prompt by reducing the number of tokens representing system prompts or tokens representing details (e.g., example information) among the tokens included in the prompt. Referring to FIG. 9, the prompt simplification module (220) can obtain a simplified prompt (42) by removing or extracting [Input Example 1: "The migration of birds is influenced by changes in weather and daylight." Output Example 1: migration, birds, influenced, weather, daylight] among the tokens included in the prompt (41).
[0139] The prompt simplification module (220) can simplify prompts using a language model or the like. For example, the prompt simplification module (220) can include a prompt conversion model, and the prompt conversion model can be a language model trained to convert input text (prompt) into text containing a smaller number of tokens while maintaining the core content of the input text. When a user-entered prompt (41) is input into the prompt conversion model, the prompt conversion model can output a simplified prompt (42).
[0140] According to one embodiment of the present disclosure, the generative model (300) may support multimodal input. For example, the generative model (300) may receive input data (e.g., images, audio, etc.) along with a prompt requesting processing of the input data (e.g., editing, summarizing, etc.), and may process the input data according to the request of the prompt.
[0141] If the generation model (300) supports multimodal input, the prompt simplification module (220) can generate an input sequence by converting input data into text and then merging it with a prompt, thereby simplifying the input sequence. At this time, the method by which the prompt simplification module (220) simplifies the input sequence may be identical to the method by which the prompt simplification module (220) simplifies the prompt as described above.
[0142] The simplified prompt (42) can be used by the KV search module (210) of the electronic device (1000) to search for keys and values. The simplified prompt (42) can be used by the electronic device (1000) to generate keys and values.
[0143] FIG. 10 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0144] An electronic device (1000) can simplify a prompt (41). The electronic device (1000) can generate keys and values using the simplified prompt (42). The electronic device (1000) can update keys and values generated for the simplified prompt (42) using the prompt (41). The sequence length of the updated keys and values may be the same as the sequence length of the keys and values generated by the simplified prompt.
[0145] An electronic device (1000) can generate a data set including at least one data sample using a prompt. Referring to FIG. 10, the electronic device (1000) can generate keys and values using a simplified prompt. The electronic device (1000) can execute a generation model using the keys and values generated using the simplified prompt and input data of the data sample. The generated keys and values can be updated using the execution result of the generation model and the output data of the data sample. For example, the KV update module (230) can update the keys and values generated for the simplified prompt (42) using at least one data sample of the data set. For example, the parameters of the generation model can also include keys (K1, K2, …, Kn) generated for the simplified prompt (42) and values (V1, V2, …, Vn) generated for the simplified prompt (42) as parameters. By performing forwarding (inference) using data samples of a data set, calculating loss, and performing backwarding, parameters of a generative model including keys (K1, K2, …, Kn) generated for a simplified prompt (42) and values (V1, V2, …, Vn) generated for the simplified prompt (42) can be updated. The sequence length of the updated keys and values can be the same as the sequence length of the keys and values corresponding to the simplified prompt area. In one embodiment, the keys and values generated for the data samples can be removed from the cache.
[0146] FIG. 11 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0147] The electronic device (1000) can obtain an output (50) corresponding to the execution result of the generation model by executing the generation model using the prompt (41) and the input data (45). The electronic device (1000) can store a pair of outputs (50) obtained using the input data (45) and the prompt (41) as data samples of a data set (550) in a working database (500). For example, a data sample in which the output (50) obtained using the prompt (41) for the input data (45) matches as the ground truth can be added to the data set (550) of the working database (500).
[0148] Referring to FIG. 11, using the prompt (41) ["Identify and extract the key terms or phrases that represent the most important concepts or ideas from the sentences provided." Input Example: "The migration of birds is influenced by changes in weather and daylight." Output Example: migration, birds, influenced, weather, daylight] and the input data (45) ["Sunset means the sun is setting."], the key and value are calculated, the generative model (300) performs the operation using the calculated key and value, and the output (50) [Sunset, means] is obtained as the result of executing the generative model. A data sample that matches the output (50) [Sunset, means] as the ground truth for the input data ["Sunset means the sun is setting."] can be added to the data set.
[0149] The KV update module (230) can perform inference using the keys and values generated using the simplified prompt (42) and the input data ["Sunset means the sun is setting."] of the data sample. The KV update module (230) can compare the result of the inference with the output data [Sunset, means] of the data sample to calculate a loss. The KV update module (230) can update the keys (K1, K2, ..., Kn) generated for the simplified prompt (42) and the values (V1, V2, ..., Vn) generated for the simplified prompt (42) using the loss. The prompts mapped to the updated keys and values can be stored in the working database (500). The sequence length of the updated keys and values can be the same as the sequence length of the keys and values corresponding to the simplified prompt area. In one embodiment, the keys and values generated for the input data and the output can be removed from the cache of keys and values.
[0150] FIG. 12 is a diagram illustrating a method for updating keys and values using a prompt according to one embodiment of the present disclosure.
[0151] The electronic device (1000) can determine whether the prompt (41) includes example information (44). If the prompt (41) includes example information (44), the electronic device (1000) can extract the example information (44) and store it as a data sample of a data set (550) in the working database (500). For example, if the example information (44) includes a pair of input data examples and output examples, the electronic device (1000) can store the input data examples and the output examples as data samples of a data set (550) in the working database (500).
[0152] Referring to FIG. 12, the output example [migration, birds, influenced, weather, daylight] of the example information (44) can be added to the data set as a data sample that matches the ground truth for the input example [The migration of birds is influenced by changes in weather and daylight."].
[0153] The KV update module (230) can perform inference using the keys and values generated using the simplified prompt (42) and the input data of the data sample [The migration of birds is influenced by changes in weather and daylight."]. The KV update module (230) can compare the result of the inference with the output data [migration, birds, influenced, weather, daylight] of the data sample to calculate a loss. The KV update module (230) can update the keys (K1, K2, …, Kn) generated for the simplified prompt (42) and the values (V1, V2, …, Vn) generated for the simplified prompt (42) using the loss. The prompts mapped to the updated keys and values can be stored in the working database (500). The sequence length of the updated keys and values can be the same as the sequence length of the keys and values corresponding to the simplified prompt area. In one embodiment, the keys and values generated for the data sample are stored in the cache of keys and values. It can be removed.
[0154] Hereinafter, with reference to the flowcharts of FIGS. 13 to 17, a method for improving the processing efficiency of a generation model according to embodiments of the present disclosure will be described. The steps included in the flowcharts of FIGS. 13 to 17 are performed by the electronic device (1000) of FIGS. 1 and 2 , and therefore, the contents previously described with reference to FIGS. 1 to 12 may be equally applied to FIGS. 13 to 17 even if omitted below.
[0155] FIG. 13 is a flowchart illustrating a method for improving the processing efficiency of a generation model according to an embodiment of the present disclosure.
[0156] Referring to FIG. 13, at step S1310, the electronic device (1000) can obtain a prompt from the user.
[0157] At step S1320, the electronic device (1000) may search for a key and value corresponding to the prompt. In one embodiment, searching for a key and value corresponding to the prompt may include searching for a key and value corresponding to a simplified prompt.
[0158] At step S1325, the electronic device (1000) can determine whether a key and value corresponding to the prompt are found. If a key and value corresponding to the prompt are found, at step S1330, the electronic device (1000) can execute a generation model using the input data and the found key and value.
[0159] If a key and value corresponding to the prompt are found, the electronic device (1000) can control the generation model to perform a calculation using intermediate calculation results (e.g., keys and values) stored in the work database. If a key and value corresponding to the prompt are not found, the electronic device (1000) can create or update a key and value and control the generation model to perform a calculation using the key and value.
[0160] If the electronic device (1000) has a prompt identical to the prompt (or a simplified prompt) in the working database, the electronic device (1000) can control the generation model to perform an operation using the keys and values corresponding to the prompt. If the electronic device (1000) has a prompt in the working database that has different details but the same intent as the prompt (or a simplified prompt), the electronic device (1000) can update the keys and values corresponding to the prompt using the prompt, and control the generation model to perform an operation using the intermediate operation results (e.g., keys and values).
[0161] If a key and value corresponding to the prompt are not found, the electronic device (1000) can obtain the key and value at step S1350. At step S1330, the electronic device (1000) can execute the generation model using the input data and the obtained key and value.
[0162] If there is no prompt identical to or similar to the prompt (or simplified prompt), the electronic device (1000) can generate keys and values for the simplified prompt and control the generation model (300) to perform operations using the generated keys and values. The specific method by which the electronic device (1000) simplifies the prompt is as described above with reference to FIG. 9.
[0163] Figure 14 is a flowchart for explaining detailed steps included in step S1350 of Figure 13.
[0164] In step S1410, the electronic device (1000) can simplify the acquired prompt. The electronic device (1000) can obtain a simplified prompt by removing some tokens from the acquired prompt and changing some tokens.
[0165] A simplified prompt may not include at least some of the details (e.g., example information about the task to be performed by the generative model). The prompt simplification module (220) may simplify the prompt by reducing the number of tokens included in the prompt, among tokens representing system prompts or tokens representing details (e.g., example information).
[0166] The electronic device (1000) can remove tokens that have a small impact on the intent or context of the prompt based on the importance of each token included in the prompt. The electronic device (1000) can simplify the prompt by extracting key tokens from among the tokens included in the prompt and reducing the number of tokens that indicate the intent. For example, the electronic device (1000) can determine the importance of tokens based on frequency of use or by using an attention mechanism, classify tokens with an importance above a certain standard as key tokens, and remove tokens with a low importance, thereby leaving only the core request portion.
[0167] The electronic device (1000) may simplify the prompt by using a language model or the like. For example, the electronic device (1000) may include a prompt conversion model, and the prompt conversion model may be a language model trained to convert the input text (prompt) into text containing a smaller number of tokens while maintaining the core content of the input text (prompt).
[0168] The simplified prompt can be used by the electronic device (1000) to search for keys and values. The simplified prompt can be used by the electronic device (1000) to generate keys and values. The specific method by which the electronic device (1000) simplifies the prompt is as described above with reference to FIG. 9.
[0169] At step S1420, the electronic device (1000) can generate a key and value using a simplified prompt. The specific method by which the electronic device (1000) generates a key and value using a simplified prompt is as described above with reference to FIG. 3.
[0170] At step S1430, the electronic device (1000) can use the prompt to update the keys and values generated using the simplified prompt. A specific method by which the electronic device (1000) uses the prompt to update the keys and values generated using the simplified prompt is as described above with reference to FIGS. 10 and 12 .
[0171] Figure 15 is a flowchart for explaining the detailed steps included in step S1430 of Figure 14.
[0172] In step S1510, if the prompt includes example information, the electronic device (1000) can extract input data and output data from the example information. In step S1520, the electronic device (1000) can obtain a second execution result of the generation model by generating a model using the extracted input data and the generated keys and values. In step S1530, the electronic device (1000) can update the generated keys and values using the second execution result and the extracted output data. If the prompt includes example information, a specific method of updating the generated keys and values using a simplified prompt is as described above with reference to FIG. 12.
[0173] Returning to FIG. 13 again, at step S1340, the electronic device (1000) can output the first execution result of the generation model.
[0174] At step S1360, the electronic device (1000) can update the key and value using the first execution result. The specific method by which the electronic device (1000) updates the key and value using the first execution result is as described above with reference to FIGS. 10 and 11.
[0175] Figure 16 is a flowchart for explaining the detailed steps included in step S1360 of Figure 13.
[0176] At step S1610, the electronic device (1000) can obtain a third execution result of the generated model by executing the generated model using the prompt and input data. At step S1620, the discovered keys and values can be adjusted using the third execution result and the first execution result. A specific method for adjusting the discovered keys and values using the third execution result and the first execution result is as described above with reference to FIGS. 10 and 11.
[0177] According to one embodiment of the present disclosure, a method for improving the processing efficiency of a generative model may be provided. The method may include a step of obtaining a prompt and a step of searching for a key and a value corresponding to the prompt. The method may include a step of executing the generative model by performing attention using input data that is the target of the prompt and the discovered key and value, if the key and value corresponding to the prompt are found. The method may include a step of outputting a first execution result of the generative model.
[0178] The method may include a step of simplifying the prompt if a key and value corresponding to the prompt are not found. The method may include a step of generating a key and a value using the simplified prompt. The method may include a step of updating the generated key and value using the prompt. In the method, the sequence length of the updated key and value may be the same as the sequence length of the generated key and value.
[0179] The step of updating the generated keys and values using the prompt may include a step of extracting input data and output data from the example information when the prompt includes example information about a task to be performed by the generation model. The step of updating the generated keys and values using the prompt may include a step of obtaining a second execution result of the generation model by executing the generation model using the extracted input data and the generated keys and values. The step of updating the generated keys and values using the prompt may include a step of adjusting the generated keys and values using the second execution result of the generation model and the extracted output data.
[0180] The method may include a step of obtaining a third execution result of the generation model by executing the generation model using the prompt and the input data. The method may include a step of adjusting the discovered key and value using the third execution result of the generation model and the first execution result.
[0181] The step of simplifying the prompt may include determining whether the prompt includes example information about a task to be performed by the generative model. If the prompt includes the example information, the step of simplifying the prompt may include obtaining a prompt from which the example information has been removed.
[0182] The method may include the step of outputting a message requesting the user to enter example information about the task to be performed by the generative model. The method may further include the step of obtaining example information about the task to be performed by the generative model.
[0183] The method may include the step of outputting one or more candidate prompts including simplified prompts. The method may include the step of receiving input indicating a prompt selected by a user. The method may include the step of executing a generative model using keys and values corresponding to the selected prompt. The method may include the step of outputting the execution result of the generative model.
[0184] According to one embodiment of the present disclosure, an electronic device may be provided. The electronic device may include a memory storing a program or at least one instruction, and at least one processor. The electronic device may obtain a prompt by the at least one processor executing the program or at least one instruction stored in the memory. The electronic device may search for a key and a value corresponding to the prompt by the at least one processor executing the program or at least one instruction stored in the memory. When the key and the value corresponding to the prompt are found by the at least one processor executing the program or at least one instruction stored in the memory, the electronic device may execute a generation model by performing attention using input data that is the target of the prompt and the found key and value. The electronic device may output a first execution result of the generation model by the at least one processor executing the program or at least one instruction stored in the memory.
[0185] The electronic device can simplify the prompt when a key and value corresponding to the prompt are not found by having the at least one processor execute the program stored in the memory or at least one instruction. The electronic device can generate a key and a value using the simplified prompt by having the at least one processor execute the program stored in the memory or at least one instruction. The electronic device can update the generated key and value using the prompt by having the at least one processor execute the program stored in the memory or at least one instruction. The sequence length of the updated key and value may be the same as the sequence length of the generated key and value.
[0186] The electronic device may, when updating the generated key and value using the prompt, extract input data and output data from the example information if the prompt includes example information about a task to be performed by the generation model. The electronic device may obtain a second execution result of the generation model by executing the generation model using the extracted input data and the generated key and value. The electronic device may adjust the generated key and value using the second execution result of the generation model and the extracted output data.
[0187] The electronic device can obtain a third execution result of the generation model by executing the generation model using the prompt and the input data. The electronic device can adjust the discovered key and value using the third execution result of the generation model and the first execution result.
[0188] The electronic device can determine whether the prompt includes exemplary information about a task to be performed by the generative model when simplifying the prompt. If the prompt includes the exemplary information, the electronic device can obtain a prompt with the exemplary information removed.
[0189] The electronic device may output a message requesting the user to input example information about the task to be performed by the generative model. The electronic device may obtain example information about the task to be performed by the generative model.
[0190] The electronic device may output one or more candidate prompts, each including a simplified prompt. The electronic device may receive an input indicating a prompt selected by the user. The electronic device may execute a generation model using keys and values corresponding to the selected prompt. The electronic device may output an execution result of the generation model.
[0191] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.
[0192] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.
[0193] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0194] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.
[0195] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A method for improving the processing efficiency of a generation model, Steps to obtain a prompt; Step of searching for keys and values corresponding to the above prompt; When a key and value corresponding to the above prompt are found, a step of executing a generation model by performing attention using the input data that is the target of the prompt and the found key and value; and A method comprising the step of outputting a first execution result of the above generation model.
2. In the first paragraph, if a key and value corresponding to the prompt are not found, a step of simplifying the prompt; Steps for generating keys and values using the above simplified prompts; and Further comprising a step of updating the generated key and value using the above prompt, A method, characterized in that the sequence length of the updated key and value is the same as the sequence length of the generated key and value.
3. In the second paragraph, the step of updating the generated key and value using the prompt is as follows: If the above prompt includes example information about the task to be performed by the generative model, a step of extracting input data and output data from the example information; A step of obtaining a second execution result of the generation model by executing the generation model using the extracted input data and the generated key and value; and A method comprising a step of adjusting the generated key and value using the second execution result of the generated model and the extracted output data.
4. In any one of paragraphs 1 to 3, A step of obtaining a third execution result of the generation model by executing the generation model using the above prompt and the input data; A method further comprising a step of adjusting the discovered keys and values using the third execution result of the above generation model and the first execution result.
5. In any one of paragraphs 2 to 4, the step of simplifying the prompt comprises: a step of determining whether the above prompt includes example information about the task to be performed by the above generative model; and A method comprising the step of obtaining a prompt from which the example information has been removed, if the prompt includes the example information.
6. In paragraph 1, A step of outputting a message requesting the user to enter example information about the task to be performed by the said generative model; and A method further comprising the step of obtaining example information about tasks to be performed by said generative model.
7. In paragraph 1, A step of outputting one or more candidate prompts including a simplified prompt; A step of receiving input indicating a prompt selected by a user; A step of executing a generation model using keys and values corresponding to the above selected prompt; and A method further comprising the step of outputting the execution result of the above generation model.
8. In electronic devices, A memory in which a program or at least one instruction is stored; and comprising at least one processor, The electronic device, wherein the at least one processor executes a program or at least one instruction stored in the memory, Get the prompt, Explore the keys and values corresponding to the above prompts, When a key and value corresponding to the above prompt are found, the generation model is executed by performing attention using the input data that is the target of the prompt and the found key and value. An electronic device that outputs the first execution result of the above generation model.
9. In the 8th paragraph, the electronic device, If no key and value corresponding to the above prompt are found, simplify the above prompt, Generate keys and values using the simplified prompts above, Update the generated keys and values using the above prompts, An electronic device, characterized in that the sequence length of the updated key and value is the same as the sequence length of the generated key and value.
10. In the 9th paragraph, the electronic device updates the generated key and value using the prompt, If the above prompt includes example information about the task to be performed by the generative model, extract input data and output data from the example information, By executing the generation model using the extracted input data and the generated key and value, a second execution result of the generation model is obtained, An electronic device characterized in that the generated key and value are adjusted using the second execution result of the generated model and the extracted output data.
11. In any one of clauses 8 to 10, the electronic device, By executing the generation model using the above prompt and the above input data, the third execution result of the generation model is obtained, An electronic device characterized in that the discovered key and value are adjusted using the third execution result of the above generation model and the first execution result.
12. In any one of paragraphs 9 to 11, the electronic device simplifies the prompt, Determine whether the above prompt contains example information about the task that the above generative model is to perform, An electronic device characterized in that, if the above prompt includes the example information, a prompt with the example information removed is obtained.
13. In any one of clauses 8 to 12, the electronic device, Output a message asking the user to enter example information about what the above generative model should do; An electronic device for obtaining example information about tasks to be performed by the above generative model.
14. In any one of clauses 8 to 13, the electronic device Outputs one or more candidate prompts that contain simplified prompts, Receives input indicating a prompt selected by the user, Execute the generated model using the keys and values corresponding to the selected prompts above, An electronic device characterized by outputting the execution result of the above generation model.
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 8 on a computer.
Citation Information
Patent Citations
Text classification method and system based on K selection strategy sparse self-attention
CN113392214A
Data processing method and device, electronic equipment and storage medium
CN116579373A
Pattern generation method and device and electronic equipment
CN116993861A
Calculation method and device of neural network model, electronic equipment and storage medium
CN117273084A
Apparatus that crushes and sorts waste insulation containers
KR102599595B1