Apparatus for speculative decoding in electronic device and method for operating same
Speculative decoding with draft models improves on-device AI efficiency by verifying small model outputs in a large model, addressing computational and memory constraints to enhance processing speed and accuracy.
Patent Information
- Application Number
- PCT/KR2025/008429
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2025-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
On-device AI systems face limitations in processing large amounts of data due to limited computational power and memory capacity, necessitating improved methods for efficient data processing.
The implementation of speculative decoding using draft models in conjunction with a target model, where draft models generate initial tokens, which are verified by the target model to determine the most suitable draft model for decoding based on probability distributions and size considerations.
This approach enhances processing efficiency and accuracy by quickly verifying small model outputs in a large model, optimizing response time and latency in AI operations.
Smart Images

Figure KR2025008429_26122025_PF_FP_ABST
Abstract
Description
Device for speculative decoding in electronic devices and method of operation thereof
[0001] The present disclosure relates to a device for performing speculative decoding in an electronic device and a method of operating the same, for example, to a device for performing speculative decoding based on an artificial intelligence (AI) model and a method of operating the same.
[0002] AI is a technology that artificially embodies human learning, reasoning, and perception abilities. This AI is evolving into generative AI, capable of producing text, images, or other media in response to prompts corresponding to user input. Large language models (LLMs) are a prime example of generative AI.
[0003] The above AI is evolving from cloud-based AI, which connects to external servers or the cloud to receive data and computation support, to on-device AI, which is installed on the device itself to provide AI services. Devices to which the above on-device AI can be applied may include devices based on a mobile environment, such as smartphones or tablet PCs. The above on-device AI can be independent of internet connection or communication status, as the device can process information on its own. However, the above on-device AI has limitations such as difficulty in hardware expansion for the applicable device, and thus, a method for processing large amounts of data based on limited computational power and / or memory capacity is required.
[0004] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.
[0005] According to one embodiment of the present disclosure, an electronic device may include a memory including one or more storage media storing instructions; and at least one processor including a processing circuit, wherein the at least one processor is configured to individually and / or collectively execute the instructions, and the electronic device performs at least one operation, wherein the at least one operation includes: outputting first tokens per draft model in response to an input of a prompt from a plurality of draft models; outputting second tokens per draft model from a target model by inputting the first tokens per draft model; and determining a target draft model to be used for speculative decoding from among the plurality of draft models based on a similarity between a first probability distribution of the first tokens per draft model and a second probability distribution of the second tokens per draft model.
[0006] According to one embodiment of the present disclosure, a non-transitory computer-readable storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when individually and / or collectively executed by at least one processor of an electronic device, may cause the electronic device to perform at least one operation including: outputting first tokens per draft model in response to an input of a prompt from a plurality of draft models; outputting second tokens per draft model from a target model by inputting the first tokens per draft model; and determining a target draft model to be used for speculative decoding from among the plurality of draft models based on a similarity between a first probability distribution of the first tokens per draft model and a second probability distribution of the second tokens per draft model.
[0007] According to one embodiment of the present disclosure, a method for executing a generative artificial intelligence (AI) model in an electronic device may be provided. The method may include: outputting first tokens for each draft model in response to a prompt input from a plurality of draft models; outputting second tokens for each draft model from a target model by taking the first tokens for each draft model as input; and determining a target draft model to be used for guess decoding from among the plurality of draft models based on a similarity between a first probability distribution of the first tokens for each draft model and a second probability distribution of the second tokens for each draft model.
[0008] According to one embodiment of the present disclosure, an electronic device may include at least one processor including a processing circuit. The electronic device may include a memory in which instructions are stored. The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to perform the following operations: obtaining an AI model including a target model and a first draft model and a second draft model smaller in size than the target model; obtaining input tokens for the AI model based on an input; identifying first output tokens of the first draft model for the input tokens and first probabilities that the first output tokens will be output from the first draft model, respectively; identifying second output tokens of the target model for the first output tokens and second probabilities that the second output tokens will be output from the target model, respectively; identifying third output tokens of the second draft model for the input tokens and third probabilities that the third output tokens will be output from the second draft model, respectively; The method may perform at least one operation including: an operation of identifying fourth output tokens of the target model for the third output tokens, and fourth probabilities that the fourth output tokens will be output from the target model, respectively; an operation of selecting one of the first draft model and the second draft model based at least in part on the first probabilities, the second probabilities, the third probabilities, and the fourth probabilities; and an operation of providing an output for the input based at least in part on the selected one draft model.
[0009] According to one embodiment of the present disclosure, an electronic device may include a memory including one or more storage media storing instructions; and at least one processor including a processing circuit. When the instructions are individually or collectively executed by the at least one processor, the electronic device may perform at least one operation including: inputting a prompt corresponding to an input into a target model, wherein the target model is relatively large in size compared to the plurality of draft models, and the plurality of draft models may have different sizes; obtaining a specific number of output tokens sequentially in response to an input of a start token based on the prompt from the plurality of draft models, wherein the start token may be input from the target model in response to an input of the prompt; performing verification on draft model-specific output tokens obtained by the plurality of draft models in the target model; and determining a first draft model to be used from among the plurality of draft models based on a result of the verification. The first probability information of the first draft model may be relatively similar to the second probability information of the target model compared to the probability information of one or more other draft models. The first probability information may be determined by a first probability value that quantifies the likelihood (e.g., accuracy) that each of the first output tokens obtained from the first draft model matches the information included in the prompt. The second probability information may be determined by a second probability value that quantifies the likelihood (e.g., accuracy) that each of the second output tokens obtained by the target model using the first output tokens as input matches the information included in the prompt.The second output tokens may be at least partially identical to the first output tokens.
[0010] According to one embodiment of the present disclosure, a non-transitory computer-readable storage medium storing at least one computer-readable instruction may be provided. The instructions, when individually and / or collectively executed by at least one processor of an electronic device, may cause the electronic device to perform at least one operation including: inputting a prompt corresponding to an input into a target model, wherein the target model is relatively large in size compared to the plurality of draft models, and the plurality of draft models may have different sizes; obtaining a specific number of output tokens sequentially in response to input of a start token based on the prompt from the plurality of draft models, wherein the start token may be input from the target model in response to input of the prompt; performing verification on draft model-specific output tokens obtained by the plurality of draft models in the target model; and determining a first draft model to be used from among the plurality of draft models based on a result of the verification. The first probability information of the first draft model may be relatively similar to the second probability information of the target model compared to the probability information of one or more other draft models. The first probability information may be determined by a first probability value that quantifies the likelihood (e.g., accuracy) that each of the first output tokens obtained from the first draft model matches the information included in the prompt. The second probability information may be determined by a second probability value that quantifies the likelihood (e.g., accuracy) that each of the second output tokens obtained by the target model using the first output tokens as input matches the information included in the prompt. The second output tokens may be at least partially identical to the first output tokens.
[0011] According to one embodiment of the present disclosure, a method for executing a generative artificial intelligence (AI) model in an electronic device may be provided. The method includes: inputting a prompt corresponding to an input into a target model, wherein the target model is relatively large in size compared to the plurality of draft models, and the plurality of draft models may have different sizes; sequentially obtaining a specific number of first output tokens in response to an input of a start token based on the prompt from the plurality of draft models, wherein the start token may be input from the target model in response to an input of the prompt; confirming first probability information that quantifies a likelihood (e.g., accuracy) that the first output tokens obtained for each draft model in the plurality of draft models match information included in the prompt; obtaining second output tokens for each draft model using the first output tokens as input from the target model; confirming second probability information that quantifies a likelihood (e.g., accuracy) that the second output tokens obtained for each draft model in the target model match information included in the prompt; And it may include an operation of determining a draft model among the plurality of draft models, in which the first probability information is relatively similar to the second probability information compared to one or more other draft models, as a first draft model, and the second output tokens may be at least partially identical to the first output tokens.
[0012] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0013] FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described herein.
[0014] FIG. 2 is a block diagram for enabling an AI model to operate in an electronic device according to one embodiment.
[0015] FIG. 3 is a block diagram of an LLM mounted on an electronic device according to one embodiment.
[0016] FIG. 4 is an operational example diagram of configuring an LLM for speculative decoding in an electronic device according to one embodiment.
[0017] FIG. 5 is an exemplary operation diagram of an electronic device equipped with an AI model based on speculative decoding according to one embodiment, in which a preferred draft model is not determined to be paired with a specific draft model included in a plurality of draft models.
[0018] FIG. 6 is an example diagram of a judgment operation for determining a preferred draft model to be paired with a target model in a group of draft models in an electronic device equipped with an AI model based on speculative decoding according to one embodiment.
[0019] FIG. 7 is an exemplary diagram of an operation of generating tokens based on speculative decoding by pairing a preferred draft model and a target model determined from a plurality of draft models in an electronic device according to one embodiment.
[0020] FIG. 8 is an exemplary diagram of an operation of generating tokens based on speculative decoding in an electronic device according to one embodiment.
[0021] FIG. 9 is a control flowchart for performing speculative decoding in an electronic device according to one embodiment.
[0022] FIG. 10 is a control flowchart for driving a generative AI model in an electronic device according to one embodiment.
[0023] FIG. 11 is a schematic diagram for driving a generative AI model in an electronic device according to one embodiment.
[0024] FIG. 12 is a block diagram of an electronic device within a network environment according to various embodiments.
[0025] FIG. 13 is a block diagram of an exemplary AI system capable of performing the operations described in this document.
[0026] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0027] FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.
[0028] Referring to FIG. 1, the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are exemplary only and do not limit the implementations described or claimed in this document. The electronic device (100) may be referred to as a mobile device, a user device, a multi-function device, a portable device, or a server.
[0029] The electronic device (100) may include components including at least one processor (110) (hereinafter, referred to as 'processor (110)'), at least one memory (120) (hereinafter, referred to as 'memory (120)'), at least one display (140) (hereinafter, referred to as 'display (140)'), at least one image sensor (150) (hereinafter, referred to as 'image sensor (150)'), at least one communication circuit (160) (hereinafter, referred to as 'communication circuit (160)'), and / or at least one sensor (170) (hereinafter, referred to as 'sensor (170)'). The components are merely exemplary. For example, the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input / output interface). For example, some components may be omitted from the electronic device (100). For example, several components may be integrated into a single component. For example, the electronic device (100) may further include at least some of the configurations and / or functions not shown. At least some of the respective components of the electronic device shown (or not shown) may be operatively, functionally, and / or electrically connected to each other.
[0030] The processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data, etc.) stored in the memory (120). The processor (110) may include a processor assembly including one or more processing circuits. The processor (110) may include any processing circuit operative to control the performance and operations of one or more components (e.g., the memory (120), the display (140), the image sensor (150), the communication circuit (160), and / or the sensor (170)) of the electronic device (100). For example, the processor (110) (e.g., the application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (110) may be implemented with multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (110) may include one or more processing circuits. For example, the processor (110) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) that is different from the first chip of the electronic device (100).
[0031] For example, the processor (110) may include a central processing unit (CPU) (111), a graphics processing unit (GPU) (112), a neural processing unit (NPU) (113), an image signal processor (ISP) (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (CP) (118), and / or a sensor interface (119). These components of the processor (110) are merely exemplary. For example, the processor (110) may further include other components. For example, some components of the processor (110) may be omitted from the processor (110). For example, some components of the processor (110) may be included as separate components of the electronic device (100) outside the processor (110). For example, some components of the processor (110) (e.g., memory controller (116)) may be included within other components (e.g., at least a portion of memory (120), an interface (e.g., available for connection to at least one component of the electronic device (100)), a display (140) and / or an image sensor (150)).
[0032] The processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in the memory (120). The CPU (111) (or central processing circuit) may be configured to control components of the processor (110) based on the execution of instructions stored in the memory (120) (e.g., volatile memory (121) and / or non-volatile memory (122)). The GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an AI model (e.g., convolution computation). The ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110). The display controller (115) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from the CPU (111), GPU (112), ISP (114), or memory (120) (e.g., volatile memory (121)) into a format suitable for the display (140). The memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121). The above storage controller (117) (or storage control circuit) may be configured to control reading data from the non-volatile memory (122) and writing data to the non-volatile memory (122).The CP (118) (communication processing circuit) may be configured to process data acquired from a component of the processor (110) into a format suitable for transmitting to another electronic device via the communication circuit (160), or to process data acquired from another electronic device via the communication circuit (160) into a format suitable for processing by a component of the processor (110). For example, the communication circuit (160) may include one or more communication circuits. The sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data on the state of the electronic device (100) and / or the state of the surroundings of the electronic device (100), acquired via the sensor (170), into a format suitable for a component of the processor (110).
[0033] The memory (120) may include one or more storage media (or one or more storage devices). For example, the memory (120) may include a memory assembly including one or more storage media. For example, the one or more storage media may include permanent memory (e.g., non-volatile memory (122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof. The memory (120) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As a non-limiting example, the cache memory may be included within the processor (110). The memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (100).
[0034] For example, the memory (120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (110). For example, the memory (120) may store instructions callable by an application programming interface (API). For example, the memory (120) may store instructions within a library.
[0035] According to an example, the electronic device (100) can execute at least one instance of an AI model. The instance may be an object corresponding to a program (or application), such as an AI model, for example. The instance may be named a replica, a pod, a container, or a virtual machine, and there is no limitation on the name thereof. The number of instances may correspond to the size of a resource (e.g., a GPU (112) or an NPU (113)), and accordingly, the number of instances may be used interchangeably with the size of the resource, or the instances may be used interchangeably with the resource.
[0036] As an example, a plurality of user requests may be input to the electronic device (100). The user requests may be associated with a service. The user request may be processed by a first instance of a first AI model, and a first processing result may be provided from the first instance of the first AI model. The first processing result may be processed by a first instance of a second AI model, and accordingly, a second processing result may be provided by the first instance of the second AI model. By serial processing of the processing results, the first instance of the M-th AI model may receive and process the N-1-th processing result. The first instance of the M-th AI model may provide the N-th processing result as a response. Accordingly, a response corresponding to the user request may be provided.
[0037] Based on the above-described process, responses corresponding to each of a plurality of user requests may be provided. Meanwhile, since processing must be performed by an instance, the time for providing responses corresponding to each of a plurality of user requests (hereinafter referred to as “response time”) may take a relatively long time. The response time may affect the latency of the instance. In order to reduce the response time, the electronic device (100) may increase the number of instances of at least one AI model, which may be referred to as scaling out. However, there may be a limit to increasing the number of instances due to hardware and / or software constraints of the electronic device (100) and / or parameter restrictions of the AI model (e.g., large language model (LLM)).
[0038] In one example, the AI model can be generated through machine learning. The machine learning can be performed, for example, on the electronic device (100) capable of operating on-device AI. The machine learning can also be performed, for example, through a separate external server (e.g., server (1208) of FIG. 12) based on a network environment (e.g., network environment (1200) of FIG. 12). In this case, the learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above.
[0039] In one example, the AI model may be an artificial neural network model written in a specified language and including multiple layers and / or operations (or computations). In one example, the AI model may include one of a feedforward neural network (FNN), a deep neural network (DNN), a convolutional neural network (CNN), a region with convolution neural network (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a Deconvolution Network (RBM), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, a Fully Convolutional Network, a long short-term memory (LSTM) Network, a Classification Network, or a combination of two or more of the above, but is not limited to the above examples. For example, an AI model can be trained on specified data, acquire input data, perform operations based on the input data, and generate output data. The AI model may additionally or alternatively include a software structure in addition to the hardware structure.
[0040] The electronic device (100) may employ speculative decoding as a means of reducing response time and / or improving operating speed in an AI model. The speculative decoding is used, for example, to operate an AI model such as LLM in an environment where constraints such as computational load or memory size exist. For example, the speculative decoding may obtain tokens corresponding to a prompt by a large model (e.g., a target model) and a small model (e.g., a draft model) that form a pair. The prompt (e.g., the prompt (250) of FIG. 2) may be, for example, a command or question that a user conveys to the AI model for a conversation with the AI model. For example, the prompt may serve as a communication window between the user and the generative AI model. The AI model must be able to accurately analyze or understand the prompt in order to provide the user with the accurate information desired. The large model may have high accuracy due to a large amount of computation and / or a large number of parameters that can be processed, but may have a slow processing speed. The small model may have high processing speed due to a small amount of computation and / or a small number of parameters that can be processed, but may have a low accuracy. By utilizing these characteristics, the speculative decoding verifies tokens quickly generated by the small model in the large model. For example, the electronic device (100) may perform a speculative decoding operation in which the small model continuously generates output tokens and the large model takes as input a certain number (e.g., a specific number (N)) of output tokens generated by the small model and verifies the accuracy of each output token.
[0041] The above-described guess decoding can be optimized by the size of the small model and / or the size of the large model. The size of the small model and / or the size of the large model may correspond to the number of parameters. The size of the small model and / or the size of the large model may be determined, for example, by the number of parameters connecting the nodes in a neural network connected to the nodes included in the corresponding model. For example, an AI model may have hundreds of millions, hundreds of billions, or trillions or more parameters. The size of the small model and / or the size of the large model may be expressed as 'large / same / small' or 'more / same / less'. For example, in the following description, the small model (e.g., the draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3) will be expressed as being 'relatively small in size' compared to the large model (e.g., the target model (310) of FIG. 3). For example, in the following description, when the large model (e.g., the target model (310) of FIG. 3) has the same size as the small model (e.g., the draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3), it will be expressed as 'the size is the same.' For example, in the following description, when the large model (e.g., the target model (310) of FIG. 3) has a relatively 'larger size' than the small model (e.g., the draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3), it will be expressed as 'the size is the same.'
[0042] For example, if the size of the large model is fixed, the optimization of the guess decoding (e.g., accuracy and / or satisfaction) may be determined by the size of the small model. For example, the small model and / or the large model may be distributed by securing various sets of input data and correct answer data of users in advance as learning data, and optimizing them for the operating environment through learning using the pre-secured learning data. The size of the small model and / or the large model may be one of the main requirements for the optimization. This is because, in an AI model, the accuracy of the result is proportional to the size of the small model and / or the large model, but the operating speed to obtain the result is inversely proportional to the size of the small model and / or the large model. In other words, as the size of the small model and / or the large model increases, the accuracy may increase, but the operating speed may slow down. In addition, in an AI model, the accuracy of the result is generally proportional to the number of parameters in the model, but the operating speed to obtain the result is inversely proportional to the number of parameters in the model. That is, as the number of parameters in the model increases, accuracy may increase, but the operation speed may slow down.
[0043] Considering the above, it is possible to define an average accuracy that satisfies various users and determine the sizes of the small model and / or large model that converge to that average. However, this may not provide sufficient satisfaction to users due to various prompts. For example, while a relatively simple task may provide satisfactory results in terms of accuracy, a relatively complex task may provide unsatisfactory results in terms of accuracy. Therefore, a model selection method that can selectively select the size of the small model (e.g., draft model) and / or large model (e.g., target model) during estimation decoding may be needed, considering accuracy and / or output satisfaction for the input.
[0044] In one example, an electronic device (100) may have multiple draft models generate a specific number (N) of initial tokens corresponding to a prompt, a target model may verify the initial tokens for each draft model, and a draft model determined from among the multiple draft models based on the verification result may be applied for speculative decoding. The draft model to be applied for the speculative decoding may be determined, for example, based on a probability distribution based on the probability values of tokens generated from the multiple draft models and the probability values of tokens generated from the target model by taking the tokens generated from the multiple draft models as input. The draft model to be applied for the speculative decoding may also be determined, for example, by considering the size of the draft model. For example, when there are multiple draft models determined based on the probability distribution, among the multiple draft models determined based on the probability distribution of output tokens, if the probability distribution is closest or the probability distribution is within a specific error range, the draft model with the smallest size may be finally determined to increase the operating speed. For example, among draft models of the same size, a draft model whose probability distribution is most similar to that of the target model, or whose probability distribution is within a certain error range relative to that of the target model, may be selected for guess decoding. For example, among draft models of the same size, a draft model with a relatively high probability distribution may be selected for guess decoding. For example, among draft models of the same size, a draft model suitable for each theme may be selected for guess decoding. In this case, the ability and / or latency of the electronic device (100) to process natural language may be optimized.
[0045] FIG. 2 is a block diagram for enabling an AI model to operate in an electronic device (e.g., the electronic device (100) of FIG. 1) according to one embodiment.
[0046] Referring to FIG. 2, the electronic device (100) may include a processor (e.g., including a processing circuit) (210) (e.g., the processor (110) of FIG. 1), a memory (230) (e.g., the memory (120) of FIG. 1), and / or an interface (IF) (e.g., including an interface circuit) (220) (e.g., the display (140) of FIG. 1). The electronic device (100) may be a device for providing a service linked to at least one AI model.
[0047] At least one of the above AI models may be based on natural language processing (NLP) technology. The NLP technology is, for example, a technology that allows the electronic device (100) to understand or process a user's input, i.e., a natural language expressed in voice and / or text (hereinafter referred to as a "prompt (250)"). The electronic device (100) can understand natural language through NLP, determine human intention based on the understood natural language, and / or convey information in a language understandable to humans. For example, the AI model can immediately process a user input entered as a prompt (250). The AI model can iterate inference to generate a token at each iteration. The NLP can learn the order of words or tokens to understand human language and predict the probability of the next word or token in a given text. The token is a basic unit for processing or understanding a prompt in the AI model. The main technologies of the above NLP include tokenization, part-of-speech tagging, syntax analysis, named entity recognition, and sentiment analysis for prompts corresponding to user input.
[0048] At least one AI model may be based on language model (LM) technology. The LM technology can predict the probability of each word or token (hereinafter collectively referred to as "token") based on a sequence of words or tokens. For example, the LM technology may predict the token that may follow a preceding token. In other words, the LM may be an AI model trained to output the most statistically appropriate token based on a prompt. For example, the electronic device (100) may include an LLM (240). The LLM (240) included in the electronic device (100) may be a single generative LLM. The LLM (240) may be a large-scale deep learning model pre-trained based on a vast amount of data. The LLM (240) may provide the ability to predict a user's intent based on a relatively small number of prompts. The LLM may be used, for example, in generative AI that generates content based on prompts input in the form of human language or text.
[0049] In one embodiment, an LLM may be referred to as an LM that includes an artificial neural network pre-trained on a large amount of data (e.g., text data). The LLM may include approximately ten times more parameters (e.g., approximately 100 billion or more parameters) than a general language model. The LLM may use a transformer artificial neural network structure based on an attention mechanism. The attention mechanism may enable the AI model to focus on important parts within the input data. The attention mechanism may be used to predict output data by predicting the extent to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of some layers of the neural network) contributes to the intermediate or final output of the neural network. For example, the RNN structure may process each element of the sequence sequentially. The RNN structure may have poor prediction performance when there is information dependency between long time-series distances. However, the attention mechanism can take into account information dependence across long time series distances by controlling the degree of weighted attention within the overall (or partial) context of the input data.
[0050] For example, a transformer may include an encoder-decoder structure. The encoder may process input data and output compressed information (e.g., an attention mechanism). The decoder may process the compressed information and output output data in token units. The encoder and / or decoder may include independent attention networks. The transformer may include a cross-attention network connecting the encoder and decoder.
[0051] For example, LLM can be trained through two operations: pre-training and fine-tuning. Pre-training may involve the process of allowing LLM to process a large amount of text data and acquire general linguistic knowledge. Pre-training may involve, for example, self-supervised learning to predict the next word using a sequence of previous words in a text sequence. Fine-tuning may involve training a large-scale language model to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. Fine-tuning may involve, for example, additional supervised learning (or adaptive learning) based on a pre-trained model using a dataset suited to the domain's purpose.
[0052] For example, the LLM can perform a task with a text input containing natural language, called a prompt (310). For example, the LLM can include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer). The term 'LLM' can refer to the neural network model itself, but can also mean a model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, a chatbot such as chatGPT can be referred to as an LLM. The 'LLM' can also include an inference engine composed of various circuits and / or executable program instructions that utilize the LLM neural network model. For example, "inputting an input prompt to the LLM" can be referred to as "inputting the input prompt to an inference engine based on the LLM."
[0053] As previously described, NLP may refer to an AI model for the electronic device (100) to understand or analyze human language, while LLM may be an AI model for predicting the next word or sentence (e.g., a subsequent token) based on given data (e.g., a preceding token). NLP is utilized for search engines, machine translation, or sentiment analysis. LLM is utilized for sentence generation, auto-completion, or voice recognition. The electronic device (100) may provide a generative AI service that uses NLP to understand a user's question, for example, and LLM to generate an appropriate answer thereto.
[0054] The processor (210) may execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the electronic device (100) that is electrically connected thereto. The processor (210) may perform various data processing or operations. As at least a part of the data processing or operations, the processor (210) may store commands or data received from other components (e.g., the I / F (220)) in a memory (230) (e.g., a volatile memory, but without limitation). As at least a part of the data processing or operations, the processor (210) may process commands or data stored in the memory (230) (e.g., a volatile memory, but without limitation). As at least a part of the data processing or operations, the processor (210) may store data resulting from processing commands or data in the memory (230) (e.g., a non-volatile memory, but without limitation). The processor (210) may include, but is not limited to, a CPU (e.g., including processing circuitry) (211) (e.g., CPU (111) of FIG. 1), an NPU (e.g., including processing circuitry) (213) (e.g., NPU (113) of FIG. 1), and / or a GPU (e.g., including processing circuitry) (215) (e.g., GPU (112) of FIG. 1), each of which includes processing circuitry. Each of the “processors,” “processing units (PUs)” or “models” may include processing circuitry and / or multiple processors. For example, the term “processor” or “model” as used herein, including in the claims, may include various processing circuitry including at least one processor, wherein one or more of the at least one processor may be individually and / or collectively configured to perform various functions described herein in a distributed manner.When "processor," "at least one processor," "model," "at least one model," and "one or more processors" are described herein as being configured to perform a plurality of functions, these terms encompass, but are not limited to, situations where one processor and / or model performs some of the recited functions and other processors and / or models perform other parts of the recited functions, and situations where a single processor and / or model can perform all of the recited functions. Furthermore, the at least one processor may comprise a combination of processors that perform the various functions described / disclosed above, and may operate in a distributed manner, for example. The at least one processor may execute program instructions to achieve or perform the various functions. Similarly, the at least one model may comprise a combination of circuits and / or processors that perform the various functions described / disclosed above, and may operate in a distributed manner, for example. The at least one processor and / or model may execute program instructions to achieve or perform the various functions.
[0055] The memory (230) may store various data used by at least one component (e.g., processor (210) and / or I / F (220)) of the electronic device (100). The data may include, for example, software (e.g., program) and input data or output data for commands related thereto. The memory (230) may include volatile memory and / or non-volatile memory. The memory (230) may include a hard disk, a ROM, a RAM (e.g., SRAM, PSRAM, or DRAM), a cache memory, and / or a register, and there is no limitation on its implementation. Some of the above-described entities (e.g., registers, but not limited thereto) may be implemented as a part of the processor (210), and there is no limitation on the form of their implementation. At least one AI model (e.g., LLM (240)) for instance execution may be stored in the memory (230).
[0056] The memory (230) can store at least one instruction. The processor (210) can execute at least one instruction stored in the memory (230). The at least one instruction, when executed by the processor (210), can cause the electronic device (100) to perform at least one operation. For example, as the at least one instruction is executed by the processor (210), at least one other component may be controlled, and / or various data processing or calculations may be performed. One operation performed by the processor (210) may mean, for example, that the operation is being performed by (or under the control of) one entity included in the processor (210) (for example, the main processor, but without limitation). That an operation is performed may mean, for example, that a particular operation is being performed by (or under the control of) multiple entities (e.g., multiple processors). That multiple operations are being performed may mean, for example, that all of the multiple operations are being performed by (or under the control of) a single entity (e.g., but not limited to, the main processor). That multiple operations are being performed may mean, for example, that some of the multiple operations are being performed by at least one entity, and some of the remaining operations are being performed by at least one other entity. At least one instruction causing the performance of one or more operations may be stored, for example, in a single memory, or may be stored distributedly in each of a plurality of memories.
[0057] In the electronic device (100), the LLM (240) may share resources (e.g., data processing or computational power) corresponding to part or all of at least one processor included in the processor (210) and / or resources (e.g., data recording area) corresponding to part or all of the memory (230). For example, the LLM (240) may be operated by at least one of the CPU (211), the NPU (213), or the GPU (215). The LLM (240) may be executed solely by the CPU (211), for example, by being allocated at least a portion of the memory (230). The LLM (240) may be executed solely by the NPU (213), for example, by being allocated at least a portion of the memory (230). The LLM (240) may be executed solely by the GPU (215), for example, by being allocated at least a portion of the memory (230). The LLM (240) can be performed by the CPU (211) and the NPU (213) in cooperation, for example, by being allocated at least a portion of the memory (230). The LLM (240) can be performed by the CPU (211) and the GPU (215) in cooperation, for example, by being allocated at least a portion of the memory (230). The LLM (240) can be performed by the NPU (213) and the GPU (215) in cooperation, for example, by being allocated at least a portion of the memory (230). The LLM (240) can be performed by the CPU (211), the NPU (213), and the GPU (215) in cooperation, for example, by being allocated at least a portion of the memory (230). The various embodiments to be described later in the present disclosure are not limited to the combination of components for performing the LLM (240) and can be implemented and / or applied based on any combination. The above LLM (240) will be described in detail with reference to the remaining drawings including FIGS. 3 and 4 below.
[0058] According to an example, the LLM (240) may include a target model (e.g., target model (310) of FIG. 3) and a plurality of draft models (e.g., draft model #1 (321-1), draft model #2 (321-2), draft model #3 (321-3), ..., draft model #N (321-M) of FIG. 3). The target model (310) may be relatively larger in size than the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). Here, M may be a positive integer. The plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have different sizes. The size of the draft models (321-1, 321-2, 321-3, ..., 321-M) may determine processing capabilities and / or latency. For example, a draft model with a larger size may generate tokens with a relatively higher probability (e.g., accuracy) of matching the information included in the prompt (250) compared to a draft model with a smaller size. However, a draft model with a larger size may have a relatively higher latency compared to a draft model with a smaller size. For example, the target model (310) may have a high accuracy due to a large amount of computational work and / or a large number of parameters that can be processed, but may have a slow processing speed, and the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have a fast processing speed due to a small amount of computational work and / or a small number of parameters, but may have low accuracy and may require supplementation. For this reason, the electronic device (100) employs an AI model based on speculative decoding, in which a target model (310) and a draft model operate in pairs. In one example, at least one target model (310) and multiple draft models (321-1, 321-2, 321-3, ...) are used.In the LLM (240) including the draft models (321-1, 321-2, 321-3, ..., 321-M), it is necessary to adaptively select a preferred draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) in order to provide speculative decoding for optimal processing power and / or delay speed. The preferred draft model may be paired with a target model for speculative decoding. The LLM (240) may determine the preferred draft model by considering the probability distribution of a specific number of output tokens generated by each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), for example, when performing speculative decoding to obtain an initial specific number (e.g., N) of tokens in response to a prompt input by a user.
[0059] According to one example, the LLM (240) can provide a prompt (250) corresponding to a user's input to the target model (310) and a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The LLM (240) can cause the target model (310) to generate a start token (or initial token) (Ts) in response to the input of the prompt (250). The LLM (240) can cause each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) to sequentially generate a specific number (e.g., 4) of output tokens based on the prompt (250) in response to the input of the start token.
[0060] According to an example, the LLM (240) can generate tokens corresponding to each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), and obtain probability values (e.g., the first probability values (q(1), q(2), q(3), ..., q(N)) of FIG. 4) which are probability information (e.g., the first probability information (411) of FIG. 4) of the generated tokens. The LLM (240) can obtain, for example, probability values of tokens (e.g., q(1-1), q(2-1), q(3-1), ..., q(N-1)) of the generated tokens corresponding to the first draft model (321-1). The LLM (240) can obtain, for example, probability values corresponding to the second draft model (321-2) (e.g., q(1-2), q(2-2), q(3-2), ..., q(N-2)). The LLM (240) can obtain, for example, probability values corresponding to the third draft model (321-3) (e.g., q(1-3), q(2-3), q(3-3), ..., q(N-3)). The LLM (240) can obtain, for example, probability values corresponding to the M-th draft model (321-M) (e.g., q(1-M), q(2-M), q(3-M), ..., q(NM)).
[0061] The above LLM can receive output tokens generated by each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) from the target model in response to the input of the start token (Ts), and perform verification on each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) based on the input tokens (e.g., output tokens generated by the draft models).
[0062] According to one example, the LLM (240) can obtain probability values (e.g., secondary probability values (p(1), p(2), p(3), ..., p(N)) of corresponding probability information (e.g., second probability information (421) of FIG. 4) of tokens generated for each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) by the target model (310). The LLM (240) can operate so that, for example, probability values (e.g., p(1-1), p(2-1), p(3-1), ..., p(N-1)) corresponding to the first draft model (321-1) are obtained by the target model (310). The LLM (240) may be operated so that, for example, probability values (e.g., p(1-2), p(2-2), p(3-2), ..., p(N-2))) corresponding to the second draft model (321-2) are obtained by the target model (310). The LLM (240) may be operated so that, for example, probability values (e.g., p(1-3), p(2-3), p(3-3), ..., p(N-3))) corresponding to the third draft model (321-3) are obtained by the target model (310). The above LLM (240) can operate so that, for example, probability values (e.g., p(1-M), p(2-M), p(3-M), .., p(NM)) corresponding to the M draft model (321-M) are obtained by the target model (310).
[0063] According to an example, the LLM (240) may determine a probability distribution for each draft model based on the probability values (e.g., the first probability information (411) of FIG. 4) determined for each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) (e.g., the first probability values (q(1), q2), q(3), ..., q(N) of FIG. 4)) of the probability information (e.g., the second probability values (p(1), p(2), p(3), ..., p(N) of FIG. 4)) of the second probability information (421) of FIG. 4) determined for the target model (310). The probability distribution may be, for example, the first probability values (q(1), q2), q(3), ..., determined by each draft model corresponding to the same output tokens. q(N)) and the secondary probability values (p(1), p(2), p(3), ..., p(N)) determined by the target model (310). The same output token can be obtained by the output tokens of each draft model (e.g., the primary output tokens (T) of FIG. 4). d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) among the output tokens of the target model (310) (e.g., the secondary output tokens (T) of FIG. 4 t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460)) may be an output token matching. The LLM (240) may determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is relatively similar to the target model (310) as a preferred draft model.
[0064] According to an example, the LLM (240) can select one or more draft models among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose output tokens are identical to the output tokens generated by the target model (310). The LLM (240) can determine a draft model having a probability distribution relatively similar to the target model among the selected one or more draft models as a preferred draft model. The probability distribution can be obtained, for example, by the first probability values (q(1), q2), q(3), ..., q(N)) determined by each draft model corresponding to the identical output tokens and the second probability values (p(1), p(2), p(3), ..., p(N)) determined by the target model (310). The identical output tokens are output tokens of each draft model (e.g., the first output tokens (T) of FIG. 4). d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) among the output tokens of the target model (310) (e.g., the secondary output tokens (T) of FIG. 4 t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460)) may be an output token matching. The LLM (240) may determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is relatively similar to the target model (310) as a preferred draft model.
[0065] The LLM (240) determines a preferred draft model among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) based on an initial specific number (e.g., N) of output tokens and a probability distribution of the specific number of output tokens in response to an input of a prompt (250), and then can generate the remaining tokens based on speculative decoding by pairing the preferred draft model with the target model (310). For example, if the LLM (240) determines a preferred draft model in response to the prompt (250), it can maintain the existing model pair (e.g., the target model (310) and the preferred draft model) without changing it until all remaining tokens are generated.
[0066] The I / F (220) may receive a prompt (250) corresponding to an input from a user, and transmit the received prompt (250) to the processor (210). The prompt (250) may be a medium that guides a generative AI (e.g., LLM (240)) to perform a task or generate a result in a desired direction. The prompt (250) may be the only window through which the user can communicate with the LLM (240). The prompt (250) needs to be clear and specific in order to obtain an answer close to the desired result from the LLM (240). The I / F (220) may receive a result processed by the LLM (240) based on the prompt (250) from the processor (210), and output a response result (260) converted into a natural language that is a form that can be recognized by humans (e.g., voice or text). The above I / F (220) can input or output natural language in the form of voice and / or text, for example, through at least one component such as a keyboard, a touch panel, a display, and / or a speaker.
[0067] FIG. 3 is a block diagram of an LLM (e.g., LLM (240) of FIG. 2) mounted on an electronic device (e.g., electronic device (100) of FIG. 1) according to one embodiment.
[0068] Referring to FIG. 3, the LLM (240) may include a target model (310) and a draft model group (320). The LLM (240) may be configured as an on-device type in the electronic device (100), for example. In this case, the target model (310) and the draft model group (320) included in the LLM (240) may be included in the electronic device (100).
[0069] The above draft model group (320) may include a plurality of draft models (e.g., draft model #1 (321-1), draft model #2 (321-2), draft model #3 (321-3), ... draft model #N (321-M) of FIG. 3). The target model (310) may be relatively larger in size than the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). Some of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have, for example, the same size. The plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have, for example, different sizes. The above draft model #1 (321-1) may be relatively larger in size than, for example, the above draft model #2 (321-2). The above draft model #2 (321-2) may be relatively larger in size than, for example, the above draft model #3 (321-3). The sizes of the draft models (321-1, 321-2, 321-3, ..., 321-M) may determine processing power and / or latency. For example, a draft model with a larger size may generate tokens with a higher probability (e.g., accuracy) of matching information contained in a prompt (e.g., prompt (250) of FIG. 2) than a draft model with a smaller size. However, a draft model with a larger size may have a relatively higher latency than a draft model with a smaller size.
[0070] According to an example, the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have the same or similar structure and / or complexity, but draft models trained specifically for different specific data sets may be used.
[0071] According to an example, the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may be draft models having the same structure, complexity, and / or number of NN parameters as the target model (310), or to which quantization is strongly applied (e.g., 4-bit quantization).
[0072] The target model (310) and the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may be provided with a prompt (250) corresponding to a user's input. The target model (310) may generate a start token (e.g., a start token (Ts) of FIG. 4) in response to the input of the prompt (250). The start token may instruct the start of token generation and verification for analysis of the prompt (250) and / or understanding of natural language through the same. The start token may be used as the first input token of the target model (310) and the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M).
[0073] The plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) can continuously generate a specific number (N) (e.g., 4) of output tokens based on the prompt (250) in response to the start token. The generation of output tokens by the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) can be performed in the same time interval. The generation of output tokens by the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) can be performed in independent time intervals. The independent time intervals can be, for example, completely distinct time intervals such that no overlapping time intervals exist. The independent time intervals can be, for example, time intervals in which overlapping time intervals partially overlap but are not completely identical. Among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), N output tokens generated consecutively in some of the draft models may be the same. Among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), N output tokens generated consecutively in some of the draft models may be different from each other. For example, the fact that the output tokens of at least two draft models are different from each other may mean, for example, that some of the output tokens generated from each of the at least two draft models are the same. For example, the fact that the output tokens of at least two draft models are different from each other may mean, for example, that the output tokens generated from each of the at least two draft models are not all the same.
[0074] The above-described plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may predict one or more candidate tokens for an input token, determine probability information for each of the one or more candidate tokens, and output one candidate token as an output token by considering the determined probability information for each candidate token. For example, the probability information may be determined by a probability value that quantifies the likelihood (e.g., accuracy) that the candidate token matches the information included in the prompt (250). The probability information of the candidate tokens may be different from each other. Even if the same candidate token is output from at least two draft models among the above-described plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), the at least two draft models may differently predict different probability information for the corresponding candidate token. At this time, a relatively high probability value generated by the draft model may indicate a high probability that the candidate token is accurate, while a relatively low probability value generated by the draft model may indicate a low probability that the candidate token is accurate. These probability values may indicate the accuracy of the tokens generated by the trained draft model.
[0075] The reason why the output tokens do not match even though the same prompt (250) is input to the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) is because there is a difference in the performance (difference in accuracy or size) of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). Despite the difference in the performance of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), the fact that the same output token is generated may be because the performance difference does not affect the prediction of the next output token by a specific input token.
[0076] The above plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) can use the output token as an input token of the next stage. Accordingly, the above plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) can initiate an operation in response to a start token, and then sequentially perform an operation of obtaining the next output token using the preceding output token as an input token until N output tokens are obtained.
[0077] The target model (310) can receive as input N output tokens output from each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The target model (310) can generate N output tokens by receiving as input N output tokens output from each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The N output tokens can be generated for each draft model. For example, the target model (310) can perform an operation of generating N output tokens by receiving as input N output tokens output from each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) in the same time period. The same time period may mean, for example, a time period that completely overlaps with each other. For example, the target model (310) may perform an operation of generating N output tokens by inputting N output tokens output by each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) in an independent time interval. The independent time interval may be, for example, a completely distinct time interval such that no overlapping time intervals exist. The independent time interval may be, for example, a time interval in which overlapping time intervals partially overlap but are not completely identical. The target model (310) may, for example, sequentially input N output tokens output from each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) and sequentially generate N output tokens corresponding to each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) at completely different time intervals. This operation may be performed by sequentially inputting N output tokens output from each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)., 321-M) is the same number of output tokens, and it can be implemented by reading and managing Kv cache (key-value cache) data corresponding to each input from memory (e.g., memory (230) of FIG. 2) for each draft model. For example, in generating tokens, Kv cache data can refer to a cache system that stores intermediate results to perform a designated task relatively quickly. The target model (310) stores the key and the corresponding value as data, and in the attention structure when the model operates, the key and value can be quickly found by pre-storing the key and value values of tokens to be operated in the future.
[0078] Among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), N output tokens generated consecutively in some of the draft models may be the same. Among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), N output tokens generated consecutively in some of the draft models may be different from each other. For example, the fact that the output tokens of at least two draft models are different from each other may mean, for example, that some of the output tokens generated from each of the at least two draft models are the same. For example, the fact that the output tokens of at least two draft models are different from each other may mean, for example, that the output tokens generated from each of the at least two draft models are not the same.
[0079] The electronic device (100) may include a predetermined component that performs a function of determining a preferred draft model to be used for speculative decoding among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The predetermined component may be, for example, an LLM (240). The predetermined component may be, for example, a specific function module included in the LLM (240). The predetermined component may be, for example, a specific function module by a combination of at least one or more of a CPU (e.g., CPU (211) of FIG. 2), an NPU (e.g., NPU (213) of FIG. 2), or a GPU (e.g., GPU (215) of FIG. 2) capable of processing other than the LLM (240) in the electronic device (100).
[0080] As an example, the predetermined component may determine a probability distribution for each draft model based on the probability values (e.g., the first probability information (411) of FIG. 4) determined for each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) (e.g., the first probability values (q(1), q2), q(3), ..., q(N) of FIG. 4)) of the probability information (e.g., the second probability values (p(1), p(2), p(3), ..., p(N) of FIG. 4)) of the second probability information (421) of FIG. 4) determined for the target model (310). The probability distribution may be, for example, the first probability values (q(1), q2), q(3), ..., q(N)) determined by each draft model in response to the same output tokens and the target The second probability values (p(1), p(2), p(3), ..., p(N)) determined by the model (310) and q(N) can be obtained. The same output token is the output tokens of each draft model (e.g., the first output tokens (T) of FIG. 4). d_out #1 , T d_out #2 , T d_out #3 , ..., Td_out #N )(440) or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) among the output tokens of the target model (310) (e.g., the secondary output tokens (T) of FIG. 4 t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460)) may be an output token matching the target model (310). The above-described component may determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is relatively similar to the target model (310) as a preferred draft model.
[0081] As an example, the LLM (240) can select one or more draft models among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose output tokens are identical to the output tokens generated by the target model (310). The LLM (240) can determine a draft model having a probability distribution relatively similar to the target model among the selected one or more draft models as a preferred draft model. The probability distribution can be obtained, for example, by the first probability values (q(1), q2), q(3), ..., q(N)) determined by each draft model corresponding to the identical output tokens and the second probability values (p(1), p(2), p(3), ..., p(N)) determined by the target model (310). The identical output tokens are output tokens of each draft model (e.g., the first output tokens (T) of FIG. 4). d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , .... Tt_in #N )(450)) among the output tokens of the target model (310) (e.g., the secondary output tokens (T) of FIG. 4 t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460)) may be an output token matching the target model (310). The above-described component may determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is relatively similar to the target model (310) as a preferred draft model.
[0082] The electronic device (100) determines a preferred draft model among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) based on an initial specific number (e.g., N) of output tokens and a probability distribution of the specific number of output tokens in response to an input of a prompt (250), and then can generate the remaining tokens based on speculative decoding by pairing the preferred draft model with the target model (310). For example, if the electronic device (100) determines a preferred draft model in response to the prompt (250), it can maintain the existing model pair (e.g., the target model (310) and the preferred draft model) without changing it until all remaining tokens are generated.
[0083] In the above description, the target model (310) and the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) are configured as on-device types in the electronic device (100), but the target model (310) may be operated by a server, and the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may be operated by the electronic device (100).
[0084] For example, the electronic device (100) may consider the operating status of the electronic device, such as the user, the application solution, the battery status, and / or the heat status, as criteria for selecting or determining a preferred draft model. The electronic device (100) may also apply criteria for selecting or determining the preferred draft model based on an AI model. In this case, the AI model may be applied within the range where the load for selecting or determining the preferred draft model does not exceed a critical level.
[0085] FIG. 4 is an exemplary operation diagram of configuring an LLM (e.g., LLM (240) of FIG. 2) for speculative decoding in an electronic device (e.g., electronic device (100) of FIG. 1) according to one embodiment. In FIG. 4, it is assumed that the m-th draft model (410) (draft model #m) is included in a plurality of draft models (e.g., a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3). Here, m is an integer greater than or equal to 1 and less than or equal to M, and M may be an integer greater than 1. The configuration and / or operation described below may be equally applied to the remaining draft models excluding the draft model #m (410) among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M).
[0086] Referring to FIG. 4, an LLM based on speculative decoding (e.g., LLM (240) of FIG. 2) may include multiple token flows. The multiple token flows may include pairs of at least one target model (420) (e.g., target model (310) of FIG. 3) and multiple draft models (e.g., multiple draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3). In the following, for the sake of convenience of explanation, one target model (420) will be described assuming this, but it should be understood that the configuration and / or operation described below may be equally applied to multiple target models. The multiple token flows may correspond to links through which output tokens are transmitted between the multiple draft models (321-1, 321-2, 321-3, ..., 321-M) and the target model (420). For example, in Fig. 4, N output tokens (T) generated by draft model #m (410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) are N input tokens (T) for the target model (420). t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) is illustrated. Here, N may be a positive integer. The draft model #m(410) may be included in the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The token flow may include a plurality of paths that connect N output ports included in the draft model #m(410) and N input ports included in the target model (420) one-to-one. The connection between the output ports of the draft model #m(410) and the input ports of the target model (420) is N output tokens (T d_out #1 , T d_out #2 , Td_out #3 , ..., T d_out #N )(440) (hereinafter referred to as “primary output tokens”) and N input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) (hereinafter referred to as “secondary input tokens”) token separation index (#x, 0 <x≤N)에 의해 식별될 수 있다. 일 예로, 상기 드래프트 모델 #m(410)과 상기 타깃 모델(420)을 연결하는 토큰 플로우 #m은 상기 드래프트 모델 #m(410)에서 제1 출력 토큰(T d_out #1 ) and the first input token (T) in the target model (420). t_in #1 ) may include a first path connecting the input ports to which the first output token (T) is to be input. d_out #1 ) are the first output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) may be one of the first input token (T t_in #1 ) are the secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) may be one of them. The remaining paths included in the token flow #m may be defined based on the token separation index as described above.
[0087] The above draft model #m (410) can input a prompt corresponding to a user input (e.g., prompt (250) of FIG. 2). The target model (420) can input a prompt corresponding to a user input. For example, the same prompt can be input to the draft model #m (410) and the target model (420). For example, the prompt input to the target model (420) can be processed or arbitrarily processed and input to the draft model #m (410). For example, the prompt input to the draft model #m (410) can be processed or arbitrarily processed and input to the target model (420).
[0088] A start token (Ts) may be input into a first input port among a plurality of input ports included in the draft model #m (410). A start token (Ts) may be input into a first input port among a plurality of input ports included in the target model (420). For example, the start token (Ts) may be a command or identifier that instructs to initiate token generation in response to a user input of a prompt (250). The start token may be generated, for example, by the target model (420) in response to an input of the prompt (250). The draft model #m (410) and / or the target model (420) may initiate token generation in response to an input of the start token (Ts).
[0089] The above draft model #m (410) generates N input tokens (T) based on the prompt (250) when the start token (Ts) is input. d_in #1 , T d_in #2 , T d_in #3 , ..., T d_in #N-1 )(430) (hereinafter referred to as “primary input tokens”) corresponding to N primary output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., Td_out #N )(440) can be created.
[0090] According to an example, the above draft model #m (410) responds to the input of the start token (Ts) by generating the first output token (hereinafter referred to as “first output token (T d_out #1 ) can be generated. The first output token (T d_out #1 ) are the first output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440). The draft model #m(410) may, for example, obtain a plurality of first candidate output tokens and probability information on the plurality of first candidate output tokens in response to the input of the start token (Ts). The probability information may include a probability value that quantifies the likelihood (e.g., accuracy) that each of the plurality of first candidate output tokens obtained from the draft model #m(410) matches the information included in the prompt (250), or may be determined by the probability value. According to an example, the draft model #m(410) selects one first candidate output token among the plurality of first candidate output tokens, which has a relatively high likelihood of matching the information included in the prompt (250), i.e., a high accuracy, as the first output token (T d_out #1 ) can be determined (e.g., see Fig. 6).
[0091] The above draft model #m(410) is the first output token (T d_out #1 ) as an input token for predicting other tokens (hereinafter referred to as “first input token (T d_in #1 ) can be used as the first input token (T d_in #1 ) are the first input tokens (T d_in #1 , T d_in #2 , T d_in #3 , ..., T d_in #N-1)(430) may be one of the above draft model #m(410), for example, the first input token (T d_in #1 ) in response to the input of the draft model #m (410), a plurality of second candidate output tokens and probability information on the plurality of second candidate output tokens can be obtained. The probability information may include a probability value (e.g., q(1), q2), q(3), ..., q(N), where N is a positive integer greater than or equal to 2) that quantifies the possibility (e.g., accuracy) that each of the plurality of second candidate output tokens obtained from the draft model #m (410) matches the information included in the prompt (250), or may be determined by the probability value. According to an example, the draft model #m (410) selects one second candidate output token among the plurality of second candidate output tokens, which has a relatively high possibility of matching the information included in the prompt (250), i.e., a high accuracy, as the second output token (T d_out #2 ) can be determined (e.g., see Fig. 6).
[0092] As described above, the draft model #m (410) can be used as a primary input token to generate the next secondary output token that follows the previously generated primary output token. In addition, the draft #m (410) can be used as a primary input token to generate the remaining primary input tokens (T d_in #3 , ..., T d_in #N ) based on probability information for the first output tokens T d_out #3 , ..., T d_out #N ) can be generated. As a result, the draft model #m (410) generates N primary input tokens (T) based on the prompt (250). d_in #1 , T d_in #2 , T d_in #3 , ..., T d_in #N-1 )(430) corresponding to N primary output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N)(440) can be created.
[0093] The above draft model #m(410) is the N primary output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) can confirm, obtain, or determine the probability information (hereinafter referred to as “first probability information (411)”) corresponding to the first probability information (411). In the following, for the purpose of convenience of explanation, the probability information will be described as “determining”, but it should not be limited thereto. The draft model #m(410) outputs the first probability information (411) as primary output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) include probability values (q(1), q2), q(3), ..., q(N), where N is a positive integer greater than or equal to 2) (hereinafter referred to as “primary probability values”), or can be determined by the probability values. The primary probability values (q(1), q2), q(3), ..., q(N)) include the primary output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) can be obtained by quantifying the probability (e.g., accuracy) that each of the above prompts (250) matches the corresponding information. As an example, the first probability values (q(1), q2), q(3), ..., q(N)) are the first input tokens (T d_in #1 , T d_in #2 , T d_in #3 , ..., T d_in #N-1 )(430) may correspond to the relatively highest probability value among the candidate output tokens obtained in response to each.
[0094] The above draft model #m (410) generates the first output tokens (T) through the corresponding token flow. d_out #1 , T d_out #2, T d_out #3 , ..., T d_out #N )(440) are the secondary input tokens (T) for the target model (420). t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450). As an example, the first output tokens (T) sequentially generated by the draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) can be sequentially input to the target model (420). As an example, the first output tokens (T) sequentially generated by the draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) can be simultaneously input to the target model (420) through a predetermined buffering. As an example, the first output tokens (T) continuously generated by the draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) Some of the output tokens may be sequentially input to the target model (420), and the remaining output tokens may be simultaneously input to the target model (420) through a predetermined buffering.
[0095] The configuration and / or operation described above for the draft model #m (410) can be equally applied to the remaining draft models among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) included in the LLM (240), and therefore, the description of the configuration and / or operation for the remaining draft models is omitted in this disclosure.
[0096] The above target model (420) generates secondary input tokens (T) based on the prompt (250) when the start token (Ts) is input. t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) can be verified. The target model (420) can be verified by the secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450), by performing verification on N secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) can be generated. The secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) are, for example, the primary output tokens (T) generated by the above draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440)(or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) can be completely matched with the above secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) are, for example, the primary output tokens (T) generated by the above draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440)(or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N)(450)) may partially match the secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) are, for example, the primary output tokens (T) generated by the above draft model #m(410) d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440)(or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) may not be completely consistent.
[0097] The target model (420) may receive the primary output tokens generated by each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) as secondary input tokens at the same time or at different times. For example, the target model (420) may generate secondary output tokens corresponding to each of the secondary input tokens provided by the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M), and may not select a draft model for which the secondary output tokens do not match the secondary input tokens (or the primary output tokens) for speculative decoding. The unselected draft model may be eliminated from being determined as a preferred draft model. For a draft model eliminated from being determined as a preferred draft model, subsequent operations may no longer be performed. In this case, one or more of the dropped draft models may not reside in memory (e.g., memory (230) of FIG. 2). For example, the dropped draft models may not use the resources of the memory (230).
[0098] According to an example, the target model (420) responds to the input of the start token (Ts) by generating secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) can obtain a plurality of candidate output tokens and probability information for the plurality of candidate output tokens. The probability information may include a probability value that quantifies the possibility (e.g., accuracy) that each of the plurality of candidate output tokens obtained for each secondary input token in the target model (420) matches the information included in the prompt (250), or may be determined by the probability value. According to an example, the target model (420) selects one candidate output token (e.g., the candidate output token with the highest probability value) among the candidate output tokens for each secondary input token that is relatively likely to match the information included in the prompt (250), that is, has a high accuracy, as the corresponding secondary input token (T t_in #n ) corresponding to the secondary output token (T t_out #n ) can be determined (e.g., see Fig. 6). Here, n can be a natural number between 1 and N. The target model (420) is the secondary input tokens (T) input by the draft model #m (410). t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) corresponding to the secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) can be determined.
[0099] The above target model (420) is the secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N)(460) can confirm, obtain, or determine the probability information (hereinafter referred to as “second probability information (421)”) corresponding to the target model (420). In the following, for the purpose of convenience of explanation, the probability information will be described as “determining”, but it should not be limited thereto. The target model (420) outputs the second probability information (421) to secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) include probability values (p(1), p(2), p(3), ..., p(N), where N is a positive integer greater than or equal to 2) (hereinafter referred to as “secondary probability values”), or can be determined by the probability values. The second probability values (p(1), p(2), p(3), ..., p(N)) may be determined by the second output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) can be obtained by quantifying the probability (e.g., accuracy) that each of the prompts (250) matches the corresponding information. As an example, the secondary probability values (p(1), p(2), p(3), ..., p(N)) are obtained by quantifying the probability (e.g., accuracy) that each of the secondary input tokens (T) matches the corresponding information contained in the prompt (250). t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450) may correspond to the relatively highest probability value among the candidate output tokens obtained in response to each.
[0100] According to an example, the LLM (240) can determine a probability distribution corresponding to the draft model #m based on the first probability values (q(1), q2), q(3), ..., q(N)) of the first probability information (411) and the second probability values (p(1), p(2), p(3), ..., p(N)) of the second probability information (421). The probability distribution can be obtained, for example, by the first probability values (q(1), q2), q(3), ..., q(N)) determined by the draft model #m (410) corresponding to the same output tokens and the second probability values (p(1), p(2), p(3), ..., p(N)) determined by the target model (420). The same output tokens are the first output tokens (T d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N )(440) (or secondary input tokens (T t_in #1 , T t_in #2 , T t_in #3 , ..., T t_in #N )(450)) among the secondary output tokens (T t_out #1 , T t_out #2 , T t_out #3 , ..., T t_out #N )(460) may be an output token matching. The LLM (240) may determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is relatively similar to the target model (420) as a preferred draft model.
[0101] According to an example, the LLM (240) can select one or more draft models among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) in which the first output tokens are the same as the second output tokens generated by the target model (420). The LLM (240) can determine a draft model having a probability distribution relatively similar to the target model among the selected one or more draft models as a preferred draft model. The probability distribution can be determined based on the first probability values (q(1), q2), q(3), ..., q(N)) and the second probability values (p(1), p(2), p(3), ..., p(N)). The first probability values (q(1), q2), q(3), ..., q(N)) are probability values determined for each of the first output tokens in the selected one or more draft models. The above secondary probability values (p(1), p(2), p(3), ..., p(N)) are probability values determined for each of the secondary output tokens in the target model (420).
[0102] For example, the LLM (240) may determine a preferred draft model among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) based on an initial specific number (e.g., N) of output tokens and a probability distribution of the specific number of output tokens in response to an input of a prompt (250), and may then generate the remaining tokens based on speculative decoding by pairing the preferred draft model with the target model (420). For example, if the LLM (240) determines a preferred draft model in response to the prompt (250), it may maintain the existing model pair (e.g., the target model (420) and the preferred draft model) without changing it until all remaining tokens are generated.
[0103] FIG. 5 is an exemplary operation diagram in which a specific draft model (e.g., draft model #m (410) of FIG. 4) included in a plurality of draft models (e.g., draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3) is not determined as a preferred draft model to be paired with a target model (e.g., target model (310) of FIG. 3) in an electronic device (e.g., electronic device (100) of FIG. 1) equipped with an AI model based on guess decoding according to one embodiment. In FIG. 5, it is assumed that the specific draft model is the m-th draft model (410) (draft model #m) included in the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). Here, m is an integer greater than or equal to 1 and less than or equal to M, and M may be an integer greater than 1. The configuration and / or operation to be described below The same can be applied to the remaining draft models except for the draft model #m (410) among the above multiple draft models (321-1, 321-2, 321-3, ..., 321-M).
[0104] Referring to FIG. 5, the draft model #m (410) responds to the input of a start token (e.g., the start token (Ts) of FIG. 4) by outputting a specific number (N) (e.g., 4) of consecutive output tokens (e.g., the first output tokens (T) of FIG. 4) based on a prompt (e.g., I go to school.) corresponding to the user's input (e.g., the prompt (250) of FIG. 2). d_out #1 , T d_out #2 , T d_out #3 , ..., T d_out #N)(440)) can be generated. Here, m may be a natural number between 1 and M. The output token generated by the draft model #m(410) may be referred to as a 'preferred output token'. For example, the draft model #m(410) generates the first output token "I(511)" using the start token (Ts) as the first input token, generates the second output token "go(513)" using the first output token "I(511)" as the second input token, generates the third output token "to(515)" using the second output token "go(513)" as the third input token, and generates the fourth output token "bed(517)" using the third output token "to(515)" as the fourth input token. Thus, the specific number of output tokens may be 'I(511)', 'go(513)', 'to(515)', and 'bed(517)'.
[0105] For example, when the start token (Ts) is input, the draft model #m (410) may obtain at least one candidate output token predicted as the first token based on the prompt (250), and determine a preferred output token from among the at least one candidate output token. The draft model #m (410) may, for example, determine a probability value for the at least one candidate output token, and determine a preferred output token based on the determined probability value. The probability value may be a numerical representation of the likelihood (e.g., accuracy) that each of the at least one candidate output tokens matches the corresponding information included in the prompt (250). The draft model #m (410) may determine a candidate output token having the highest probability value from among the at least one candidate output token as a preferred output token. For example, the draft model #m (410) may use the third generated output token 'to' (515) as an input token, and obtain "bed", "school", and "the" as candidate output tokens. The above draft model #m (410) can determine “80%”, “10%”, and “10%” as probability values (q(x)) (510) corresponding to the obtained candidate output tokens “bed”, “school”, and “the”, respectively. The above draft model #m (410) can determine “bed”, which has the highest probability value (q(x)) (510), among the obtained candidate output tokens “bed”, “school”, and “the”, as the preferred output token. Here, the preferred output token is determined corresponding to one input token, but the remaining preferred output tokens “I”, “go”, or “to” can also be determined in the same manner.
[0106] The target model (420) can perform a verification operation on a specific number (N) (e.g., 4) of output tokens generated for each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) in response to an input of a start token (e.g., a start token (Ts) of FIG. 4). For example, the target model (420) can perform a verification process of calculating a correlation between the initial token and the N input token values provided by each of the N input tokens, and generating N tokens with the highest probability. For example, the target model (420) can perform a verification process for each of the draft models (321-1, 321-2, 321-3, ..., 321-M). In the following, it will be described that the target model (420) performs a verification operation on one draft model (e.g., the mth draft model (410), where m is a natural number between 1 and M), but substantially the same verification operation can be performed on the remaining draft models.
[0107] According to an example, the target token (420) may generate output tokens corresponding to tokens output from each of a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) and input in a batch form. The target model (420) may use the first output tokens generated by the draft model as input tokens to generate second output tokens.
[0108] For example, the target model (420) may generate preferred output tokens (531, 533, 535, 537, 539) corresponding to input tokens (521, 523, 525, 527) using the start token (Ts) and “I (511)”, “go (513)”, “to (515)”, or “bed (517)” generated by the draft model #m (410). The preferred output tokens generated by the target model (420) may be, for example, “I (531)”, “go (533)”, “to (535)”, “school (537)”, and “. (539)”. The target model (420) can determine output token probabilities (p(“I”), p(“go”), p(“to”), p(“bed”), p(“.”)) corresponding to each of the preferred output tokens (531, 533, 535, 537, 539). For example, the target model (420) can determine “30%”, “50%”, and “20%” as probability values (p(x)) (520) corresponding to candidate output tokens “bed”, “school”, and “the” obtained for the input token “to (525)”, respectively.
[0109] Since the input tokens of the target model (420), “I(511)”, “go(513)”, “to(515)”, and “bed(517)”, and the output tokens of the target model (420), “I(521)”, “go(523)”, “to(525)”, and “school(527)”, do not match, the draft model #m(410) may be eliminated from being selected as the preferred draft model. However, if the input tokens and the output tokens of the target model (420) match as “I(521)”, “go(523)”, “to(525)”, and “school(527)”, the draft model #m(410) may be a candidate to be selected as the preferred draft model.
[0110] As described above, all tokens generated from the target model (420) and each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have their own probability values. The probability values of the tokens for each draft model obtained by the target model (420) and the probability values of the tokens obtained from the corresponding draft model may be compared, and the draft model with the largest number of identical tokens may be selected. The sum of the probability distributions for the selected draft models may be obtained for each draft model. The draft model having the smallest difference between the sum of the probability distributions obtained for each draft model and the sum of the probability distributions obtained by the target model (420) (e.g., various matrices such as the difference of the squares or the difference of the absolute values) may be determined as the preferred draft model.
[0111] FIG. 6 is an exemplary diagram of a judgment operation for determining a preferred draft model to be paired with a target model (e.g., target model (310) of FIG. 3) in a draft model group (e.g., draft model group (320) of FIG. 3) in an electronic device (e.g., electronic device (100) of FIG. 1) equipped with an AI model (e.g., LLM (240) of FIG. 2) based on guess decoding) according to one embodiment. In FIG. 6, it is assumed that the draft model group (320) includes three draft models (e.g., draft model #1 (321-1), draft model #2 (321-2), and draft model #3 (321-3) of FIG. 3). The draft model group (320) may be, for example, a set of target draft models that may be determined as preferred draft models. Although the configuration and / or operation described below is not included in the above draft model group (320), it can be equally applied to the remaining draft models included in the above plurality of draft models (321-1, 321-2, 321-3, ..., 321-M).
[0112] Referring to Fig. 6, draft model #1 (321-1) sequentially obtained four output tokens (T1-1, T1-2, T1-3, T1-4) of “the”, “school”, “bus”, and “is” in response to a prompt (e.g., “The school bus is ~”). The draft model #1 (321-1) can determine the first probability values (P1-1, P1-2, P1-3, P1-4) corresponding to the obtained output tokens “the”, “school”, “bus”, and “is” as “0.1”, “0.5”, “0.1”, and “0.3”, respectively.
[0113] The target model (310) can obtain four output tokens (T1'-1, T1'-2, T1'-3, T1'-4), namely "the", "school", "bus", and "is", by taking as input the tokens (T1-1, T1-2, T1-3, T1-4) output by the draft model #1 (321-1), namely "the", "school", "bus", and "is". Since the output tokens generated by the draft model #1 (321-1) and the output tokens generated by the target model (310) match, the draft model #1 (321-1) can be determined as a candidate that can be selected as a preferred draft model. The target model (310) above can determine the probability values (P1'-1, P1'-2, P1'-3, P1'-4) corresponding to each of the output tokens “the”, “school”, “bus”, and “is” generated by using the output tokens of the draft model #1 (321-1) as input as “0.2”, “0.3”, “0.4”, and “0.6”.
[0114] Draft model #2 (321-2) sequentially acquired four output tokens (T2-1, T2-2, T2-3, T2-4) of “the”, “school”, “bus”, and “is” in response to a prompt (e.g., “The school bus is ~”). Draft model #2 (321-2) can determine the first probability values (P2-1, P2-2, P2-3, P2-4) corresponding to the acquired output tokens “the”, “school”, “bus”, and “is” as “0.25”, “0.35”, “0.35”, and “0.62”, respectively.
[0115] The target model (310) can obtain four output tokens (T2'-1, T2'-2, T2'-3, T2'-4) of "the", "school", "bus", and "is" by taking as input the tokens "the", "school", "bus", and "is" output by the draft model #2 (321-2). Since the output tokens generated by the draft model #2 (321-2) and the output tokens generated by the target model (310) match, the draft model #2 (321-2) can be determined as a candidate that can be selected as the preferred draft model. The target model (310) above can determine the probability values (P2'-1, P2'-2, P2'-3, P2'-4) corresponding to each of the output tokens “the”, “school”, “bus”, and “is” generated by using the output tokens of the draft model #2 (321-2) as input as “0.3”, “0.3”, “0.35”, and “0.5”.
[0116] Draft model #3 (321-3) sequentially acquired four output tokens (T3-1, T3-2, T3-3, T3-4) of “the”, “boy”, “bus”, and “building” in response to a prompt (e.g., “The school bus is ~”). Draft model #3 (321-3) can determine the first probability values (P3-1, P3-2, P3-3, P3-4) corresponding to the acquired output tokens “the”, “boy”, “bus”, and “building” as “0.1”, “0.5”, “0.2”, and “0.3”, respectively.
[0117] The target model (310) can obtain four output tokens (T3'-1, T3'-2, T3'-3, T3'-4) of "the", "school", "bus", and "goes" by taking as input the tokens "the", "boy", "bus", and "building" output by the draft model #3 (321-3). Since the output tokens generated by the draft model #3 (321-3) and the output tokens generated by the target model (310) do not match, the target model (310) can decide to exclude the draft model #3 (321-3) from the preferred draft models. The target model (310) above can determine the probability values (P3'-1, P3'-2, P3'-3, P3'-4) corresponding to each of the output tokens “the”, “school”, “bus”, and “goes” generated by using the output tokens of the draft model #3 (321-3) as input as “0.2”, “0.3”, “0.4”, and “0.1”.
[0118] The electronic device (100) may include a predetermined component that performs a function of determining a preferred draft model among the three draft models (e.g., draft model #1 (321-1), draft model #2 (321-2), draft model #3 (321-3)). The predetermined component may be, for example, an LLM (e.g., LLM (240) of FIG. 2). The predetermined component may be, for example, a specific function module included in the LLM (240). The predetermined component may be, for example, a specific function module by a combination of at least one or more of a CPU (e.g., CPU (211) of FIG. 2), an NPU (e.g., NPU (213) of FIG. 2), or a GPU (e.g., GPU (215) of FIG. 2) capable of processing other than the LLM (240) in the electronic device (100).
[0119] For example, the predetermined component may perform a verification process of comparing the probability values of the same tokens “the”, “school”, “bus”, and “is” output from the target model (310) and a plurality of candidate draft models (e.g., draft model #1 (321-1) and draft model #2 (321-2)) in speculative decoding. The plurality of candidate draft models may be, for example, draft model #1 (321-1) and draft model #2 (321-2) among three draft models (321-1, 321-2, 321-3) that generated the same output tokens as the target model (310).
[0120] For example, the verification operation for the draft model #1 (321-1) is performed by comparing the probability values (P1-1, P1-2, P1-3, P1-4) of “0.1”, “0.5”, “0.1”, and “0.3” determined by the draft model #1 (321-1) for the same output tokens (T1-1, T1-2, T1-3, T1-4) of “the”, “school”, “bus”, and “is”, respectively, with the probability values (P1'-1, P1'-2, P1'-3, P1-4) of “0.2”, “0.3”, “0.4”, and “0.6” determined by the target model (310) for the same output tokens (T1'-1, T1'-2, T1'-3, T1'-4) of “the”, “school”, “bus”, and “is”, respectively. Square each difference value (① (0.2-0.1) 2 , ② (0.3-0.5) 2 , ③ (0.4-0.1) 2 ,④ (0.6-0.3) 2 ) and the sum of the above square values (①+②+③+④) can be obtained as 0.23 as a probability distribution corresponding to the above draft model #1 (321-1).
[0121] For example, the verification operation for the draft model #2 (321-2) is performed by determining the probability values (P2-1, P2-2, P2-3, P2-4) of “the”, “school”, “bus”, and “is”, which are the same output tokens (T2-1, T2-2, T2-3, T2-4) as “0.25”, “0.35”, “0.35”, and “0.62”, and the probability values (P2'-1, P2'-2, P2'-3, P2-4) of “the”, “school”, “bus”, and “is”, which are the same output tokens (T2'-1, T2'-2, T2'-3, T2'-4) as “0.3”, “0.3”, “0.35”, and “0.62”, which are the same output tokens (T2'-1, T2'-2, T2'-3, T2'-4) as “the”, “school”, “bus”, and “is”, which are the same output tokens (T2'-1, T2'-2, T2'-3, T2'-4) as “0.3”, “0.3”, “0.35”, and Square each of the difference values of “0.5” (① (0.3-0.25) 2, ② (0.3-0.35) 2 , ③ (0.35-0.35) 2 ,④ (0.5-0.62) 2 ) and the sum of the above square values (①+②+③+④) can be obtained as 0.0194 as a probability distribution corresponding to the above draft model #2 (321-2).
[0122] For example, the verification operation for the draft model #2 (321-2) may determine the candidate draft model with the smallest probability distribution difference among the probability distributions obtained for a plurality of candidate draft models as the preferred draft model. The smallest probability distribution means that the draft model is determined to be most similar to the target model (310). According to the example above, since the probability distribution corresponding to draft model #2 (321-2) has a relatively small value of '0.0194' compared to the probability distribution corresponding to draft model #1 (321-1), '0.23', the draft model #2 (321-2) may be paired with the target model (310) to perform the guess decoding operation.
[0123] FIG. 7 is an exemplary diagram of an operation of generating tokens based on guess decoding by pairing a preferred draft model determined from a plurality of draft models (e.g., a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3) and a target model (e.g., a target model (310) of FIG. 3) in an electronic device (e.g., an electronic device (100) of FIG. 1) according to one embodiment.
[0124] Referring to FIG. 7, the preferred draft model (710) can take as input the target model (720) and the same prompt (780) corresponding to the user input (e.g., the prompt (250) of FIG. 2). The preferred draft model (710) can be determined by a specific number (e.g., M) of initial tokens generated by a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) and the probability values of the initial tokens.
[0125] The preferred draft model (710) may generate a specific number of tokens (t1, t2, t3) sequentially from the initial tokens based on the prompt (250). At this time, the number of tokens generated by the preferred draft model (710) may be the same as or different from the number (M) of the initial tokens. Since the operation of the preferred draft model (710) generating the specific number of tokens (t1, t2, t3) is substantially the same as the operation for generating the initial token, further description thereof is omitted. The preferred draft model (710) may provide the generated tokens (t1, t2, t3) as inputs to the target model (720) (operation 740).
[0126] The target model (720) can generate tokens (t1', t2', t3') corresponding to tokens (t1, t2, t3) input by the preferred draft model (710). The tokens (t1, t2, t3) generated by the preferred draft model (710) can be verified through comparison with the tokens (t1', t2', t3') generated by the target model (720). If there is a token among the tokens (t1, t2, t3) generated by the preferred draft model (710) that does not match the tokens (t1', t2', t3') generated by the target model (720), the corresponding token (e.g., t3) can be corrected to a token predicted to be correct (e.g., t3'_correct) (operation 750).
[0127] Information related to the above error correction may be managed by a buffer (730). The buffer (730) may request an update of the preferred draft model (710) when the buffer size exceeds a threshold level (operation 770).
[0128] Information related to the above error correction may be provided to the preferred draft model (710) (operation 760). The information related to the error correction may include information on the target token to be corrected and / or the token to be corrected. The preferred draft model (710) may perform correction on the corresponding token based on the information related to the error correction. The preferred draft model (710) may resume generation of consecutive tokens using the corrected token as a new input token.
[0129] FIG. 8 is an exemplary diagram of an operation of generating tokens based on guess decoding in an electronic device (e.g., the electronic device (100) of FIG. 1), according to one embodiment.
[0130] Referring to FIG. 8, the preferred draft model (710) can take as input the same prompt (e.g., prompt (250) of FIG. 2) corresponding to the target model (720) and the user input. The preferred draft model (710) can be determined by a specific number (e.g., M) of initial tokens generated by multiple draft models (321-1, 321-2, 321-3, ..., 321-M) and the probability values of the initial tokens.
[0131] The preferred draft model (710) may generate a specific number (e.g., five) of tokens (711, 713, 715, 717, 719) (e.g., “I”, “go”, “to, “bed”, “.”) based on the prompt (250) and transmit the generated tokens to the target model (720). The target model (720) may perform verification (721, 723, 725, 727) on the specific number (e.g., five) of tokens (711, 713, 715, 717, 719) (e.g., “I”, “go”, “to, “bed”, “.”) input from the preferred draft model (710), and may select one of the specific number (e.g., five) of tokens (711, 713, 715, 717, 719). It can be recognized that an error correction is required for the token (717) “bed.” The target model (720) can request the preferred draft model (710) to correct the token (717) “bed” in which the error occurred to “the.”
[0132] The preferred draft model (710) above can correct the last token (717), “bed,” to “the” based on the error correction-related information. The preferred draft model (710) can generate consecutive tokens using the error-corrected token as a new input token and input them into the target model (720).
[0133] The preferred draft model (710) can generate a specific number (e.g., five) of tokens (731, 733, 735, 737, 739) (e.g., “cam”, “pus”, “.”, “It”, “was”) consecutive to the token (e.g., the corrected token, “the”) previously input to the target model (720) based on the prompt (250) and transfer the generated tokens to the target model (720). The target model (720) can perform verification (741, 743, 745, 747, 749) on the specific number (e.g., five) of tokens (731, 733, 735, 737, 739) (e.g., “cam”, “pus”, “.”, “It”, “was”) input from the preferred draft model (710).
[0134] The preferred draft model (710) and target model (720) of the above pair can perform a speculative decoding operation until the generation (e.g., execution) of all tokens corresponding to the prompt (250) is completed.
[0135] FIG. 9 is a control flowchart for performing speculative decoding in an electronic device (e.g., the electronic device (100) of FIG. 1), according to one embodiment.
[0136] Referring to FIG. 9, the electronic device (100) may, in operation 910, receive a prompt (e.g., prompt (250) of FIG. 2) input by a user. The prompt (250) may be, for example, a command or question that the user transmits to the LLM (e.g., LLM (240) of FIG. 2) mounted on the electronic device (100) for conversation with the LLM. For example, the prompt (250) may serve as a communication window between the user and the LLM (240). The LLM (240) mounted on the electronic device (100) must be able to accurately analyze or understand the prompt (250) in order to provide the user with accurate information desired by the user. The above prompt (250) may be input as a plurality of draft models (e.g., a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3) included in the LLM (240) and at least one target model (e.g., a target model (310) of FIG. 3). The target model (310) may be relatively larger in size than the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have different sizes.
[0137] The electronic device (100) may, at operation 920, determine a preferred draft model to be used for speculative decoding among a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). In one example, the electronic device (100) may, in response to a start token, cause the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) to obtain a specific number (N) (e.g., 4) of first output tokens in succession based on the prompt (250). The start token may be generated by the target model (310) in response to an input of the prompt. The electronic device (100) may be operable to check first probability information that quantifies the likelihood (e.g., accuracy) that the first output tokens obtained for each draft model from the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) match the information included in the prompt (250).
[0138] The electronic device (100) may be operable to cause the target model (310) to obtain second output tokens for each draft model using the first output tokens as input. The second output tokens may, for example, be at least partially identical to the first output tokens. The second output tokens may, for example, be substantially identical to the second output tokens.
[0139] The electronic device (100) may be operable to check second probability information that quantifies the likelihood (e.g., accuracy) that the second output tokens obtained for each draft model from the target model (310) match the information included in the prompt (250). The electronic device (100) may be operable to determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) in which the first probability information is relatively similar to the second probability information compared to one or more other draft models as a preferred draft model.
[0140] According to one example, the electronic device (100) may be configured to determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) having a probability distribution similar to the target model (310)) as a preferred draft model. The probability distribution may be probability information determined by the probability values of identical output tokens (hereinafter referred to as “third output tokens”) that are overlapped and included in the first output tokens and the second output tokens, respectively, corresponding to each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)). As an example, the probability distribution may be obtained by squaring the difference value between the first probability value and the second probability value, and adding up the squared values calculated corresponding to the third output tokens. The first probability value may be determined by the corresponding draft model for each of the third output tokens. The second probability value can be determined by the target model (310) for each of the third output tokens. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select a draft model having a relatively smaller size from among the multiple preferred draft models. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select one of the multiple preferred draft models by considering status information of the electronic device (100). The status information can include information indicating a heating state and / or a battery charging state.
[0141] According to an example, the electronic device (100) may be operable to select one or more draft models among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) whose output tokens are identical to the target model (310) as at least one candidate draft model. The electronic device (100) may be operable to determine a draft model having a probability distribution similar to the target model (310) among the at least one selected candidate draft model as a preferred draft model. The probability distribution may be probability information determined by probability values determined by the at least one selected candidate draft model and probability values determined by the target model (310) for third output tokens corresponding to the identical output tokens. As an example, the probability distribution may be obtained by squaring a difference value between a first probability value and a second probability value and adding up the squared values calculated corresponding to the third output tokens. The first probability value can be determined by the draft model for each of the third output tokens. The second probability value can be determined by the target model (310) for each of the third output tokens. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select a draft model having a relatively small (or smallest) size among the multiple preferred draft models. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select one among the multiple preferred draft models by considering status information of the electronic device (100). The status information can include information indicating a heating state or a charging state of a battery.
[0142] The electronic device (100), at operation 930, may operate to generate the remaining tokens corresponding to the prompt (250) based on a guess decoding method by pairing the determined preferred draft model and the target model (310). For example, after determining the preferred draft model, the electronic device (100) may operate to verify the remaining tokens based on the prompt (250) in units of a specific number by pairing the preferred draft model and the target model (310). The electronic device (100), at operation 930, may operate to generate a result value responding to the prompt (250) based on a verification result for all output tokens corresponding to the prompt (250). The electronic device (100) may operate to output the generated result value. The result value output by the electronic device (100) may be confirmed by a user.
[0143] FIG. 10 is a control flowchart for driving a generative AI model in an electronic device (e.g., the electronic device (100) of FIG. 1) according to one embodiment.
[0144] Referring to FIG. 10, the electronic device (100) may operate, in operation 1010, to input a prompt (e.g., prompt (250) of FIG. 2) input by a user into a plurality of draft models (e.g., a plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3) included in an LLM (240) and at least one target model (e.g., target model (310) of FIG. 3). The target model (310) may be relatively larger in size than the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M). The plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) may have different sizes. The above prompt (250) may be, for example, a command or question that the user transmits to the LLM (240) installed in the electronic device (100) for conversation with the LLM (240) (e.g., the LLM (240) of FIG. 2). For example, the prompt (250) may serve as a communication window between the user and the LLM (240). Therefore, the electronic device (100) needs to implement the LLM (240) so that it can accurately analyze or understand the prompt (250) in order to accurately provide desired information.
[0145] The electronic device (100) may, in operation 1020, operate to cause each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) to generate a specific number (e.g., N) of first output tokens in succession based on the prompt (250) in response to a start token generated by the target model (310). As an example, the successive output tokens generated by each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) are as illustrated in FIG. 6. The start token may be generated by the target model (310) in response to an input of the prompt (250). The electronic device (100) may be operable to check first probability information that quantifies the likelihood (e.g., accuracy) that the first output tokens obtained for each draft model from the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) match the information included in the prompt (250).
[0146] The electronic device (100), in operation 1030, may operate to cause the target model (310) to obtain second output tokens for each draft model using the first output tokens as input. The second output tokens may, for example, be at least partially identical to the first output tokens. The second output tokens may, for example, be substantially identical to the second output tokens. The electronic device (100) may operate to check second probability information that quantifies the likelihood (e.g., accuracy) that the second output tokens obtained for each draft model in the target model (310) match the information included in the prompt (250).
[0147] The electronic device (100), in operation 1030, may determine one draft model included in the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) as a preferred draft model based on the first probability information and the second probability information. For example, the electronic device (100) may operate to determine a draft model in which the first probability information is relatively similar to the second probability information compared to one or more other draft models as the preferred draft model.
[0148] According to one example, the electronic device (100) may be configured to determine a draft model among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)) having a probability distribution similar to the target model (310)) as a preferred draft model. The probability distribution may be probability information determined by the probability values of identical output tokens (hereinafter referred to as “third output tokens”) that are overlapped and included in the first output tokens and the second output tokens, respectively, corresponding to each of the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M)). As an example, the probability distribution may be obtained by squaring the difference value between the first probability value and the second probability value, and adding up the squared values calculated corresponding to the third output tokens. The first probability value may be determined by the corresponding draft model for each of the third output tokens. The second probability value can be determined by the target model (310) for each of the third output tokens. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select a draft model having a relatively smaller size from among the multiple preferred draft models. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select one of the multiple preferred draft models by considering status information of the electronic device (100). The status information can include information indicating a heating state or a battery charging state.
[0149] According to an example, the electronic device (100) may be operable to select one or more draft models among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose output tokens are the same as the target model (310) as at least one candidate draft model. The electronic device (100) may be operable to determine a draft model having a probability distribution similar to the target model (310) among the at least one selected candidate draft model as a preferred draft model. The probability distribution may be probability information determined by probability values determined by the at least one selected candidate draft model and probability values determined by the target model (310) for third output tokens corresponding to the same output tokens. As an example, the probability distribution may be obtained by squaring a difference value between a first probability value and a second probability value and adding up the squared values calculated corresponding to the third output tokens. The first probability value can be determined by the draft model for each of the third output tokens. The second probability value can be determined by the target model (310) for each of the third output tokens. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select a draft model having a relatively smaller size among the multiple preferred draft models. For example, if the electronic device (100) determines that there are multiple preferred draft models, it can operate to select one of the multiple preferred draft models by considering status information of the electronic device (100). The status information can include information indicating a heating state or a charging state of a battery.
[0150] According to one example, if there are multiple draft models among the plurality of draft models (321-1, 321-2, 321-3, ..., 321-M) whose probability distribution is the same as or similar to the target model (310), the electronic device (100) may determine a draft model having a larger number of tokens identical to the target model (310) as the preferred draft model, regardless of the difference in the probability distribution.
[0151] The electronic device (100), in operation 1040, may operate to generate the remaining tokens corresponding to the prompt (250) based on a guess decoding method by pairing the determined preferred draft model and the target model (310). For example, after determining the preferred draft model, the electronic device (100) may operate to verify the remaining tokens based on the prompt (250) in units of a specific number by pairing the preferred draft model and the target model (310). The electronic device (100) may operate to generate a result value responding to the prompt (250) based on the verification results for all output tokens corresponding to the prompt (250).
[0152] The electronic device (100) may operate to output the generated result value in operation 1050. The result value output by the electronic device (100) may be confirmed by the user.
[0153] For example, the electronic device (100) may determine whether there is a prompt that is identical to or similar to an input prompt among the prompts for which a response result was output by the operation of a generative AI model that was performed previously. If there is a corresponding prompt that is identical to or similar to an input prompt, the electronic device (100) may determine the draft model used for speculative decoding for the corresponding prompt as the preferred draft model for speculative decoding for the input prompt.
[0154] According to one example, when multiple prompts are input, the electronic device (100) can omit the operation of determining a preferred draft model and process the multiple prompts simultaneously using a preset draft model.
[0155] FIG. 11 is a configuration diagram for driving a generative AI model in an electronic device (e.g., the electronic device (100) of FIG. 1) according to one embodiment.
[0156] Referring to FIG. 11, the electronic device (100) may include at least one processor (e.g., AP) (1120) (e.g., including a processing circuit) for driving a generative AI model, an input driving module (1121) (e.g., including various circuits and / or executable program instructions), an input framework module (1122), an NPU (1123) (e.g., including a processing circuit), a token accuracy determination and BD update module (1128), and / or a token determination module (1129). The input framework module (1122) may include an input processing module (1122-1) or an output processing module (1122-2), each of which may include various processing circuits. The NPU (1123) may include at least one target model (1124) (e.g., target module (310) of FIG. 3) or multiple draft modules (e.g., multiple draft modules (321-1, 321-2, 321-3, ..., 321-M) of FIG. 3). The NPU (1123) may include, for example, draft model #1 (1125), draft model #2 (1126), or draft model #3 (1127).
[0157] The electronic device (100) may include a display (1110). The display (1110) may display a screen (1111) for conversation with a user. The user may input a command to be transmitted to the LLM (240) operating on the AP (1120) for conversation with the LLM (240) or a prompt (250 of FIG. 2) corresponding to a question (reference numeral 1113) through the screen (1111). The display (1110) may be controlled by the input driving module (1121) included in the AP (1120). The prompt (250) corresponding to the user's input may be input to the input processing module (1122-1) included in the input framework module (1122) under the control of the input driving module (1121) (1130).
[0158] The above input processing module (1122-1) can convert a prompt in natural language, such as voice and / or text, into machine language, which is a language that can be recognized within the electronic device (100).
[0159] The NPU (1123) can receive (1140) the machine language converted by the input processing module (1122-1) and analyze the prompt (250) converted into the machine language to generate tokens. The NPU (1123) can generate tokens corresponding to the prompt (250) based on a guess decoding method.
[0160] According to an example, each of the draft model #1 (1125), the draft model #2 (1126), and the draft model #3 (1127) included in the NPU (1123) may obtain a specific number (e.g., N) of consecutive first output tokens in response to an input of a start token based on the prompt (250). The start token may be generated by the target model (1124) in response to an input of the prompt (250). Each of the draft model #1 (1125), the draft model #2 (1126), and the draft model #3 (1127) may determine first probability information that quantifies the likelihood (e.g., accuracy) that the first output tokens match the information included in the prompt (250). Each of the above draft model #1 (1125), the above draft model #2 (1126), and the above draft model #3 (1127) can input the first probability information determined corresponding to the first output tokens and / or each of the first output tokens into the target model (1124).
[0161] The target model (1124) can input first output tokens for each draft model generated by each of the draft model #1 (1125), the draft model #2 (1126), and the draft model #3 (1127), and can obtain second output tokens for each draft model based on the prompt (250). The target model (1124) can determine second probability information that quantifies the likelihood (e.g., accuracy) that the second output tokens for each draft model match the information included in the prompt (250).
[0162] The target model (1124) may provide information (1150) regarding the token generation result to the token accuracy determination and DB update module (1128). The information (1150) regarding the token generation result may include first output tokens and first probability information for each draft model, and second output tokens and second probability information for each draft model. The information (1150) regarding the token generation result may also be provided to the output processing module (1122-2) included in the input framework module (1122) (1180). The output processing module (1122-2) may output the information (1150) regarding the token generation result as a processing result corresponding to the input (1130) (1190). The display (1110) may display the processing result on the screen (1111) based on the output (1190) of the output processing module (1122-2) (1115).
[0163] The token accuracy determination and DB update module (1128) can analyze the information (1150) regarding the token generation result to determine the accuracy of the first output tokens for each draft. The token accuracy determination and DB update module (1128) can update the related data of the DB based on the accuracy of the determined first output tokens for each draft. The token accuracy determination and DB update module (1128) can provide information (1160) regarding the accuracy of the determined first output tokens for each draft to the token determination module (1129).
[0164] The token determination module (1129) may select one of the draft model #1 (1125), the draft model #2 (1126), and the draft model #3 (1127) as a preferred draft model based on information about the accuracy of the first output tokens for each determined draft provided from the token accuracy determination and DB update module (1128) (1170).
[0165] Among the above draft models #1 (1125), #2 (1126), and #3 (1127), the draft model selected as the preferred draft model by the token determination module (1129) is paired with the target model (1124) to generate or analyze the remaining tokens based on the guess decoding method.
[0166] FIG. 12 is a block diagram of an electronic device (1201) (e.g., the electronic device (100) of FIG. 1) within a network environment (1200) according to various embodiments.
[0167] Referring to FIG. 12, in a network environment (1200), an electronic device (1201) may communicate with an electronic device (1202) via a first network (1298) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1204) or a server (1208) via a second network (1299) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1201) may communicate with the electronic device (1204) via the server (1208). According to one embodiment, the electronic device (1201) may include a processor (1220), a memory (1230), an input module (1250), an audio output module (1255), a display module (1260), an audio module (1270), a sensor module (1276), an interface (1277), a connection terminal (1278), a haptic module (1279), a camera module (1280), a power management module (1288), a battery (1289), a communication module (1290), a subscriber identification module (1296), or an antenna module (1297). In some embodiments, the electronic device (1201) may omit at least one of these components (e.g., the connection terminal (1278)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1276), camera module (1280), or antenna module (1297)) may be integrated into a single component (e.g., display module (1260)).
[0168] The processor (1220) may control at least one other component (e.g., hardware or software component) of the electronic device (1201) connected to the processor (1220) by executing, for example, software (e.g., program (1240)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1220) may store commands or data received from other components (e.g., sensor module (1276) or communication module (1290)) in volatile memory (1232), process the commands or data stored in volatile memory (1232), and store result data in non-volatile memory (1234). According to one embodiment, the processor (1220) may include a main processor (1221) (e.g., a central processing unit or an application processor) or a secondary processor (1223) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1221). For example, when the electronic device (1201) includes the main processor (1221) and the secondary processor (1223), the secondary processor (1223) may be configured to use less power than the main processor (1221) or to be specialized for a given function. The secondary processor (1223) may be implemented separately from the main processor (1221) or as a part thereof.
[0169] The auxiliary processor (1223) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1260), the sensor module (1276), or the communication module (1290)) of the electronic device (1201), for example, on behalf of the main processor (1221) while the main processor (1221) is in an inactive (e.g., sleep) state, or together with the main processor (1221) while the main processor (1221) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1223) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1280) or a communication module (1290)). In one embodiment, the auxiliary processor (1223) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1201) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1208)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0170] The memory (1230) can store various data used by at least one component (e.g., the processor (1220) or the sensor module (1276)) of the electronic device (1201). The data can include, for example, software (e.g., the program (1240)) and input data or output data for commands related thereto. The memory (1230) can include a volatile memory (1232) or a non-volatile memory (1234).
[0171] The program (1240) may be stored as software in memory (1230) and may include, for example, an operating system (1242), middleware (1244), or an application (1246).
[0172] The input module (1250) can receive commands or data to be used in a component of the electronic device (1201) (e.g., a processor (1220)) from an external source (e.g., a user) of the electronic device (1201). The input module (1250) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0173] The audio output module (1255) can output audio signals to the outside of the electronic device (1201). The audio output module (1255) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0174] The display module (1260) can visually provide information to an external party (e.g., a user) of the electronic device (1201). The display module (1260) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1260) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0175] The audio module (1270) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1270) can acquire sound through the input module (1250), output sound through the sound output module (1255), or an external electronic device (e.g., electronic device (1202)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1201).
[0176] The sensor module (1276) can detect the operating status (e.g., power or temperature) of the electronic device (1201) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1276) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0177] The interface (1277) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1201) with an external electronic device (e.g., the electronic device (1202)). In one embodiment, the interface (1277) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0178] The connection terminal (1278) may include a connector through which the electronic device (1201) may be physically connected to an external electronic device (e.g., the electronic device (1202)). According to one embodiment, the connection terminal (1278) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0179] The haptic module (1279) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1279) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0180] The camera module (1280) can capture still images and videos. According to one embodiment, the camera module (1280) may include one or more lenses, image sensors, image signal processors, or flashes.
[0181] The power management module (1288) can manage power supplied to the electronic device (1201). According to one embodiment, the power management module (1288) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0182] A battery (1289) may power at least one component of the electronic device (1201). In one embodiment, the battery (1289) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0183] The communication module (1290) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1201) and an external electronic device (e.g., electronic device (1202), electronic device (1204), or server (1208)), and the performance of communication through the established communication channel. The communication module (1290) may operate independently from the processor (1220) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1290) may include a wireless communication module (1292) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1294) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1204) via a first network (1298) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1299) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1292) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1296) to verify or authenticate the electronic device (1201) within a communication network such as the first network (1298) or the second network (1299).
[0184] The wireless communication module (1292) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1292) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1292) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1292) may support various requirements specified in the electronic device (1201), an external electronic device (e.g., the electronic device (1204)), or a network system (e.g., the second network (1299)). According to one embodiment, the wireless communication module (1292) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0185] The antenna module (1297) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1297) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1297) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1298) or the second network (1299), may be selected from the plurality of antennas by, for example, the communication module (1290). A signal or power may be transmitted or received between the communication module (1290) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1297).
[0186] According to various embodiments, the antenna module (1297) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0187] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0188] According to one embodiment, commands or data may be transmitted or received between the electronic device (1201) and an external electronic device (1204) via a server (1208) connected to a second network (1299). Each of the external electronic devices (1202, or 1204) may be the same or a different type of device as the electronic device (1201). According to one embodiment, all or part of the operations executed in the electronic device (1201) may be executed in one or more of the external electronic devices (1202, 1204, or 1208). For example, when the electronic device (1201) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1201) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1201). The electronic device (1201) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1201) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (1204) may include an Internet of Things (IoT) device. The server (1208) may be an intelligent server utilizing machine learning and / or a neural network.In one embodiment, an external electronic device (1204) or server (1208) may be included in the second network (1299). The electronic device (1201) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0189] Figure 13 is a block diagram of an exemplary AI system (1300) capable of performing the operations described in this document. The AI system (1300) may be a generative AI system, but will be referred to as the "AI system (1300)" hereinafter.
[0190] Referring to FIG. 13, the AI system (1300) may include a User Query / Response Interface (1310) (e.g., I / F (220) of FIG. 2) (hereinafter, referred to as 'I / F (1310)'), an AI framework (1320), a generative AI model (1330), a database (1340), or an Application / Service Component (1350).
[0191] The above I / F (1310) can receive input (e.g., data acquired or generated by an electronic device (e.g., the electronic device (100) of FIG. 1 or the electronic device (1201) of FIG. 12) (hereinafter referred to as 'electronic device (100)') or user input, etc.). The data acquired or generated by the electronic device (100) can include image or video data generated using a processor (e.g., the processor (110) of FIG. 1 or the processor (1220) of FIG. 12), values transmitted through a sensor or sensor hub (e.g., external illuminance, an angle of the terminal, a display (e.g., the display (140) of FIG. 1) or the temperature of the electronic device (100), display (140) size or expansion / reduction information, a captured image of an image sensor (e.g., the image sensor (150) of FIG. 1), etc.). The user input may be in the form of natural language, touch coordinates or stylus coordinates obtained through a touch panel or digitizer included in the display (140), images, and / or videos. In addition, context information may also be transmitted when the user input is transmitted. The context information may include various additional information at the time of the user input. The additional information may include, for example, information on the application currently being used by the user or information on the user's location. In addition, the user input may also be in the form of a mixture of the above-described natural language, images, sounds, and context information. In addition, the user input may also be in the form of a non-natural language, such as selecting a menu. The I / F (1310) may provide the user with the result of analyzing the AI system (1300) and / or the input as an output. The output may be in the form of natural language or a specific content. The output may also be provided in the form of an action requested by the user. The output may also be provided in the form of a specific value specified by the user. The above I / F (1310) can output the results of the generative AI system (1300) to the user.The output may be in natural language or in a specific content format. The output may also be provided in the form of a user-requested action, etc.
[0192] The AI framework (1320) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query. For example, the AI framework (1320) can include a prompt design component (1321), an API / plug-in management component (1323), or an output modification component (or refiner component) (1325).
[0193] The user input received from the I / F (1310) may be transmitted to the prompt design component (1321). The prompt design component (1321) may be used to generate a prompt (e.g., the prompt (250) of FIG. 2) suitable for inputting the user input into the generative AI model (1330) (e.g., LLM, LVM (large vision model), or LMM (large multimodal model)). The prompt design component (1321) may be an AI component that uses a machine learning algorithm or a neural network to develop a better prompt (250) over time. The prompt design component (1321) may access a knowledge component including user preference data, a prompt library, and prompt examples based on the user input to generate the prompt (250), and transmit the generated prompt (250) to the generative AI model (1330).
[0194] The above API / plug-in management component (1323) may communicate with external information when there is a request for additional information when transmitting user input as input to the generative AI model (1330). The API / plug-in management component (1323) establishes a channel that can communicate with the outside of the AI interface through the API, and allows access to various data sources (e.g., knowledge repositories (1345)) through the established channel. For example, the knowledge repositories (1345) may store user preference data (1343) and / or prompt libraries (1341). If an action that performs a user input as a final step, rather than an intermediate result, must be performed in an application or service, the API / plug-in management component (1323) may request the action to the application / service component (1350) through the API. Information obtained from the outside may be used to generate a prompt (250) in the prompt design component (1321) together with the user input, or may be passed as an input to the generative AI model (1330).
[0195] The output adjustment component (1325) (also referred to as a refiner component) can fine-tune or reprocess the results output from the generative AI model (1330). The output adjustment component (1325) can verify, for example, whether the content generated through the generative AI model (1330) is irrelevant, does not contain biased content, or does not contain harmful content. The output adjustment component (1325) can also determine to what extent the content matches the result desired by the user and, if additional processing is required, can proceed with the process. The output adjustment component (1325) can additionally configure and provide the user with hints to avoid unwanted output.
[0196] The generative AI model (1330) may generally refer to an AI neural network that creates new types of data based on user input information. The generative AI model (1330) may include a model that generates images and / or a model that generates language. The model that generates images may include, for example, a generative adversarial network (GAN) or a variational autoencoder (VAE). The model that generates images may be, for example, a diffusion-based AI model that uses a VAE and a transformer structure. The model that generates language may be a model trained to output the most statistically appropriate output value based on an input value. Representative examples thereof include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM as an AI model (1330) that can recognize various types of data input, such as text, images, voices, and videos, and generate new data corresponding thereto.
[0197] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.
[0198] According to an example, the electronic device (100) may include a memory (230) including one or more storage media for storing instructions. The electronic device (100) may include at least one processor (210) including a processing circuit. When the instructions are individually or collectively executed by the at least one processor (210), the instructions may cause the electronic device (100) to perform at least one operation. The at least one operation may include an operation of outputting first tokens (440) for each draft model in response to an input of a prompt (250) in a plurality of draft models (320 or 410). The at least one operation may include an operation of outputting second tokens (460) for each draft model in a target model (310 or 420) by inputting the first tokens (450) for each draft model. The at least one operation may include an operation of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 410) based on a similarity between a first probability distribution of the first tokens (440 or 450) for each draft model and a second probability distribution of the second tokens (460) for each draft model.
[0199] In one example, the at least one action may include inputting the prompt (250) into the target model (310 or 420) and the plurality of draft models (320 or 410).
[0200] In one example, the at least one action may include outputting a start token (or initial token) (Ts) in response to an input of the prompt (250) in the target model (310 or 420).
[0201] According to an example, the at least one operation may be performed in response to an input of the start token (Ts) to output the first tokens (440) for each draft model.
[0202] According to one example, the at least one operation may include an operation of performing the guess decoding by the target model (720) and the target draft model (710) when the target draft model (710) is determined.
[0203] In one example, the first tokens (440 or 450) for each draft model may be substantially identical to the second tokens (460) for each draft model.
[0204] According to one example, the at least one operation may include an operation of determining the first probability distribution by quantifying the probability values of the first tokens (440) to be generated for each draft model.
[0205] According to one example, the at least one operation may include an operation of determining the second probability distribution by quantifying the probability values of the second tokens (460) to be generated for each draft model.
[0206] In one example, the at least one operation may include determining the target draft model (710) based on a size of the plurality of draft models (320 or 410), the size being determined by the number of parameters in the draft model.
[0207] According to one example, the at least one operation may include an operation of determining the target draft model (710) based on status information of the electronic device (100). The status information may include information indicating a heating state or a charging state of a battery.
[0208] According to one example, the size of the target model (310) - the size of the target model (310) is determined by the number of parameters in the target model (310) - may be configured to be larger than the size of the plurality of draft models (320) - the size of each draft model (320) is determined by the number of parameters in the corresponding draft model.
[0209] According to one example, the sizes of the plurality of draft models (320) may be configured differently.
[0210] According to one example, the at least one operation may include an operation of excluding the specific draft model (321-3) from candidates that can be determined as the target draft model (710) in response to the fact that the first tokens (T3-1, T3-2, T3-3, T3-4) per draft model output from the specific draft model (321-3) included in the plurality of draft models (320) are not identical to the second tokens (T3'-1, T3'-2, T3'-3, T3'-4) per draft model output from the target model (320 or 410).
[0211] According to one example, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a part of at least one processor of the electronic device (100), may cause the electronic device (100) to perform at least one operation. The at least one operation may include an operation (1020) of outputting first tokens (440) for each draft model in response to an input of a prompt (250) in a plurality of draft models (320 or 410). The at least one operation may include an operation (1030) of outputting second tokens (460) for each draft model in a target model (310 or 420) by inputting the first tokens (450) for each draft model. The at least one operation may include an operation (1030) of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 410) based on a similarity between a first probability distribution of the first tokens (440 or 450) for each draft model and a second probability distribution of the second tokens (460) for each draft model.
[0212] In one example, the at least one action may include inputting the prompt (250) into the target model (310 or 420) and the plurality of draft models (320 or 410).
[0213] In one example, the at least one action may include outputting a start token (or initial token) (Ts) in response to an input of the prompt (250) in the target model (310 or 420).
[0214] In one example, the at least one operation may include an operation of outputting first tokens (440) for each draft model in response to an input of the start token (Ts).
[0215] According to one example, the at least one operation may include an operation of performing the guess decoding by the target model (720) and the target draft model (710) when the target draft model (710) is determined.
[0216] In one example, the first tokens (440 or 450) for each draft model may be substantially identical to the second tokens (460) for each draft model.
[0217] According to one example, the at least one operation may include an operation of determining the first probability distribution by quantifying the probability values of the first tokens (440) to be generated for each draft model.
[0218] According to one example, the at least one operation may include an operation of determining the second probability distribution by quantifying the probability values of the second tokens (460) to be generated for each draft model.
[0219] In one example, the at least one operation may include determining the target draft model (710) based on the size of the plurality of draft models (320 or 410), the size of each draft model (320) being determined by the number of parameters in the corresponding draft model.
[0220] According to one example, the at least one operation may include an operation of determining the target draft model (710) based on status information of the electronic device (100). The status information may include information indicating a heating state or a charging state of a battery.
[0221] According to one example, the size of the target model (310) - the size of the target model (310) is determined by the number of parameters in the target model (310) - may be configured to be larger than the sizes of the plurality of draft models (320), and the sizes of the plurality of draft models (320) - the size of each draft model (320) is determined by the number of parameters in the corresponding draft model - may be configured to be different from each other.
[0222] According to one example, the at least one operation may include an operation of excluding the specific draft model (321-3) from candidates that can be determined as the target draft model (710) in response to the fact that the first tokens (T3-1, T3-2, T3-3, T3-4) per draft model output from the specific draft model (321-3) included in the plurality of draft models (320) are not identical to the second tokens (T3'-1, T3'-2, T3'-3, T3'-4) per draft model output from the target model (320 or 410).
[0223] According to an example, a method of executing a generative artificial intelligence (AI) model (240) in an electronic device (100) may be provided. The method may include an operation (1020) of outputting first tokens (440) for each draft model in response to an input of a prompt (250) in a plurality of draft models (320 or 410). The method may include an operation (1030) of outputting second tokens (460) for each draft model in a target model (310 or 420) by inputting the first tokens (450) for each draft model. The method may include an operation (1030) of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 420) based on a similarity between a first probability distribution of the first tokens (440 or 450) for each draft model and a second probability distribution of the second tokens (460) for each draft model.
[0224] In one example, the method may include inputting the prompt (250) into the target model (310 or 420) and the plurality of draft models (320 or 410).
[0225] In one example, the method may include an operation of outputting a start token (or initial token) (Ts) in response to an input of the prompt (250) in the target model (310 or 420).
[0226] According to an example, the method may include an operation of outputting first tokens (440) for each draft model in response to an input of the start token (Ts).
[0227] According to one example, the method may include an operation of performing the guess decoding by the target model (720) and the target draft model (710) when the target draft model (710) is determined.
[0228] In one example, the first tokens (440 or 450) for each draft model may be substantially identical to the second tokens (460) for each draft model.
[0229] According to one example, the method may include an operation of determining the first probability distribution by quantifying the probability values of the first tokens (440) to be generated for each draft model.
[0230] According to one example, the method may include an operation of determining the second probability distribution by quantifying the probability values of the second tokens (460) to be generated for each draft model.
[0231] According to one example, the method may include an operation of determining the target draft model (710) based on the sizes of the plurality of draft models (320 or 410), where the size of each draft model (710) is determined by the number of parameters in the corresponding draft model.
[0232] According to one example, the method may include an operation of determining the target draft model (710) based on status information of the electronic device (100). The status information may include information indicating a heating state or a charging state of a battery.
[0233] According to one example, the size of the target model (310) - the size of the target model (310) is determined by the number of parameters in the target model (310) - may be configured to be larger than the sizes of the plurality of draft models (320) - the size of each draft model (320) is determined by the number of parameters in the corresponding draft model - and the sizes of the plurality of draft models (320) may be configured to be different from each other.
[0234] According to one example, the method may include an operation of excluding the specific draft model (321-3) from candidates that can be determined as the target draft model (710) in response to the fact that the first tokens (T3-1, T3-2, T3-3, T3-4) per draft model output from a specific draft model (321-3) included in the plurality of draft models (320) are not identical to the second tokens (T3'-1, T3'-2, T3'-3, T3'-4) per draft model output from the target model (320 or 410).
[0235] According to an example, the electronic device (100) may include at least one processor (210). The electronic device (100) may include a memory (230) in which instructions are stored. The instructions, when individually or collectively executed by the at least one processor (210), may cause the electronic device (100) to perform at least one operation. The at least one operation may include an operation of obtaining an AI model (240) including a target model (310 or 420) and a first draft model (320) and a second draft model (410) that are smaller in size than the target model (310 or 420). The at least one operation may include an operation of obtaining input tokens for the AI model (240) based on a user input. The at least one operation may include an operation of checking first output tokens of the first draft model (320) for the input tokens, and first probabilities that the first output tokens will be output from the first draft model (320), respectively. The at least one operation may include an operation of checking second output tokens of the target model (310 or 420) for the first output tokens, and second probabilities that the second output tokens will be output from the target model (310 or 420), respectively. The at least one operation may include an operation of checking third output tokens of the second draft model (410) for the input tokens, and third probabilities that the third output tokens will be output from the second draft model (410), respectively. The at least one operation may include an operation of checking fourth output tokens of the target model (310 or 420) for the third output tokens, and fourth probabilities that the fourth output tokens will be output from the target model (310 or 420), respectively.The at least one operation may include selecting one of the first draft model (320) and the second draft model (410) based at least in part on the first probabilities, the second probabilities, the third probabilities, and the fourth probabilities. The at least one operation may include providing an output for the user input based at least in part on the selected one draft model.
[0236] In one example, the at least one operation may include selecting the first draft model (320) from among the first draft model (320) and the second draft model (410) if a difference between at least some of the first probabilities and at least some of the corresponding second probabilities is less than a difference between at least some of the third probabilities and at least some of the corresponding fourth probabilities.
[0237] In one example, the at least one operation may include selecting the second draft model (410) from among the first draft model (320) and the second draft model (410) if a difference between at least some of the first probabilities and at least some of the corresponding second probabilities is greater than a difference between at least some of the third probabilities and at least some of the corresponding fourth probabilities.
[0238] In one example, the at least one operation may include selecting the one draft model based on a difference between a first probability that the same output token will be output from the first draft model (320) and a second probability that the same output token will be output from the target model (310 or 420), when the first output tokens and the second output tokens contain the same output token.
[0239] In one example, the at least one operation may include selecting the one draft model based on the square of the difference between the first probability and the second probability, if the first output tokens and the second output tokens contain the same output token.
[0240] In one example, the at least one operation may include selecting the one draft model based on a probability that the different output token will be output from the target model (310 or 420) if the different output token that is not included in the first output tokens is included in the second output tokens.
[0241] In one example, the at least one operation may include selecting the one draft model based on the square of the probability that the different output token that is not included in the first output tokens is output from the target model (310 or 420) if the different output token is included among the second output tokens.
[0242] In one example, the at least one operation may include an operation of verifying the first output tokens by repeating the operation of making the output of the first draft model (320) for the input tokens into the input of the first draft model (320) a plurality of times.
[0243] In one example, the at least one operation may include an operation of verifying the third output tokens by repeating the operation of making the output of the second draft model (410) for the input tokens again as input to the second draft model (410) multiple times.
[0244] In one example, the at least one operation may include changing at least some of the tokens to be output from the selected one draft model to at least some of the second output tokens based on a difference between at least some of the first probabilities and at least some of the corresponding second probabilities.
[0245] In one example, the at least one operation may include selecting the one draft model further based on the number of identical tokens included in the first output tokens and the second output tokens, and the number of identical tokens included in the third output tokens and the fourth output tokens, when a difference between at least some of the first probabilities and at least some of the corresponding second probabilities and a difference between at least some of the third probabilities and at least some of the corresponding fourth probabilities are equal or similar.
[0246] According to an example, the first draft model (320) may have a different size than the second draft model (410).
[0247] In one example, the at least one operation may include selecting the one draft model further based on the sizes of the first draft model (320) and the second draft model (410) when a difference between at least some of the first probabilities and at least some of the corresponding second probabilities and a difference between at least some of the third probabilities and at least some of the corresponding fourth probabilities are equal or similar.
[0248] According to one example, the at least one operation may include selecting the one draft model further based on state information of the electronic device (100) when a difference between at least some of the first probabilities and at least some of the corresponding second probabilities and a difference between at least some of the third probabilities and at least some of the corresponding fourth probabilities are the same or similar.
[0249] According to one example, the status information of the electronic device (100) may include the heating status of at least a part of the electronic device (100) or the charging status of the battery.
[0250] In one example, the at least one operation may include generating the input tokens based on the user input and the target model (310 or 420).
[0251] According to an example, the operating speed of the target model (310 or 420) may be slower than the operating speed of the first draft model (320) and the operating speed of the second draft model (410).
[0252] According to an example, the number of parameters of the target model (310 or 420) may be greater than the number of parameters of the first draft model (320) and the number of parameters of the second draft model (410).
[0253] In one example, the complexity of the target model (310 or 420) may be greater than the complexity of the first draft model (320) and the complexity of the second draft model (410).
[0254] In one example, quantization may be applied more strongly to the first draft model (320) and the second draft model (410) than to the target model (310 or 420).
[0255] As an example, the AI model (240) may include a large language model (LLM).
[0256] It should be understood that the embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to a specific embodiment, but rather to encompass various modifications, equivalents, or substitutes of the embodiment. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the item, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0257] The term "module" used in one embodiment of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0258] One embodiment of the present document may be implemented as software including one or more instructions stored in a storage medium (e.g., memory (230)) readable by a machine (e.g., electronic device (100)). For example, a processor (e.g., CPU (211), NPU (213), GPU (215) of a device (e.g., electronic device (100)) can call at least one command among one or more commands stored from a storage medium and execute it. This enables the device to operate to perform at least one function according to the called at least one command. The one or more commands may include code generated by a compiler or code executable by an interpreter. A storage medium readable by the device may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' only means that the storage medium is a tangible device and does not include a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is temporarily stored in the storage medium.
[0259] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0260] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (100), A memory (230) including one or more storage media for storing instructions; and At least one processor (210) comprising a processing circuit, Here, at least one processor (210) is configured to individually and / or collectively execute the instructions, and cause the electronic device (100) to perform at least one operation, At least one of the above actions: An operation of outputting first tokens (440) for each draft model in response to input of a prompt (250) in multiple draft models (320 or 410); An operation of inputting the first tokens (450) for each draft model in the target model (310 or 420) and outputting the second tokens (460) for each draft model; and An operation of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 410) based on the similarity between the first probability distribution of the first tokens (440 or 450) for each draft model and the second probability distribution of the second tokens (460) for each draft model. An electronic device (100) comprising:
2. In paragraph 1, At least one processor (210) individually or collectively: An action of inputting the above prompt (250) into the target model (310 or 420) and the plurality of draft models (320 or 410); An operation of outputting a start token (Ts) in response to the input of the prompt (250) in the target model (310 or 420); and An operation of outputting the first tokens (440) for each draft model in response to the input of the above start token (Ts). An electronic device (100) configured to perform.
3. In paragraph 1 or 2, At least one processor (210) individually or collectively: Based on the determination of the target draft model (710), the target model (720) and the guess decoding by the target draft model (710) An electronic device (100) configured to perform.
4. In any one of paragraphs 1 to 3, At least one processor (210) individually or collectively: An operation of determining the first probability distribution by quantifying the probability values of the first tokens (440) to be generated for each draft model; and An operation of determining the second probability distribution by quantifying the probability values of the second tokens (460) to be generated for each of the above draft models. An electronic device (100) configured to perform.
5. In any one of paragraphs 1 to 4, At least one processor (210) individually or collectively: The target draft model (710) is determined based on the size of the plurality of draft models (320 or 410) - the size is determined by the number of parameters in the draft model - or An operation of determining the target draft model (710) based on the status information of the electronic device (100) - the status information includes information indicating the heating status or the charging status of the battery. An electronic device (100) configured to perform.
6. In any one of paragraphs 1 to 5, The size of the target model (310) - the size of the target model (310) is determined by the number of parameters in the target model (310) - is configured to be larger than the sizes of the plurality of draft models (320) - the size of each draft model (320) is determined by the number of parameters in the corresponding draft model. The sizes of the above plurality of draft models (320) are configured differently, in an electronic device (100).
7. In paragraph 1, At least one processor (210) individually or collectively: An operation of excluding the specific draft model (321-3) from candidates that can be determined as the target draft model (710) in response to the fact that the first tokens (T3-1, T3-2, T3-3, T3-4) for each draft model output from a specific draft model (321-3) included in the plurality of draft models (320) are not identical to the second tokens (T3'-1, T3'-2, T3'-3, T3'-4) for each draft model output from the target model (320 or 410) An electronic device (100) configured to perform.
8. In a method performed in an electronic device (100), An operation (1020) of outputting first tokens (440) for each draft model in response to input of a prompt (250) in multiple draft models (320 or 410); An operation (1030) of inputting the first tokens (450) for each draft model in the target model (310 or 420) and outputting the second tokens (460) for each draft model; and An operation (1030) of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 420) based on the similarity between the first probability distribution of the first tokens (440 or 450) for each draft model and the second probability distribution of the second tokens (460) for each draft model A method comprising:
9. In paragraph 8, An operation (1010) of inputting the above prompt (250) into the target model (310) and the plurality of draft models (320); An operation (1020) of outputting a start token (Ts) in response to the input of the prompt (250) in the target model (310); and An operation of outputting the first tokens (440) for each draft model in response to the input of the above start token (Ts). A method comprising:
10. In paragraph 8 or 9, An operation of performing the guess decoding by the target model (720) and the target draft model (710) based on the determination of the target draft model (710) A method comprising:
11. In any one of paragraphs 8 to 10, An operation of determining the first probability distribution by quantifying the probability values of the first tokens (440) to be generated for each draft model; and An operation of determining the second probability distribution by quantifying the probability values of the second tokens (460) to be generated for each of the above draft models. A method comprising:
12. In any one of paragraphs 8 to 11, The target draft model (710) is determined based on the size of the plurality of draft models (320 or 410) - the size is determined by the number of parameters in the draft model - or An operation of determining the target draft model (710) based on the status information of the electronic device (100) - the status information includes information indicating the heating status or the charging status of the battery. A method comprising:
13. In any one of paragraphs 8 to 12, The size of the target model (310) - the size of the target model (310) is determined by the number of parameters in the target model (310) - is configured to be larger than the sizes of the plurality of draft models (320) - the size of each draft model (320) is determined by the number of parameters in the corresponding draft model. A method in which the sizes of the above plurality of draft models (320) are configured differently.
14. In paragraph 8, A method comprising an operation of excluding the specific draft model (321-3) from candidates that can be determined as the target draft model (710) in response to the fact that the first tokens (T3-1, T3-2, T3-3, T3-4) per draft model output from a specific draft model (321-3) included in the plurality of draft models (320) are not identical to the second tokens (T3'-1, T3'-2, T3'-3, T3'-4) per draft model output from the target model (320 or 410).
15. In a storage medium (230) storing at least one instruction readable by a computer, The at least one instruction, when executed by at least a part of at least one processor (210) of the electronic device (100), causes the electronic device (100) to perform at least one operation; At least one of the above actions: An operation (1020) of outputting first tokens (440) for each draft model in response to input of a prompt (250) in multiple draft models (320 or 410); An operation (1030) of inputting the first tokens (450) for each draft model in the target model (310 or 420) and outputting the second tokens (460) for each draft model; and An operation (1030) of determining a target draft model (710) to be used for guess decoding among the plurality of draft models (320 or 410) based on the similarity between the first probability distribution of the first tokens (440 or 450) for each draft model and the second probability distribution of the second tokens (460) for each draft model A recording medium (230) including:
Citation Information
Patent Citations
Method for making ceramics having embroidery
KR1020220043356A
Similarity based per item model selection for medical imaging
US11288797B2
Dynamic accuracy-based deployment and monitoring of machine learning models in provider networks
US20190156247A1
Modifying artificial intelligence models using model fragments
US20200311561A1
Deep-learning model creation recommendations
US20210133558A1