Electronic device and method for generating sentence by using trained model, and non-transitory computer-readable storage medium

WO2026168722A1PCT designated stage Publication Date: 2026-08-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-08-13

Smart Images

  • Figure KR2025021036_13082026_PF_FP_ABST
    Figure KR2025021036_13082026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: a memory for storing instructions; and at least one processor. When executed individually or collectively by the at least one processor, the instructions can instruct the electronic device to: generate a first set of input sequences; obtain, from a trained model, output sequences inferred from the first set of input sequences; identify an output sequence from among the output sequences on the basis of verification of the first set of input sequences and the output sequences; and generate, on the basis of identifying the output sequence, a second set of input sequences of which the number of sequences is less than the number of sequences of the first set of input sequences. Each of some input sequences from among the second set of input sequences can include a second number of tokens, which is greater than a first number of tokens of each of the first set of input sequences and corresponds to a second setting value different from a first setting value.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for generating sentences using a trained model

[0001] The following descriptions relate to an electronic device, a method, and a non-transient computer-readable storage medium for generating sentences using a trained model.

[0002] With the advancement of electronic devices, technological development related to electronic devices equipped with artificial intelligence (AI) technology is underway. Electronic devices equipped with AI technology can provide various services to users. For example, electronic devices equipped with AI technology can provide a response to an input prompt by performing natural language processing on the input prompt. For example, the response may include a sentence.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the generation of a first set of input sequences. Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model stored in the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the training model to obtain output sequences inferred from the first set of input sequences by simultaneously providing the first set of input sequences to the training model. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to identify an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to generate a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences based on identifying the output sequence. Each of some of the input sequences in the second set of input sequences may include a second number of tokens according to a second setting value that is greater than the first number and different from the first setting value.

[0005] A method performed by an electronic device may include an operation of generating a first set of input sequences. Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model stored in the electronic device. The method may include an operation of obtaining output sequences inferred from the first set of input sequences from the trained model by simultaneously providing the first set of input sequences to the trained model. The method may include an operation of identifying an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences. The method may include an operation of generating a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences based on identifying the output sequences. Each of some of the second set of input sequences may include a second number of tokens according to a second set value that is greater than the first number and different from the first set value.

[0006] A non-transient computer-readable storage medium may store one or more programs comprising instructions that cause the electronic device to generate a first set of input sequences when executed individually or collectively by at least one processor of the electronic device. Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model stored in the electronic device. The non-transient computer-readable storage medium may store one or more programs comprising instructions that cause the electronic device to obtain output sequences inferred from the first set of input sequences from the trained model by simultaneously providing the first set of input sequences to the trained model when executed individually or collectively by the at least one processor. The above non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device to identify an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences when executed individually or collectively by the at least one processor. The above non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device to generate a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences based on identifying the output sequences when executed individually or collectively by the at least one processor. Each of some input sequences of the second set of input sequences may include a second number of tokens according to a second setting value that is greater than the first number and different from the first setting value.

[0007] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0008] Figure 2 illustrates an example of a neural network running on an electronic device.

[0009] FIG. 3 is a simplified block diagram of an exemplary electronic device.

[0010] Figure 4a illustrates an example of a trained model including a large language model (LLM).

[0011] Figure 4b illustrates examples of candidate tokens sampled from a draft token.

[0012] FIG. 4c illustrates an example of a method for generating an output sequence by providing an input sequence to a trained model.

[0013] FIG. 5 illustrates an example of a method for providing a function for generating multiple sentences without using self-speculative decoding (SSD).

[0014] FIG. 6a illustrates an example of a method for providing a function for generating multiple sentences using an SSD.

[0015] FIG. 6b illustrates an example of a method for performing verification in multi-sentence generation using SSD.

[0016] FIG. 6c illustrates an example of a case where the generation of some sentences is terminated in multi-sentence generation using an SSD.

[0017] FIG. 7 illustrates an example of a method for dynamically adjusting a setting value to be used for sampling each of the input sequences by redistributing resources for sequences for which sentence generation has ended to the input sequences while generating sentences using a trained model.

[0018] FIGS. 8 to 11 illustrate examples of a method for dynamically adjusting a setting value to be used for sampling each of the input sequences by redistributing resources for a sequence for which sentence generation has ended to the input sequences.

[0019] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of other embodiments. A singular expression may include a plural expression unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art described in this disclosure. Terms used in this disclosure that are defined in a general dictionary may be interpreted as having the same or similar meaning as they have in the context of the relevant technology, and are not to be interpreted in an ideal or overly formal sense unless explicitly defined in this disclosure. In some cases, even terms defined in this disclosure are not to be interpreted to exclude the embodiments of this disclosure.

[0020] In the various embodiments of the present disclosure described below, a hardware-based approach is described as an example. However, since the various embodiments of the present disclosure include techniques using both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0021] Additionally, in this disclosure, expressions of "greater than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled; however, this is merely for the purpose of expressing an example and does not exclude descriptions of "greater than" or "less than." Conditions described as "greater than" may be replaced with "greater than," conditions described as "less than" may be replaced with "less than," and conditions described as "greater than and less than" may be replaced with "greater than and less than." Furthermore, "A" to "B" below refer to at least one of the elements from A (including A) to B (including B).

[0022] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0024] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0025] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0026] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0027] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0028] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0029] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0030] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0031] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0032] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0033] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0034] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0036] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0038] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0039] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0041] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0043] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0045] Figure 2 illustrates an example of a neural network running on an electronic device.

[0046] Referring to FIG. 2, a neural network (200) can be obtained from a set of parameters stored in memory (e.g., memory (130) of FIG. 1) by an electronic device (101). For example, the neural network (200) may be an example of a model stored in memory (130). For example, the set of parameters may be included in the composition information of the model stored in memory (130).

[0047] For example, the neural network (200) may include multiple layers. For example, the neural network (200) may include an input layer (210), one or more hidden layers (220), and an output layer (230). The input layer (210) may correspond to a vector and / or matrix representing input data of the neural network (200). For example, the vector representing input data may have elements corresponding to the number of nodes included in the input layer (210). For example, the elements included in the matrix representing input data may correspond to each of the nodes included in the input layer (210). Signals generated at each of the nodes within the input layer (210) by the input data may be transmitted from the input layer (210) to the hidden layers (220). The output layer (230) may generate output data of the neural network (200) based on one or more signals received from the hidden layers (220). For example, the output data may correspond to a vector and / or matrix having elements corresponding to the number of nodes included in the output layer (230).

[0048] For example, first nodes included in a specific layer among a plurality of layers included in the neural network (200) may correspond to a weighted sum of at least one of the second nodes of the preceding layer of the specific layer within a sequence of the plurality of layers. For example, the electronic device (101) may identify a weight to be applied to at least one of the second nodes from a set of parameters stored in memory (130). Training the neural network (200) may include the operation of changing and / or determining one or more weights related to the weighted sum.

[0049] For example, one or more hidden layers (220) may be located between the input layer (210) and the output layer (230) and may convert input data transmitted through the input layer (210) into a value that is easy to predict. The input layer (210), one or more hidden layers (220), and the output layer (230) may include multiple nodes. One or more hidden layers (220) may be convolution filters or fully connected layers in a convolutional neural network (CNN), or various types of filters or layers grouped based on special functions or features. In one example, one or more hidden layers (220) may be layers based on a recurrent neural network (RNN) in which the output value is input back into the hidden layer at the current time. The neural network (200) may form a deep neural network by including a number of hidden layers (220). Training the above deep neural network may be referred to as deep learning. Among the nodes of the neural network (200), the nodes included in the hidden layers (220) may be referred to as hidden nodes.

[0050] For example, nodes included in the input layer (210) and one or more hidden layers (220) may be connected to each other through connecting lines having connection weights, and nodes included in the hidden layer and output layer (230) may also be connected to each other through connecting lines having connection weights. Tuning and / or training the neural network (200) may mean changing the connection weights between nodes included in each of the layers included in the neural network (200) (e.g., input layer (210), one or more hidden layers (220), and output layer (230)). For example, tuning of the neural network (200) may be performed based on supervised learning and / or unsupervised learning.

[0051] For example, in the state of acquiring a neural network (200), the electronic device (101) may identify weights corresponding to connecting lines connecting an input layer (210), one or more hidden layers (220), and / or an output layer (230) stored in memory (e.g., memory (130)). The electronic device (101) may acquire weighted sums based on connecting lines sequentially along a plurality of layers of the neural network (200) (e.g., the input layer (210), the one or more hidden layers (220), and the output layer (230)) in order to acquire output data from the neural network (200) based on the identified weights. The acquired weighted sums may be stored in at least one processor (e.g., processor (120)) and / or memory (130) of the electronic device (101). For example, the electronic device (101) can repeatedly update the weighted sum stored in memory (130) as it sequentially obtains the weighted sum along a plurality of layers.

[0052] Each of the multiple layers of the neural network (200) may have an independent data type and / or precision. For example, if the connecting lines between a first layer and a second layer among the multiple layers have weights based on a first data type for representing a floating-point number, the electronic device (101) may obtain weighted sums based on the first data type from numerical values ​​and weights corresponding to the nodes of the first layer. In the above example, if the connecting lines between a second layer and a third layer among the multiple layers have weights based on a second data type for representing an integer number, the electronic device (101) may obtain weighted sums based on the second data type from the obtained weighted sums and weights based on the second data type.

[0053] For example, if multiple layers have different data types, the electronic device (101) can obtain weighted sums corresponding to each of the multiple layers based on the different data types by using at least one processor (e.g., processor (120)). As the electronic device (101) accesses the memory (130) based on the weighted sums obtained based on the different data types, the bandwidth of the memory (130) can be utilized more efficiently. For example, as the bandwidth of the memory (130) is utilized more efficiently, the electronic device (101) can obtain output data more quickly from the neural network (200) based on the multiple layers.

[0054] FIG. 3 is a simplified block diagram of an exemplary electronic device.

[0055] Referring to FIG. 3, the electronic device (301) may be one of various forms of electronic devices, such as a laptop, smartphones having various form factors (e.g., bar-type smartphones, foldable (multi-foldable) type smartphones, or sliderable (or rollable) type smartphones), a tablet, a cellular phone, and other similar computing devices. The components, their relationships, and their functions illustrated in FIG. 3 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (301) may be referred to as a mobile device, a user device, a multi-functional device, a portable device, or a server. The electronic device (301) of FIG. 3 may be an example of the electronic device (101) of FIG. 1.

[0056] The electronic device (301) may include at least one processor (310) and memory (320). The components (e.g., at least one processor (310), memory (320)) are merely exemplary. For example, the electronic device (301) may include other components. For example, the other components may include a power management integrated circuitry (PMIC) or a rechargeable battery. For example, the other components may include a microphone for receiving voice signals from the outside. For example, the other components may include a display for displaying visual information. For example, the other components may include a speaker for outputting auditory information. For example, some components may be omitted from the electronic device (301). For example, some components may be integrated into a single component.

[0057] At least one processor (310) may be implemented as one or more integrated circuitry (IC) chips and may perform various data processing operations. At least one processor (310) may include at least one electrical circuit and may process instructions (or programs, data, etc.) stored in memory (320) individually or collectively in a distributed manner. At least one processor (310) may include a processor assembly including one or more processing circuits. At least one processor (310) may include any processing circuit that is operational to control the performance and operations of one or more components (e.g., memory (320)) of the electronic device (301). For example, at least one processor (310) (e.g., application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or a set of chips). For example, at least one processor (310) may be implemented with multiple cores (or multiple core circuits), multiple chips, or multiple sets of chips. For example, at least one processor (310) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example, but not limited to, at least one processor (310) may include a first processor (e.g., including a processing circuit) included in a first chip and a second processor (e.g., including a processing circuit) included in a second chip different from the first chip.

[0058] For example, at least one processor (310) may include a central processing unit (CPU) (311) (e.g., including processing circuits) and a neural processing unit (NPU) (312) (e.g., including processing circuits). The components of at least one processor (310) (e.g., CPU (311) and NPU (312)) are merely exemplary.

[0059] At least one processor (310) may cause other components of the electronic device (301) to perform various operations by executing instructions stored in memory (320). For example, the CPU (311) (or central processing circuit (311)) may be configured to control other components of at least one processor (310) (e.g., NPU (312)) based on the execution of instructions stored in memory (320). For example, the NPU (312) may be configured to generate (or infer) output data (or output sequence) using a trained model (345) from input data (or, input sequence) obtained from the CPU (311).

[0060] For example, at least one processor (310) may include at least a part of the processor (120) of FIG. 1 or correspond to at least a part of the processor (120) of FIG. 1.

[0061] Memory (320) may include one or more storage media (or one or more storage devices). For example, memory (320) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, flash memory, permanent memory such as ROM (read-only memory), semi-permanent memory such as RAM (random access memory), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (320) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (301). As an example, but not limited to, the cache memory may be included within at least one processor (310). The memory (320) can be fixedly embedded in the electronic device (301) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (301).

[0062] For example, memory (320) may store one or more software applications, such as an operating system (or system software application), a firmware software application, a driver software application, a plugin (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software application. For example, the one or more software applications may include instructions executable by at least one processor (310). For example, memory (320) may store instructions that can be called by an application programming interface (API). For example, memory (320) may store instructions within a library.

[0063] For example, the memory (320) may include at least a portion of the memory (130) of FIG. 1 or correspond to at least a portion of the memory (130) of FIG. 1.

[0064] Referring to FIG. 3, programs installed on an electronic device (301) may be included in any one of different layers, including an application layer (330), a framework layer (340), and / or a hardware abstraction layer (HAL) (350), depending on the target. For example, within the hardware abstraction layer (350), programs designed to target the hardware of the electronic device (301) (e.g., NPU (312)) (e.g., modules, or drivers) may be included. The framework layer (340) may be referred to as a generative framework layer (or generative AI framework) in that it includes one or more programs for providing sentence generation services using a trained model (345). For example, the layers illustrated in FIG. 3 are logically (or for convenience of explanation) separated, and may not imply that the address space of memory (320) is separated by said layers.

[0065] For example, within the framework layer (340), programs designed to target at least one of the hardware abstraction layer (350) and / or the application layer (330) may be included. Programs included in the framework layer (340) may provide an application programming interface (API) that is executable (or invokeable) based on other programs.

[0066] As a non-limiting example, the programs included in the framework layer (340) may include a self-speculative decoding module. For example, the self-speculative decoding module may be used to generate a sentence stream (or sentence) using a trained model (345). For example, the self-speculative decoding module may improve the speed of sentence generation by using multiple tokens rather than a single token when performing a single operation of the trained model (345). For example, the self-speculative decoding module may be implemented based on a serial method or a parallel method. As a non-limiting example, the serial method may include an early exit. For example, the early exit may utilize some of the multiple layers included in the trained model (345). In the present disclosure, the self-speculative decoding module based on the parallel method may be used. For example, the self-conjecturing decoding module based on the above parallel technique can generate draft candidate tokens at once and perform verification between the candidate tokens inferred from the draft candidate tokens and the draft candidate tokens.

[0067] As a non-limiting example, the programs included in the framework layer (340) may include a management module that stores configuration values ​​to be used for sampling. For example, the management module may store and manage configuration values ​​according to the maximum number of tokens available for each sentence (e.g., X) for the maximum number of tokens that can be input to the trained model (345) when generating multiple sentences. For example, if the trained model (345) is AR-64 and the number of sentences to be generated according to multi-sentence generation is 8, the maximum number (N) may be 64 and the maximum number (X) may be 8. In this case, the maximum number (X) may represent the maximum number of tokens included in a single input sequence. In the above example, the number of input tokens of the trained model (345) may be 64, the number of input sequences included in one set may be 8, and the number of tokens in each input sequence may be 8. In the above example, each input sequence is exemplified as including a number of tokens corresponding to the same setting value, but the present disclosure is not limited thereto. For example, some of the input sequences may use a first setting value, and other parts of the input sequences may use a second setting value. For example, the management module may store setting values ​​that require fewer than the number of tokens available in a single input sequence as candidate values. For example, the candidate values ​​may represent setting values ​​that can be used as setting values ​​for a single input sequence. In the above example, if a single input sequence may include 8 tokens, the candidate values ​​available as setting values ​​for the input sequence may include {1}, {2}, and {3}. Specific details regarding the candidate values ​​available as setting values ​​for the input sequence are exemplified and described below with reference to Table 3.

[0068] As a non-limiting example, the programs included in the framework layer (340) may include a distribution module for allocating resources of a sequence in which sentence generation has been terminated (or, completed, ended, finished) to sequences prior to the termination of sentence generation. For example, the distribution module may be referred to as a resource allocation module or an allocation module. For example, the distribution module may identify resources used in the sequence in which sentence generation has been terminated among the sequences output from a previous run of the trained model (345) (or sequences inferred from input sequences, input sequences). As a non-limiting example, the resources may be referred to as the number of tokens contained in the terminated sequence. For example, the distribution module may distribute (or allocate) the identified resources to at least some of the sequences to be input in the next run. In the present disclosure, distributing (or allocating) resources to sequences may indicate that the number of tokens included (or used) in the sequence is changed (or increased). For example, assume a case where there exists a first sequence containing 8 tokens, a second sequence containing 8 tokens, and a third sequence containing 8 tokens in a first run of the trained model (345). For example, when sentence generation of a fourth sequence inferred from the first sequence is finished, the distribution module may identify that the resource is 8. For example, the distribution module may distribute the identified resource (e.g., 8) to the second sequence and the third sequence, respectively. Accordingly, the second sequence may contain 12 tokens in the second run immediately following the first run of the trained model (345). Additionally, the third sequence may contain 12 tokens in the second run of the trained model (345).In the above example, the resource (e.g., 8) is illustrated as being distributed evenly among the sequences, but the present disclosure is not limited thereto. For example, the resource (e.g., 8) may be distributed differentially. Specific examples of the function of the distribution module may be referenced below in FIG. 7 and FIG. 8 through 11.

[0069] As a non-limiting example, the programs included in the framework layer (340) may include a setting module for selecting a setting value to be used for the sampling. For example, the setting module may select a setting value to be used for the operation of the trained model (345) by adjusting or maintaining the setting value to be used for the sampling. In the present disclosure, adjusting or maintaining the setting value may be included in distributing resources to the sequence. In other words, the setting module may be included in at least a part of the distribution module, or the distribution module and the setting module may be implemented as a single module. For example, specific examples of the function of the setting module may be referenced below in FIGS. 7 and FIGS. 8 through 11.

[0070] For example, within the application layer (330), a program designed to target a user of the wearable device (103) may be included. Programs included in the application layer (330) may include a software application for providing a sentence generation service through an artificial intelligence model (e.g., a trained model (345)). For example, a software application for providing a sentence generation service may be referred to as a multi-stream application or a smart reply application. For example, programs included in the application layer (330) (e.g., a software application) may call an API to cause the execution of a function supported by programs included in the framework layer (340). By example, without limitation, the function may include a function for generating multiple sentences (or a multi-sentence generation function). For example, the multi-sentence generation function may include a function for generating and outputting multiple sentences in response to a single prompt (or input prompt, input sentence).

[0071] For example, the memory (320) may include a trained model (345). For example, the trained model (345) may be composed of multiple neural network layers (e.g., the neural network (200) of FIG. 2). Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of operations of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized by the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the trained model (345) during the learning process is reduced or minimized.Artificial neural networks may include deep neural networks, such as, for example, CNN (convolutional neural network), R-CNN (region with convolution neural network), RPN (region proposal network), RNN (recurrent neural network), S-DNN (stacking-based deep neural network), S-SDNN (state-space dynamic neural network), deconvolution network, RBM (restricted Boltzmann machine), DBN (deep belief network), BRDNN (bidirectional recurrent deep neural network), fully convolutional network, classification network, LSTM (long short-term memory) network, vision transformer, diffusion, GAN (generative adversarial network), or deep q-networks, but are not limited to the examples mentioned above.

[0072] An electronic device (301) according to the present disclosure may use a trained model (345) to recommend / execute / infer a response to input data (e.g., a prompt). In the present disclosure, using the trained model (345) may be referred to as driving the trained model (345). For example, depending on the driving of the trained model (345), an output sequence (or output token) may be inferred (or generated) from an input sequence (or input token). As an example without limitation, the electronic device (301) may load the trained model (345) stored in memory (320) (or non-volatile memory in memory (320)) to perform inference (or store it in volatile memory in memory (320)) and perform inference using the loaded trained model (345). At least one processor (310) can perform a preprocessing process on the input data to convert it into a form suitable for use as input to the trained model (345). The trained model (345) can be created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using multiple training data by a learning algorithm, thereby creating a predefined behavioral rule or a trained model (345) set to perform a desired characteristic (or purpose). Inference prediction is a technique for logically reasoning and predicting by judging information, and may include knowledge-based reasoning, optimization prediction, preference-based planning, recommendation, etc.

[0073] For example, the trained model (345) may include an LLM for generating a response using an input prompt. Specific details regarding this may be referenced in FIG. 4a below.

[0074] Figure 4a illustrates an example of a trained model including a large language model (LLM).

[0075] FIG. 4a illustrates an example of a trained model (345) including an LLM (403). Referring to FIG. 4a, the trained model (345) may include an embedding layer (401) and a large language model (LLM) (403). The trained model (345) may include (or be implemented as) an LLM (403), which is an artificial neural network-based language model that has learned a large amount of text data through prior training. The LLM (403) may include relatively more parameters (e.g., more than 10 billion) than a conventional general language model. The LLM (403) may use a transformer artificial neural network structure based on an attention mechanism.

[0076] An attention mechanism is a technique that helps an artificial intelligence model focus on important parts within input data. An attention mechanism can be used to predict output data by predicting the extent to which some parts of time-series input data (e.g., input data such as voice or video, or input data of some layers of a neural network) contribute to the intermediate or final output of the neural network. While a recurrent neural network (RNN) structure that processes each element of a sequence sequentially suffers from degraded prediction performance when there is information dependency over long time-series distances, an attention mechanism can account for information dependency over long time-series distances by controlling the degree of weighted attention within the overall context (or part thereof) of the input data. For example, tokens processed through an attention mask may not be processed by the trained model (345) or may be ignored by the trained model (345).

[0077] A transformer can be composed of an encoder-decoder structure. The encoder processes input data to output compressed information (e.g., contextual representation), and the decoder processes the compressed information to output data in token units. Each of the encoder and decoder may include an independent attention network and a cross-attention network connecting the encoder and decoder.

[0078] For example, the training of the LLM (403) may include pre-training and / or fine-tuning. Pre-training is a process of enabling the LLM (403) to acquire general language knowledge using a large amount of text data, and may include, for example, self-supervised learning that predicts the next word using the previous word sequence of a text sequence. Fine-tuning is a process of training the LLM (403) to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task, and the LLM may be further supervised (or adaptive) using a dataset suitable for the domain purpose based on the pre-trained model. The LLM (403) may perform tasks with text input containing natural language called a prompt.

[0079] For example, fine-tuning can be omitted during training of the LLM (403). To improve the performance of the task desired by the user, the prompts input to the LLM (403) can be controlled. Examples of the task and / or guides for performing the task can be added to the prompts, such as in-context learning or zero-shot / few-shot learning. Examples of publicly available LLMs (403) include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer).

[0080] In the present disclosure, LLM (403) may refer to the trained model (345) itself. However, the present disclosure is not limited thereto. For example, LLM (403) may refer to a model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM (403)-based chatbot or LLM (403)-based translator such as ChatGPT may also refer to 'LLM'. 'LLM' may include an inference engine using a neural network model. For example, "inputting an input prompt into the LLM" may mean "inputting an input prompt into an inference engine based on the LLM". For example, "output of the LLM for the input prompt" may refer to the output information of the last neural network layer of the LLM (403) (or output information modified through additional processing) obtained when the input prompt is input into an inference engine based on the LLM (403). In the above example, it is described that an input prompt (or prompt) is input, but the present disclosure is not limited thereto. For convenience of explanation, it is described below that a token is input to the trained model (345).

[0081] For example, the embedding layer (401) may be used to convert tokens input to the trained model (345) into embedding values. For example, each of the embedding values ​​may be defined as a vector. For example, the trained model (345) may convert the acquired tokens into embedding values ​​using the embedding layer (401). However, the present disclosure is not limited thereto. For example, the trained model (345) may not include the embedding layer (401). For example, the embedding layer (401) may be located outside the trained model (345) and may provide (or transmit) the embedding values ​​converted from the tokens to the trained model (345). For convenience of explanation, the present disclosure assumes that the trained model (345) includes the embedding layer (401), but the present disclosure is not limited thereto.

[0082] In the present disclosure, the trained model (345) may be an autoregressive model (AR model). The autoregressive model may mean a model that predicts data to be used at the next time point (or next operation, next operation) using data used at the previous time point (or previous operation, previous operation).

[0083] For example, the trained model (345) may be implemented based on configuration information of the trained model (345). For example, the configuration information may include at least one of graph parameters or weight parameters of the trained model (345). For example, the graph parameters may represent operations, variables, and / or connection relationships of the trained model (345). For example, the weight parameters may include weights assigned to a plurality of nodes of the trained model (345) and / or connections between the plurality of nodes. For example, the length (or size) of an input sequence that can be input to the trained model (345) may be determined (or defined) according to the configuration information. For example, the input sequence may be referred to as an input token, input data, input value, or input token sequence. For example, the input sequence may be defined as a plurality of tokens. For example, the number of the plurality of tokens of the input sequence can be determined by the graph parameter (and / or the weight parameter).

[0084] For example, the trained model (345) may be used to generate (or infer) an output sequence from the input sequence. For example, the output sequence may be referred to as an output token, output data, output value, or output token sequence. For example, the output sequence may be defined as a plurality of tokens. For example, the number of the plurality of tokens in the output sequence may correspond (or be the same as) the number of the plurality of tokens in the input sequence.

[0085] Referring to the above description, the trained model (345) stored (or included) in the electronic device (301) of the present disclosure may be implemented by pre-set (or designed, defined) configuration information (e.g., graph parameters). Accordingly, the maximum number (N) of inputable tokens in the trained model (345) may be fixed (or specified). For example, depending on the maximum number (N) of inputable tokens of the trained model (345), which is an autoregressive model, the trained model (345) may be referred to as AR-N. In one example, if there is input data of a number of tokens exceeding the maximum number of inputable tokens, the tokens divided according to the maximum number (N) among the input data may be input to the trained model (345). To process the input data, the trained model (345) may be driven multiple times. In another example, if there is input data with a number of tokens less than the maximum number of inputtable tokens, arbitrary tokens may be input to the trained model (345) along with the tokens of the input data. In this case, processing of the arbitrary tokens may be ignored by providing additional information to the trained model (345). As a non-limiting example, the additional information may be referred to as a mask or an attention mask.

[0086] Figure 4b illustrates examples of candidate tokens sampled from a draft token.

[0087] FIG. 4b illustrates an example of a method for obtaining candidate tokens from draft tokens by performing sampling of draft tokens. In FIG. 4b, when generating a sentence using a trained model (345), one valid token (410) and two draft tokens (420) are shown, but the present disclosure is not limited thereto. For example, the number of draft tokens associated with the valid token (410) may be one or three or more.

[0088] For example, the valid token (410) may be one of the tokens included in the received prompt based on the execution of a software application for a sentence generation service. As an example without limitation, the valid token (410) may be the last token among the tokens of the prompt. Or, for example, one of the tokens of the output sequence generated (or inferred) by providing the tokens of the input sequence to the trained model (345) may be used as the valid token (410). For example, the one token among the tokens of the output sequence may be the approved token (or the last token among the approved tokens) based on verification between the output sequence and the input sequence. For example, the valid token (410) may be different (or distinguishable) from the tokens generated (or used) for self-inferential decoding (e.g., draft token, candidate token, guess token).

[0089] For example, draft tokens (420) may be tokens to be placed after (or consecutive with) a valid token (410). For example, among the tokens (or words) included in the sentence to be generated, a token to be placed after (or consecutive with) the current valid token (410) may be referred to as a draft token. In the example of FIG. 4b, the draft tokens (420) may include a first draft token (d1) (421) and a second draft token (d2) (422). For example, the first draft token (d1) (421) may represent a token to be placed after the valid token (410). For example, the second draft token (d2) (422) may represent a token to be placed after the first draft token (d1) (421).

[0090] For example, each of the draft tokens (420) may be sampled according to a set value. For example, the set value may represent a parameter to be used for sampling. For example, the set value may include at least one factor. For example, sampling the draft tokens may represent determining (or selecting) one or more candidate tokens among the tokens available as draft tokens.

[0091] For example, tokens (430) available as a first draft token (d1) (421) can be identified. For example, the tokens (430) may be tokens accessible for sentence generation in the trained model (345) (or tokens stored in memory (320). As a non-limiting example, the number of tokens (430) may be 100,000. For example, the tokens (430) may include token (d1_a) (431), token (d1_b) (432), token (d1_c) (433), and token (d1_d) (434). For example, the tokens (430) may be sorted according to the logit value (or probability value) of the token. As a non-limiting example, the tokens (430) may be arranged in the order of token (d1_a) (431), token (d1_b) (432), token (d1_c) (433), and token (d1_d) (434) according to descending order of logit values. In other words, the logit value of token (d1_a) (431) may have the maximum value among the tokens (430), and the logit value of token (d1_b) (432) may be lower than the logit value of token (d1_a) (431). In the present disclosure, the logit value may represent the probability of being used as a specific token (e.g., a candidate token, or a token inferred by the trained model (345)) among the tokens (e.g., tokens (430)).

[0092] For example, tokens (440) available as second draft tokens (d2) (422) can be identified. For example, the tokens (440) may be tokens accessible for sentence generation in the trained model (345) (or tokens stored in memory (320). By example, without limitation, the number of tokens (440) may be 100,000. In one example, the tokens (440) may correspond to (or be identical to) the tokens (430). For example, the tokens (440) may include token (d2_a) (441), token (d2_b) (442), token (d2_c) (443), and token (d2_d) (444). For example, the tokens (440) may be sorted according to the logit value (or probability value) of the token. As an example that is not limited, the tokens (440) may be arranged in descending order of logit value as token (d2_a) (441), token (d2_b) (442), token (d2_c) (443), and token (d2_d) (444). In other words, the logit value of token (d2_a) (441) may have the maximum value among the tokens (440), and the logit value of token (d2_b) (442) may be lower than the logit value of token (d2_a) (441).

[0093] For example, a setting value may represent a parameter for determining the number of candidate tokens to be generated by performing sampling of draft tokens. For example, the setting value may include at least one factor. For example, the setting value may be defined by the symbol ({}). For example, if the setting value includes two factors, the setting value may be defined as {fc1, fc2}. For example, fc1 may indicate a first factor to be used for sampling the first draft token (d1) (421), and fc2 may indicate a second factor to be used for sampling the second draft token (d2) (422). The number of at least one factor of the setting value may be equal to the number of draft tokens (e.g., 2). As a non-limiting example, the setting value may be defined as {fc}. Or, as a non-limiting example, the setting value may be defined as {fc1, fc2, fc3}.

[0094] In the example of FIG. 4b, the first setting value (429-1) may be {1, 1}. Depending on the first setting value (429-1), the first draft token (421) and the second draft token (422) may be sampled, respectively. For example, the candidate token sampled from the first draft token (421) may be the token (d1_a) (431). For example, the candidate token sampled from the second draft token (422) may be the token (d2_a) (441). Referring to the above, each of the factors of the first setting value (429-1) may be used to determine the number of candidate tokens corresponding to the factor among the tokens available as draft tokens. At this time, the number of candidate tokens corresponding to the factor may be determined according to the order of magnitude of the logit value (or probability value) (or top-k (k is the factor)) among the tokens available as draft tokens. For example, among the tokens (430) of the first draft token (421), the token (d1_a) (431) may be determined as a candidate token. For example, among the tokens (440) of the second draft token (422), the token (d2_a) (441) may be determined as a candidate token. In the above example, sampling by the first setting value (429-1) of {1, 1} is exemplified, but when the factor of the setting value is 1, sampling may not actually occur because one candidate token is determined for one draft token.

[0095] In the example of FIG. 4b, the second setting value (429-2) may be {3, 2}. Depending on the second setting value (429-2), the first draft token (421) and the second draft token (422) may each be sampled. For example, candidate tokens sampled from the first draft token (421) may include token (d1_a) (431), token (d1_b) (432), and token (d1_c) (433). For example, candidate tokens sampled from the second draft token (422) may include token (d2_a) (441) and token (d2_b) (442). Referring to the foregoing, each of the factors of the second setting value (429-2) may be used to determine the number of candidate tokens corresponding to the factor among the tokens available as draft tokens. At this time, the number of candidate tokens corresponding to the factor may be determined according to the order of the magnitude of the logit value (or probability value) among the tokens available as draft tokens (or top-k). For example, among the tokens (430) of the first draft token (421), token (d1_a) (431), token (d1_b) (432), and token (d1_c) (433) may be determined as candidate tokens. For example, among the tokens (440) of the second draft token (422), token (d2_a) (441) and token (d2_b) (442) may be determined as candidate tokens.

[0096] In the above example, two tokens (e.g., token (d2_a) (441) and token (d2_b) (442)) may be determined as candidate tokens of the second draft token (422) according to the second setting value (429-2). The second setting value (429-2) may be generated for each of the tokens determined according to the second factor (e.g., 2) among a plurality of factors (e.g., token (d2_a) (441) and token (d2_b) (442)), and the tokens determined according to the first factor (e.g., 3) (e.g., token (d1_a) (431), token (d1_b) (432), and token (d1_c) (433)). In the above example, the candidate tokens of the second draft token (422) may include six tokens. For example, a first pair of tokens (d2_a) (441) and token (d2_b) (442) may be generated for token (d1_a) (431), a second pair of tokens (d2_a) (441) and token (d2_b) (442) may be generated for token (d1_b) (432), and a third pair of tokens (d2_a) (441) and token (d2_b) (442) may be generated for token (d1_c) (433).

[0097] The number of candidate tokens generated according to the setting value can be calculated according to the following mathematical formula.

[0098]

[0099] In mathematical formula 1, the above θ represents the total number of candidate tokens generated according to the setting value, M represents the number of draft tokens, and fcj represents the i-th factor of the setting value. In the example of FIG. 4b, the total number of candidate tokens generated according to the second setting value (429-2) which is {3, 2} ( ) can be 9 (=3+3*2). The total number of candidate tokens generated according to the second setting value (429-2) which is {3, 2} ( ) can be exemplified according to a sample tree method, as in the example (449) of FIG. 4b. The setting value of the present disclosure may be referenced as a branch when used in the sample tree method.

[0100] Referring to the above description, when generating a sentence using the trained model (345), the sentence generation speed (or the inference speed of the trained model (345)) can be improved by using candidate tokens sampled from the draft token. For example, when one draft token (or one candidate token) is used, if the draft token inferred from the one draft token using the trained model (345) is rejected during verification, the one draft token can be input into the trained model (345), inferred, and verified again. However, when multiple candidate tokens for a specific draft token are used, the probability that one of the multiple candidate tokens inferred from the multiple candidate tokens using the trained model (345) will be approved during verification can be increased. In other words, the probability of approval during verification is increased, and the sentence generation speed can be improved.

[0101] A specific example of a method for generating a sentence using a trained model (345) according to self-conjectural decoding using multiple candidate tokens can be referenced in FIG. 4c below.

[0102] FIG. 4c illustrates an example of a method for generating an output sequence by providing an input sequence to a trained model.

[0103] FIG. 4c illustrates an example of how a trained model (345) generates (or infers) an output sequence from an input sequence. In FIG. 4c, for convenience of explanation, the maximum number (N) of inputable tokens of the trained model (345) is 32 (or AR-32), but the present disclosure is not limited thereto. Also, in FIG. 4c, for convenience of explanation, the setting value to be used for sampling draft tokens is assumed to be {3, 2}.

[0104] Referring to FIG. 4c, the input sequence may include valid tokens (v0) (410), candidate tokens (450), and guess tokens (460). In the present disclosure, a guess token may represent a token used as an input to generate (or infer) a plurality of tokens when running a trained model (345) based on self-guessing decoding. For example, the guess token may be referred to as a forecast token. If a guess token inferred (or generated) from an input guess token during the run of the trained model (345) at a specific time point is associated with a candidate token approved in verification, the inferred guess token may be used as a draft token during the run of the trained model (345) at the next time point.

[0105] For example, the candidate tokens (450) may include candidate tokens (451) sampled from the first draft token (d1) (e.g., the first draft token (d1) (421) of FIG. 4b). For example, the candidate tokens (451) may include candidate token (d1_a) (451-1), candidate token (d1_b) (451-2), and candidate token (d1_c) (451-3). For example, the candidate tokens (450) may include candidate tokens (452) sampled from the second draft token (d2) (e.g., the second draft token (d2) (422) of FIG. 4b). For example, candidate tokens (452) may include candidate token (d2_a) (452-1), candidate token (d2_b) (452-2), candidate token (d2_a) (452-3), candidate token (d2_b) (452-4), candidate token (d2_a) (452-5), and candidate token (d2_b) (452-6). For example, candidate token (d2_a) (452-1), candidate token (d2_a) (452-3), and candidate token (d2_a) (452-5) may have the same value (e.g., logit value). For example, candidate token (d2_b) (452-2), candidate token (d2_b) (452-4), and candidate token (d2_b) (452-6) may have the same value (e.g., logit value).

[0106] For example, the guess tokens (460) may include guess tokens (f_x, f_y) (461), guess tokens (f_x, f_y) (462), and guess tokens (f_x, f_y) (463). For example, the guess tokens (f_x, f_y) (461) may be associated with the valid token (v0) (410). For example, the guess tokens (f_x, f_y) (462) may be associated with the candidate token (d1_a) (451-1). For example, the guess tokens (f_x, f_y) (463) may be associated with the candidate token (d2_b) (452-6).

[0107] For example, guess tokens (460) can be generated based on valid tokens (v0) (410) and candidate tokens (450). For example, the number of guess tokens (460) can be determined based on the number of valid tokens (v0) (410) and the number of candidate tokens (450). For example, the number of guess tokens (460) (e.g., 20) can be defined as the product of the number of valid tokens (v0) (410) (e.g., 1) and the number of candidate tokens (450) (e.g., 9) (e.g., 10) and the number of draft tokens (e.g., 2).

[0108] In the above example, the number of tokens in the input sequence may be 30 (= 1+9+20). For example, the number of tokens in the input sequence may be less than or equal to the maximum number of tokens that can be input to the trained model (345) (e.g., 32). Any tokens corresponding to the difference (e.g., 2) between the maximum number of tokens that can be input to the trained model (345) (e.g., 32) and the number of tokens in the input sequence (e.g., 30) may be additionally included in the input sequence, and said any tokens may be set to be ignored by a mask. In FIG. 4c, said any tokens are not shown for convenience of explanation.

[0109] As a non-limiting example, the electronic device (301) may provide additional information when providing the input sequence to the trained model (345). For example, the additional information may include an attention mask (or mask) for setting the arbitrary tokens of the input sequence to be ignored. Or, for example, the additional information may include position information (e.g., position ids) for indicating the position of each token of the input sequence. Or, for example, the additional information may include information related to a previous operation (e.g., a key value cache) to increase the computation speed of the trained model (345) (or to reduce unnecessary operations). For example, the additional information may change depending on the input sequence to be input and the result of the previous operation. For example, the additional information (e.g., a key value cache) may refer to a system that stores intermediate results to perform a specified task relatively quickly. For convenience of explanation, the provision of the above additional information when the trained model (345) is operated is omitted below.

[0110] For example, the electronic device (301) can generate an output sequence using a trained model (345). For example, the output sequence can be inferred from the input sequence. For example, the output sequence may include valid tokens (v1) (470), candidate tokens (480), and guess tokens (490). For example, the valid token (v1) (470) can be inferred from the valid token (v0) (410). For example, the candidate tokens (480) can be inferred from the candidate tokens (450). For example, the guess tokens (490) can be inferred from the guess tokens (460).

[0111] For convenience of explanation in this disclosure, the output sequence is described as tokens, but the disclosure is not limited thereto. For example, the output sequence may be output values ​​(e.g., logit values). In this case, the output values ​​of the output sequence may be decoded into tokens corresponding to the output values. As a non-limiting example, the output values ​​may be decoded into tokens corresponding to the output values ​​by a tokenizer included in the trained model (345). Hereinafter, it is assumed that the output sequence contains tokens instead of output values.

[0112] For example, the candidate tokens (480) may include candidate tokens (481) that can be used as a second draft token (d2) inferred (predicted) from a first draft token (d1). For example, the candidate tokens (481) may include a candidate token (d2_1) (481-1), a candidate token (d2_2) (481-2), and a candidate token (d2_3) (481-3). For example, the candidate tokens (480) may include candidate tokens (482) that can be used as a third draft token (d3) inferred (predicted) from a second draft token (d2). For example, candidate tokens (482) may include candidate token (d3_1) (482-1), candidate token (d3_2) (482-2), candidate token (d3_1) (482-3), candidate token (d3_2) (482-4), candidate token (d3_1) (482-5), and candidate token (d3_2) (482-6). For example, candidate token (d3_1) (482-1) and candidate token (d3_2) (482-2) may be associated with candidate token (d2_1) (481-1). For example, candidate token (d3_1) (482-3) and candidate token (d3_2) (482-4) may be associated with candidate token (d2_2) (481-2). For example, candidate token (d3_1) (482-5) and candidate token (d3_2) (482-6) can be associated with candidate token (d2_3) (481-3).

[0113] For example, the guess tokens (490) may include guess tokens (f2_v1, f3_v1) (491), guess tokens (f3_d21, f4_d21) (492), and guess tokens (f4_d32, f5_d32) (493). For example, the guess tokens (f2_v1, f3_v1) (491) may be associated with the valid token (v1) (470). For example, the guess tokens (f3_d21, f4_d21) (492) may be associated with the candidate token (d2_1) (481-1). For example, the guess tokens (f4_d32, f5_d32) (493) may be associated with the candidate token (d3_2) (482-6).

[0114] In the present disclosure, symbols indicating tokens (e.g., v0, v1, d1_a, d2_a, f2_v1) may be defined according to the following rules.

[0115] For example, with respect to the character of the symbol, 'v' may indicate (or represent) a valid token, 'd' may indicate (or represent) a draft token or a candidate token for a draft token, and 'f' may indicate (or represent) a guess token. For example, with respect to the character of the symbol and a consecutive number, '0' may indicate (or represent) a priority position, '1' may indicate (or represent) a position following '0', and '2' may indicate (or represent) a position following '1'. In the example of FIG. 4c, the valid token (v1) (470) may indicate a token inferred from the valid token (v0) (410) to be placed after the valid token (v0) (410). For example, the candidate token (d1_a) (451-1) may indicate a candidate for a draft token that is likely to be placed after the valid token (v0) (410).

[0116] For example, the character following the underscore (_) of the symbol in the candidate token may indicate (or represent) the candidate token sampled from the draft token. In the example of FIG. 4c, the candidate token (d1_a) (451-1) may be an example of a candidate token sampled from the draft token (d1). For example, the number following the underscore (_) of the symbol in the candidate token may indicate (or represent) the candidate token inferred from the sampled candidate token. In the example of FIG. 4c, the candidate token (d2_1) (481-1) may be an example of a candidate token inferred by the model (345) trained from the candidate token (d1_a) (451-1).

[0117] For example, the character following the underscore (_) of the symbol in the guess token may indicate different arbitrary values. In the example of FIG. 4c, the guess token (f_x) of the guess tokens (461) and the guess token (f_y) of the guess tokens (461) may be different values. Also, in the example of FIG. 4c, the guess token (f_x) of the guess tokens (461) and the guess token (f_x) of the guess tokens (462) may be the same value.

[0118] For example, the characters and numbers following the underscore (_) of the symbol in the guess token may indicate (or represent) the token associated with the guess token. In the example of FIG. 4c, the guess tokens (491) including the guess token (f2_v1) and the guess token (f3_v1) may be associated with the valid token (v1) (470). In the example above, the guess token (f2_v1) may be the token to be placed after the valid token (v1) (470), and the guess token (f3_v1) may be the token to be placed after the guess token (f2_v1). Also, in the example of FIG. 4c, the guess tokens (492) including the guess token (f3_d21) and the guess token (f4_d21) may be associated with the candidate token (d2_1) (481-1). Additionally, in the example of FIG. 4c, the guess tokens (493), including the guess token (f4_d32) and the guess token (f5_d32), can be associated with the candidate token (d3_2) (482-6).

[0119] Referring to FIG. 4c, each of the tokens of the output sequence can be inferred (or generated) from a corresponding token among the tokens of the input sequence.

[0120] For example, a valid token (v1) (470) can be inferred (or generated) from a valid token (v0) (410). For example, the valid token (v1) (470) may be a token that is placed after the valid token (v0) (410) and may be a token that does not require verification (or a token approved without verification).

[0121] For example, candidate token (d2_1) (481-1) can be inferred (or generated) from candidate token (d1_a) (451-1). For example, candidate token (d2_2) (481-2) can be inferred (or generated) from candidate token (d1_b) (451-2). For example, candidate token (d2_3) (481-3) can be inferred (or generated) from candidate token (d1_c) (451-3). For example, candidate token (d3_1) (482-1) can be inferred (or generated) from candidate token (d2_a) (452-1). For example, candidate token (d3_2) (482-2) can be inferred (or generated) from candidate token (d2_b) (452-2). For example, candidate token (d3_1) (482-3) can be inferred (or generated) from candidate token (d2_a) (452-3). For example, candidate token (d3_2) (482-4) can be inferred (or generated) from candidate token (d2_b) (452-4). For example, candidate token (d3_1) (482-5) can be inferred (or generated) from candidate token (d2_a) (452-5). For example, candidate token (d3_2) (482-6) can be inferred (or generated) from candidate token (d2_b) (452-6).

[0122] For example, a guess token (f2_v1) can be inferred (or generated) from the guess token (f_x) of the guess tokens (461). For example, a guess token (f3_v1) can be inferred (or generated) from the guess token (f_y) of the guess tokens (461). For example, a guess token (f3_d21) can be inferred (or generated) from the guess token (f_x) of the guess tokens (462). For example, a guess token (f4_d21) can be inferred (or generated) from the guess token (f_y) of the guess tokens (462). For example, a guess token (f4_d32) can be inferred (or generated) from the guess token (f_x) of the guess tokens (463). For example, the guess token (f5_d32) can be inferred (or generated) from the guess token (f_y) of the guess tokens (463).

[0123] For example, the electronic device (301) may perform verification of the input sequence and the output sequence of the trained model (345). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). As described above, the valid token (v1) (470) inferred from the valid token (v0) (410) may be considered valid.

[0124] For example, the electronic device (301) can compare each of the candidate tokens (451) sampled from the valid token (v1) (470) and the first draft token (d1). For example, the electronic device (301) can sequentially compare the valid token (v1) (470) with the candidate token (d1_a) (451-1), candidate token (d1_b) (451-2), and candidate token (d1_c) (451-3) of the candidate tokens (451). For example, if the valid token (v1) (470) is identical to the candidate token (d1_a) (451-1), the candidate token (d1_a) (451-1) among the candidate tokens (451) may be accepted, and the remaining candidate tokens, candidate token (d1_b) (451-2) and candidate token (d1_c) (451-3), may be rejected. This is because only one of the candidate tokens (451) of the first draft token (d1) can be used as the first draft token (d1). However, the present disclosure is not limited thereto. For example, the valid token (v1) (470) may be different from all the tokens of the candidate tokens (451) of the first draft token (d1). If candidate token (d1_a) (451-1) is approved among candidate tokens (451), candidate token (d1_b) (451-2) and candidate token (d1_c) (451-3) may be rejected immediately upon comparison (or without verification) with the remaining candidate tokens. Alternatively, for example, if the valid token (v1) (470) is different from candidate token (d1_a) (451-1), candidate token (d1_a) (451-1) among candidate tokens (451) is rejected, and the electronic device (301) may compare the valid token (v1) (470) with candidate token (d1_b) (451-2). Additionally, if the valid token (v1) (470) is also different from the candidate token (d1_b) (451-2), the candidate token (d1_b) (451-2) among the candidate tokens (451) is also rejected, and the electronic device (301) can compare the valid token (v1) (470) with the candidate token (d1_c) (451-3).For convenience of explanation, it is assumed below that candidate token (d1_a) (451-1) among candidate tokens (451) is approved.

[0125] For example, the electronic device (301) may identify a candidate token (d2_1) (481-1) inferred from candidate token (d1_a) (451-1) as an approved candidate token when candidate token (d1_a) (451-1) among candidate tokens (451) is identical to a valid token (v1) (470). In this case, the remaining candidate tokens among candidate tokens (451), namely candidate token (d1_b) (451-2) and candidate token (d1_c) (451-3), are rejected, and the candidate token and speculative token associated with the rejected tokens may be rejected. For example, candidate token (d2_2) (481-2) inferred from candidate token (d1_b) (451-2), candidate token (d2_a) (452-3) and candidate token (d2_b) (452-4) among candidate tokens (452), candidate token (d3_1) (482-3) and candidate token (d3_2) (482-4) among candidate tokens (482), some of the guess tokens among the guess tokens (460), and some of the guess tokens among the guess tokens (490) may be rejected.

[0126] For example, the electronic device (301) can compare the candidate token (d2_1) (481-1) with some of the candidate tokens (452) sampled from the second draft token (d2). For example, the electronic device (301) can compare the candidate token (d2_1) (481-1) sequentially with the candidate token (d2_a) (452-1) and the candidate token (d2_b) (452-2) among the candidate tokens (452). For example, if the candidate token (d2_1) (481-1) is identical to the candidate token (d2_a) (452-1), the candidate token (d2_a) (452-1) may be accepted, and the candidate token (d2_b) (452-2) may be rejected. This is because only one of the candidate tokens (e.g., candidate token (d2_a) (452-1) and candidate token (d2_b) (452-2)) associated with candidate token (d1_a) (451-1) (or candidate token (d2_1) (481-1)) among the candidate tokens (452) of the second draft token (d2) can be used as the second draft token (d2). However, the present disclosure is not limited thereto. For example, candidate token (d2_1) (481-1) may be different from both candidate token (d2_a) (452-1) and candidate token (d2_b) (452-2).

[0127] For example, the electronic device (301) can identify the last approved candidate token among the approved candidate tokens in the output sequence as the candidate token (d2_1) (481-1) when the candidate token (d2_a) (452-1) and the candidate token (d2_b) (452-2) are different from the candidate token (d2_1) (481-1). For example, the last approved candidate token (d2_1) (481-1) can be used as a valid token (v0) in the next drive for sentence generation using the trained model (345). Additionally, the electronic device (301) can identify the guess tokens (492) associated with the candidate token (d2_1) (481-1) in the output sequence. For example, the guess token (f3_d21) of the guess tokens (492) can be used as the first draft token (d1) in the next drive, and the guess token (f4_d21) can be used as the second draft token (d2) in the next drive.

[0128] In another example, the electronic device (301) can identify the candidate token (d3_1) (482-1) inferred from the candidate token (d2_a) (452-1) as the accepted candidate token when the candidate token (d2_a) (452-1) is identical to the candidate token (d2_1) (481-1). For example, the electronic device (301) can identify the last accepted candidate token among the accepted candidate tokens in the output sequence as the candidate token (d3_1) (482-1). For example, the last accepted candidate token (d3_1) (482-1) can be used as the valid token (v0) in the next drive for sentence generation using the trained model (345). Additionally, the electronic device (301) can identify the guess tokens associated with the candidate token (d3_1) (482-1) in the output sequence. For example, the guess tokens (not shown) associated with the candidate token (d3_1) (482-1) can be used as the first draft token (d1) and the second draft token (d2) in the following drive.

[0129] For example, the electronic device (301) may add the tokens approved according to the verification to at least part of the sentence to be generated. For example, assume that the electronic device (301) identifies the valid token (v1) (470) and the candidate token (d2_1) (481-1) of the output sequence as the valid tokens during the verification of the operation of the trained model (345) exemplified in FIG. 4c. In this case, the sentence to be generated may include the word corresponding to the valid token (v0) (410), the word corresponding to the valid token (v1) (470), and the word corresponding to the candidate token (d2_1) (481-1). Afterward, the electronic device (301) may perform the next operation of the trained model (345) to determine the word following the word corresponding to the candidate token (d2_1) (481-1).

[0130] Referring to FIGS. 4a through 4c, the electronic device (301) may use multiple tokens (e.g., multiple candidate tokens and multiple guess tokens) in self-guessing decoding when generating sentences using a trained model (345). Accordingly, the sentence generation speed may be improved compared to the case where a single token (e.g., a single draft token) is used. The trained model (345) stored in the electronic device (301) may use a setting value to be used for sampling. For example, the minimum number of valid tokens, candidate tokens, and guess tokens according to the setting value may be defined as shown in Table 1 below.

[0131] Minimum number of setting value tokens Minimum number of setting value tokens{1}4{3, 2}30{2}6{3, 3}39{3}8{4, 1}27{4}10{4, 2}39{5}12{4, 3}51{6}14{4, 4}63{7}16{5, 1}33{8}18{5, 2}48{9}20{5, 3}63{10}22{5, 4}78{11}24{5, 5}93{12}26{6, 1}39{13}28{6, 2}57{14}30{6, 3}75{15}32{6, 4}93{1, 1}9{6, 5}111{2, 1}15{2, 2, 2}60{2, 2}21{3, 2, 1}64{3, 1}21…

[0132] The minimum number of tokens according to the setting values ​​of Table 1 above can be calculated according to the method for determining the number of guessing tokens of FIG. 4b, Equation 1, and FIG. 4c. For example, in a trained model (345) having a maximum number of inputable tokens greater than or equal to the minimum number of tokens of Table 1 above, the corresponding setting value may be used. In the example of Table 1 above, in a trained model (345) (e.g., AR-32) with a maximum number of 32, setting values ​​of {1} to {15}, {1, 1}, {2, 1}, {2, 2}, {3, 1}, {3, 2}, or {4, 1} may be used. Or, in the example of Table 1 above, in a trained model (345) (e.g., AR-4) with a maximum number of 4, setting value of {1} may be used. Alternatively, in the example of Table 1 above, for a trained model (345) (e.g., AR-8) with a maximum number of 8, setting values ​​of {1}, {2}, or {3} may be used. Alternatively, in the example of Table 1 above, for a trained model (345) (e.g., AR-16) with a maximum number of 16, setting values ​​of {1} to {7}, or {2, 1} may be used.

[0133] The examples in Table 1 above are merely illustrative for convenience of explanation and the present disclosure is not limited thereto. For example, the setting values ​​may include not only {1, 1} but also {1, 2}. For example, the setting values ​​available in the trained model (345) which is AR-32 may be further exemplified as shown in the table below.

[0134] Minimum number of configuration value tokens … … {1, 1}9{1, 2}12{1, 3}15{1, 4}18{1, 5}21{1, 6}24{1, 7}27{1, 8}30{2, 1}15{2, 2}21{2, 3}27{3, 1}21{3, 2}30{4, 1}27

[0135] For example, an electronic device (301) storing (or including) a trained model (345) that is AR-32 may use a predefined setting value according to the constraints of the trained model (345). For example, the predefined setting value may be determined by considering the optimal acceptance rate based on meaningful tokens (e.g., valid tokens, candidate tokens, and guess tokens) included in the input token. Referring to the example in FIG. 4c, the acceptance rate may be determined based on whether the candidate tokens of the preceding draft token are accepted, since if the candidate tokens of the preceding draft token (e.g., first draft token (d1)) are rejected, all candidate tokens of the subsequent draft token (e.g., second draft token (d2)) are rejected. However, if the setting value (or the factor for sampling) for the preceding draft token (e.g., the first draft token (d1)) is increased in order to increase the approval rate of candidate tokens for the preceding draft token (e.g., the first draft token (d1)), the approval rate of the subsequent draft token (e.g., the second draft token (d2)) may decrease. Accordingly, the sentence generation speed (or the approval rate for multiple draft tokens) may not be substantially improved.

[0136] In the above example, the approval rate is exemplified, but the present disclosure is not limited thereto. For example, the approval rate may be referred to as a token generation rate or a token rate. For example, the token generation rate may be defined as the ratio of the number of tokens approved in verification among the tokens inferred (or generated) as a result of running the trained model (345) to the number of runs of the trained model (345). For example, if 6 tokens among the inferred tokens are approved while the trained model (345) runs 5 times, the token generation rate may be 1.2.

[0137] To minimize the trade-off described above, the electronic device (301) may use a setting value having an averagely high token generation rate for draft tokens for driving the trained model (345). For example, the electronic device (301) may use a setting value of {3, 2} for driving the trained model (345) which is AR-32. For example, the setting value may be determined according to a predefined LUT (look-up table) (e.g., Table 1, Table 2). As an example without limitation, the token generation rate of the trained model (345) according to the setting value may be pre-calculated (or determined, defined). In this case, the token generation rate pre-calculated may be a value expected based on the results of an experiment (or test) on the trained model (345).

[0138] In the example of FIG. 4c, an example is illustrated in which an electronic device (301) generates a single sentence using a trained model (345), but the present disclosure is not limited thereto. For example, the electronic device (301) may generate multiple sentences in response to a single input (or prompt) based on the execution of a function for generating multiple sentences. For example, the electronic device (301) may use multiple input sequences as input tokens provided to the trained model (345) to generate the multiple sentences. Input sequences used as inputs in a single operation of the trained model (345) may be referred to as a set of input sequences. In other words, a specific set of input sequences may include multiple input sequences, and the multiple input sequences may be used to generate the multiple sentences.

[0139] Figures 5 and 6a through 6c below describe specific examples of a method for providing a function for generating multiple sentences using a trained model (345).

[0140] FIG. 5 illustrates an example of a method for providing a function for generating multiple sentences without using self-speculative decoding (SSD).

[0141] FIG. 5 illustrates examples (501, 502) of a method in which an electronic device (301) provides a function for generating multiple sentences without using an SSD. In the present disclosure, not using an SSD can be represented as the electronic device (301) including one token within one stream to generate one sentence using a trained model (345). In other words, the electronic device (301) can generate (or infer) one sentence by providing one stream (or one token) to the trained model (345).

[0142] Example (501) illustrates an example of inferring (or generating) output sequences using input sequences in a first drive of a trained model (345). Referring to Example (501), the electronic device (301) may generate input sequences of a first set (510). By example, without limitation, the first set (510) may include a plurality of input sequences (511, 512, 513). For example, the first set (510) may include a first input sequence (s1) (511), a second input sequence (s2) (512), a third input sequence (s3), a fourth input sequence (s4), a fifth input sequence (s5), a sixth input sequence (s6), a seventh input sequence (s7), and an eighth input sequence (s8) (513). For example, the first set (510) of input sequences may be referred to as an input sequence set, a sequence set, a stream set, a batch set, or a string set. Each input sequence within the first set (510) may be referred to as a sequence, a stream, a batch, or a string. For example, the first input sequence (511) may be referred to as a first sequence, a first stream, a first batch, or a first string. The number of input sequences included in the first set (510) may correspond to the number of sentences to be generated during multi-sentence generation. For example, the electronic device (301) may identify the number of sentences to be generated for multi-sentence generation. In the example of FIG. 5, the number of sentences to be generated may be 8. Accordingly, the electronic device (301) may generate a first set (510) containing 8 input sequences.

[0143] For example, if an SSD is not used, each sequence of the first set (510) may include one token. In this case, the one token included in each sequence may be a valid token. For example, the first input sequence (s1) (511) may include one token. For example, the second input sequence (s2) (512) may include one token. For example, the token included in each sequence of the first set (510) may be one of the tokens included in the prompt for multi-sentence generation. As a non-limiting example, the token included in each sequence of the first set (510) may be the last token among the tokens of the prompt. For example, the prompt may represent a voice signal received based on the execution of a software application for a sentence generation service (or the execution of a function for multi-sentence generation of the software application). As a non-limiting example, the trained model (345) exemplified in FIG. 5 may be AR-8 because 8 tokens are input.

[0144] For example, the electronic device (301) may provide the generated input sequences (511, 512, 513) to the trained model (345). For example, the electronic device (301) may use the trained model (345) to obtain output sequences inferred from the input sequences. Referring to the example (501) of FIG. 5, the electronic device (301) may generate output sequences of a first set (520). By example, without limitation, the first set (520) may include a plurality of output sequences (521, 522, 523). For example, the first set (520) may include a first output sequence (s1) (521), a second output sequence (s2) (522), a third output sequence (s3), a fourth output sequence (s4), a fifth output sequence (s5), a sixth output sequence (s6), a seventh output sequence (s7), and an eighth output sequence (s8) (523).

[0145] For example, an output sequence can be inferred (or generated) from an input sequence. In other words, an output sequence can be inferred from a corresponding input sequence. In the example (501) of FIG. 5, a first output sequence (s1) (521) can be inferred from a first input sequence (s1) (511), a second output sequence (s2) (522) can be inferred from a second input sequence (s2) (512), and an eighth output sequence (s8) (523) can be inferred from an eighth input sequence (s8) (513).

[0146] For example, the electronic device (301) may perform a second drive following the first drive after generating input sequences of a second set (520) from input sequences of a first set (510) in the first drive. This may be because, according to the first drive, not all sentences are completed (or, the generation of all sentences is not finished). As a non-limiting example, the electronic device (301) may identify the end of sentence generation based on identifying whether a token within each sequence corresponds to a reference token for sentence generation. Specific details regarding the second drive may be referenced in Example (502).

[0147] Example (502) illustrates an example of inferring (or generating) output sequences using input sequences in the first drive of the trained model (345). Referring to Example (502), the electronic device (301) may generate input sequences of a second set (530). By example, without limitation, the second set (530) may include a plurality of input sequences (531, 532, 533). For example, the second set (530) may include a first input sequence (s1) (531), a second input sequence (s2) (532), a third input sequence (s3), a fourth input sequence (s4), a fifth input sequence (s5), a sixth input sequence (s6), a seventh input sequence (s7), and an eighth input sequence (s8) (533). As a non-limiting example, each of the input sequences (531, 532, 533) of the second set (530) may correspond to each of the output sequences (521, 522, 523) of the first set (520). For example, the first input sequence (s1) (531) may be identical to (or correspond to) the first output sequence (s1) (521). For example, the second input sequence (s2) (532) may be identical to (or correspond to) the second output sequence (s2) (522). For example, the eighth input sequence (s8) (533) may be identical to (or correspond to) the eighth output sequence (s8) (528). In the present disclosure, an input sequence corresponding to an output sequence may indicate that it is a sequence containing the same token and identification information of the token (e.g., token id).

[0148] For example, the number of input sequences included in the second set (530) may correspond to the number of sentences to be generated during multi-sentence generation. However, the present disclosure is not limited thereto. For example, if sentence generation for some sequences is terminated in the first drive of example (501), the number of input sequences included in the second set (530) of example (502) may be less than the number of sentences to be generated during multi-sentence generation. For example, if sentence generation for some sequences is terminated in the first drive, in the second drive following the first drive, some sequences may be excluded from the sequence to be input, or a dummy sequence replaced from some sequences may be used as the sequence to be input.

[0149] In example (502), as in example (501), if an SSD is not used, each sequence of the second set (530) may include one token. In this case, the one token included in each sequence may be a valid token. For example, the first input sequence (531) may include one token. For example, the second input sequence (532) may include one token.

[0150] For example, the electronic device (301) may provide the generated input sequences (531, 532, 533) to the trained model (345). For example, the electronic device (301) may use the trained model (345) to obtain output sequences inferred from the input sequences. Referring to the example (502) of FIG. 5, the electronic device (301) may generate output sequences of a second set (540). By example, without limitation, the second set (540) may include a plurality of output sequences (541, 542, 543). For example, the second set (540) may include a first output sequence (s1) (541), a second output sequence (s2) (542), a third output sequence (s3), a fourth output sequence (s4), a fifth output sequence (s5), a sixth output sequence (s6), a seventh output sequence (s7), and an eighth output sequence (s8) (543).

[0151] For example, an output sequence can be inferred (or generated) from an input sequence. In other words, an output sequence can be inferred from a corresponding input sequence. In the example (502) of FIG. 5, a first output sequence (s1) (541) can be inferred from a first input sequence (s1) (531), a second output sequence (s2) (542) can be inferred from a second input sequence (s2) (532), and an eighth output sequence (s8) (543) can be inferred from an eighth input sequence (s8) (533).

[0152] Referring to the above description, in order to increase the speed of sentence generation when generating multiple sentences even when an SSD is not used, the electronic device (301) can provide (or simultaneously provide) multiple sequences to the trained model (345). Accordingly, the electronic device (301) can infer multiple sequences or generate multiple sentences through a single operation without running the trained model (345) multiple times.

[0153] In FIG. 5, an example is provided where an SSD is not used, but an SSD may be used to improve the approval rate (or token generation rate). Specific examples of how to perform multiple sentence generation when an SSD is used may be referenced in FIG. 6a to 6c below.

[0154] FIG. 6a illustrates an example of a method for providing a function for generating multiple sentences using an SSD.

[0155] FIG. 6a illustrates an example of a method in which an electronic device (301) uses an SSD to provide a function for generating multiple sentences. In the present disclosure, using an SSD may be represented as the electronic device (301) including a plurality of tokens within a single stream for generating a single sentence using a trained model (345). For example, the plurality of tokens may include valid tokens, candidate tokens (or draft tokens), and guess tokens. Specific details regarding this may be referenced in FIG. 4b and FIG. 4c. In other words, the electronic device (301) can generate (or infer) a single sentence by providing a trained model (345) with a single stream containing a plurality of tokens.

[0156] Referring to FIG. 6a, an example is illustrated in which output sequences are inferred (or generated) using input sequences in the operation of a trained model (345). Referring to FIG. 6a, an electronic device (301) may generate input sequences of a first set (610). By example, without limitation, the first set (610) may include a plurality of input sequences (611, 612, 613). For example, the first set (610) may include a first input sequence (s1) (611), a second input sequence (s2) (612), a third input sequence (s3), a fourth input sequence (s4), a fifth input sequence (s5), a sixth input sequence (s6), a seventh input sequence (s7), and an eighth input sequence (s8) (613). The number of input sequences included in the first set (610) may correspond to the number of sentences to be generated during multi-sentence generation. For example, the electronic device (301) can identify the number of sentences to be generated for multi-sentence generation. In the example of FIG. 6a, the number of sentences to be generated may be 8. Accordingly, the electronic device (301) can generate a first set (610) containing 8 input sequences.

[0157] For example, when using an SSD, each sequence of the first set (610) may include multiple tokens. For example, the first input sequence (s1) (611) may include eight tokens. For example, the second input sequence (s2) (612) may include eight tokens. For example, the eighth input sequence (s8) (613) may include eight tokens. Although not shown in FIG. 6a, the third input sequence (s3), the fourth input sequence (s4), the fifth input sequence (s5), the sixth input sequence (s6), and the seventh input sequence (s7) may each include eight tokens.

[0158] In the example of FIG. 6a, the setting value used for sampling in each sequence may be {3}. Depending on the setting value ({3}), each sequence may include one valid token, three candidate tokens, and four guess tokens. In the example of FIG. 6a, the tokens included in the first input sequence (s1) (611) may include a valid token (v0), three candidate tokens (d1_a, d1_b, d1_c), and four guess tokens (f_x). Additionally, for example, the tokens included in the second input sequence (s2) (612) may include a valid token (v0), three candidate tokens (d1_a, d1_b, d1_c), and four guess tokens (f_x). Additionally, for example, the tokens included in the eighth input sequence (s8) (613) may include a valid token (v0), three candidate tokens (d1_a, d1_b, d1_c), and four guess tokens (f_x).

[0159] For example, among the tokens included in each sequence of the first set (610), a valid token (e.g., v0) may include one of the tokens included in the prompt for multi-sentence generation. As a non-limiting example, the tokens included in each sequence of the first set (610) may include the last token among the tokens of the prompt. For example, the prompt may represent a voice signal received based on the execution of a software application for a sentence generation service (or the execution of a function for multi-sentence generation of the software application). As a non-limiting example, the trained model (345) exemplified in FIG. 6a may be AR-64 because a total of 64 (=8*8) tokens are input. For convenience of explanation, a trained model that is AR-64 is exemplified below.

[0160] For example, the electronic device (301) may provide the generated input sequences (611, 612, 613) to the trained model (345). For example, the electronic device (301) may use the trained model (345) to obtain output sequences inferred from the input sequences. Referring to FIG. 6a, the electronic device (301) may generate output sequences of a first set (620). By example, without limitation, the first set (620) may include a plurality of output sequences (621, 622, 623). For example, the first set (620) may include a first output sequence (s1) (621), a second output sequence (s2) (622), a third output sequence (s3), a fourth output sequence (s4), a fifth output sequence (s5), a sixth output sequence (s6), a seventh output sequence (s7), and an eighth output sequence (s8) (623).

[0161] For example, an output sequence can be inferred (or generated) from an input sequence. In other words, an output sequence can be inferred from a corresponding input sequence. In the example of FIG. 6a, a first output sequence (s1) (621) can be inferred from a first input sequence (s1) (611), a second output sequence (s2) (622) can be inferred from a second input sequence (s2) (612), and an eighth output sequence (s8) (623) can be inferred from an eighth input sequence (s8) (613).

[0162] For example, the first output sequence (s1) (621) may include a valid token (v1), three candidate tokens (d2_1, d2_2, d2_3), and four guess tokens (f2_v1, f3_d21, f3_d22, f3_d23). For example, the valid token (v1) of the first output sequence (s1) (621) may be inferred from the valid token (v0) of the first input sequence (s1) (611). For example, the candidate token (d2_1) of the first output sequence (s1) (621) may be inferred from the candidate token (d1_a) of the first input sequence (s1) (611). For example, a candidate token (d2_2) of the first output sequence (s1) (621) can be inferred from a candidate token (d1_b) of the first input sequence (s1) (611). For example, a candidate token (d2_3) of the first output sequence (s1) (621) can be inferred from a candidate token (d1_c) of the first input sequence (s1) (611). For example, a guess token (f2_v1) of the first output sequence (s1) (621) can be inferred from the first guess token (f_x) among the guess tokens of the first input sequence (s1) (611). For example, a guess token (f3_d21) of the first output sequence (s1) (621) can be inferred from the second guess token (f_x) among the guess tokens of the first input sequence (s1) (611). For example, the guess token (f3_d22) of the first output sequence (s1) (621) can be inferred from the third guess token (f_x) among the guess tokens of the first input sequence (s1) (611). For example, the guess token (f3_d23) of the first output sequence (s1) (621) can be inferred from the fourth guess token (f_x) among the guess tokens of the first input sequence (s1) (611).

[0163] In FIG. 6a, an example is provided in which an electronic device (301) infers output sequences from input sequences using a trained model (345). For example, the electronic device (301) may use an NPU (312) configured to perform inference using the trained model (345). For example, the electronic device (301) may perform verification of the input sequences and output sequences of the trained model (345). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). Specific details regarding the verification may be referenced below in FIG. 6b.

[0164] FIG. 6b illustrates an example of a method for performing verification in multi-sentence generation using SSD.

[0165] FIG. 6b illustrates examples (601, 602) of a method in which an electronic device (301) generates an output sequence inferred from an input sequence using a trained model (345) and then performs verification on the input sequence and the output sequence. Example (601) illustrates verification of a first set of input sequences and a first set of output sequences of a first drive of the trained model (345). Example (602) illustrates verification of a second set of input sequences and a second set of output sequences of a second drive of the trained model (345) that is rightly subsequent to (or immediately subsequent to) the first drive.

[0166] In Examples (601) and (602), for convenience of explanation, a first input sequence (s1) among the input sequences and a first output sequence (s1) among the output sequences are illustrated, but the present disclosure is not limited thereto. In FIG. 6b, for convenience of explanation, the case where the maximum number (N) of inputable tokens of the trained model (345) is 64 (or AR-32) is illustrated, but the present disclosure is not limited thereto. Also, in FIG. 6b, for convenience of explanation, the case where the setting value used for sampling the draft token of each sequence is {3} is assumed.

[0167] Referring to example (601), the first input sequence (s1) (611) may include a valid token (v0) (611-1), candidate tokens (611-2, 611-3, 611-4), and guess tokens (611-5, 611-6, 611-7, 611-8). For example, the candidate tokens (611-2, 611-3, 611-4) may include candidate tokens sampled from the first draft token (d1) (e.g., the first draft token (d1) (421) of FIG. 4b). For example, the candidate tokens may include candidate token (d1_a) (611-2), candidate token (d1_b) (611-3), and candidate token (d1_c) (611-4). For example, the guess tokens (611-5, 611-6, 611-7, 611-8) may include guess token (f_x) (611-5), guess token (f_x) (611-6), guess token (f_x) (611-7), and guess token (f_x) (611-8). For example, the guess token (f_x) (611-5) may be associated with the valid token (v0) (611-1). For example, the guess token (f_x) (611-6) may be associated with the candidate token (d1_a) (611-2). For example, the guess token (f_x) (611-7) may be associated with the candidate token (d1_b) (611-3). For example, a guess token (f_x) (611-8) can be associated with a candidate token (d1_c) (611-4). For example, guess tokens (611-5, 611-6, 611-7, 611-8) can be generated based on valid tokens (v0) (611-1) and candidate tokens (611-2, 611-3, 611-4). For example, the number of guess tokens (611-5, 611-6, 611-7, 611-8) can be determined based on the number of valid tokens (v0) (611-1) and the number of candidate tokens (611-2, 611-3, 611-4).For example, the number of guess tokens (611-5, 611-6, 611-7, 611-8)) (e.g., 4) can be defined as the product of the number of valid tokens (v0) (611-1) (e.g., 1) and the number of candidate tokens (611-2, 611-3, 611-4) (e.g., 3) (e.g., 4) and the number of draft tokens (e.g., 1).

[0168] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, a first output sequence (s1) (621) can be inferred from a first input sequence (s1) (611). For example, the first output sequence (s1) (621) may include a valid token (v1) (621-1), candidate tokens (621-2, 621-3, 621-4), and guess tokens (621-5, 621-6, 621-7, 621-8). For example, the valid token (v1) (621-1) can be inferred from a valid token (v0) (611-1). For example, candidate tokens (621-2, 621-3, 621-4) can be inferred from candidate tokens (611-2, 611-3, 611-4). For example, guess tokens (621-5, 621-6, 621-7, 621-8) can be inferred from guess tokens (611-5, 611-6, 611-7, 611-8).

[0169] For example, candidate tokens (621-2, 621-3, 621-4) may include candidate tokens that can be used as a second draft token (d2) inferred (predicted) from a first draft token (d1). For example, candidate tokens may include candidate token (d2_1) (621-2), candidate token (d2_2) (621-3), and candidate token (d2_3) (621-4). For example, guess tokens (621-5, 621-6, 621-7, 621-8) may include guess token (f2_v1) (621-5), guess token (f3_d21) (621-6), guess token (f3_d22) (621-7), and guess token (f3_d23) (621-8). For example, the guess token (f2_v1) (621-5) can be associated with the valid token (v1) (621-1). For example, the guess token (f3_d21) (621-6) can be associated with the candidate token (d2_1) (621-2). For example, the guess token (f3_d22) (621-7) can be associated with the candidate token (d2_2) (621-3). For example, the guess token (f3_d23) (621-8) can be associated with the candidate token (d2_3) (621-4).

[0170] For example, the electronic device (301) may perform verification on the first input sequence (s1) (611) and the first output sequence (s1) (621) of the trained model (345). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). As described above, the valid token (v1) (621-1) inferred from the valid token (v0) (611-1) may be considered as approved.

[0171] For example, the electronic device (301) can compare each of the candidate tokens (611-2, 611-3, 611-4) sampled from the valid token (v1) (621-1) and the first draft token (d1). For example, the electronic device (301) can sequentially compare the valid token (v1) (621-1) with the candidate token (d1_a) (611-2), candidate token (d1_b) (611-3), and candidate token (d1_c) (611-4) of the candidate tokens (611-2, 611-3, 611-4). For example, if the valid token (v1) (621-1) is identical to the candidate token (d1_a) (611-2), among the candidate tokens (611-2, 611-3, 611-4), the candidate token (d1_a) (611-2) may be accepted, and the remaining candidate tokens, candidate token (d1_b) (611-3) and candidate token (d1_c) (611-4), may be rejected. This is because only one of the candidate tokens (611-2, 611-3, 611-4) of the first draft token (d1) can be used as the first draft token (d1). However, the present disclosure is not limited thereto. For example, the valid token (v1) (621-1) may be different from all of the candidate tokens (611-2, 611-3, 611-4) of the first draft token (d1). If candidate token (d1_a) (611-2) among the candidate tokens (611-2, 611-3, 611-4) is approved, candidate token (d1_b) (611-3) and candidate token (d1_c) (611-4) may be rejected immediately upon comparison (or, without verification) with the remaining candidate tokens. Alternatively, for example, if the valid token (v1) (621-1) is different from the candidate token (d1_a) (611-2), the candidate token (d1_a) (611-2) among the candidate tokens (611-2, 611-3, 611-4) is rejected, and the electronic device (301) can compare the valid token (v1) (621-1) with the candidate token (d1_b) (611-3).Additionally, if the valid token (v1) (621-1) is also different from the candidate token (d1_b) (611-3), the candidate token (d1_b) (611-3) among the candidate tokens (611-2, 611-3, 611-4) is also rejected, and the electronic device (301) can compare the valid token (v1) (621-1) with the candidate token (d1_c) (611-4). For convenience of explanation, it is assumed below that the candidate token (d1_b) (611-3) among the candidate tokens (611-2, 611-3, 611-4) is accepted.

[0172] For example, the electronic device (301) may identify the candidate token (d2_2) (621-3) inferred from the candidate token (d1_b) (611-3) as the approved candidate token when the candidate token (d1_b) (611-3) among the candidate tokens (611-2, 611-3, 611-4) is identical to the valid token (v1) (621-1). In this case, the remaining candidate tokens among the candidate tokens (611-2, 611-3, 611-4), namely candidate token (d1_a) (611-2) and candidate token (d1_c) (611-4), are rejected, and the candidate token and speculative token associated with the rejected tokens may be rejected. For example, candidate token (d2_1)(621-2) inferred from candidate token (d1_a)(611-2), candidate token (d2_3)(621-4) inferred from candidate token (d1_c)(611-4), some of the guess tokens (611-5, 611-6, 611-8) among the guess tokens (611-5, 611-6, 611-7, 611-8), and some of the guess tokens (621-5, 621-6, 621-7, 621-8) among the guess tokens (621-5, 621-6, 621-8) may be rejected.

[0173] For example, the electronic device (301) can identify the last approved candidate token among the approved candidate tokens in the first output sequence (s1) (621) as the candidate token (d2_2) (621-3). For example, the last approved candidate token (d2_2) (621-3) can be used as the valid token (v2) of the first input sequence (s1) (631) in the next drive for sentence generation using the trained model (345) (e.g., the second drive of example (602). Additionally, the electronic device (301) can identify the guess token (f3_d22) (621-7) associated with the candidate token (d2_2) (621-3) in the first output sequence (s1) (621). For example, the guess token (f3_d22) (621-7) can be used as the first draft token (d3) in the following drive.

[0174] For example, the electronic device (301) may add the tokens approved according to the verification as at least part of the sentence to be generated. For example, assume that the electronic device (301) identifies the valid token (v1) (621-1) and the candidate token (d2_2) (621-3) of the first output sequence (s1) (621) as the valid tokens in the verification of the operation of the trained model (345) exemplified in the example (601) of FIG. 6a. In this case, the sentence to be generated may include the word corresponding to the valid token (v0) (611-1), the word corresponding to the valid token (v1) (621-1), and the word corresponding to the candidate token (d2_2) (621-3). Afterward, the electronic device (301) may perform the next operation of the trained model (345) to determine the word following the word corresponding to the candidate token (d2_2) (621-3). For example, the sentence to be generated may be a sentence generated using a first input sequence (s1) (611) (and / or a first output sequence (s1) (621)). For example, the electronic device (301) may generate eight sentences by performing the above operations for each of the input sequences (611, 612, 613) in the first set (610).

[0175] Referring to example (602), the first input sequence (s1) (631) may be included in a second set of input sequences. For example, the second set may include eight input sequences. For example, the first input sequence (s1) (631) may include valid tokens (v2) (631-1), candidate tokens (631-2, 631-3, 631-4), and guess tokens (631-5, 631-6, 631-7, 631-8). For example, the candidate tokens (631-2, 631-3, 631-4) may include candidate tokens sampled from the first draft token (d3) (e.g., the first draft token (d1) (421) in FIG. 4b). For example, candidate tokens may include candidate token (d3_a) (631-2), candidate token (d3_b) (631-3), and candidate token (d3_c) (631-4). For example, guess tokens (631-5, 631-6, 631-7, 631-8) may include guess token (f_x) (631-5), guess token (f_x) (631-6), guess token (f_x) (631-7), and guess token (f_x) (631-8). For example, guess token (f_x) (631-5) may be associated with valid token (v2) (631-1). For example, guess token (f_x) (631-6) may be associated with candidate token (d3_a) (631-2). For example, the guess token (f_x) (631-7) can be associated with the candidate token (d3_b) (631-3). For example, the guess token (f_x) (631-8) can be associated with the candidate token (d3_c) (631-4). For example, the guess tokens (631-5, 631-6, 631-7, 631-8) can be generated based on the valid token (v2) (631-1) and the candidate tokens (631-2, 631-3, 631-4). For example, the number of guess tokens (631-5, 631-6, 631-7, 631-8) can be determined based on the number of valid tokens (v2) (631-1) and the number of candidate tokens (631-2, 631-3, 631-4).For example, the number of guess tokens (631-5, 631-6, 631-7, 631-8)) (e.g., 4) can be defined as the product of the number of valid tokens (v0) (631-1) (e.g., 1) and the number of candidate tokens (631-2, 631-3, 631-4) (e.g., 3) (e.g., 4) and the number of draft tokens (e.g., 1).

[0176] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, a first output sequence (s1) (641) can be inferred from a first input sequence (s1) (631). For example, the first output sequence (s1) (641) can be included in a second set of output sequences. For example, the first output sequence (s1) (641) can include a valid token (v3) (641-1), candidate tokens (641-2, 641-3, 641-4), and guess tokens (641-5, 641-6, 641-7, 641-8). For example, the valid token (v4) (641-1) can be inferred from the valid token (v3) (611-1). For example, candidate tokens (641-2, 641-3, 641-4) can be inferred from candidate tokens (631-2, 631-3, 631-4). For example, guess tokens (641-5, 641-6, 641-7, 641-8) can be inferred from guess tokens (631-5, 631-6, 631-7, 631-8).

[0177] For example, candidate tokens (641-2, 641-3, 641-4) may include candidate tokens that can be used as a second draft token (d4) inferred (predicted) from a first draft token (d3). For example, candidate tokens may include candidate token (d4_1) (641-2), candidate token (d4_2) (641-3), and candidate token (d4_3) (641-4). For example, guess tokens (641-5, 641-6, 641-7, 641-8) may include guess token (f4_v3) (641-5), guess token (f5_d41) (641-6), guess token (f5_d42) (641-7), and guess token (f5_d43) (641-8). For example, the guess token (f4_v3) (641-5) can be associated with the valid token (v3) (641-1). For example, the guess token (f5_d41) (641-6) can be associated with the candidate token (d4_1) (641-2). For example, the guess token (f5_d42) (641-7) can be associated with the candidate token (d4_2) (641-3). For example, the guess token (f5_d43) (641-8) can be associated with the candidate token (d4_3) (641-4).

[0178] For example, the electronic device (301) can perform verification on the first input sequence (s1) (631) and the first output sequence (s1) (641) of the trained model (345). For example, the electronic device (301) can identify the candidate token (d4_3) (641-4) inferred from the candidate token (d3_c) (631-4) as the approved candidate token if the candidate token (d3_c) (631-4) among the candidate tokens (631-2, 631-3, 631-4) is identical to the valid token (v3) (641-1). In this case, among the candidate tokens (641-2, 641-3, 641-4), the remaining candidate tokens, candidate token (d4_a) (641-2) and candidate token (d4_b) (641-3), are rejected, and the candidate token and guess token associated with the rejected tokens may be rejected. For example, candidate token (d4_1)(641-2) inferred from candidate token (d3_a)(631-2), candidate token (d4_2)(641-3) inferred from candidate token (d3_b)(631-3), some of the guess tokens (631-5, 631-6, 631-7) among the guess tokens (631-5, 631-6, 631-7, 631-8), and some of the guess tokens (641-5, 641-6, 641-7, 641-8) among the guess tokens (641-5, 641-6, 641-7) among the guess tokens (641-5, 641-6, 641-7) among the guess tokens (641-8).

[0179] For example, the electronic device (301) can identify the last approved candidate token among the approved candidate tokens in the first output sequence (s1) (641) as the candidate token (d4_3) (641-4). For example, the last approved candidate token (d4_3) (641-4) can be used as the valid token of the first input sequence (s1) in the next drive for sentence generation using the trained model (345). Additionally, the electronic device (301) can identify the guess token (f5_d43) (641-8) associated with the candidate token (d4_3) (641-4) in the first output sequence (s1) (641). For example, the guess token (f5_d43) (641-8) can be used as the first draft token in the next drive.

[0180] For example, the electronic device (301) may add the tokens approved according to the verification as at least part of the sentence to be generated. For example, assume that the electronic device (301) identifies the valid token (v3) (641-1) and the candidate token (d4_3) (641-4) of the first output sequence (s1) (641) as the valid tokens in the verification of the operation of the trained model (345) exemplified in the example (602) of FIG. 6a. In this case, the sentence to be generated may include the word corresponding to the valid token (v0) (611-1), the word corresponding to the valid token (v1) (621-1), the candidate token (d2_2) (621-3) (or the word corresponding to the valid token (v2) (631-1)), the word corresponding to the valid token (v3) (641-1), and the word corresponding to the candidate token (d4_3) (641-4). Afterward, the electronic device (301) may perform the following operation of the trained model (345) to determine the word following the word corresponding to the candidate token (d4_3) (641-4). For example, the sentence to be generated may be a sentence generated using the first input sequence (s1) (631) (and / or the first output sequence (s1) (641)). For example, the electronic device (301) may generate eight sentences by performing the above operations for each of the input sequences in the second set (e.g., the first input sequence (s1) (631)).

[0181] Referring to FIGS. 6a and 6b, the electronic device (301) can generate a set of output sequences from a set of input sequences and perform verification. FIGS. 6a and 6b illustrate cases where the sequences within each set have not yet completed sentence generation (or have not finished sentence generation), but the present disclosure is not limited thereto. For example, while generating a plurality of sentences, some sentences may be generated faster than others. In this case, inference and verification of the sequences for the sentences that have been generated may not be performed (or, the performance may be bypassed, or the performance may be stopped). Specific details regarding this may be referenced in FIG. 6c below.

[0182] FIG. 6c illustrates an example of a case where the generation of some sentences is terminated in multi-sentence generation using an SSD.

[0183] FIG. 6c illustrates examples (651, 652) of cases where the electronic device (301) has finished generating sentences for a set of output sequences inferred in each of the first drive of the trained model (345) and the second drive following the first drive, and for some of the output sequences within the set. Example (651) illustrates an example of the first drive of the trained model (345). Example (652) illustrates an example of the second drive of the trained model (345). In FIG. 6c, for convenience of explanation, it is assumed that each sequence contains 8 tokens, generates a total of 8 sentences (e.g., s1 to s8), and that the trained model (345) is AR-64.

[0184] Referring to Example (651), the electronic device (301) can generate (or infer) output sequences of the first set (620) by providing input sequences of the first set (610) to a trained model (345). In Example (651), the electronic device (301) can identify one or more approved candidate tokens for each of the output sequences of the first set (620) based on verification of the output sequences of the first set (620) and the input sequences of the first set (610), and can identify whether the last approved candidate token among the one or more candidate tokens corresponds to a reference token. As a non-limiting example, the word corresponding to the reference token may include a period (.), or an exclamation mark (!), or a question mark (?). However, the present disclosure is not limited thereto. For example, the electronic device (301) can identify that sentence generation of the sequence is terminated when the last approved candidate token corresponds to the reference token. In the following example (651), the electronic device (301) assumes that among the output sequences of the first set (620), the second output sequence (s2), the fourth output sequence (s4), the seventh output sequence (s7), and the eighth output sequence (s8) have finished generating sentences.

[0185] Referring to example (652), the electronic device (301) can generate (or infer) output sequences of a second set (670) by providing input sequences of a second set (660) to a trained model (345). Referring to example (652), the input sequences of the second set (660) may include a first input sequence (s1) (661), a third input sequence (s3) (663), a fifth input sequence (s5) (665), and a sixth input sequence (s6) (666). In the example (652) of FIG. 6c, for convenience of explanation, the second input sequence (s2) (662), the fourth input sequence (s4) (664), the seventh input sequence (s7) (667), and the eighth input sequence (s8) (668) are shown, but the second set (660) may not include the second input sequence (s2) (662), the fourth input sequence (s4) (664), the seventh input sequence (s7) (667), and the eighth input sequence (s8) (668). This may be because the second input sequence (s2) (662), the fourth input sequence (s4) (664), the seventh input sequence (s7) (667), and the eighth input sequence (s8) (668) are sequences in which sentence generation was terminated in a previous drive (e.g., the first drive). Since the second set (660) includes a first input sequence (s1) (661), a third input sequence (s3) (663), a fifth input sequence (s5) (665), and a sixth input sequence (s6) (666), the second set (670) may include a first output sequence (s1) (671), a third output sequence (s3) (673), a fifth output sequence (s5) (675), and a sixth output sequence (s6) (676). The second set (670) may also not include a second output sequence (s2) (672), a fourth output sequence (s4) (674), a seventh output sequence (s7) (677), and an eighth output sequence (s8) (678), just like the second set (660).In other words, the trained model (345) may not generate (or infer) the second output sequence (s2) (672), the fourth output sequence (s4) (674), the seventh output sequence (s7) (677), and the eighth output sequence (s8) (678).

[0186] Referring to the above description, the electronic device (301) may input dummy tokens for the resources (e.g., 32 tokens) for the second input sequence (s2) (662), the fourth input sequence (s4) (664), the seventh input sequence (s7) (667), and the eighth input sequence (s8) (668) instead of using them to substantially generate sentences. Additionally, the electronic device (301) (or at least one processor (310), CPU (311)) may provide the KV cache (or KV value) for the terminated sequence as a dummy value to the trained model (345) as it identifies the sequence for which sentence generation has ended. This may be because the electronic device (301) uses a fixed setting value to be used for sampling each sequence within the set of input sequences. In some drives (e.g., the second drive), resources for some sequences (e.g., the second input sequence (s2) (662), the fourth input sequence (s4) (664), the seventh input sequence (s7) (667), and the eighth input sequence (s8) (668)) of the set of input sequences to be input to the trained model (345) may be wasted. Additionally, when generating multiple sentences, since the output for all sentences (e.g., eight sentences) is provided to the user after all sentences have been generated, the ability to improve the sentence generation speed may be limited if resources for already completed sequences are not utilized as described above.

[0187] In the following, the present disclosure may dynamically adjust the setting value to be used for sampling each of the sequences (or sets of sequences) for sentences when generating multiple sentences using a trained model (345). The present disclosure may allocate (or distribute) the resources (e.g., tokens) used for generating some sentences to other sequences for generating other sentences when the generation of some sentences is finished (or completed) while generating multiple sentences using the trained model (345). Allocating resources to a sequence may include adjusting the setting value of the sequence. This may be because the number of tokens (e.g., candidate tokens and guess tokens) included in the sequence is adjusted when using the adjusted setting value. Accordingly, the present disclosure may shorten the completion time (or generation time) for generating all sentences by minimizing wasted resources when generating multiple sentences. For example, as illustrated in the example of FIG. 6c, if dummy tokens (and dummy values) are used without allocating the resources of a terminated sequence for other sequences, the utilization of the NPU (312) (or the trained model (345)) may be reduced. In contrast, the present disclosure, as described below, can increase the utilization of the NPU (312) (or the trained model (345)) by allocating the resources for other sequences. In other words, the present disclosure can improve the acceptance rate (or token generation rate) of candidate tokens and improve sentence generation speed by utilizing optimized setting values ​​of sequences for every run of the trained model (345).

[0188] FIG. 7 illustrates an example of a method for dynamically adjusting a setting value to be used for sampling each of the input sequences by redistributing resources for sequences for which sentence generation has ended to the input sequences while generating sentences using a trained model.

[0189] Figure 7 illustrates an example of a flow of operations for dynamically adjusting setting values ​​to be used for sampling while generating sentences using a trained model.

[0190] At least some of the above methods of FIG. 7 may be performed by the electronic device (301) of FIG. 3. For example, at least some of the above methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (301). In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0191] In operation (700), the electronic device (301) may execute a software application for providing a sentence generation service. For example, the electronic device (301) may execute the software application based on receiving user input for an icon (or visual object) representing the software application. In the above example, the software application for providing the sentence generation service is described, but the present disclosure is not limited thereto. The software application may include a software application for generating an output using a trained model (345) including an LLM (403).

[0192] In operation (705), the electronic device (301) may generate a prompt. For example, the electronic device (301) may receive input from a user based on the execution of the software application. For example, the input may include a touch input regarding the display of the electronic device (301) or a voice input regarding the microphone of the electronic device (301). For example, the electronic device (301) may generate the prompt indicating the input. For example, the prompt may include at least one token.

[0193] In operation (710), the electronic device (301) can identify a defined setting value regarding the number of sentences and the trained model (345). For example, the electronic device (301) can identify the number of sentences to be generated based on the execution (or activation) of the multi-sentence generation function of the software application. By example, without limitation, the number of sentences to be generated may be 8. In this case, it is assumed that the maximum number of tokens that can be input to the trained model (345) is 64 (i.e., AR-64). For example, the electronic device (301) may distribute resources uniformly to the sequence for each sentence to ensure that the generation of all sentences (e.g., 8) is completed as quickly as possible during multi-sentence generation. Accordingly, each sequence may contain 8 tokens. The electronic device (301) can identify the set value defined according to the configuration information (e.g., graph parameters) of the trained model (345) when generating a sequence using eight tokens. For example, the electronic device (301) can identify the set value defined for the trained model (345) from the LUT (e.g., Table 1 above). As an example without limitation, the LUT may be stored in the electronic device (301) (or memory (320)). For example, the set value may be used to generate an input sequence containing tokens having a number less than or equal to the maximum number of torques that can be input to the trained model (345). For example, the set value may be used for sampling draft tokens.

[0194] In operation (715), the electronic device (301) may generate a set of input sequences. For example, the electronic device (301) may generate a set of input sequences corresponding to the number of sentences to be generated during multi-sentence generation (e.g., 8). As an example without limitation, the electronic device (301) may generate a first set containing 8 input sequences. Each of the input sequences of the first set, to be input (or provided) upon the initial operation of the trained model (345), may include tokens (e.g., 8 tokens) generated based on the same first setting value (e.g., {3}). Each of the input sequences of the first set may include valid tokens, candidate tokens, and guess tokens. For example, the electronic device (301) may identify the last token among at least one token of the prompt as a valid token. For example, the electronic device (301) may identify the number of draft tokens associated with the valid token according to the setting value. For example, the electronic device (301) can generate candidate tokens sampled from draft tokens according to the first setting value defined with respect to the trained model (345). For example, the electronic device (301) can generate guess tokens according to the valid tokens and candidate tokens.

[0195] As a non-limiting example, if the valid token of the input sequence is the last token of the prompt, the electronic device (301) may use any token (or any value) as the draft token. If the valid token of the input sequence is the last token of the prompt, the operation of the trained model (345) using the input sequence may be referred to as the initial operation.

[0196] As an example that is not limited, if the number of tokens in the input sequence is less than the maximum number of tokens that can be input to the trained model (345), the input sequence may include any additional tokens. For example, the number of any tokens may correspond to the difference between the maximum number of tokens that can be input to the trained model (345) and the number of tokens in the input sequence.

[0197] In operation (720), the electronic device (301) may provide a set of input sequences to the trained model (345). For example, the electronic device (301) may provide the first set of input sequences to the trained model (345) for the initial operation.

[0198] In operation (725), the electronic device (301) may generate a set of output sequences. For example, the electronic device (301) may generate a first set of output sequences inferred from the first set of input sequences using a trained model (345). For example, each of the output sequences of the first set may be inferred (or generated) from each of the input sequences of the first set. For example, the output sequence may include output values ​​(e.g., logit values). Or, for example, the output sequence may include tokens. For example, the tokens of the output sequence may be generated as they are decoded by a tokenizer from the output values. For convenience of explanation, it is assumed below that the output sequence includes tokens. For example, each of the output sequences of the first set may include a valid token inferred from a valid token of a corresponding input sequence among the input sequences of the first set, candidate tokens inferred from candidate tokens of the corresponding input sequence, and guess tokens inferred from guess tokens of the corresponding input sequence. As a non-limiting example, the output sequence may include a token corresponding to any token of the input sequence. For convenience of explanation, the description of any token of the input sequence and the output sequence is omitted below.

[0199] In operation (730), the electronic device (301) can perform verification. For example, the electronic device (301) can perform verification on each of the first set of input sequences input to the trained model (345) and each of the first set of output sequences output from the trained model (345). Specific details regarding the verification may be referenced in FIG. 4c and FIG. 6b described above.

[0200] In operation (735), the electronic device (301) can identify whether there is a terminated sequence. For example, the electronic device (301) can identify whether there is a terminated sequence among the first set of output sequences. For example, the electronic device (301) can identify whether a criterion for the termination (or completion) of sentence generation is satisfied for each of the first set of output sequences. For example, the criterion may include a correspondence between a candidate token approved in the verification among the candidate tokens of the output sequence and a criterion token for the termination (or completion) of sentence generation. As a non-limiting example, the word corresponding to the criterion token may include a period (.), or an exclamation mark (!), or a question mark (?). However, the present disclosure is not limited thereto. If the criterion is satisfied, the electronic device (301) can generate a completed sentence.

[0201] Alternatively, for example, the above criteria may include receiving user input for the termination of sentence generation. The electronic device (301) may generate an incomplete sentence if the above criteria are satisfied.

[0202] In operation (735), the electronic device (301) may perform operation (715) again if there is no terminated sequence among the first set of output sequences. In this case, when performing operation (715) again, the electronic device (301) may generate a set of input sequences by reusing the first set value. Alternatively, in operation (735), the electronic device (301) may perform operation (740) again if there is a terminated sequence among the first set of output sequences.

[0203] In operation (740), the electronic device (301) can identify whether sentence generation for all sequences has ended. For example, the electronic device (301) can identify whether all of the output sequences of the first set are terminated sequences.

[0204] In operation (740), the electronic device (301) may perform operation (760) if all of the output sequences of the first set are terminated sequences. Alternatively, in operation (740), the electronic device (301) may perform operation (745) if at least some of the sequences of the first set of output sequences are not terminated.

[0205] In operation (745), the electronic device (301) can identify the number of tokens of a terminated sequence. For example, the electronic device (301) can identify the number of tokens of a terminated sequence among the output sequences of the first set. In one example, if one sequence among the output sequences of the first set is terminated, the device can identify eight tokens of the terminated sequence. Or, in the example, if two sequences among the output sequences of the first set are terminated, the device can identify sixteen tokens of the terminated sequences. In operation (745), identifying the number of tokens of a terminated sequence may be referred to as identifying the resources of the terminated sequence.

[0206] In operation (750), the electronic device (301) can identify the number of remaining sentences. For example, the electronic device (301) can identify the number of remaining sentences excluding the sentences of the terminated sequence from the number of sentences to be generated according to multi-sentence generation (e.g., 8). In the above example, if there are 2 terminated sequences, the number of remaining sentences may be 6.

[0207] In operation (755), the electronic device (301) may select a setting value to be used for at least some of the input sequences based on the token generation rate. For example, the electronic device (301) may select a setting value to be used for at least some of the second set of input sequences to be provided to the trained model (345) in a second drive following the first drive of the trained model (345), based on the token generation rate. In the present disclosure, selecting a setting value may include adjusting or maintaining the setting value. For example, the token generation rate may be defined for each setting value. For example, a first setting value may have a first token generation rate. For example, a second setting value different from the first setting value may have a second token generation rate. By example, without limitation, the second token generation rate may be the same as the first token generation rate. Or, by example, without limitation, the second token generation rate may be different from the first token generation rate. For example, the token generation rate defined according to the setting value can be defined as shown in the table below.

[0208] Configuration Values ​​Draft Number of Tokens Minimum Number of Tokens Token Generation Rate {3}18 1.15{4}110 1.19{5}112 1.21{6}114 1.22...{15}132 1.24...{32}164 1.24{1, 1}29 1.12{1, 2}212 1.12{1, 3}215 1.12...{3, 2}230 1.28...

[0209] Table 3 described above can be referenced as a LUT (look-up table) that is defined for a trained model (345) and indicates the number of tokens included in the input sequence and the number of draft tokens included in the input sequence.

[0210] For example, the electronic device (301) may select a setting value to be applied to at least some of the input sequences of the second set based on a token generation rate defined (or tested, preset, calculated) according to a setting value. For example, the number of input sequences of the second set may be less than the number of input sequences of the first set. This may be because sentence generation based on at least some of the sequences of the input sequences of the first set has ended. In the above example, if sentence generation of two input sequences among the input sequences of the first set has ended, the number of input sequences of the second set may be 6. For example, the electronic device (301) may select (or adjust) a setting value to be used in some of the input sequences of the second set as a second setting value different from the first setting value. At this time, the electronic device (301) may select (or maintain) the first setting value to be used in some other input sequences that are distinct (or different) from some of the input sequences of the second set of input sequences. However, the present disclosure is not limited thereto. For example, the electronic device (301) may select (or adjust) the setting value to be used in all of the input sequences of the second set of input sequences to a second setting value that is different from the first setting value. In the present disclosure, when selecting the setting value to be used in at least some of the input sequences of the second set of input sequences, the electronic device (301) may select a setting value that generates all sentences (e.g., 8 sentences) as quickly as possible (or is likely to generate all sentences (e.g., 8 sentences) as quickly as possible). In other words, the electronic device (301) may allocate (or distribute) the resources of a sequence in which sentence generation has ended to a sequence in which sentence generation has not ended (or sentence generation is in progress).Specific details regarding a method for selecting setting values ​​to generate all sentences to be generated as quickly as possible when generating multiple sentences (or a method for allocating or distributing resources) can be exemplified below with reference to FIGS. 8 to 11.

[0211] After operation (755), the electronic device (301) may perform operation (715) again. In operation (715) performed again after operation (755), the electronic device (301) may generate a set of input sequences. For example, the electronic device (301) may generate the second set of input sequences. For example, the electronic device (301) may generate each of the input sequences of the second set according to the first setting value. For example, the electronic device (301) may generate each of the other input sequences that are distinct (or different) from the input sequences of the second set according to the second setting value. For example, the number of tokens included in the input sequence generated according to the first setting value may be different from the number of tokens included in the input sequence generated according to the second setting value. The number of tokens included in the input sequence generated according to the second setting value above may be greater than the number of tokens included in the input sequence generated according to the first setting value above.

[0212] After the re-executed operation (715), operations (720) through (740) may be performed again. For example, in the re-executed operation (740), the electronic device (301) assumes that sentence generation for all sequences has been completed.

[0213] In operation (760), the electronic device (301) can output sentences. For example, the electronic device (301) can output sentences generated according to the operations described above. For example, the sentences to be output may be sentences generated according to multiple sentence generation (e.g., 8 sentences).

[0214] For example, each of the above sentences may include a word corresponding to a valid token and a word corresponding to an approved candidate token of an output sequence output in each run of the trained model (345). As an example without limitation, the electronic device (301) may output the sentences as auditory information through the speaker of the electronic device (301). As an example without limitation, the electronic device (301) may output the sentences as visual information through the display of the electronic device (301).

[0215] FIGS. 8 to 11 illustrate examples of a method for dynamically adjusting a setting value to be used for sampling each of the input sequences by redistributing resources for a sequence for which sentence generation has ended to the input sequences.

[0216] FIGS. 8 through 11 illustrate examples of a method for dynamically adjusting a setting value to be used for sampling by redistributing resources for a sequence in which sentence generation is terminated in a specific drive to input sequences to be generated in a drive following said specific drive, while the electronic device (301) sequentially drives a trained model (345). For example, FIG. 8 illustrates an example of a first drive of the trained model (345) when the electronic device (301) performs at least some of the operations of FIG. 5. For example, FIG. 9 illustrates an example of a second drive of the trained model (345) that follows (or is subsequent to) said first drive of the electronic device (301) when the electronic device (301) performs at least some of the operations of FIG. 5. For example, FIG. 10 illustrates an example of a third drive that follows (or follows) the second drive of the trained model (345) when the electronic device (301) performs at least some of the operations of FIG. 5. For example, FIG. 11 illustrates an example of a fourth drive that follows (or follows) the third drive of the trained model (345) when the electronic device (301) performs at least some of the operations of FIG. 5.

[0217] In FIGS. 8 to 11, for convenience of explanation, it is assumed that the trained model (345) is AR-64 and the number of sentences to be generated according to the function for multi-sentence generation is 8. In other words, the maximum number of tokens that can be input in one run of the trained model (345) may be 64. However, the present disclosure is not limited to the examples described above. For example, the trained model (345) may be AR-32. Or, for example, the number of sentences to be generated may be less than 8 (e.g., 4) or more than 8 (e.g., 16).

[0218] Referring to FIG. 8, the electronic device (301) can generate a first set (810) of input sequences for sentences to be generated based on the execution of a function for generating multiple sentences. For example, the electronic device (301) can identify that the number of sentences to be generated during multi-sentence generation is 8 in order to generate the first set (810) of input sequences. For example, the electronic device (301) can identify a first set (810) containing 8 input sequences to be used to generate 8 sentences. For example, the electronic device (301) can uniformly allocate (or distribute) 64 tokens, which is the number of tokens that can be simultaneously input (or provided) to the trained model (345), to the 8 input sequences. Accordingly, the electronic device (301) can identify that the number of tokens to be included in each input sequence is 8. For example, the electronic device (301) can identify a setting value that causes the token generation rate to have a maximum value for each input sequence that may contain eight tokens. For example, the setting value may be defined for a trained model (345). Referring to Table 3, among the setting values ​​(or candidate values) requiring a number of eight or fewer tokens, the electronic device (301) can select (or adjust, identify, determine) a first setting value (e.g., {3}) having a maximum token generation rate as the setting value to be used in each of the first set of input sequences. For example, the electronic device (301) can generate the first set of input sequences (810) based on the first setting value (e.g., {3}).

[0219] In the example of FIG. 8, the first set (810) may include a first input sequence (s1) (811), a second input sequence (s2) (812), a third input sequence (s3) (813), a fourth input sequence (s4) (814), a fifth input sequence (s5) (815), a sixth input sequence (s6) (816), a seventh input sequence (s7) (817), and an eighth input sequence (s8) (818).

[0220] For example, the first input sequence (s1) (811) may include valid tokens (v0) (811-1), candidate tokens (811-2, 811-3, 811-4), and guess tokens (811-5, 811-6, 811-7, 811-8) according to the first setting value (e.g., {3}). For example, the candidate tokens (811-2, 811-3, 811-4) may include candidate tokens sampled from the first draft token (d1) (e.g., the first draft token (d1) (421) of FIG. 4b). For example, the candidate tokens may include candidate token (d1_a) (811-2), candidate token (d1_b) (811-3), and candidate token (d1_c) (811-4). For example, the guess tokens (811-5, 811-6, 811-7, 811-8) may include guess token (f_x) (811-5), guess token (f_x) (811-6), guess token (f_x) (811-7), and guess token (f_x) (811-8). For example, the guess token (f_x) (811-5) may be associated with the valid token (v0) (811-1). For example, the guess token (f_x) (811-6) may be associated with the candidate token (d1_a) (811-2). For example, the guess token (f_x) (811-7) may be associated with the candidate token (d1_b) (811-3). For example, the guess token (f_x) (811-8) can be associated with the candidate token (d1_c) (811-4).

[0221] In FIG. 8, tokens included in the first input sequence (s1) (811) are illustrated for convenience of explanation, but the present disclosure is not limited thereto. For example, the content regarding the tokens included in each of the input sequences of the first set (810) may be substantially the same as the content regarding the tokens of the first input sequence (s1) (811). For examples regarding the number of tokens, the number of draft tokens, the setting value for sampling, and the token generation rate for the input sequences of the first set (810), the table below may be referenced.

[0222] Input sequences Settings Number of value tokens Token generation rate s1{3}81.15s2{3}81.15s3{3}81.15s4{3}81.15s5{3}81.15s6{3}81.15s7{3}81.15s8{3}81.15

[0223] Referring to Table 4 above, in the first drive, the input sequences of the first set (810) may include the same number of tokens according to the same setting value. This allows the electronic device (301) to equally assign tokens that can be assigned (or input) to the trained model (345) so that in the first drive (or initial drive), the possibility of sentences (e.g., 8 sentences) to be generated according to the multi-sentence generation function being generated simultaneously (or substantially simultaneously) is improved (or the deviation in generation time between sentences is minimized).

[0224] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, the electronic device (301) can generate a first set (820) of output sequences inferred from a first set (810). For example, the first set (820) may include a first output sequence (s1) (821), a second output sequence (s2) (822), a third output sequence (s3) (823), a fourth output sequence (s4) (824), a fifth output sequence (s5) (825), a sixth output sequence (s6) (826), a seventh output sequence (s7) (827), and an eighth output sequence (s8) (828). Each of the output sequences of the first set (820) may include eight tokens because they are inferred from each of the input sequences of the first set (810).

[0225] For example, the first output sequence (s1) (821) can be inferred from the first input sequence (s1) (811). For example, the first output sequence (s1) (821) may include a valid token (v1) (821-1), candidate tokens (821-2, 821-3, 821-4), and guess tokens (821-5, 821-6, 821-7, 821-8). For example, the valid token (v1) (821-1) can be inferred from the valid token (v0) (811-1). For example, the candidate tokens (821-2, 821-3, 821-4) can be inferred from the candidate tokens (811-2, 811-3, 811-4). For example, guess tokens (821-5, 821-6, 821-7, 821-8) can be inferred from guess tokens (811-5, 811-6, 811-7, 811-8). For example, candidate tokens (821-2, 821-3, 821-4) may include candidate token (d2_1)(821-2), candidate token (d2_2)(821-3), and candidate token (d2_3)(821-4). For example, the guess tokens (821-5, 821-6, 821-7, 821-8) may include the guess token (f2_v1) (821-5), the guess token (f3_d21) (821-6), the guess token (f3_d22) (821-7), and the guess token (f3_d23) (821-8). For example, the guess token (f2_v1) (821-5) may be associated with the valid token (v1) (821-1). For example, the guess token (f3_d21) (821-6) may be associated with the candidate token (d2_1) (821-2). For example, the guess token (f3_d22) (821-7) may be associated with the candidate token (d2_2) (821-3). For example, the guess token (f3_d23) (821-8) can be associated with the candidate token (d2_3) (821-4).

[0226] In FIG. 8, tokens included in the first output sequence (s1) (821) are shown for convenience of explanation, but the present disclosure is not limited thereto. For example, the content regarding the tokens included in each of the output sequences of the first set (820) may be substantially the same as the content regarding the tokens of the first output sequence (s1) (821).

[0227] For example, the electronic device (301) may perform verification on a first set (810) of input sequences and a first set (820) of output sequences of a trained model (345). For example, the electronic device (301) may perform verification on a first input sequence (s1) (811) and a first output sequence (s1) (821). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). Specific details regarding the verification may be referenced in FIG. 4c and FIG. 6b described above. For convenience of explanation, the electronic device (301) assumes that, based on a comparison of each of the candidate tokens (811-2, 811-3, 811-4) sampled from the valid token (v1) (821-1) and the first draft token (d1), the candidate token (d1_b) (811-3) among the candidate tokens (811-2, 811-3, 811-4) is approved. For example, the electronic device (301) may identify the last approved candidate token among the approved candidate tokens in the first output sequence (s1) (821) as the candidate token (d2_2) (821-3). In the above example, since the number of draft tokens is 1, the approved candidate token among the candidate tokens (811-2, 811-3, 811-4) of the first draft token (d1) is the last approved candidate token, but the present disclosure is not limited thereto. For example, if the number of draft tokens is 2 or more, the last approved candidate token may be one of the multiple approved tokens. For example, the last approved candidate token (d2_2) (821-3) may be used as the valid token (v0) of the first input sequence (s1) in the second drive of FIG. 9, which is the next drive for sentence generation using the trained model (345).Additionally, the electronic device (301) can identify a guess token (f3_d22) (821-7) associated with a candidate token (d2_2) (821-3) in the first output sequence (s1) (821). For example, the guess token (f3_d22) (821-7) can be used as the first draft token (d1) in the second drive. In the above example, it is described assuming that the first output sequence (s1) (821) is not a sequence in which sentence generation has ended, but the present disclosure is not limited thereto.

[0228] For example, the electronic device (301) can identify whether each of the output sequences of the first set (820) is a sequence in which sentence generation has ended. For example, the electronic device (301) can identify whether a criterion for the termination of sentence generation is satisfied for each of the output sequences of the first set (820). For example, the electronic device (301) can identify whether the last accepted candidate token (e.g., candidate token (d2_2) (821-3)) among the tokens of the first output sequence (s1) (821) corresponds to the criterion token (899). For example, the electronic device (301) can determine that the first output sequence (s1) (821) is a sequence in which sentence generation has ended if the last accepted candidate token (e.g., candidate token (d2_2) (821-3)) among the tokens of the first output sequence (s1) (821) corresponds to the criterion token (899).

[0229] Additionally, for example, the electronic device (301) can determine whether all output sequences of the first set (820) are terminated sequences. In the example of FIG. 8, for convenience of explanation, it is assumed that among the output sequences of the first set (820), the first output sequence (s1) (821) is a sequence in which sentence generation has ended, and the remaining output sequences are sequences in which sentence generation has not ended (or sentence generation is in progress). For example, the electronic device (301) can perform the operation of FIG. 9 based on identifying that not all output sequences of the first set (820) are sequences in which sentence generation has ended.

[0230] Referring to FIG. 9, the electronic device (301) can identify the number of tokens of the first output sequence (s1) (821) that has ended among the output sequences of the first set (820) of FIG. 8. For example, the electronic device (301) can identify 8 tokens of the first output sequence (s1) (821) as redistributable (or reallocable) resources. For example, the electronic device (301) can identify the number of remaining sentences. For example, the electronic device (301) can identify 7 remaining sentences, as the generation of one sentence among the 8 sentences has ended. For example, the electronic device (301) can identify the output sequences of the first set (820) for the remaining sentences. For example, the output sequences of the first set (820) for the above remaining sentences may include a second output sequence (s2) (822), a third output sequence (s3) (823), a fourth output sequence (s4) (824), a fifth output sequence (s5) (825), a sixth output sequence (s6) (826), a seventh output sequence (s7) (827), and an eighth output sequence (s8) (828).

[0231] For example, the electronic device (301) may select a setting value to be used in at least some of the input sequences based on the token generation rate. For example, the electronic device (301) may identify a second set (910) comprising seven input sequences corresponding to the second output sequence (s2) (822), the third output sequence (s3) (823), the fourth output sequence (s4) (824), the fifth output sequence (s5) (825), the sixth output sequence (s6) (826), the seventh output sequence (s7) (827), and the eighth output sequence (s8) (828) of the first set (820). For example, the second set (910) may include a second input sequence (s2) (912), a third input sequence (s3) (913), a fourth input sequence (s4) (914), a fifth input sequence (s5) (915), a sixth input sequence (s6) (916), a seventh input sequence (s7) (917), and an eighth input sequence (s8) (918). In FIG. 9, a first input sequence (s1) (911) is shown, but the first input sequence (s1) (911) may not be included in the second set (910). This may be because the first output sequence (s1) (821) corresponding to the first input sequence (s1) (911) is a sequence in which sentence generation has ended.

[0232] For example, the electronic device (301) may recognize that when resources (e.g., 8 tokens) for the first output sequence (s1) (821) are distributed evenly among the sequences for the remaining sentences, the number of tokens for each of the second input sequence (s2) (912), the third input sequence (s3) (913), the fourth input sequence (s4) (914), the fifth input sequence (s5) (915), the sixth input sequence (s6) (916), the seventh input sequence (s7) (917), and the eighth input sequence (s8) (918) of the second set (910) increases from 8 to 9. However, referring to Table 3 above, even with 9 tokens, the electronic device (301) may identify (or recognize) that it is not possible to select a setting value having a token generation rate higher than the first setting value (e.g., {3}). In other words, when using a method that distributes resources evenly to input sequences for residual sentences, a token generation rate higher than the current token generation rate (e.g., 1.15) (or the average value of the token generation rates of the sequences) (or the average value of the token generation rates of the sequences) may not be provided. As a more specific example, the electronic device (301) can identify a first token generation rate (e.g., 1.15) according to the first setting value (e.g., {3}), a second token generation rate (e.g., 1.19) according to the second setting value (e.g., {4}), and a third token generation rate (e.g., 1.10) according to the third setting value (e.g., {2}). For example, the electronic device (301) may determine the second set value to be used in some of the input sequences of the second set (910) based on the third token generation rate which is lower than the first token generation rate and the second token generation rate which is higher than the first token generation rate. In the above example, the third token generation rate is described as being lower than the first token generation rate, but the present disclosure is not limited thereto. For example, the third token generation rate may be equal to the first token generation rate.At this time, the case where the third token generation rate is the same as the first token generation rate may include cases where the number of tokens allocated at the same setting value (e.g., {3}) is different (e.g., 8 and 9).

[0233] The electronic device (301) may perform differential resource distribution when equal resource distribution is not possible. For example, the electronic device (301) may use a previously used setting value (e.g., the first setting value being {3}) for some of the input sequences of the second set (910), and may use a second setting value different from the first setting value for other input sequences of the second set (910) that are different from the input sequences said to be. For example, the electronic device (301) may further allocate resources (e.g., 8 tokens) for the first output sequence (s1) (821) to four input sequences, each with two tokens. Referring to the example in FIG. 9, the electronic device (301) may additionally assign two tokens to each of the second input sequence (s2) (912), the third input sequence (s3) (913), the fourth input sequence (s4) (914), and the fifth input sequence (s5) (915) of the second set (910). Accordingly, the electronic device (301) may assign 10 tokens to each of the second input sequence (s2) (912), the third input sequence (s3) (913), the fourth input sequence (s4) (914), and the fifth input sequence (s5) (915), and assign 8 tokens to each of the sixth input sequence (s6) (916), the seventh input sequence (s7) (917), and the eighth input sequence (s8) (918). Examples of the number of tokens, the number of draft tokens, the setting value for sampling, and the token generation rate for the input sequences of the second set (910) may be referenced in the table below.

[0234] Input sequences Settings Number of value tokens Token generation rate s1---s2{4}101.19s3{4}101.19s4{4}101.19s5{4}101.19s6{3}81.15s7{3}81.15s8{3}81.15

[0235] Referring to Table 5 above, in the second drive, some of the input sequences of the second set (910) (e.g., second output sequence (s2) (912), third output sequence (s3) (913), fourth output sequence (s4) (914), fifth output sequence (s5) (915)) may contain 10 tokens according to the second setting value (e.g., {4}), and other some of the input sequences of the second set (910) (e.g., sixth output sequence (s6) (916), seventh output sequence (s7) (917), eighth output sequence (s8) (918)) may contain 8 tokens according to the first setting value (e.g., {3}). This allows the electronic device (301) to assign tokens that can be assigned (or input) to the trained model (345) so that, in the second drive, the generation speed of sentences (e.g., 8 sentences) to be generated according to the multi-sentence generation function is improved, and the possibility of sentences (e.g., 8 sentences) being generated simultaneously (or substantially simultaneously) (or the deviation in generation time between sentences is minimized) is improved.

[0236] For example, the second input sequence (s2) (912) may include valid tokens (v0) (912-1), candidate tokens (912-2, 912-3, 912-4, 912-5), and guess tokens (912-6, 912-7, 912-8, 912-9, 912-10) according to the second setting value (e.g., {4}). For example, the candidate tokens (912-2, 912-3, 912-4, 912-5) may include candidate tokens sampled from the first draft token (d1) (e.g., the first draft token (d1) (421) of FIG. 4b). For example, candidate tokens may include candidate token (d1_a) (912-2), candidate token (d1_b) (912-3), candidate token (d1_c) (912-4), and candidate token (d1_d) (912-5). For example, guess tokens (912-6, 912-7, 912-8, 912-9, 912-10) may include guess token (f_x) (912-6), guess token (f_x) (912-7), guess token (f_x) (912-8), guess token (f_x) (912-9), and guess token (f_x) (912-10). For example, guess token (f_x) (912-6) may be associated with valid token (v0) (912-1). For example, the guess token (f_x) (912-7) can be associated with the candidate token (d1_a) (912-2). For example, the guess token (f_x) (912-8) can be associated with the candidate token (d1_b) (912-3). For example, the guess token (f_x) (912-9) can be associated with the candidate token (d1_c) (912-4). For example, the guess token (f_x) (912-10) can be associated with the candidate token (d1_d) (912-5).

[0237] The content for each of the third input sequence (s3) (913), the fourth input sequence (s4) (914), and the fifth input sequence (s5) (915) can be substantially the same as the content for the second input sequence (s2) (912). For example, each of the third input sequence (s3) (913), the fourth input sequence (s4) (914), and the fifth input sequence (s5) (915) may contain 10 tokens.

[0238] For example, the sixth input sequence (s6) (916) may include a valid token (v0), a candidate token (d1_a), a candidate token (d1_b), a candidate token (d1_c), and guess tokens (f_x) according to the first setting value (e.g., {3}). The content regarding the tokens of the sixth input sequence (s6) (916) may be referenced substantially identically to the content regarding the first input sequence (s1) (811) of FIG. 8.

[0239] The content for each of the seventh input sequence (s7) (917) and the eighth input sequence (s8) (918) can be substantially the same as the content for the sixth input sequence (s6) (916). For example, each of the seventh input sequence (s7) (917) and the eighth input sequence (s8) (918) may contain eight tokens.

[0240] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, the electronic device (301) can generate a second set (920) of output sequences inferred from a second set (910). For example, the second set (920) may include a second output sequence (s2) (922), a third output sequence (s3) (923), a fourth output sequence (s4) (924), a fifth output sequence (s5) (925), a sixth output sequence (s6) (926), a seventh output sequence (s7) (927), and an eighth output sequence (s8) (928). In FIG. 9, a first output sequence (s1) (921) is shown, but the first output sequence (s1) (921) may not be included in the second set (920). This may be because the first output sequence (s1) (821) corresponding to the first output sequence (s1) (921) is a sequence in which sentence generation has ended.

[0241] For example, the second output sequence (s2) (922) can be inferred from the second input sequence (s2) (912). For example, the second output sequence (s2) (922) may include a valid token (v1) (922-1), candidate tokens (922-2, 922-3, 922-4, 922-5), and guess tokens (922-6, 922-7, 922-8, 922-9, 922-10). For example, the valid token (v1) (922-1) can be inferred from the valid token (v0) (912-1). For example, candidate tokens (922-2, 922-3, 922-4, 922-5) can be inferred from candidate tokens (912-2, 912-3, 912-4, 912-5). For example, guess tokens (922-6, 922-7, 922-8, 922-9, 922-10) can be inferred from guess tokens (912-6, 912-7, 912-8, 912-9, 912-10). For example, candidate tokens (922-2, 922-3, 922-4, 922-5) may include candidate token (d1_a) (922-2), candidate token (d1_b) (922-3), candidate token (d1_c) (922-4), and candidate token (d1_d) (922-5). For example, guess tokens (922-6, 922-7, 922-8, 922-9, 922-10) may include guess token (f2_v1) (922-6), guess token (f3_d21) (922-7), guess token (f3_d22) (922-8), guess token (f3_d23) (922-9), and guess token (f3_d24) (922-10). For example, the guess token (f2_v1) (922-6) can be associated with the valid token (v1) (922-1). For example, the guess token (f3_d21) (922-7) can be associated with the candidate token (d2_1) (922-2). For example, the guess token (f3_d22) (922-8) can be associated with the candidate token (d2_2) (922-3). For example, the guess token (f3_d23) (922-9) can be associated with the candidate token (d2_3) (922-4).For example, the guess token (f3_d24) (912-10) can be associated with the candidate token (d1_d) (912-5).

[0242] Additionally, for example, the sixth output sequence (s6) (926) can be inferred from the sixth input sequence (s6) (916). The contents of the tokens of the sixth output sequence (s6) (926) can be substantially referenced to the contents of the first output sequence (s1) (821) of FIG. 8.

[0243] For example, the electronic device (301) may perform verification on a second set (910) of input sequences and a second set (920) of output sequences of a trained model (345). For example, the electronic device (301) may perform verification on a second input sequence (s2) (912) and a second output sequence (s2) (922). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). Specific details regarding the verification may be referenced in FIG. 4c and FIG. 6b described above. For convenience of explanation, the electronic device (301) assumes that, based on a comparison of each of the candidate tokens (912-2, 912-3, 912-4, 912-5) sampled from the valid token (v1) (922-1) and the first draft token (d1), the candidate token (d1_d) (912-5) among the candidate tokens (912-2, 912-3, 912-4, 912-5) is approved. For example, the electronic device (301) may identify the last approved candidate token among the approved candidate tokens in the second output sequence (s2) (912) as candidate token (d2_4) (922-5). In the above example, since the number of draft tokens is 1, the approved candidate token among the candidate tokens (912-2, 912-3, 912-4, 912-5) of the first draft token (d1) is the last approved candidate token, but the present disclosure is not limited thereto. For example, if the number of draft tokens is 2 or more, the last approved candidate token may be one of the multiple approved tokens. For example, the last approved candidate token (d2_4) (922-5) may be used as the valid token (v0) of the second input sequence (s2) in the third drive of FIG. 10, which is the next drive for sentence generation using the trained model (345).Additionally, the electronic device (301) can identify a guess token (f3_d24) (922-10) associated with a candidate token (d2_4) (922-5) in the second output sequence (s2) (922). For example, the guess token (f3_d24) (922-10) can be used as the first draft token (d1) in the third drive. In the above example, it is described assuming that the second output sequence (s2) (922) is not a sequence in which sentence generation has ended, but the present disclosure is not limited thereto.

[0244] Additionally, the electronic device (301) can perform verification on the sixth input sequence (s6) (916) and the sixth output sequence (s6) (926). For convenience of explanation, it is assumed below that the candidate token (d1_a) among the candidate tokens (d1_a, d1_b, d1_c) of the sixth input sequence (s6) (916), which are sampled from the valid token (v1) and the first draft token (d1), is approved based on comparison. For example, the electronic device (301) can identify the last approved candidate token among the approved candidate tokens in the sixth output sequence (s6) (926) as the candidate token (d2_1) of the sixth output sequence (s6) (926). For example, the last approved candidate token (d2_1) can be used as the valid token (v0) of the sixth input sequence (s6) (1016) in the third drive of FIG. 10, which is the next drive for sentence generation using the trained model (345). Additionally, the electronic device (301) can identify the guess token (f3_d21) associated with the candidate token (d2_1) in the sixth output sequence (s6) (926). For example, the guess token (f3_d24) can be used as the first draft token (d1) in the third drive.

[0245] For example, the electronic device (301) can identify whether each of the output sequences of the second set (920) is a sequence in which sentence generation has ended. For example, the electronic device (301) can identify whether a criterion for the termination of sentence generation is satisfied for each of the output sequences of the second set (920). For example, the electronic device (301) can identify whether the last accepted candidate token (e.g., candidate token (d2_4) (922-5)) among the tokens of the second output sequence (s2) (922) corresponds to the criterion token (999). For example, the electronic device (301) can determine that the second output sequence (s2) (922) is a sequence in which sentence generation has ended if the last accepted candidate token (e.g., candidate token (d2_4) (922-5)) among the tokens of the second output sequence (s2) (922) corresponds to the criterion token (999). For example, the electronic device (301) can identify whether the last approved candidate token (e.g., candidate token (d2_1)) among the tokens of the sixth output sequence (s6) (926) corresponds to the reference token (999). For example, if the last approved candidate token (e.g., candidate token (d2_1)) among the tokens of the sixth output sequence (s6) (926) corresponds to the reference token (999), the electronic device (301) can determine the sixth output sequence (s6) (926) as the sequence where sentence generation has ended.

[0246] Additionally, for example, the electronic device (301) can determine whether all output sequences of the second set (920) are terminated sequences. In the example of FIG. 9, for convenience of explanation, it is assumed that among the output sequences of the second set (920), the second output sequence (s2) (922) and the third output sequence (s3) (923) are sequences in which sentence generation has ended, and the remaining output sequences are sequences in which sentence generation has not ended (or sentence generation is in progress). In the above example, sequences with a large number of tokens included in the input sequence (e.g., the second output sequence (s2) (922) and the third output sequence (s3) (923)) are described as having sentence generation ended, but the present disclosure is not limited thereto. A sequence with a large number of tokens included in the input sequence is more likely to have sentence generation terminated (or, to have tokens included in the sentence generated), but a sequence with a large number of tokens included in the input sequence may not necessarily have sentence generation terminated faster than a sequence with a small number of tokens included in the input sequence. For example, the electronic device (301) can perform the operation of FIG. 10 based on identifying that not all output sequences of the second set (920) are sequences in which sentence generation has terminated.

[0247] Referring to FIG. 10, the electronic device (301) can identify the number of tokens of the second output sequence (s2) (922) and the third output sequence (s3) (923) that have been terminated among the output sequences of the second set (920) of FIG. 9. For example, the electronic device (301) can identify 20 tokens, which is the number of tokens of the second output sequence (s2) (922) and the third output sequence (s3) (923), as redistributable (or reallocable) resources. For example, the electronic device (301) can identify the number of remaining sentences. For example, the electronic device (301) can identify 5 remaining sentences, as the generation of 3 sentences out of 8 sentences has been terminated. For example, the electronic device (301) can identify the output sequences of the second set (920) for the remaining sentences. For example, the output sequences of the second set (920) for the above remaining sentences may include a fourth output sequence (s4) (924), a fifth output sequence (s5) (925), a sixth output sequence (s6) (926), a seventh output sequence (s7) (927), and an eighth output sequence (s8) (928).

[0248] For example, the electronic device (301) may select a setting value to be used in at least some of the input sequences based on a token generation rate. For example, the electronic device (301) may identify a third set (1010) comprising five input sequences corresponding to the fourth output sequence (s4) (924), the fifth output sequence (s5) (925), the sixth output sequence (s6) (926), the seventh output sequence (s7) (927), and the eighth output sequence (s8) (928) of the second set (920). For example, the third set (1010) may include the fourth input sequence (s4) (1014), the fifth input sequence (s5) (1015), the sixth input sequence (s6) (1016), the seventh input sequence (s7) (1017), and the eighth input sequence (s8) (1018). In FIG. 10, a first input sequence (s1) (1011), a second input sequence (s2) (1012), and a third input sequence (s3) (1013) are shown, but the first input sequence (s1) (1011), the second input sequence (s2) (1012), and the third input sequence (s3) (1013) may not be included in the third set (1010). This may be because the first output sequence (s1) (821), the second output sequence (s2) (922), and the third output sequence (s3) (923) corresponding to the first input sequence (s1) (1011), the second input sequence (s2) (1012), and the third input sequence (s3) (1013) are sequences in which sentence generation has ended.

[0249] For example, if the electronic device (301) allocates resources (e.g., 20 tokens) for the second output sequence (s2) (921) and the third output sequence (s3) (922) evenly to the sequences for the remaining sentences, the number of tokens for each of the fourth input sequence (s4) (1014), the fifth input sequence (s5) (1015), the sixth input sequence (s6) (1016), the seventh input sequence (s7) (1017), and the eighth input sequence (s8) (1018) of the third set (1010) can be increased by four. Referring to Table 3 above, the electronic device (301) can identify the setting value of each of the fourth input sequence (s4) (1014) and the fifth input sequence (s5) (1015) as a setting value (e.g., {6}) changed from the second setting value (e.g., {4}). In the above example, the number of tokens included in each of the fourth input sequence (s4) (1014) and the fifth input sequence (s5) (1015) can be increased from 10 to 14. Additionally, the electronic device (301) can identify the setting value of each of the sixth input sequence (s6) (1016), the seventh input sequence (s7) (1017), and the eighth input sequence (s8) (1018) as a setting value (e.g., {5}) changed from the first setting value (e.g., {3}). In the above example, the number of tokens included in each of the sixth input sequence (s6) (1016), the seventh input sequence (s7) (1017), and the eighth input sequence (s8) (1018) can be increased from 8 to 12. In the above example, the difference (e.g., 4) between the number of tokens (e.g., 14) of the fourth input sequence (s4) (1014) and the number of tokens (e.g., 10) of the fourth input sequence (s4) (914) may be the same as the difference (e.g., 4) between the number of tokens (e.g., 12) of the sixth input sequence (s6) (1016) and the number of tokens (e.g., 8) of the sixth input sequence (s6) (916).

[0250] The method described above may be a method of evenly distributing resources of a terminated sequence. However, the method described above may allow the 6th input sequence (s6) (1016), 7th input sequence (s7) (1017), and 8th input sequence (s8) (1018), which correspond to the 6th input sequence (s6) (916), 7th input sequence (s7) (917), and 8th input sequence (s8) (918) that used relatively fewer resources in the 2nd drive, to still use relatively fewer resources in the 3rd drive. In this case, the possibility that all sentences (e.g., 8 sentences) are generated simultaneously (or substantially simultaneously) when generating multiple sentences (or that the deviation in generation time between sentences is minimized) may be reduced. Therefore, instead of the method of evenly distributing as described above, the electronic device (301) may use a method of differentially distributing. Specific details related to this are exemplified in FIG. 10.

[0251] For example, the electronic device (301) may use a third setting value (e.g., {5}) for some of the input sequences of the third set (1010) and a fourth setting value (e.g., {6}) for other input sequences of the third set (1010) that are different from said input sequences. Referring to the example of FIG. 10, the electronic device (301) may additionally assign two tokens to each of the fourth input sequence (s4) (1014) and the fifth input sequence (s5) (1015) of the third set (1010). Additionally, for example, the electronic device (301) may additionally assign four tokens to the sixth input sequence (s6) (1016) of the third set (1010). Accordingly, the electronic device (301) may assign 12 tokens to each of the fourth input sequence (s4) (1014), the fifth input sequence (s5) (1015), and the sixth input sequence (s6) (1016). Additionally, for example, the electronic device (301) may assign 6 additional tokens to each of the seventh input sequence (s7) (1017) and the eighth input sequence (s8) (1018) of the third set (1010). Accordingly, the electronic device (301) may assign 14 tokens to each of the seventh input sequence (s7) (1017) and the eighth input sequence (s8) (1018). For examples of the number of tokens, the number of draft tokens, the setting value for sampling, and the token generation rate for the input sequences of the third set (1010), the following table may be referenced.

[0252] Input sequences Settings Number of value tokens Token generation rate s1---s2---s3---s4{5}121.21s5{5}121.21s6{5}121.21s7{6}141.22s8{6}141.22

[0253] Referring to Table 6 above, in the third drive, some of the input sequences of the third set (1010) (e.g., fourth output sequence (s4) (1014), fifth output sequence (s5) (1015), sixth output sequence (s6) (1016)) may contain 12 tokens according to the third setting value (e.g., {5}), and other some of the input sequences of the third set (1010) (e.g., seventh output sequence (s7) (1017), eighth output sequence (s8) (1018)) may contain 14 tokens according to the fourth setting value (e.g., {6}). This allows the electronic device (301) to assign tokens that can be assigned (or input) to the trained model (345) so that, in the third drive, the generation speed of sentences (e.g., 8 sentences) to be generated according to the multi-sentence generation function is improved, and the possibility of sentences (e.g., 8 sentences) being generated simultaneously (or substantially simultaneously) (or the deviation in generation time between sentences is minimized) is improved.

[0254] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, the electronic device (301) can generate a third set (1020) of output sequences inferred from a third set (1010). For example, the third set (1020) may include a fourth output sequence (s4) (1024), a fifth output sequence (s5) (1025), a sixth output sequence (s6) (1026), a seventh output sequence (s7) (1027), and an eighth output sequence (s8) (1028). In FIG. 10, a first output sequence (s1) (1021), a second output sequence (s2) (1022), and a third output sequence (s3) (1023) are shown, but the first output sequence (s1) (1021), the second output sequence (s2) (1022), and the third output sequence (s3) (1023) may not be included in the third set (1020). This may be because the sequences corresponding to the first output sequence (s1) (1021), the second output sequence (s2) (1022), and the third output sequence (s3) (1023) (e.g., the first output sequence (s1) (821), the second output sequence (s2) (922), and the third output sequence (s3) (923)) are sequences where sentence generation has ended.

[0255] For example, the electronic device (301) may perform verification on a third set (1010) of input sequences and a third set (1020) of output sequences of a trained model (345). For example, the electronic device (301) may perform verification on a fourth input sequence (s4) (1014) and a fourth output sequence (s4) (1024). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). Specific details regarding the verification may be referenced in FIG. 4c and FIG. 6b described above.

[0256] For example, the electronic device (301) can identify whether each of the output sequences of the third set (1020) is a sequence in which sentence generation has ended. For example, the electronic device (301) can identify whether a criterion for the termination of sentence generation has been satisfied for each of the output sequences of the third set (1020). For example, the electronic device (301) can identify whether the last approved candidate token among the tokens of each of the output sequences of the third set (1020) corresponds to the criterion token (1099). Additionally, for example, the electronic device (301) can determine whether all of the output sequences of the third set (1020) are sequences in which sentence generation has ended. In the example of FIG. 10, for convenience of explanation, it is assumed that among the output sequences of the third set (1020), the sixth output sequence (s6) (1026), the seventh output sequence (s7) (1027), and the eighth output sequence (s8) (1028) are sequences where sentence generation has ended, and the remaining output sequences are sequences where sentence generation has not ended (or sentence generation is in progress). For example, the electronic device (301) can perform the operation of FIG. 11 based on identifying that not all of the output sequences of the third set (1020) are sequences where sentence generation has ended.

[0257] Referring to FIG. 11, the electronic device (301) can identify the number of tokens of the terminated sixth output sequence (s6) (1026), seventh output sequence (s7) (1027), and eighth output sequence (s8) (1028) among the output sequences of the third set (1020) of FIG. 10. For example, the electronic device (301) can identify 40 (= 12+14+14), which is the number of tokens of the sixth output sequence (s6) (1026), seventh output sequence (s7) (1027), and eighth output sequence (s8) (1028), as a redistributable (or reallocable) resource. For example, the electronic device (301) can identify the number of remaining sentences. For example, the electronic device (301) can identify 2 remaining sentences, as the generation of 6 sentences out of 8 sentences has ended. For example, the electronic device (301) can identify output sequences of a third set (1020) for the remaining sentences. For example, the output sequences of the third set (1020) for the remaining sentences may include a fourth output sequence (s4) (1024) and a fifth output sequence (s5) (1025).

[0258] For example, the electronic device (301) may select a setting value to be used in at least some of the input sequences based on a token generation rate. For example, the electronic device (301) may identify a fourth set (1110) comprising two input sequences corresponding to the fourth output sequence (s4) (1024) and the fifth output sequence (s5) (1025) of the third set (1020). For example, the fourth set (1110) may include the fourth input sequence (s4) (1114) and the fifth input sequence (s5) (1115). In FIG. 11, the first input sequence (s1) (1111), the second input sequence (s2) (1112), the third input sequence (s3) (1113), the sixth input sequence (s6) (1116), the seventh input sequence (s7) (1117), and the eighth input sequence (s8) (1118) are shown, but the first input sequence (s1) (1111), the second input sequence (s2) (1112), the third input sequence (s3) (1113), the sixth input sequence (s6) (1116), the seventh input sequence (s7) (1117), and the eighth input sequence (s8) (1118) may not be included in the fourth set (1110). This may be because the first output sequence (s1) (821), second output sequence (s2) (922), third output sequence (s3) (923), sixth output sequence (s6) (1026), seventh output sequence (s7) (1027), and eighth output sequence (s8) (1028) corresponding to the first input sequence (s1) (1111), second input sequence (s2) (1112), third input sequence (s3) (1113), sixth input sequence (s6) (1116), seventh input sequence (s7) (1117), and eighth input sequence (s8) (1118) are sequences where sentence generation has ended.

[0259] For example, the electronic device (301) may further increase the number of tokens for each of the fourth input sequence (s4) (1114) and fifth input sequence (s5) (1115) of the fourth set (1110) when allocating resources (e.g., 40 tokens) for the sixth output sequence (s6) (1026), the seventh output sequence (s7) (1027), and the eighth output sequence (s8) (1028) evenly to the sequences for the remaining sentences. Referring to Table 3 described above, the electronic device (301) may identify the setting value of each of the fourth input sequence (s4) (1114) and the fifth input sequence (s5) (1115) as the fifth setting value (e.g., {2, 3}). The above fifth setting value (e.g., {2, 3})) may be a setting value (or candidate values) that has the maximum token generation rate among setting values ​​(or candidate values) that require a number lower than the maximum number of tokens that can be included in each sequence (e.g., 32) when the maximum number of tokens that can be input to the trained model (345) (e.g., 64) is evenly distributed among two sequences.

[0260] For example, the electronic device (301) may use the fifth setting value (e.g., {2, 3}) for each of the input sequences of the fourth set (1110). Referring to the example in FIG. 11, the electronic device (301) may additionally assign 18 tokens to each of the fourth input sequence (s4) (1114) and the fifth input sequence (s5) (1115) of the fourth set (1110). Accordingly, the electronic device (301) may assign 30 tokens to each of the fourth input sequence (s4) (1114) and the fifth input sequence (s5) (1115). For examples of the number of tokens, the number of draft tokens, the setting value for sampling, and the token generation rate for the input sequences of the fourth set (1110), the following table may be referenced.

[0261] Input sequences Settings Number of value tokens Token generation rate s1---s2---s3---s4{2, 3}301.28s5{2, 3}301.28s6---s7---s8---

[0262] Referring to Table 7 above, in the fourth drive, the input sequences of the fourth set (1110) (e.g., fourth output sequence (s4) (1114), fifth output sequence (s5) (1115)) may include 30 tokens according to the fifth setting value (e.g., {2, 3}). This allows the electronic device (301) to assign (or input) tokens to the trained model (345) so that, in the fourth drive, the generation speed of sentences (e.g., 8 sentences) to be generated according to the multi-sentence generation function is improved, and the possibility of sentences (e.g., 8 sentences) being generated simultaneously (or substantially simultaneously) (or the deviation in generation time between sentences is minimized) is improved.

[0263] For example, the electronic device (301) can generate output sequences using a trained model (345). For example, the electronic device (301) can generate a fourth set (1120) of output sequences inferred from a fourth set (1110). For example, the fourth set (1120) may include a fourth output sequence (s4) (1124) and a fifth output sequence (s5) (1125). In FIG. 11, the first output sequence (s1) (1121), the second output sequence (s2) (1122), the third output sequence (s3) (1123), the sixth output sequence (s6) (1126), the seventh output sequence (s7) (1127), and the eighth output sequence (s8) (1128) are shown, but the first output sequence (s1) (1121), the second output sequence (s2) (1122), the third output sequence (s3) (1123), the sixth output sequence (s6) (1126), the seventh output sequence (s7) (1127), and the eighth output sequence (s8) (1128) may not be included in the fourth set (1120). This may be because the sequences corresponding to the first output sequence (s1) (1121), the second output sequence (s2) (1122), the third output sequence (s3) (1123), the sixth output sequence (s6) (1126), the seventh output sequence (s7) (1127), and the eighth output sequence (s8) (1128) (e.g., the first output sequence (s1) (821), the second output sequence (s2) (922), the third output sequence (s3) (923), the sixth output sequence (s6) (1026), the seventh output sequence (s7) (1027), the eighth output sequence (s8) (1028)) are sequences where sentence generation has ended.

[0264] For example, the electronic device (301) may perform verification on a fourth set (1110) of input sequences and a fourth set (1120) of output sequences of a trained model (345). For example, the electronic device (301) may perform verification on a fourth input sequence (s4) (1114) and a fourth output sequence (s4) (1124). As an example without limitation, the electronic device (301) may perform the verification using a CPU (311). Or, for example, the electronic device (301) may perform the verification using a CPU (311) and an NPU (312) (or another trained model for verification). Specific details regarding the verification may be referenced in FIG. 4c and FIG. 6b described above.

[0265] For example, the electronic device (301) can identify whether each of the output sequences of the fourth set (1120) is a sequence in which sentence generation has ended. For example, the electronic device (301) can identify whether a criterion for the termination of sentence generation has been satisfied for each of the output sequences of the fourth set (1120). For example, the electronic device (301) can identify whether the last approved candidate token among the tokens of each of the output sequences of the fourth set (1120) corresponds to the criterion token (1199). Additionally, for example, the electronic device (301) can determine whether all of the output sequences of the fourth set (1120) are sequences in which sentence generation has ended. In the example of FIG. 11, for convenience of explanation, it is assumed that among the output sequences of the fourth set (1120), the fourth output sequence (s4) (1126) and the fifth output sequence (s5) (1127) are sequences in which sentence generation has ended. For example, the electronic device (301) can perform the operation (760) of FIG. 7 based on identifying that all of the output sequences of the fourth set (1120) are sequences in which sentence generation has ended.

[0266] For example, the electronic device (301) can output sentences generated in response to an input prompt in a function for generating multiple sentences. For example, the sentences to be output may include a sentence generated using a first output sequence (s1) (821), a sentence generated using a second output sequence (s2) (922), a sentence generated using a third output sequence (s3) (923), a sentence generated using a fourth output sequence (s4) (1124), a sentence generated using a fifth output sequence (s5) (1125), a sentence generated using a sixth output sequence (s6) (1026), a sentence generated using a seventh output sequence (s7) (1027), and a sentence generated using an eighth output sequence (s8) (1028).

[0267] Referring to the above description, the operations exemplified in FIGS. 8 to 11 are described as being performed sequentially, but the present disclosure is not limited thereto. In other words, the present disclosure is not limited to the order of operations exemplified in FIGS. 8 to 11, but can assign tokens that can be assigned (or input) to a trained model (345) so that the generation speed of sentences (e.g., 8 sentences) to be generated according to the multi-sentence generation function is improved, and the possibility of sentences (e.g., 8 sentences) being generated simultaneously (or the deviation in generation time between sentences is minimized) is improved.

[0268] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0269] As described above, the electronic device (301) may include a memory (320) that stores instructions and includes one or more storage media. The electronic device (301) may include at least one processor (310) that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to generate a first set of input sequences. Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model (345) stored in the electronic device (301). When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the trained model (345) to obtain output sequences inferred from the first set of input sequences by simultaneously providing the first set of input sequences to the trained model (345). When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to identify an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences. When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to generate a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences based on identifying the output sequences. Each of the input sequences in the second set of input sequences may include a second number of tokens according to a second setting value that is greater than the first number and different from the first setting value.

[0270] According to one embodiment, each of the first set of input sequences may be used to generate sentences based on the execution of a function for multi-sentence generation. The identified output sequence among the output sequences may be a sequence used to generate a sentence in which sentence generation is completed based on the verification.

[0271] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to execute a software application for providing a sentence generation service using the trained model (345). When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to execute the function for the multiple sentence generation of the software application. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the electronic device (301) to generate a prompt indicating a received voice input based on the executed function. Each of the first set of input sequences may be used to generate the sentences indicating a response to the prompt.

[0272] According to one embodiment, each of the other input sequences that are different from the other input sequences among the second set of input sequences may include the first number of tokens according to the first setting value.

[0273] According to one embodiment, each of the other input sequences that are different from the other input sequences among the second set of input sequences may include the second number of tokens according to the second setting value.

[0274] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of input sequences to determine the second set of input sequences to be used in each of the second set of input sequences, based on identifying the output sequence, so that the tokens of the output sequence are uniformly distributed among the second set of input sequences.

[0275] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be caused to determine the second setting value based on a look-up table (LUT) defined with respect to the trained model (345) and indicating a token generation rate according to the number of tokens included in the input sequence and the number of draft tokens included in the input sequence.

[0276] According to one embodiment, the token generation rate may be defined as the ratio of the number of tokens approved in verification among the tokens inferred as a result of the operation of the trained model (345) to the number of operations of the trained model (345).

[0277] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be caused to identify the number of sentences set for the function for generating multiple sentences. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be caused to generate the first set of input sequences, each containing the first number of tokens according to the first set value, which are determined based on the number of sentences.

[0278] According to one embodiment, the number of sequences in the first set may correspond to the number of sentences. The number of sequences in the second set may be less than the number of sentences.

[0279] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be caused to identify each of the second setting value and the third setting value, which are available as setting values ​​to be used in each of the parts of the second set of input sequences. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be caused to identify each of the first token generation rate of the trained model (345) according to the first setting value, the second token generation rate of the trained model (345) according to the second setting value, and the third token generation rate of the trained model (345) according to the third setting value. When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the setting value to be determined as the second setting value based on the third token generation rate lower than the first token generation rate and the second token generation rate higher than the first token generation rate.

[0280] According to one embodiment, the number of input tokens of the trained model (345) generated according to the second set of input sequences and the second setting value may be less than the maximum number of tokens that can be input to the trained model (345). The number of input tokens of the trained model (345) generated according to the second set of input sequences and the third setting value may be less than the maximum number of tokens that can be input to the trained model (345).

[0281] According to one embodiment, the output sequences inferred from the first set of input sequences may be the first set of output sequences. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of input sequences to be provided to the trained model (345) simultaneously, thereby causing the second set of output sequences inferred from the second set of input sequences to be obtained from the trained model (345). When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of output sequences to identify one or more output sequences for which sentence generation has ended, based on verification of the second set of input sequences and the second set of output sequences. When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause to generate a third set of input sequences having a number of sequences less than the number of sequences of the second set of input sequences based on identifying the one or more output sequences. Each of the input sequences of the third set of input sequences may include a third number of tokens according to a third setting value that is greater than the second number and different from the second setting value. Each of the other input sequences of the third set of input sequences that are different from the input sequences may include a fourth number of tokens according to a fourth setting value that is greater than the second number and different from the second setting value and the third setting value.

[0282] According to one embodiment, each of the other input sequences that differ from the some input sequences among the second set of input sequences may include the first number of tokens according to the first set value. The some input sequences among the third set of input sequences may correspond to some of the some input sequences among the second set of input sequences. The other some input sequences among the third set of input sequences may correspond to some of the other some input sequences among the second set of input sequences. The difference between the third number and the second number may be the same as the difference between the fourth number and the first number.

[0283] According to one embodiment, the output sequences inferred from the first set of input sequences may be the output sequences of the first set. Each of the other input sequences among the second set of input sequences that differ from some of the input sequences may include the first number of tokens according to the first set value. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of input sequences to be provided to the trained model (345) simultaneously, thereby causing the second set of output sequences inferred from the second set of input sequences to be obtained from the trained model (345). When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of output sequences to be identified based on verification of the second set of input sequences and the second set of output sequences. When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause to generate a third set of input sequences having a number of sequences less than the number of sequences of the second set of input sequences based on identifying the one or more output sequences. The third set of input sequences may correspond to some of the input sequences of the second set of input sequences and some of the other input sequences of the second set of input sequences. Each of the third set of input sequences may include a third number of tokens according to a third setting value that is greater than the second number and different from the second setting value.

[0284] According to one embodiment, the output sequences inferred from the first set of input sequences may be the first set of output sequences. When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the second set of input sequences to be provided to the trained model (345) simultaneously, thereby causing the second set of output sequences inferred from the second set of input sequences to be obtained from the trained model (345). When the instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may cause the completion of sentence generation for each of the second set of output sequences to be identified based on verification of the second set of input sequences and the second set of output sequences. When the above instructions are executed individually or collectively by the at least one processor (310), the electronic device (301) may be generated based on the first set of input sequences, the first set of output sequences, and the second set of output sequences, and may be caused to output sentences representing a response to one (a) prompt.

[0285] According to one embodiment, each of the input sequences of the first set may include one valid token, one or more draft tokens and one or more candidate tokens generated according to the first setting value, and one or more speculative tokens according to the one or more candidate tokens. The sum of the number of valid tokens, the number of the one or more candidate tokens, and the number of the one or more speculative tokens may be the first number.

[0286] According to one embodiment, the at least one processor (310) may include a neural processing unit (NPU) including a processing circuit and a central processing unit (CPU) including a processing circuit. The NPU may be configured to drive the trained model (345). The CPU may be configured to generate input sequences, perform the verification, and determine a setting value to be used in the input sequences.

[0287] A method performed by an electronic device (301) as described above may include an operation of generating a first set of input sequences. Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model (345) stored in the electronic device (301). The method may include an operation of obtaining output sequences inferred from the first set of input sequences from the trained model (345) by simultaneously providing the first set of input sequences to the trained model (345). The method may include an operation of identifying an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences. The method may include an operation of generating a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences based on identifying the output sequences. Each of the input sequences in the second set of input sequences may include a second number of tokens according to a second setting value that is greater than the first number and different from the first setting value.

[0288] As described above, a non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device (301) to generate a first set of input sequences when executed individually or collectively by at least one processor (310) of the electronic device (301). Each of the first set of input sequences may include a first number of tokens according to a first set value defined for a trained model (345) stored in the electronic device (301). The non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device (301) to obtain output sequences inferred from the first set of input sequences from the trained model (345) by simultaneously providing the first set of input sequences to the trained model (345) when executed individually or collectively by the at least one processor (310). The above non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device (301) to identify an output sequence among the output sequences based on verification of the first set of input sequences and the output sequences when executed individually or collectively by the at least one processor (310). The above non-transient computer-readable storage medium may store one or more programs including instructions that cause the electronic device (301) to generate a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences when executed individually or collectively by the at least one processor (310) based on identifying the output sequence.Each of the input sequences in the second set of input sequences may include a second number of tokens according to a second setting value that is greater than the first number and different from the first setting value.

[0289] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0290] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0291] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” may each include any one of the items listed together in the corresponding phrase, or any combination thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another corresponding component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0292] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0293] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0294] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0295] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, Memory that stores instructions and includes one or more storage media; and It includes at least one processor comprising a processing circuit, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Generating a first set of input sequences, wherein each of the first set of input sequences includes a first number of tokens according to a first setting value defined for a trained model stored in the electronic device; By simultaneously providing the first set of input sequences to the trained model, output sequences inferred from the first set of input sequences are obtained from the trained model; Based on the verification of the first set of input sequences and the output sequences, an output sequence is identified among the output sequences; and Based on identifying the above output sequence, causing to generate a second set of input sequences having a number of sequences less than the number of sequences of the first set of input sequences, and Each of the input sequences of the second set of input sequences comprises a second number of tokens that is greater than the first number and corresponds to a second setting value different from the first setting value. Electronic device.

2. In Claim 1, Each of the above first set of input sequences is used to generate sentences based on the execution of a function for multi-sentence generation, and The identified output sequence among the above output sequences is a sequence used to generate a sentence in which sentence generation is completed based on the verification, Electronic device.

3. In Claim 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Execute a software application for providing a sentence generation service using the above-mentioned trained model; Executing the function for generating the multiple sentences of the above software application; and Based on the above-mentioned executed function, cause to generate a prompt indicating the received voice input, and Each of the above first set of input sequences is used to generate the sentences representing a response to the prompt, Electronic device.

4. In Claim 1, Each of the other input sequences among the second set of input sequences that are different from the other input sequences includes the first number of tokens according to the first setting value. Electronic device.

5. In Claim 1, Each of the other input sequences among the second set of input sequences that are different from the other input sequences includes the second number of tokens according to the second setting value. Electronic device.

6. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on identifying the output sequence, causing the setting value to be used in each of the subsets of the second set of input sequences to be determined as the second setting value, so that the tokens of the output sequence are uniformly distributed among the subsets of the second set of input sequences. Electronic device.

7. In Claim 6, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Causing the above setting value to be determined as the second setting value based on a LUT (look-up table) defined for the above-trained model and indicating a token generation rate according to the number of tokens included in the input sequence and the number of draft tokens included in the input sequence, Electronic device.

8. In Claim 7, The above token generation rate is defined as the ratio of the number of tokens approved in verification among the tokens inferred as a result of running the trained model to the number of times the trained model is run, Electronic device.

9. In Claim 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identifying the number of sentences set with respect to the function for generating the above multiple sentences; and Causing to generate the first set of input sequences, each comprising the first number of tokens according to the first setting value, which are determined based on the number of the above sentences. Electronic device.

10. In Claim 9, The number of sequences in the first set above corresponds to the number of sentences, and The number of sequences in the second set above is less than the number of sentences above. Electronic device.

11. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identifying each of the second setting value and the third setting value that are available as setting values ​​to be used in each of the partial input sequences among the second set of input sequences; Identifying each of the first token generation rate of the trained model according to the first setting value, the second token generation rate of the trained model according to the second setting value, and the third token generation rate of the trained model according to the third setting value; and Causing the setting value to be determined as the second setting value based on the third token generation rate lower than the first token generation rate and the second token generation rate higher than the first token generation rate. Electronic device.

12. In Claim 11, The number of input tokens of the trained model generated according to the input sequences of the second set and the second setting value is less than the maximum number of tokens that can be input to the trained model, and The number of input tokens of the trained model generated according to the input sequences of the second set and the third setting value is less than the maximum number of tokens that can be input to the trained model. Electronic device.

13. In Claim 1, The output sequences inferred from the input sequences of the first set are the output sequences of the first set, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By simultaneously providing the second set of input sequences to the trained model, a second set of output sequences inferred from the second set of input sequences are obtained from the trained model; Based on verification of the input sequences of the second set and the output sequences of the second set, one or more output sequences among the output sequences of the second set for which sentence generation has ended are identified; and Based on identifying one or more of the above output sequences, causing to generate a third set of input sequences having a number of sequences less than the number of sequences of the second set of input sequences, and Each of the input sequences of the third set of input sequences comprises a third number of tokens that is greater than the second number and corresponds to a third setting value different from the second setting value, and Each of the other input sequences among the input sequences of the third set that are different from the other input sequences is greater than the second number and includes a fourth number of tokens according to a fourth setting value different from the second setting value and the third setting value. Electronic device.

14. In Claim 13, Each of the other input sequences among the second set of input sequences that are different from some of the input sequences includes the first number of tokens according to the first setting value, and Some of the input sequences among the input sequences of the third set correspond to some of the input sequences among the input sequences of the second set, and Among the input sequences of the third set, the other some input sequences correspond to some of the other some input sequences of the second set, and The difference between the third number and the second number is the same as the difference between the fourth number and the first number. Electronic device.

15. In Claim 1, The output sequences inferred from the first set of input sequences are the first set of output sequences, and Each of the other input sequences among the second set of input sequences that are different from some of the input sequences includes the first number of tokens according to the first setting value, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By simultaneously providing the second set of input sequences to the trained model, a second set of output sequences inferred from the second set of input sequences are obtained from the trained model; Based on verification of the input sequences of the second set and the output sequences of the second set, one or more output sequences among the output sequences of the second set are identified; and Based on identifying one or more of the above output sequences, causing to generate a third set of input sequences having a number of sequences less than the number of sequences of the second set of input sequences, and The input sequences of the third set above correspond to some of the input sequences of the second set above and some of the other input sequences of the second set above, and Each of the input sequences of the third set above comprises a third number of tokens that is greater than the second number and corresponds to a third setting value different from the second setting value. Electronic device.