Electronic device and method for enhancing diversity of response to user utterance
By combining general and sub-prompts based on domain, user, and device information, the electronic device enhances response diversity and personalization, addressing the limitations of uniform responses in existing systems.
Patent Information
- Application Number
- PCT/KR2024/018424
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-17
AI Technical Summary
Existing electronic devices struggle to provide diverse and personalized responses to user utterances due to the limitations of large-scale language models, often resulting in uniform and limited responses that do not account for individual user characteristics or device-specific features.
The electronic device employs a multi-prompt system that combines a general prompt with sub-prompts tailored to the domain, user information, and device characteristics to generate a final prompt for input into a large language model, enhancing response diversity and personalization.
This approach allows for a wide range of responses that reflect individual user preferences and device capabilities, improving user experience while reducing query costs by minimizing the number of tokens processed by the language model.
Smart Images

Figure KR2024018424_17072025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for enhancing the diversity of responses to user utterances
[0001] The present disclosure relates to an electronic device and method for enhancing the diversity of responses to user utterances.
[0002] Electronic devices may support the ability to provide responses to user utterances. For example, when a user utters a user utterance to trigger an action on the electronic device, the electronic device may receive the user utterance. The electronic device may extract command text from the user utterance and input a prompt based on the extracted command text into a large-scale language model to provide a response.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art in connection with the present disclosure.
[0004] An electronic device is provided. The electronic device may include at least one processor including a processing circuit. The electronic device may include a memory storing instructions. The instructions, when executed by the at least one processor, may cause the electronic device to extract a command text from a user utterance, determine a domain corresponding to the command text based on the command text, determine a first sub-prompt related to the domain, determine at least one second sub-prompt based on information stored in the memory or an external electronic device, generate a final prompt to be input into a large language model (LLM) based at least in part on the command text, the first sub-prompt, and the at least one second sub-prompt, and input the final prompt into the large language model, thereby obtaining response information for the user utterance.
[0005] A method performed by an electronic device is provided. The method may include extracting a command text from a user utterance. The method may include determining a domain corresponding to the command text based on the command text. The method may include determining a first sub-prompt related to the domain. The method may include determining at least one second sub-prompt based on information stored in the memory or an external electronic device. The method may include generating a final prompt to be input into a large language model (LLM) based at least in part on the command text, the first sub-prompt, and the at least one second sub-prompt. The method may include obtaining response information for the user utterance by inputting the final prompt into the large language model.
[0006] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0007] Figure 2 illustrates a state in which an exemplary electronic device provides a response according to a user's speech.
[0008] Figure 3a is a block diagram illustrating components of an exemplary electronic device.
[0009] Figure 3b is a block diagram showing a prompt database.
[0010] Figure 4 is a flow chart showing the operation of an exemplary electronic device that obtains response information according to a user's speech.
[0011] Figure 5 illustrates the operation of determining a domain using a large-scale language model.
[0012] Figure 6 illustrates an exemplary process for generating a final prompt by combining a general prompt and a first sub-prompt.
[0013] Figure 7 illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a third sub-prompt.
[0014] Figure 8 illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a fourth sub-prompt.
[0015] Figure 9a illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a fourth sub-prompt.
[0016] Figure 9b illustrates an exemplary process of generating another final prompt by further combining a third sub-prompt with the final prompt generated through the process of Figure 9a.
[0017] Figure 9c illustrates the response provided via the final prompt of Figure 9a.
[0018] Figure 9d illustrates the response provided via the final prompt of Figure 9b.
[0019] Figure 10a shows a prompt of an electronic device according to a comparative example.
[0020] FIG. 10b illustrates a prompt of an electronic device according to one embodiment.
[0021] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0022] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0023] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0024] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0025] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0026] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0027] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0028] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0029] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0030] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0031] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0032] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) to an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0033] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0034] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0035] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0036] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0037] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0038] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0039] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0040] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0041] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0042] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0043] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0044] Figure 2 illustrates a state in which an exemplary electronic device provides a response according to a user's speech.
[0045] Referring to FIG. 2, an exemplary electronic device (101) may be configured to provide an action based on a user utterance (210) using an artificial intelligence model. The artificial intelligence model is a technology for understanding human language and processing actions according to human language using a machine learning algorithm. When a user utterance (210) is input into the electronic device (101), the electronic device (101) may extract the user utterance (210) as text, determine the intent of the user utterance (210) based on the extracted text, and provide a response (220) corresponding to the determined intent. For example, when a user provides a user utterance (210) commanding the transmission of a text message in order to send a text message with specific content to a specific person, the electronic device (101) may be configured to generate a text message corresponding to the specific content to the specific person based on the user utterance (210) and transmit the generated text message.
[0046] The exemplary electronic device (101) can be implemented in various ways. For example, the electronic device (101) can include a smart phone (201). The smart phone (201) can have a bar type structure, but is not limited thereto. For example, the smart phone (201) can also have a foldable type and / or rollable type structure. For example, the electronic device (101) can include an electronic device such as a tablet (202), a laptop, a computer, a wearable device (e.g., a smart watch (203), wireless earphones), and / or a home appliance (e.g., a speaker device (204), a TV (205), a robot), and in addition, can be implemented as various electronic devices.
[0047] An exemplary electronic device (101) can acquire commands that cause specified actions through a large language model (LLM). The large language model, an artificial intelligence model trained with text data, can be used to provide a response (220) to a user utterance (210).
[0048] LLM can be referred to as a language model composed of an artificial neural network pre-trained on a large amount of text data. Compared to existing general language models, LLM can contain more than 10 times more parameters (for example, more than 100 billion parameters). LLM can use a transformer artificial neural network structure based on the attention mechanism. The attention mechanism is a technology that helps an artificial intelligence model focus (attention) on important parts of input data. The attention mechanism can be used to predict output data by predicting the degree to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of some layers of a neural network) contributes to the intermediate or final output of the neural network. The recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration (attention) within the overall (or partial) context of the input data.
[0049] For example, an LLM or transformer may include an encoder-decoder structure. The encoder may process input data to output compressed information (e.g., a contextual representation), and the decoder may process the compressed information to output token-based data. Each encoder and decoder may include an independent attention network, and a cross-attention network connecting the encoder and decoder may be included.
[0050] For example, LLM can be trained in two stages: pre-training and fine-tuning. Pre-training involves training the LLM to process large amounts of text data and acquire general linguistic knowledge. For example, this could involve self-supervised learning, such as predicting the next word using a sequence of previous words in a text sequence. Fine-tuning involves training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. Additional supervised learning (or adaptive learning) can be performed on a pre-trained model using a dataset tailored to the domain's purpose. The LLM can perform tasks using text inputs containing natural language, called prompts. For example, the LLM can include Bidirectional Encoder Representations from Transformer (BERT) and generative pre-trained transformers (GPT). For example, fine-tuning can be omitted during LLM training. The prompts input to the LLM can be controlled to improve the performance of the task desired by the user. In a similar way to in-context learning or zero-shot / few-shot learning, you can provide additional examples of tasks and / or guidance on performing the task in the prompt.
[0051] The term "LLM" can refer to the neural network model itself, but it can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot such as ChatGPT can also be referred to as an LLM. "LLM" can also include an inference engine that utilizes an LLM neural network model. For example, "inputting an input prompt into an LLM" can refer to "inputting an input prompt into an LLM-based inference engine." For example, "the output of the LLM for an input prompt" can refer to the output information (or output information modified through further processing) of the last neural network layer of the LLM obtained when the input prompt is input into an LLM-based inference engine.
[0052] According to one embodiment, the electronic device (101) can obtain a response (220) provided to the user by generating a prompt based on a user utterance (210) and inputting the generated prompt into a large-scale language model. If the electronic device (101) provides a limited number of responses (220) within a predefined set of responses, only responses (220) that do not reflect the characteristics of the electronic devices (101) that can be implemented in various ways and the user's personal characteristics are provided. If a large number of prompts are input to provide various responses (220), the number of tokens to be processed through the large-scale language model increases, which causes an increase in query cost.
[0053] An exemplary electronic device (101) may be configured to utilize at least one sub-prompt based on unique information to provide a response (220) to a user utterance (210). For example, the electronic device (101) may be configured to generate a final prompt to be input into a large-scale language model by combining a sub-prompt corresponding to a domain according to the user utterance (210), a sub-prompt according to the characteristics of the electronic device (101), and / or a sub-prompt according to the characteristics of the user with a common prompt. By generating a prompt by combining sub-prompts based on various pieces of information, the electronic device (101) may provide a variety of responses that match individual characteristics, rather than a uniform and limited response. Hereinafter, an electronic device (101) for providing a variety of responses to a user utterance (210) will be described.
[0054] Figure 3a is a block diagram illustrating components of an exemplary electronic device. Figure 3b is a block diagram illustrating a prompt database.
[0055] Referring to FIG. 3A, an exemplary electronic device (101) may include at least one processor (310) and memory (320).
[0056] According to an exemplary embodiment, the memory (320) may store data such as instructions for the operation of the electronic device (101), basic programs, applications, and setting information. The memory (320) may be configured as volatile memory, non-volatile memory, or a combination of volatile memory and non-volatile memory. The memory (320) may store a prompt database (321), user account information (322), and / or electronic device information (323).
[0057] According to an exemplary embodiment, at least one processor (310) may control the overall operation of the electronic device (101). For example, at least one processor (310) may cause the electronic device (101) to provide a response according to a user's speech by executing instructions stored in the memory (320). At least one processor (120) may include a processing circuit. At least one processor (310) may include, but is not limited to, an application processor (AP) (e.g., a central processing unit (CPU)) and / or a communication processor (CP) (e.g., a modem). At least one processor (310) may include, but is not limited to, a graphics processing unit (e.g., a GPU), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a wireless-fidelity (Wi-Fi) chip, and a Bluetooth chip. ® It may include a chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, connectivity chips, a sensor controller, a touch controller, a finger-print sensor controller, a display drive integrated circuit (DDI), an audio CODEC chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on chip (SoC), an integrated circuit (IC), or similar circuits.
[0058] According to an exemplary embodiment, at least one processor (310) may extract command text based on a user utterance. At least one processor (310) may determine a domain based on the extracted command text. At least one processor (310) may determine at least one sub-prompt. At least one processor (310) may generate a final prompt to be input into a large-scale language model (330) by combining the determined at least one sub-prompt with a general prompt.
[0059] According to an exemplary embodiment, at least one processor (310) may include a voice assistant (311), a prompt retriever (312), and / or a prompt generator (313). The voice assistant (311) may be used to extract command text based on a user utterance and to guide a response of the electronic device (101). The voice assistant (311) may execute an application (e.g., a speech recognition application) to extract the command text, determine a domain based on the command text, and / or execute an application (e.g., a text application) to execute an action corresponding to the response. The prompt retriever (312) may be used to determine a domain based on the extracted command text. The prompt generator (313) may be used to determine at least one sub-prompt and to combine the determined at least one sub-prompt with a general prompt to generate a final prompt to be input into a large-scale language model (330). At least one of the voice assistant (311), the prompt retriever (312), and the prompt generator (313) may be a logically distinct logic circuit within at least one processor (310), or may be a separate circuit physically separated from at least one processor (310).
[0060] An exemplary electronic device (101) may include an input device (340) and / or an output device (350). The input device (340) may be used to receive a user speech. For example, the input device (340) may include, but is not limited to, a microphone. The input device (340) may be configured to receive the user speech and provide data corresponding to the received user speech to at least one processor (310) (e.g., a voice assistant (311)). The output device (350) may be used to provide a response to the user speech. For example, the output device (350) may include a display (351) and / or a speaker (352). The display (351) may be configured to display the response as visual information. The speaker (352) may be configured to display the response as auditory information.
[0061] Referring to FIG. 3B, the prompt database (321) may store prompts used to generate prompts input into a large-scale language model (e.g., the large-scale language model (330) of FIG. 3A). According to an exemplary embodiment, the prompt database (321) may store a general prompt (371), a first sub-prompt (372), and a second sub-prompt (373).
[0062] According to an exemplary embodiment, a common prompt (371) may be a prompt related to the definition of the basic concept of a response. For example, a common prompt (371) may include defined prompts such as "You are a voice assistant" or "You are a personal tutor who teaches English." In terms of its relationship to the basic concept of a response, a common prompt (371) may be referred to as a system prompt.
[0063] According to an exemplary embodiment, the first sub-prompt (372) and the second sub-prompt (373) may be used to determine a sub-prompt (e.g., a second prompt) to be combined with the general prompt (371) (e.g., a first prompt). For example, the first sub-prompt (372) may be related to a domain according to the user's utterance. For example, the second sub-prompt (373) may include a third sub-prompt (373a), a fourth sub-prompt (373b), and / or a fifth sub-prompt (373c). The third sub-prompt (373a) may be related to user information. The fourth sub-prompt (373b) may be related to information of the electronic device (101). The fifth sub-prompt (373c) may be related to other information (e.g., persona information, language information). Descriptions of each of the sub-prompts are described below.
[0064] According to an exemplary embodiment, the electronic device (101) may generate a prompt to be input into the large-scale language model (330) based at least in part on a command text extracted from a user utterance, a first sub-prompt (372), and at least one second sub-prompt (373). The electronic device (101) may determine a first prompt corresponding to a general prompt (371), and may determine a second prompt based on the command text, the first sub-prompt (372), and at least one second sub-prompt (373). The electronic device (101) may be configured to generate a final prompt to be input into the large-scale language model (330) by combining the first prompt and the second prompt. Since the prompt to be input into the large-scale language model (330) is generated at least in part based on domain information, user information, and / or electronic device information (323), the electronic device (101) may provide a variety of responses that match individual characteristics, rather than a uniform and limited response.
[0065] Referring back to FIG. 3A, data used to generate a prompt may be stored in the memory (320), but is not limited thereto. According to an exemplary embodiment, the electronic device (101) may also use data stored in an external electronic device (301) (e.g., a server) to provide a response. The external electronic device (301) may store a prompt database (302), user account information (303), and / or electronic device information (304). Based on identifying a user utterance, the electronic device (101) may obtain at least one sub-prompt from the prompt database (302) stored in the external electronic device (301) via the wireless communication circuit (360). At least some of the prompt database (321), the user account information (322), and the electronic device information (323) may be stored in the memory (320), and the remaining some may be stored in the external electronic device (301).
[0066] In the above description, it has been described that at least one processor (310) of the electronic device (101) obtains a response to a user utterance by inputting a prompt to the large-scale language model (330), but is not limited thereto. According to an exemplary embodiment, the electronic device (101) may provide the user utterance to the external electronic device (301) via the wireless communication circuit (360) to provide the response. The processor (305) of the external electronic device (301) may generate a response based on the user utterance provided from the electronic device (101). For example, the processor (305) may include a prompt retriever (306) and a prompt generator (307). The processor (305) may generate a prompt corresponding to the user utterance and input the prompt to the large-scale language model (308), thereby obtaining response information corresponding to the user utterance. The external electronic device (301) may transmit the response information to the electronic device (101). The electronic device (101) can provide the user with response information provided from an external electronic device (301). The descriptions of the electronic device (101) for generating the final prompt, described below, can be substantially equally applied to the external electronic device (301).
[0067] Figure 4 is a flowchart illustrating the operation of an exemplary electronic device that acquires response information based on a user's utterance. Figure 5 illustrates the operation of determining a domain using a large-scale language model. Figure 6 illustrates an exemplary process of generating a final prompt by combining a general prompt and a first sub-prompt.
[0068] The operations described in FIG. 4 may be operations performed by an electronic device (e.g., an electronic device (101) of FIG. 3A) when instructions stored in a memory (e.g., a memory (320) of FIG. 3A) are executed by a processor (e.g., at least one processor (310) of FIG. 3A).
[0069] Referring to FIG. 4, at operation 401, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to extract command text from user utterance.
[0070] According to an exemplary embodiment, at least one processor (310) may receive a voice command by a user utterance. The user may provide the user utterance to cause an operation of the electronic device (101). For example, if the user provides a user utterance such as “call mom,” the at least one processor (310) may be configured to obtain the user utterance through an input device (e.g., the input device (340) of FIG. 3A) and extract a command text from the user utterance. The at least one processor (310) may extract the command text by preprocessing the voice command and extracting features. The command text extraction operation may be performed by, but is not limited to, a voice assistant providing speech-to-text (STT) conversion (e.g., the voice assistant (311) of FIG. 3A).
[0071] At operation 403, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to determine a domain corresponding to the command text.
[0072] Referring to FIG. 5, a large-scale language model (330) may be used to identify a domain corresponding to the intent of a command text among one or more predefined domains (510). According to an exemplary embodiment, at least one processor (310) may generate a prompt for determining a domain based on the extracted command text. The at least one processor (310) may input the prompt to the large-scale language model (330). The prompt may include a domain list including a description of each of the one or more domains (510) and a command requesting domain information by referring to the description of each of the one or more domains (510). The at least one processor (310) may determine the domain based on the domain information output by the large-scale language model (330). For example, to determine a domain using the large-scale language model (330), the at least one processor (310) may input the prompt described in [Table 1] below to the large-scale language model (330).
[0073] The following are the domains that the electronic device can support. Domain list: - Call: This plugin provides the functions for phone calls, get call logs... - Camera: This plugin provides the functions for camera mode... Refer to the domain list to find out which domain the user's speech belongs to.
[0074] For example, if the user utterance is "Call Mom", the large-scale language model (330) can identify the domain corresponding to the user utterance as the phone domain (511) by referring to the description of the action for the function of the phone domain (511). For example, if the user utterance is related to the function of the camera, the large-scale language model (330) can identify the domain corresponding to the user utterance as the camera domain (512) by referring to the description of the action for the function of the camera domain (512). However, the present invention is not limited thereto. For example, the domain can be determined using at least one of a deep neural network model, a machine learning model, or TF-IDF (Term Frequency - Inverse Document Frequency) configured to output the domain when a command text extracted from the user utterance is input.
[0075] Referring again to FIG. 4, at operation 405, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to determine a first sub-prompt corresponding to the domain.
[0076] According to an exemplary embodiment, at least one processor (310) may be configured to determine a first sub-prompt, which is used to generate a final prompt to be input into a large-scale language model (330). The first sub-prompt may be related to a domain determined through the large-scale language model (330). The first sub-prompt may be referred to as a domain prompt from the perspective of being determined in relation to a domain. The operation of determining the first sub-prompt may be performed by, but is not limited to, a prompt retriever (e.g., prompt retriever (312) of FIG. 3A ).
[0077] Referring to FIG. 6, at least one processor (310) may be configured to obtain a first sub-prompt (372) related to a determined domain from a prompt database (e.g., the prompt database (321, 302) of FIG. 3A). The first sub-prompt (372) may include one or more prompts (611, 612, 613, 614) for each of one or more domains.
[0078] According to an exemplary embodiment, the domain may include, but is not limited to, clock, web, call, and / or message. The first sub-prompt (372) may include a time-related prompt (611), a web-related prompt (612), a call-related prompt (613), and / or a message-related prompt (614). For example, the time-related prompt (611) may include a prompt defined as, "When telling the time, tell me only the hour and minute, excluding the seconds." For example, the web-related prompt (612) may include a prompt defined as, "When telling me the search results, tell me the source of the search results as well." For example, the phone-related prompt (613) may include a prompt defined as, "When telling me the phone number, use the recipient's name to generate a response." For example, a prompt (614) related to a message may include a defined prompt such as, "When you originally sent the text to the recipient, write a response in the past tense stating that you sent the text." A first sub-prompt (372) predefined for each of one or more domains may be combined with a general prompt (601) to generate a final prompt (605).
[0079] According to an exemplary embodiment, at least one processor (310) may determine a first sub-prompt (372) related to the determined domain. For example, if the determined domain is a telephone, at least one processor (310) may determine a telephone-related prompt (613) as a sub-prompt used in generating the final prompt (605). Based on determining the telephone-related prompt (613) as a sub-prompt, at least one processor (310) may obtain the telephone-related prompt (613) from the prompt database (321). At least one processor (310) may generate the final prompt (605) to be input to the large-scale language model (330) by combining the first prompt corresponding to the general prompt (605) and the second prompt related to the telephone-related prompt (613). Since the final prompt (605) to be input into the large-scale language model (330) is generated based at least in part on the phone-related prompt (613), the final prompt may include a prompt to generate a response using the recipient's name, according to the predefined phone-related prompt (613).
[0080] According to an exemplary embodiment, the first sub-prompt (372) may include one or more functional functions (e.g., function, API, deep link) corresponding to the domain and / or a description of the functional function. For example, a prompt (611) related to the message domain may include text including a function name and a comment corresponding to the function name (e.g., “exec_show_recent_message() / Show recently arrived messages” or “exec_delete_message(sMessage) / Delete sMessage”). The comment may include text separated by a delimeter such as “ / ”. The comment may include text surrounded by designated characters such as “ / *” and “* / ”.
[0081] Referring again to FIG. 4 , at operation 407, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to determine at least one second sub-prompt (e.g., the second sub-prompt (373) of FIG. 3B ).
[0082] According to an exemplary embodiment, at least one processor (310) may be configured to determine at least one second sub-prompt used to generate a final prompt to be input to the large-scale language model (330). The at least one second sub-prompt may include at least one of a third sub-prompt (e.g., the third sub-prompt (373a) of FIG. 3b), a fourth sub-prompt (e.g., the fourth sub-prompt (373b) of FIG. 3b), or a fifth sub-prompt (e.g., the fifth sub-prompt (373c) of FIG. 3b).
[0083] According to an exemplary embodiment, the third sub-prompt may be related to user information. For example, the user information may include, but is not limited to, the user's age information, the user's gender information, the user's preference information, and / or the user's emotional information. The third sub-prompt may be referred to as a user prompt in terms of its relationship to user information. The third sub-prompt is described below with reference to FIG. 7. When the third sub-prompt is combined with the final prompt, the user's characteristics may be reflected in the response output through the final prompt.
[0084] According to an exemplary embodiment, the fourth sub-prompt may be related to electronic device information. For example, the electronic device information may include, but is not limited to, information on the type of the electronic device and / or information on an output device (350) included in the electronic device (101). The fourth sub-prompt may be referred to as a device prompt in terms of its relationship to the electronic device information. The fourth sub-prompt is described below with reference to FIG. 8. When the fourth sub-prompt is combined with the final prompt, the characteristics of the electronic device (101) may be reflected in the response output through the final prompt.
[0085] According to an exemplary embodiment, the fifth sub-prompt may relate to other information. For example, the other information may include various information in addition to the aforementioned domain, electronic device information, and user information. For example, it may include, but is not limited to, persona information and / or language information.
[0086] For example, if a fifth sub-prompt related to persona information is combined to generate a final prompt, the persona information may be reflected in the response output by the final prompt. If the persona information is set to kindness and confidence as the basic persona, the response generated through the large-scale language model (330) may be generated based on kindness and confidence according to the persona information. The electronic device (101) may provide a system that can change the persona settings according to the user's settings.
[0087] For example, if a fifth sub-prompt related to language information is combined to generate a final prompt, the language information may be reflected in the response output by the final prompt. At least one processor (310) may identify the user's language through the user's speech and select a prompt based on the identified language. For example, if the user's speech is in Korean, the final prompt may be generated by considering Korean grammar, characteristics of spoken and written language, cultural characteristics of Korean, punctuation, etc. Other information may include various types of information in addition to persona information and language information.
[0088] At operation 409, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to generate a final prompt based at least in part on the command text, the first sub-prompt, and the at least one second sub-prompt.
[0089] According to an exemplary embodiment, at least one processor (310) may be configured to generate a final prompt to be input into a large-scale language model (330) by combining a first prompt and a second prompt. The first prompt may correspond to a general prompt. The second prompt may be determined based on the first sub-prompt and at least one second sub-prompt. If a user utterance is required for generating a response, the at least one processor (310) may also add the user utterance to the final prompt. The operation of generating the final prompt may be performed by, but is not limited to, a prompt generator (e.g., prompt generator (313) of FIG. 3A ).
[0090] According to an exemplary embodiment, the second prompt to be combined with the first prompt may include a first sub-prompt and at least one second sub-prompt. At least one processor (310) may select at least one of a third sub-prompt, a fourth sub-prompt, or a fifth sub-prompt. For example, if the third sub-prompt is selected, the first sub-prompt and the third sub-prompt may be determined as the second sub-prompt to be combined with the first prompt. For example, if the third sub-prompt and the fourth sub-prompt are selected, the first sub-prompt, the third sub-prompt, and the fourth sub-prompt may be determined as the second sub-prompt to be combined with the first prompt. The at least one second sub-prompt may be selected as predetermined by the user, but is not limited thereto.
[0091] At operation 411, the instructions, when executed individually or collectively by at least one processor (310), may cause the electronic device (101) to obtain response information for the user utterance by inputting the final prompt into a large-scale language model (330).
[0092] According to an exemplary embodiment, at least one processor (310) may be configured to obtain response information for a user utterance by inputting a final prompt into a large-scale language model (330), and provide a response based on the obtained response information. The obtained response information may be displayed through an output device (e.g., an output device (350) of FIG. 3A). For example, at least one processor (310) may control a display (e.g., a display (351) of FIG. 3A) to display a response based on the response information as visual information. For example, at least one processor (310) may control a speaker (e.g., a speaker (352) of FIG. 3A) to output a response based on the response information as auditory information. The operation of providing a response may be performed by, but is not limited to, a voice assistant (311).
[0093] According to an exemplary embodiment, since the final prompt input to the large-scale language model (330) is generated by a combination of a general prompt and one or more sub-prompts related to various information, the diversity of responses provided by the electronic device (101) can be enhanced. For example, if a response is generated using only general prompts, it is difficult to reflect domain information, user information, electronic device information, and other information, and thus a uniform and limited response may be provided. According to an exemplary embodiment, the electronic device (101) can provide an enhanced user experience by providing a response that reflects domain information, user information, electronic device information, and / or other information.
[0094] Figure 7 illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a third sub-prompt.
[0095] Referring to FIG. 7, at least one processor (310) may be configured to obtain a third sub-prompt (373a) related to user information from a prompt database (e.g., prompt database (321, 302) of FIG. 3a).
[0096] According to an exemplary embodiment, the user information may indicate personal characteristics of a user of an electronic device (e.g., the electronic device (101) of FIG. 3A). For example, the user information may include, but is not limited to, user age information, user gender information, user preference information, and / or user emotion information. The third sub-prompt (373a) may include prompts (711, 712) related to user gender information, prompts (713, 714) related to user age information, prompts (715, 716, 717) related to user preference information, and / or prompts related to user emotion. For example, user account information (e.g., user account information (322, 303) of FIG. 3A) may be stored in a memory (e.g., memory (320) of FIG. 3A) or an external electronic device (e.g., external electronic device (301) of FIG. 3A). At least one processor (310) can obtain user information from user account information (322, 303).
[0097] According to an exemplary embodiment, the user information may determine the tone of a response based on response information output through a large-scale language model (e.g., the large-scale language model (330) of FIG. 3A). For example, the user's age information may indicate whether the user is a child or an adult. If the user's age information indicates a child, the prompt (713) related to the user's age information may include a prompt defined as, "Use informal speech rather than formal speech, and use simple and friendly expressions that even children can understand," so as to generate a friendly response. If the user's age information indicates an adult, the prompt (714) related to the user's age information may include a prompt defined as, "Use formal speech overall, and respond simply and clearly," so as to generate a serious response.
[0098] In an exemplary embodiment, the user's gender information may indicate whether the user is male or female. For example, a preferred tone of voice may be surveyed based on the user's gender, and a prompt related to the user's gender information may be defined so that the survey results are reflected in the response. If the survey results indicate that men prefer a gentle tone of voice and women prefer a polite tone of voice, a prompt related to the user's gender information may be defined based on the survey results. For example, if the user's gender information indicates that the user is male, the prompt (711) related to the user's gender information may include a prompt defined as, "As a result of reading the user's information from the user account information, the user's gender is {male}. Male users prefer a gentle tone of voice, so please respond in a gentle tone." so that a gentle tone of voice can be generated. For example, if the user's gender information indicates a female, the prompt (712) related to the user's gender information may include a prompt defined as "As a result of reading the user's information from the user account information, the user's gender is {female}. Female users prefer a polite tone of response, so please respond in a polite tone." so as to generate a polite tone of response.
[0099] According to an exemplary embodiment, prompts (715, 716, 717) related to the user's preference information may be defined based on the user's preset preferences. For example, if the user's preference information indicates a formal tone, the prompt (715) related to the user's preference information may include a defined prompt such as, "Please respond with polite expressions and keep your responses simple and clear." For example, if the user's preference information indicates a casual tone, the prompt (716) related to the user's preference information may include a defined prompt such as, "Please use a friendly and natural conversational style." and / or a defined prompt such as, "Please use expressions and language used in everyday conversation." For example, if the user's preference information indicates a witty tone, the prompt (717) related to the user's preference information may include a defined prompt such as, "Please use a humorous and entertaining conversational style." and / or a defined prompt such as, "You may use trendy words and memes in your responses."
[0100] According to an exemplary embodiment, a prompt related to a user's emotion may be defined based on the identified user's emotion. For example, the electronic device (101) may include a sensor for acquiring data related to the user's emotion. The sensor may include, but is not limited to, a heart rate sensor for identifying emotions by measuring a heart rate, a skin resistance sensor for identifying emotions by measuring the electrical resistance of the skin, a temperature sensor for identifying emotions by measuring body temperature, a voice recognition sensor for identifying emotions by measuring the tone, rate, and intensity of the user's voice, and / or a facial recognition sensor for identifying emotions by measuring the user's facial expression. For example, if the user's emotion information indicates an emotion requiring comfort, such as sadness, anger, or fear, the prompt related to the user's emotion may include a prompt defined as, "The emotion analysis result for the user's utterance indicates that {comfort} is needed. Please respond in a warm and serious tone." The user's emotion information may be identified from the user's utterance. For example, prompts related to the user's emotions might include defined prompts such as, "The utterance interpretation result is {inquiry about research paper}, so please respond in a polite tone."
[0101] In addition to the exemplary embodiments described above, the third sub-prompts (373a) related to user information may include various embodiments.
[0102] According to an exemplary embodiment, at least one processor (310) may determine a third sub-prompt (373a) related to user information. The at least one processor (310) may determine the first sub-prompt (372) and the third sub-prompt (373a) related to the domain as second prompts, and may generate a final prompt (705) to be input into the large-scale language model (330) by combining the determined second prompt with the first prompt corresponding to the general prompt (601). For example, the at least one processor (310) may generate the final prompt (705) by combining the general prompt (601), the phone-related prompt (613), and the user age-related prompt (714). Since the final prompt (705) to be input into the large-scale language model (330) is generated based at least in part on the third sub-prompt (373a), a response output by inputting the final prompt into the large-scale language model (330) may include personal characteristics of the user. An electronic device (101) according to an exemplary embodiment can enhance the user experience by providing a response that reflects the user's personal characteristics.
[0103] Figure 8 illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a fourth sub-prompt.
[0104] Referring to FIG. 8, at least one processor (310) may be configured to obtain a fourth sub-prompt (373b) related to electronic device information from a prompt database (e.g., prompt database (321, 302) of FIG. 3a).
[0105] According to an exemplary embodiment, the electronic device information may include information related to structural characteristics and / or information related to functional characteristics of the electronic device (e.g., the electronic device (101) of FIG. 3A). For example, the electronic device information may include, but is not limited to, at least one of type information of the electronic device and / or output device information included in the electronic device. The fourth sub-prompt (373b) may include prompts (811, 812, 813, 814) related to type information of the electronic device and / or prompts related to output device information (815, 816) included in the electronic device. For example, the electronic device information may be defined in the form of system properties or system features of the electronic device (101) in a memory (e.g., the memory (320) of FIG. 3A) or an external electronic device (e.g., the external electronic device (301) of FIG. 3A).
[0106] According to an exemplary embodiment, the type information of the electronic device may indicate the type of the electronic device. As described above, the electronic device (101) may include electronic devices such as a smart phone, a tablet, a laptop, a computer, and / or a TV. The output device information included in the electronic device may indicate what type of output device (350) the electronic device (101) includes. For example, the electronic device (101) may include, but is not limited to, a display for outputting visual information (e.g., the display (351) of FIG. 3A) and / or a speaker for outputting auditory information (e.g., the speaker (352) of FIG. 3A).
[0107] For example, when the electronic device (101) includes a display (351), the type information of the electronic device may include information related to the size of the screen displayed through the display (351). The display (351) included in the electronic device (101) may have a screen of different sizes depending on the type of the electronic device. In the case of a smart watch worn on a user's wrist, the display (351) may be included to display a screen of a relatively small size. In the case of a smart phone carried by a user, the display (351) may be included to display a screen of a relatively large size. For example, when the type information of the electronic device indicates a smart watch, the prompt (812) related to the type of the electronic device may include a defined prompt such as, "The screen of the smart watch is small, so make a sentence as concise as possible within 100 characters." For example, if the type information of the electronic device indicates a smart phone, a prompt (813) related to the type of the electronic device may include a prompt defined as, "Smart phones can scroll down, so create a sentence of less than 1,000 characters."
[0108] For example, the output device information included in the electronic device may indicate what type of output device (350) the electronic device (101) includes. Since the form of providing a response to a user's utterance varies depending on the type of the output device (350), the third sub-prompt (373a) may include a prompt based on the type of the output device (350) included in the electronic device (101). For example, if the output device information included in the electronic device indicates a display (351), the prompt (815) related to the output device information included in the electronic device may include a prompt defined as, "For a device with a screen, respond with an emoticon that goes well with the response." For example, if the output device information included in the electronic device indicates only a speaker (352) excluding the display (351), the prompt (816) related to the output device information included in the electronic device may include a prompt defined as, "For a device without a screen, explain the response as detailed and long as possible."
[0109] In addition to the exemplary embodiments described above, the fourth sub-prompts (373b) may include various embodiments.
[0110] According to an exemplary embodiment, at least one processor (310) may determine a fourth sub-prompt (373b) related to electronic device information. At least one processor (310) may determine the first sub-prompt (372) and the fourth sub-prompt (373b) related to the domain as second prompts, and may generate a final prompt (805) to be input into the large-scale language model (330) by combining the determined second prompt with the first prompt corresponding to the general prompt (601). For example, at least one processor (310) may generate the final prompt (805) by combining the general prompt (601), the phone-related prompt (613), and the type-related prompt (812). Since the final prompt (805) to be input into the large-scale language model (330) is generated based at least in part on the fourth sub-prompt (373b), the response output by the final prompt (805) being input into the large-scale language model (330) may include characteristics of the electronic device (101). The electronic device (101) according to the exemplary embodiment may enhance the user experience by providing a response that reflects the characteristics of the electronic device (101).
[0111] Figure 9a illustrates an exemplary process for generating a final prompt by combining a general prompt, a first sub-prompt, and a fourth sub-prompt. Figure 9b illustrates an exemplary process for generating another final prompt by further combining a third sub-prompt with the final prompt generated through the process of Figure 9a. Figure 9c illustrates a response provided through the final prompt of Figure 9a. Figure 9d illustrates a response provided through the final prompt of Figure 9b.
[0112] As described above, the exemplary electronic device (101) may be configured to generate a final prompt (901) by combining a first prompt corresponding to a general prompt (601) and a second prompt determined by one or more sub-prompts. The generated final prompt (901) may be input into a large-scale language model (330) to generate a response to a user utterance. The response to the user utterance may be provided to the user via an output device (350). For example, the response may be provided as visual information via a display (351) and / or as auditory information via a speaker (352).
[0113] Referring to FIG. 9A, the first sub-prompt (372) and the fourth sub-prompt (373b) may be combined with the general prompt (601) to generate the final prompt (901). For example, at least one processor (310) may determine the first sub-prompt (372) corresponding to the domain as a phone-related prompt (613) based on command text extracted from the user's utterance. For example, the phone-related prompt (613) may include a defined prompt such as, "You perform the role of making and receiving calls to the user. When calling someone, create a response using the recipient's name." If at least one second sub-prompt is predetermined as the fourth sub-prompt (373b) related to electronic device information, at least one processor (310) may determine the fourth sub-prompt (373b). For example, if the electronic device (101) is a smart watch, the prompt (812) related to the electronic device type information may include a prompt defined as, “The screen of the smart watch is small, so please make the sentence as brief as possible within 100 characters.”
[0114] According to an exemplary embodiment, at least one processor (310) may determine the second prompt as the determined first sub-prompt (372) and the fourth sub-prompt (373b), and may generate the final prompt (901) by combining the first prompt and the second prompt corresponding to the general prompt (601). For example, the general prompt (601) corresponding to the definition of the basic concept may include a defined prompt such as “You are an artificial intelligence assistant that carries out the user’s request.” At least one processor (310) may generate the final prompt (901) by combining the command text, the general prompt (601), the first sub-prompt (372) (e.g., a prompt related to a phone (613)), and the fourth sub-prompt (373b) (e.g., a prompt related to information about the type of electronic device (812)).
[0115] Referring to FIG. 9B, the first sub-prompt (372), the third sub-prompt (373a), and the fourth sub-prompt (373b) may be combined with the general prompt (601) to generate a final prompt (901). For example, at least one processor (310) may additionally combine the third sub-prompt (373a) with the first final prompt (901) generated through the process illustrated in FIG. 9A to generate a second final prompt (902) to be input into the large-scale language model (330). For example, if the user's age information indicates a child, the prompt (713) related to the user's information may include a prompt defined as, "Use informal speech rather than formal speech, but use expressions that are easy and friendly enough for children to understand." At least one processor (310) can generate a final prompt (901) by combining command text, a general prompt (601), a first sub-prompt (372), a third sub-prompt (373a), and a fourth sub-prompt (373b).
[0116] Referring to FIG. 9c, the exemplary electronic device (101) may be configured to provide a user with a response (920) output in response to an input of a second final prompt (902). For example, the electronic device (101) may provide the response (920) via a display (351) and / or a speaker (352).
[0117] FIG. 9C illustrates an example of a response (920) provided via an output device (350) when a first final prompt (901) generated through the process of FIG. 9A is input into a large-scale language model (e.g., a large-scale language model (330) of FIG. 3A). Referring to FIG. 9C, at least one processor (e.g., at least one processor (310) of FIG. 3A) can generate a first final prompt (e.g., the first final prompt (901) of FIG. 9C) by combining a command text, a general prompt (e.g., the general prompt (601) of FIG. 9C), a first sub-prompt (e.g., the first sub-prompt (372) of FIG. 9C), and a fourth sub-prompt (e.g., the fourth sub-prompt (373b) of FIG. 9C). At least one processor (310) can obtain a response (920) to a user utterance (910) by inputting a first final prompt (901) into a large-scale language model (330). For example, if the user utterance (910) is “Call Mom,” the at least one processor (310) can receive the user utterance (910) via an input device (e.g., an input device (340) of FIG. 3A) (e.g., a microphone). The at least one processor (310) can obtain a response (920) such as “Call Mom’s contact number” by inputting the first final prompt (901), which is generated by combining text, a general prompt (601), a first sub-prompt (372), and a fourth sub-prompt (373b), into the large-scale language model (330). By the first sub-prompt (372), a response (920) including the recipient's name (e.g., "Mom") can be obtained. By the fourth sub-prompt (373b), a response (920) including a brief sentence of less than 100 characters can be obtained. The electronic device (101) can display the response (920) as visual information on the display (351) and / or output it as auditory information through the speaker (352).
[0118] FIG. 9d illustrates an example of a response (930) provided via an output device (350) when a second final prompt (902) generated through the process of FIG. 9b is input into a large-scale language model (e.g., the large-scale language model (330) of FIG. 3a). Referring to FIG. 9d, at least one processor (e.g., at least one processor (310) of FIG. 3a) can generate a second final prompt (e.g., the second final prompt (902) of FIG. 9b) by combining a command text, a general prompt (e.g., the general prompt (601) of FIG. 9b), a first sub-prompt (e.g., the first sub-prompt (372) of FIG. 9b), a third sub-prompt (e.g., the third sub-prompt (373a) of FIG. 9b), and a fourth sub-prompt (e.g., the fourth sub-prompt (373b) of FIG. 9b). At least one processor (310) can obtain a response (930) to a user utterance (910) by inputting a second final prompt (902) into a large-scale language model (330). For example, if the user utterance (910) is “Call Mom,” the at least one processor (310) can receive the user utterance (910) via an input device (e.g., an input device (340) of FIG. 3A) (e.g., a microphone). The at least one processor (310) can obtain a response (930) such as “I’ll call Mom!” by inputting a first final prompt (901) generated by combining text, a general prompt (601), a first sub-prompt (372), a third sub-prompt (373a), and a fourth sub-prompt (373b) into the large-scale language model (330). By the third sub-prompt (373a), a response (930) including informal speech can be obtained.
[0119] According to an exemplary embodiment, the electronic device (101) can enhance the diversity of responses to a user utterance (910) by generating a final prompt (901) using one or more sub-prompts based on various information. As the user's personal characteristics and / or the characteristics of the electronic device (101) are reflected in the prompts for generating responses, a variety of responses can be provided. According to an exemplary embodiment, the user experience can be enhanced.
[0120] Fig. 10a illustrates a prompt of an electronic device according to a comparative example. Fig. 10b illustrates a prompt of an electronic device according to one embodiment.
[0121] According to one embodiment of the present disclosure, the electronic device (101) can reduce query costs by reducing the number of tokens to be processed through a large-scale language model (330) by using at least one sub-prompt.
[0122] Referring to FIG. 10A, an electronic device according to a comparative example may be referred to as an electronic device that generates a final prompt (1001) without using sub-prompts based on various information. In the case of the electronic device according to the comparative example, a final prompt (1001) including a relatively large number of tokens may be used to enhance the diversity of responses to user utterances. For information to be reflected in the response, the final prompt (1001) may include a plurality of conditional statements (if statements). For example, prompts for each domain, each type of electronic device, and / or each user may be included in the final prompt (1001) in the form of conditional statements. The cost for a large-scale language model (e.g., the large-scale language model (330) of FIG. 3A) may be determined based on the number of tokens in the prompt. Since the electronic device according to the comparative example uses a final prompt (1001) including a relatively large number of tokens, the query cost may increase.
[0123] Referring to FIG. 10B, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 3A) may be referred to as an electronic device that generates a final prompt (1002) using sub-prompts based on various information. In the case of the electronic device (101) according to one embodiment, since the final prompt is generated by combining one or more sub-prompts with a general prompt, the diversity of responses to a user utterance may be enhanced, and the number of tokens of the final prompt (1002) may be reduced. For example, since the electronic device (101) determines a first sub-prompt (372) related to a domain, a third sub-prompt (373a) related to user information, and / or a fourth sub-prompt (373b) related to electronic device information, the final prompt (1002) may not be configured in the form of a conditional statement. Unlike the electronic device according to the comparative example, the electronic device (101) according to one embodiment can utilize a final prompt (1002) containing relatively few tokens because it combines only specific sub-prompts related to user information and / or electronic device information. Since the electronic device (101) according to one embodiment utilizes a final prompt (1002) containing relatively few tokens, query costs can be reduced.
[0124] An electronic device (101) is provided. The electronic device (101) may include at least one processor (310). The electronic device (101) may include a memory (320) that stores instructions. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to extract a command text from a user utterance, determine a domain corresponding to the command text based on the command text, determine a first sub-prompt related to the domain, determine at least one second sub-prompt based on information stored in the memory (320) or the external electronic device (301), generate a final prompt to be input into a large language model (330) (LLM) based at least in part on the command text, the first sub-prompt, and the at least one second sub-prompt, and input the final prompt into the large language model (330), thereby obtaining response information for the user utterance.
[0125] According to an exemplary embodiment, the instructions, when executed individually or collectively by the at least one processor (310), may cause the electronic device (101) to determine the domain by generating a prompt for determining the domain based on the command text and inputting the prompt for determining the domain into the large-scale language model (330).
[0126] According to an exemplary embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain the domain information from the command text using an artificial intelligence model.
[0127] According to an exemplary embodiment, the at least one second sub-prompt may include a third sub-prompt related to user information stored in the memory (320) or in the external electronic device (301).
[0128] According to an exemplary embodiment, the user information may include at least one of the user's age information, the user's gender information, the user's preference information, or the user's emotion information.
[0129] According to an exemplary embodiment, the user information may be configured to determine the tone of a response to be output based on the response information.
[0130] According to an exemplary embodiment, the at least one second sub-prompt may include a fourth sub-prompt related to the electronic device information stored in the memory (320) or the external electronic device (301).
[0131] According to an exemplary embodiment, the electronic device information may include at least one of information on the type of the electronic device or information on an output device (350) included in the electronic device (101).
[0132] According to an exemplary embodiment, the output device (350) may further include a display (351) for providing visual information. The type information of the electronic device may include information related to the size of the screen displayed through the display (351).
[0133] According to an exemplary embodiment, the at least one second sub-prompt may include a fifth sub-prompt related to at least one of persona information or language information stored in the memory (320) or in the external electronic device (301).
[0134] According to an exemplary embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate the final prompt by combining a first prompt corresponding to a common prompt stored in the memory (320) or the external electronic device (301), and a second prompt based on the command text, the first sub-prompt, and the at least one second sub-prompt.
[0135] According to an exemplary embodiment, the electronic device (101) may further include a display (351) for providing visual information. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to control the display (351) to display a response based on the response information as the visual information based on the acquisition of the response information.
[0136] According to an exemplary embodiment, the electronic device (101) may further include a speaker (352) for providing auditory information. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to control the speaker (352) to output a response based on the response information as the auditory information based on the acquisition of the response information.
[0137] According to an exemplary embodiment, the memory (320) or the external electronic device (301) may include a database that stores at least one of the electronic device information or the user's account information of the electronic device (101).
[0138] A method performed by an electronic device (101) is provided. The method may include extracting a command text from a user utterance. The method may include determining a domain corresponding to the command text based on the command text. The method may include determining a first sub-prompt related to the domain. The method may include determining at least one second sub-prompt based on information stored in the memory (320) or an external electronic device (301). The method may include generating a final prompt to be input into a large language model (LLM) (330) based at least in part on the command text, the first sub-prompt, and the at least one second sub-prompt. The method may include obtaining response information for the user utterance by inputting the final prompt into the large language model (330).
[0139] According to an exemplary embodiment, the method may further include the action of generating a prompt for determining the domain based on the command text. The method may further include the action of determining the domain by inputting the prompt for determining the domain into the large-scale language model (330).
[0140] According to an exemplary embodiment, the at least one second sub-prompt may include a third sub-prompt related to user information stored in the memory (320) or in the external electronic device (301).
[0141] According to an exemplary embodiment, the at least one second sub-prompt may include a fourth sub-prompt related to the electronic device information stored in the memory (320) or the external electronic device (301).
[0142] According to an exemplary embodiment, the at least one second sub-prompt may include a fifth sub-prompt related to at least one of persona information or language information stored in the memory (320) or in the external electronic device (301).
[0143] According to an exemplary embodiment, the method may further include generating the final prompt by combining a first prompt corresponding to a common prompt stored in the memory (320) or the external electronic device (301), and a second prompt based on the command text, the first sub-prompt, and the at least one second sub-prompt.
[0144] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0145] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0146] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0147] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (120) (e.g., the processor (120)) of a machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0148] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as a memory (130) of a manufacturer's server, an application store's server, or a relay server.
[0149] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, At least one processor comprising a processing circuit; and A memory comprising one or more storage media storing instructions, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Extract command text from user utterances, Based on the above command text, determine a domain corresponding to the above command text, Determine the first sub-prompt related to the above domain, Based on information stored in said memory or external electronic device, determining at least one second sub-prompt, Generating a final prompt that is input to a large-scale language model, based at least in part on the command text, the first sub-prompt and the at least one second sub-prompt; By inputting the final prompt above into the large-scale language model, it causes the response information for the user utterance to be obtained. Electronic devices.
2. In paragraph 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Generate a prompt to determine the domain based on the above command text, By inputting the above prompt for determining the above domain into the above large-scale language model, causing it to determine the above domain, Electronic devices.
3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Using an artificial intelligence model, causing the domain information to be obtained from the command text. Electronic devices.
4. In any one of paragraphs 1 to 3, At least one of the second sub-prompts above, comprising a third sub-prompt related to user information stored within said memory or within said external electronic device; Electronic devices.
5. In paragraph 4, The above user information is: Containing at least one of the user's age information, the user's gender information, the user's preference information, or the user's emotion information. Electronic devices.
6. In paragraph 5, The above user information is: configured to determine the tone of the response to be output based on the above response information, Electronic devices.
7. In any one of paragraphs 1 to 6, At least one of the second sub-prompts above, comprising a fourth sub-prompt related to the electronic device information stored in the memory or the external electronic device; Electronic devices.
8. In paragraph 7, The above electronic device information is, At least one of the type information of the electronic device, or the output device information included in the electronic device, Electronic devices.
9. In paragraph 8, The above output device is, Further comprising a display for providing visual information, Information on the type of the above electronic device, Contains information related to the size of the screen displayed through the above display. Electronic devices.
10. In any one of paragraphs 1 to 9, At least one of the second sub-prompts above, comprising a fifth sub-prompt related to at least one of persona information or language information stored in said memory or in said external electronic device; Electronic devices.
11. In any one of paragraphs 1 to 10, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: causing the final prompt to be generated by combining a first prompt corresponding to a general prompt, stored in the memory or in the external electronic device, and a second prompt based on the command text, the first sub-prompt and the at least one second sub-prompt; Electronic devices.
12. In any one of paragraphs 1 to 11, Further comprising a display for providing visual information, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the acquisition of the above response information, causing the display to be controlled so as to display a response based on the above response information as the visual information. Electronic devices.
13. In any one of paragraphs 1 to 12, Further comprising a speaker for providing auditory information; The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the acquisition of the above response information, causing the speaker to be controlled to output a response based on the above response information as the auditory information. Electronic devices.
14. In any one of paragraphs 1 to 13, The above memory or the above external electronic device, A database comprising at least one of the electronic device information or the account information of the user of the electronic device, Electronic devices.
15. In a method performed by an electronic device, The act of extracting command text from user utterance; An action of determining a domain corresponding to the command text based on the command text; An action to determine a first sub-prompt related to the above domain; An action of determining at least one second sub-prompt based on information stored in said memory or external electronic device; An operation of generating a final prompt to be input to a large-scale language model, based at least in part on said command text, said first sub-prompt and said at least one second sub-prompt; and An operation for obtaining response information for the user utterance by inputting the final prompt into the large-scale language model, method.
Citation Information
Patent Citations
Incremental speech input interface with real time feedback
EP3640938A1
Response generation apparatus and method of the same
JP2023158992A
Door opening and closing device for semiconductor or display manufacturing facility to prevent gaps from occuring
KR1020230060578A
Motorcycle battery controller with protective electrical connections
KR1020240009183A
KR20230071045A