Electronic device for prompting and method for operating same
The electronic device addresses the variability in generative model outputs by determining if training is needed and generating hierarchically organized prompt-response pairs to optimize prompts, resulting in improved response quality.
Patent Information
- Application Number
- PCT/KR2024/019274
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-12
AI Technical Summary
The quality of output from generative models varies significantly based on the quality of prompts input, which can be affected by the user's knowledge and experience, leading to inconsistent results.
An electronic device and method that determine if training of a generative model is required for a user utterance, and if so, generate hierarchically organized prompt-response pairs to improve prompt quality before inputting the prompt to the trained generative model.
This approach enhances the quality of responses generated by the generative model by ensuring that prompts are optimized based on user inputs, leading to more consistent and effective output.
Smart Images

Figure KR2024019274_12062025_PF_FP_ABST
Abstract
Description
Electronic device for prompting and method of operation thereof
[0001] Embodiments of the present invention relate to an electronic device for prompting and a method of operating the same.
[0002] Generative models (e.g., generative artificial intelligence (AI)) have recently gained popularity, and are being utilized in areas such as article writing and blogging. Because the quality of the output from a generative model varies depending on the prompts inputted into the model, prompt engineering technology is attracting attention.
[0003] To utilize a generative model, users must input prompts (or text corresponding to the prompts) into the model. Since users write their own natural language prompts, the quality of the prompts varies depending on their knowledge and experience, which can also impact the quality of the output produced by the generative model.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0005] An operating method of an electronic device according to one embodiment may include an operation of receiving a user utterance. The operating method may include an operation of determining whether training of a generative model is required to obtain a response corresponding to the utterance. If training of the generative model is required, the operating method may include an operation of generating hierarchically organized prompt-response pairs based on the utterance. The operating method may include an operation of training the generative model based on the hierarchically organized prompt-response pairs. The operating method may include an operation of generating a prompt corresponding to the utterance. The operating method may include an operation of inputting the prompt to a trained generative model to obtain a response corresponding to the utterance.
[0006] An electronic device according to one embodiment may include one or more processors. The electronic device may include a memory storing instructions. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to receive a user utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to determine whether training of a generative model is required to obtain a response corresponding to the utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to generate hierarchically organized prompt-response pairs based on the utterance if training of the generative model is required. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to train the generative model based on the hierarchically organized prompt-response pairs. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to generate a prompt corresponding to the utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to input the prompt to a trained generative model, thereby obtaining a response corresponding to the utterance.
[0007] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0008] FIG. 2 is a schematic block diagram of an electronic device according to one embodiment.
[0009] FIG. 3 is a diagram illustrating prompting of an electronic device according to one embodiment.
[0010] FIG. 4 illustrates an example of hierarchically structured prompt-response pairs according to one embodiment.
[0011] FIGS. 5 to 7 are drawings for explaining the output of an electronic device according to one embodiment.
[0012] FIG. 8 illustrates a flowchart of a method of operating an electronic device according to one embodiment.
[0013] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.
[0014]
[0015] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment.
[0016] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0017] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134).
[0018] According to one embodiment, the processor (120) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (120) may include one or more processors. For example, the processor (120) may include a combination of one or more processors such as a CPU, a GPU, an MPU, an AP, and a CP.
[0019] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0020] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0021] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto.
[0022] According to one embodiment, the memory (130) may include one or more memories. The instructions stored in the memory (130) may be stored in a single memory. The instructions stored in the memory (130) may be divided and stored in a plurality of memories. The instructions stored in the memory (130) may be individually or collectively executed by the processor (120) to cause the electronic device (101) (e.g., the electronic device (201) of FIG. 2) to perform and / or control the prompting described with reference to FIGS. 2 to 8. The instructions stored in the memory (130) may be individually or collectively executed by a plurality of processors to cause the electronic device (101) (e.g., the electronic device (201) of FIG. 2) to perform and / or control the prompting described with reference to FIGS. 2 to 8. According to one embodiment, the memory (130) may include volatile memory (132) or non-volatile memory (134).
[0023] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0024] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0025] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0026] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0027] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0028] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a depth sensor, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0029] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0030] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0031] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0032] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0033] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0034] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0035] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0036] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0037] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0038] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0039] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0040] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0041]
[0042] FIG. 2 is a schematic block diagram of an electronic device according to one embodiment.
[0043] In one embodiment, the electronic device (201) can generate a prompt for utilizing a generative model (e.g., generative artificial intelligence (AI)). A generative model is a subfield of machine learning and artificial intelligence and may include a neural network model (e.g., a language model) that analyzes given data and generates new data or content based on the data. The prompt may be data for transmitting a user's input or request to the generative model. The generative model can analyze the prompt and generate an appropriate response accordingly. A properly written prompt can help the generative model output the content desired by the user.
[0044] In one embodiment, the electronic device (201) may train a generative model prior to generating prompts. The electronic device (201) may perform curriculum learning on the generative model. Curriculum learning may support the generative model to learn progressively more complex concepts and skills. The electronic device (201) may utilize hierarchically structured prompt-response pairs for curriculum learning of the generative model. The hierarchically structured prompt-response pairs may be generated by the electronic device (201). During curriculum learning, sequential training of the generative model may be performed, starting from the lowest-level prompt-response pair (e.g., the lowest difficulty level) among the hierarchically structured prompt-response pairs. The generative model may address new and / or more complex concepts based on learning from prompts in previous levels. In the initial stages (e.g., during the training stage using the lowest-level prompts), the generative model may focus on understanding words or simple phrases, and in later stages, the generative model may move on to understanding more complex grammar, sentence structure, and higher-level concepts.
[0045] According to one embodiment, the electronic device (201) may first generate hierarchically structured prompt-response pairs for curriculum learning of a generative model. The hierarchically structured prompt-response pairs may be hierarchically structured so that the generative model can correctly respond to prompts of progressively (or stepwise) higher difficulty. The difficulty of the prompts may be defined based on the complexity of the prompts (e.g., grammar, sentence structure), the difficulty of the concepts included in the prompts, and / or the degree of semantic understanding required to utilize the prompts. Sequential learning of the generative model may be performed starting from the prompt-response pairs of the lowest level (e.g., lowest difficulty) among the hierarchically structured prompt-response pairs. The hierarchically structured prompt-response pairs may contribute to the generative model being able to respond to various tasks.
[0046] In one embodiment, the electronic device (201) may not simply segment a single, long, and complex user input into multiple actions when generating prompt-response pairs. The electronic device (201) may construct a hierarchical relationship when generating the prompt-response pairs. Based on the hierarchically structured prompt-response pairs, the electronic device (201) may train a generative model prior to performing a complex task.
[0047] In one embodiment, a generative model trained by the electronic device (201) can return an appropriate response in response to a user input requesting the performance of a complex task (e.g., a user input requesting multiple tasks, a user utterance consisting of a complex sentence structure, and / or a user input including rarely used terms). The electronic device (201) can enhance the user experience of utilizing the generative model.
[0048] In one embodiment, the electronic device (201) may use an artificial intelligence model to generate a prompt.
[0049] According to one embodiment, the artificial intelligence model (or AI neural network) may include various foundation models such as a language model, a code model, an image model, and / or other artificial intelligence neural network models. The artificial intelligence model may include a large language model (LLM) and / or a large vision model (LVM). For convenience of explanation, the present disclosure will describe the large language model (LLM) and / or the large vision model (LVM) as examples.
[0050] The AI model that can be used in this disclosure may include an LLM, an AI neural network-based language model that has learned a large amount of text data through pre-training. The LLM may contain a relatively larger number of parameters (e.g., approximately 10 billion or more) than existing general language models. The LLM may utilize a transformer AI neural network structure based on an attention mechanism.
[0051] In one embodiment, the training of the LLM may include pre-training and / or fine-tuning. Pre-training may involve training the LLM to acquire general language knowledge using a large amount of text data. For example, pre-training may involve self-supervised learning, which predicts the next word in a text string using a previous word string. Fine-tuning may involve training the LLM to be suitable for a specific domain (e.g., chatbot, AI assistant, translation, summary generation, question answering) and / or task. Fine-tuning may involve further training (e.g., supervised learning, adaptive learning) the LLM using a dataset corresponding to the specific domain and / or task based on the pre-trained model. The LLM may perform a task based on text input containing natural language, referred to as a prompt.
[0052] In one embodiment, fine-tuning can be omitted in LLM learning. Users can control the prompts provided to the LLM to improve performance on a desired task. For example, users can control whether the prompts provide additional examples of tasks and / or guidance for performing the task, such as in-context learning, zero-shot learning, and / or few-shot learning. Publicly available LLMs include Bidirectional Encoder Representations from Transformer (BERT) and generative pre-trained transformer (GPT).
[0053] The term "LLM" can refer to the language neural network model itself, but it can also refer to the model of an LLM-based application (e.g., chatbot, AI assistant, translation, summary generation, text classification, sentence generation). For example, an LLM-based chatbot like ChatGPT or an LLM-based translator can also be referred to as "LLM."
[0054] "LLM" may include an inference engine utilizing the LLM neural network model. For example, "inputting an input prompt to the LLM" may mean "inputting the input prompt to an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may mean the output information of the last neural network layer of the LLM obtained when the input prompt is input to the LLM-based inference engine, and / or the output information modified through additional processing.
[0055] The attention mechanism is a technique that allows an AI model to focus (attention) on important parts of input data. The attention mechanism can be used to predict output data by predicting the extent to which a portion of time-series input data (e.g., time-series input data such as voice or video, or input data of some layers of a neural network) contributes to the output of the intermediate layers and / or the final output of the neural network. While a recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, may exhibit poor prediction performance when there is information dependence between long time-series distances, the attention mechanism can account for information dependence between long time-series distances by controlling the level of weight concentration (attention) within the entire and / or partial context of the input data. A transformer can be configured as an encoder-decoder structure. The encoder can process the input data and output compressed information (e.g., a contextual representation). The decoder can process compressed information and output data in token units. Each of the encoder and the decoder can include an independent attention network, and can further include a cross-attention network connecting the encoder and the decoder. Referring to FIG. 2, according to one embodiment, an electronic device (201) (e.g., the electronic device (101) of FIG. 1) and a server (202) (e.g., the server (108) of FIG. 1) can be connected via a local area network (LAN), a wide area network (WAN), a value added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof.The electronic device (201) and the server (202) can communicate with each other using a wired communication method or a wireless communication method (e.g., wireless LAN (WiFi), Bluetooth, Bluetooth low energy, ZigBee, WFD (WiFi direct), UWB (ultra wide band), infrared communication (IrDA, infrared data association), NFC (near field communication)). The server (202) can manage a generative model (e.g., 203 of FIG. 3).
[0056] According to one embodiment, the electronic device (201) may be implemented as at least one of a smartphone, a tablet personal computer, a mobile phone, a speaker (e.g., an AI speaker), a video phone, an e-book reader, a desktop personal computer, a laptop personal computer, a netbook computer, a workstation, a server, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device.
[0057] According to one embodiment, the electronic device (201) may include one or more processors (220) (e.g., the processor (120) of FIG. 1). The electronic device (201) may include a memory (230) (e.g., the memory (130) of FIG. 1). The processor (220) (e.g., an application processor) may access the memory (230) to execute instructions. The memory (230) may store various data used by at least one component (e.g., the processor (220)) of the electronic device (201). The processor (220) may execute an automatic speech recognition module (ASR module) (221), a prompt planner (222), a prompt generation module (223), and a prompt management module (224).
[0058] According to one embodiment, the processor (220) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (220) may include one or more processors. For example, the processor (220) may include a combination of one or more processors such as a CPU, a GPU, an MPU, an AP, and a CP.
[0059] According to one embodiment, the memory (230) may include one or more memories. The instructions stored in the memory (230) may be stored in a single memory. The instructions stored in the memory (230) may be divided and stored in multiple memories. The instructions stored in the memory (230) may be individually or collectively executed by the processor (220) to cause the electronic device (201) to perform and / or control the prompting described with reference to FIGS. 2 to 8. The instructions stored in the memory (230) may be individually or collectively executed by multiple processors to cause the electronic device (201) to perform and / or control the prompting described with reference to FIGS. 2 to 8.
[0060] According to one embodiment, the ASR module (221) can convert data related to voice input (e.g., utterance) received by the electronic device (201) into text data.
[0061] According to one embodiment, the prompt planner (222) may determine whether training of a generative model is required based on the converted text data. The prompt planner (222) may determine whether training of the generative model is required to obtain a response corresponding to the utterance. The prompt planner (222) may utilize parameters including the perplexity of the sentence (e.g., text data) corresponding to the utterance, the complexity of the sentence, and / or the uncertainty of the sentence.
[0062] In one embodiment, perplexity may be a value indicating how novel the input data is compared to the training data of a generative model. For example, for a generative model trained on data from the news domain, the utterance "Turn on KBS" may have a perplexity value close to 0, and the utterance "Let's go skiing" may have a perplexity value close to 1. The closer the perplexity value is to 0, the more similar the utterance is to the probability distribution of the generative model. The perplexity value can be calculated using Equation 1.
[0063]
[0064] [Mathematical Formula 1]
[0065]
[0066]
[0067] In Equation 1, H(p) can be the entropy of the probability distribution. Equation 1 is merely an example to aid understanding and is not limited thereto. It can be modified, applied, or expanded in various ways. The lower the perplexity value, the higher the similarity between the training data and input data (e.g., utterances) of the generative model.
[0068] In one embodiment, the complexity of a sentence may be a value indicating how diverse and difficult the linguistic and syntactic elements (e.g., compound sentences or complex sentences, simple sentences, independent clauses, or dependent clauses) within the sentence are. The complexity of a sentence may be calculated by considering the diversity of vocabulary, sentence length, grammatical structure, and / or semantic connectivity. For example, the closer the complexity value of a sentence is to 1, the more complex the sentence may be.
[0069] In one embodiment, the uncertainty of a sentence may be a value representing the degree of uncertainty regarding the output of a generative model when input data of a type different from the training data of the generative model is input. For example, for a generative model trained on weather domain data, the utterance "Call me" may have an uncertainty value close to 1. The lower the uncertainty value, the more similar the input data (e.g., the utterance) is to the training data of the generative model.
[0070] According to one embodiment, there may be a difference in the configuration of the hierarchically organized prompt-response pairs depending on parameters (e.g., parameters including perplexity of a sentence (e.g., text data) corresponding to an utterance, sentence complexity, and / or sentence uncertainty). For example, when a sentence (e.g., a sentence corresponding to a user utterance) has a low perplexity value, a low sentence complexity value, and a low sentence uncertainty value, the generative model may have a high understanding of the user utterance. Accordingly, the electronic device (201) (e.g., the prompt generation module (223)) may not perform training of the generative model or may simply configure the prompt-response pairs used for training the generative model (e.g., configure them so that only a small number of hierarchies exist). For example, if a sentence (e.g., a sentence corresponding to a user utterance) has a high perplexity value, a high complexity value, and a high uncertainty value, a generative model may have a low understanding of the user utterance. Therefore, the electronic device (201) (e.g., a prompt generation module (223)) may densely structure prompt-response pairs (e.g., structure them so that multiple layers exist) for training the generative model.
[0071] According to one embodiment, the prompt generation module (223) may generate hierarchically organized prompt-response pairs based on utterances when training of a generative model is required. The hierarchically organized prompt-response pairs may be hierarchically organized so that the generative model can correctly respond to prompts of progressively (or stepwise) increasing difficulty. The difficulty of a prompt may be defined based on the complexity of the prompt (e.g., grammar, sentence structure), the difficulty of the concepts included in the prompt, and / or the degree to which semantic understanding is required to utilize the prompt.
[0072] According to one embodiment, the prompt generation module (223) can generate hierarchically structured prompt-response pairs based on an actor-critic algorithm of hierarchical reinforcement learning. An actor of hierarchical reinforcement learning can generate prompt-response pairs for each layer. A critic of hierarchical reinforcement learning can evaluate the prompt-response pairs for each layer generated by the actor. The actor and the critic may be trained together. The actor and the critic can generate hierarchically structured prompt-response pairs so that the generative model can perform intermediate goals (e.g., sub-goals) (e.g., sequentially outputting responses corresponding to hierarchically structured prompts) toward a final goal (e.g., outputting a response corresponding to a user utterance).
[0073] In one embodiment, the prompt generation module (223) may generate hierarchically structured prompt-response pairs based on templates corresponding to utterances. The templates may be stored in the memory (230). The prompt generation module (223) may select and utilize templates corresponding to utterances from the memory (230).
[0074] According to one embodiment, the prompt generation module (223) can generate hierarchically structured prompt-response pairs based on a syntactic analysis of a sentence corresponding to an utterance. The prompt generation module (223) can perform a syntactic analysis of a sentence (e.g., decomposing sentence components). The prompt generation module (223) can replace a sentence component with another word (or phrase). The other word replaced by the prompt generation module (223) may belong to a higher concept than the word included in the sentence. The other word replaced by the prompt generation module (223) may be a word that the generative model uses with a high frequency. The prompt generation module (223) can generate hierarchically structured prompt-response pairs by replacing and / or expanding sentence components based on a syntactic analysis of the sentence.
[0075] According to one embodiment, the prompt management module (224) can manage the hierarchically structured prompt-response pairs generated by the prompt generation module (223). The prompt management module (224) can store the hierarchically structured prompt-response pairs. The prompt generation module (223) can use the prompt-response pairs stored in the prompt management module (224) to generate new prompt-response pairs. For example, the prompt generation module (223) can replace a prompt-response pair of a specific layer among the prompt-response pairs stored in the prompt management module (224) to generate new prompt-response pairs. When a multi-turn conversation with the user is ongoing, the prompt generation module (223) can actively use the prompt-response pairs stored in the prompt management module (224). The prompt management module (224) can also store failed prompt-response pairs.
[0076] In one embodiment, the prompt management module (224) may provide feedback on hierarchically structured prompt-response pairs generated by the prompt generation module (223). The feedback provided by the prompt management module (224) to the prompt generation module (223) may occur during the training (e.g., curriculum learning) process of the generative model. Curriculum learning of the generative model will be described below.
[0077] According to one embodiment, the processor (220) can train (e.g., learn a curriculum) a generative model based on hierarchically structured prompt-response pairs. Curriculum learning can assist the generative model in learning progressively more complex concepts and skills. During curriculum learning, sequential training of the generative model can be performed, starting from the lowest level (e.g., lowest difficulty) of the hierarchically structured prompt-response pairs. The generative model can handle new and / or more complex concepts based on learning the prompts of the previous levels. In the initial stages (e.g., during the training stage using the lowest level prompts), the generative model can focus on understanding words or simple phrases, and in later stages, the generative model can move on to understanding complex grammar, sentence structures, and higher-level concepts.
[0078] According to one embodiment, the processor (220) may select a first prompt-first response pair, which is the lowest layer among the hierarchically structured prompt-response pairs. The processor (220) may input the first prompt into a generative model to obtain an output response. The prompt management module (224) may verify a correspondence between the output response and the first response. The prompt management module (224) may provide the correspondence (e.g., feedback) between the output response and the first response to the prompt generation module (223). If the output response and the first response correspond, the processor (220) may input a second prompt, which is one level higher than the first prompt, to the generative model. If the output response and the first response do not correspond, the processor (220) may debug (e.g., replace, supplement, reinforce) the first prompt. The processor (220) may repeat the above operations to train the generative model. Training of a generative model can be completed when the generative model outputs the correct response to the prompt at the top layer of the hierarchically structured prompt-response pairs.
[0079] In one embodiment, the processor (220) may generate a prompt to be input into the trained generative model. The processor (220) may input the prompt into the trained generative model to obtain a response corresponding to the utterance. The electronic device (201) may provide an output corresponding to the response to the user.
[0080] According to one embodiment, the electronic device (201) may utilize hierarchically structured prompt-response pairs to perform an action corresponding to a single user input (e.g., an utterance). As described above, the hierarchically structured prompt-response pairs may be composed of relatively short and easily understandable prompts, rather than a single complex prompt (e.g., see FIG. 4 ). A single complex prompt may be considerably long due to various constraints such as tone, context, and length. A single complex and long prompt may be costly, considering the current token-based charging system. A single complex and long prompt may require a large amount of review when an appropriate response is not output.
[0081] In one embodiment, the electronic device (201) may utilize prompt-response pairs composed of relatively short and easily understandable prompts instead of a single complex prompt. The prompt-response pairs utilized by the electronic device (201) may be hierarchically structured so that the generative model can correctly respond to prompts of progressively (or stepwise) higher difficulty. Sequential training of the generative model may be performed starting from the prompt-response pair of the lowest level (e.g., the lowest difficulty) among the hierarchically structured prompt-response pairs. The hierarchically structured prompt-response pairs may contribute to the generative model's ability to respond to various tasks.
[0082] In one embodiment, a generative model trained by the electronic device (201) can return an appropriate response in response to a user input requesting the performance of a complex task (e.g., a user input requesting multiple tasks, consisting of a complex sentence structure, and / or including rarely used terms). The electronic device (201) can enhance the user experience of utilizing the generative model.
[0083]
[0084] FIG. 3 is a diagram illustrating prompting of an electronic device according to one embodiment, and FIG. 4 illustrates examples of hierarchically structured prompt-response pairs according to one embodiment.
[0085] Referring to FIG. 3, according to one embodiment, an electronic device (201) (e.g., the electronic device (101) of FIG. 1) can return an appropriate response to a user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”). The user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) may be an utterance requesting multiple tasks (e.g., TV on, netflix start). The user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) may be an utterance including a conditional statement (e.g., “When the sun goes down”). The user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) may be an utterance including a term (e.g., TV) that a generative model (e.g., a generative model that does not perform learning for controlling external electronic devices) rarely uses. User utterances (e.g., “When the sun goes down, I turn on the TV and start Netflix”) may be utterances that require training of a generative model to obtain a corresponding response.
[0086] According to one embodiment, in operation 311, the electronic device (201) may receive a user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”).
[0087] According to one embodiment, in operation 313, the ASR module (221) may convert a user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) into text data and output the text data (e.g., “When the sun goes down, turn on the TV and start Netflix”) to the prompt planner (222).
[0088] According to one embodiment, in operation 315, the prompt planner (222) may determine whether training of a generative model is required to obtain a response corresponding to the utterance, and output the determination result (e.g., a value between 0 and 1) to the prompt generation module (223). As described above, a user utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) may be an utterance that requires training of a generative model to obtain a corresponding response.
[0089] According to one embodiment, in operation 317, the prompt generation module (223) may generate hierarchically structured prompt-response pairs (e.g., 400 of FIG. 4) and transmit a first prompt (e.g., 411 of FIG. 4), which is the lowest layer among the hierarchically structured prompt-response pairs, to the generative model (203). Referring to FIG. 4, the hierarchically structured prompt-response pairs (400) generated by the prompt generation module (223) may be confirmed. The hierarchically structured prompt-response pairs (400) may include a first prompt-first response pair (410), a second prompt-second response pair (420), and a third prompt-third response pair (430). The layers of the hierarchically structured prompt-response pairs are not limited to three as in FIG. 4. The first prompt (e.g., 411 in FIG. 4) may be used for few-shot prompting. However, prompts at a higher level than the first prompt (e.g., the second prompt (421), the third prompt (431)) may not be used for few-shot prompting. The hierarchically structured prompt-response pairs (400) may be designed to take into account a case where the generative model (203) correctly performs the action corresponding to 'Turn off Wi-Fi', but fails to perform the action corresponding to 'Turn on the TV'. The prompt generation module (223) may also structure the prompt-response pairs in other ways (e.g., 'Electronic devices can be turned on and off.' -> 'TV is an electronic device.' > 'TV can be turned on and off.').
[0090] According to one embodiment, in operation 319, the generative model (203) may transmit an output response to a first prompt (e.g., 411 of FIG. 4) to an electronic device (201) (e.g., a prompt management module (224)). The prompt management module (224) may verify a correspondence between the output response (e.g., the output response to the first prompt (e.g., 411 of FIG. 4) of the generative model (203)) and the first response (e.g., 412 of FIG. 4).
[0091] According to one embodiment, the prompt management module (224) may provide a correspondence (e.g., feedback) between the output response and the first response (e.g., 412 of FIG. 4) to the prompt generation module (223).
[0092] According to one embodiment, in operation 321, if the output response and the first response correspond, the prompt generation module (223) may input a second prompt (e.g., 421 of FIG. 4) that is one level higher than the first prompt (e.g., 411 of FIG. 4) into the generative model (203). If the output response and the first response do not correspond, the prompt generation module (223) may debug (e.g., replace, supplement, reinforce) the first prompt (e.g., 411 of FIG. 4).
[0093] In one embodiment, the curriculum method can selectively compensate for a failed step by utilizing the result of an intermediate process (e.g., the output response of the generative model (203) to the first prompt (e.g., 411 in FIG. 4)). The output of the generative model for the same input is not always the same. That is, if it fails to obtain the correct output by inputting a single complex prompt, it may be difficult to segment and debug the complex prompt. If the output response of the generative model (203) in operation 319 does not correspond to the first response (e.g., 412 in FIG. 4), the generative model (203) may not recognize the actions 'turn on' and 'turn off'. The electronic device (201) can efficiently debug the prompt by performing debugging of the prompt only in the necessary parts.
[0094] According to one embodiment, in operation 321, if the output response and the first response do not correspond, the prompt generation module (223) may debug (e.g., replace, supplement, reinforce) the first prompt (e.g., 411 of FIG. 4) and transmit the debugged first prompt (not shown) to the generative model (203).
[0095] According to one embodiment, in operation 323, the generative model (203) may transmit an output response to the debugged first prompt to an electronic device (201) (e.g., a prompt management module (224)). The prompt management module (224) may verify a correspondence between the output response and the debugged first response.
[0096] According to one embodiment, the processor (220) may train the generative model (203) by repeating the above operations. Training of the generative model (203) may be completed when the generative model outputs a correct response (e.g., a response corresponding to 432 of FIG. 4) to a prompt (e.g., 431 of FIG. 4) of the highest layer among the hierarchically structured prompt-response pairs (e.g., 400 of FIG. 4).
[0097] In one embodiment, at step 325, the prompt generation module (223) may generate a prompt to be input into the trained generative model (e.g., 203) and transmit the generated prompt to the generative model (e.g., 203). The generated prompt may be generated based on an utterance (e.g., “When the sun goes down, turn on the TV and start Netflix”) or may be generated utilizing one of the prompt-response pairs stored in the prompt management module (224).
[0098] According to one embodiment, in operation 327, the prompt management module (224) may obtain a response corresponding to the utterance from a generative model (e.g., 203). The electronic device (201) may provide an output corresponding to the response to the user. The output provided by the electronic device (201) to the user will be described below.
[0099]
[0100] FIGS. 5 to 7 are drawings for explaining the output of an electronic device according to one embodiment.
[0101] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2) can provide an output to a user corresponding to a response (e.g., a response output by a generative model in response to a user utterance).
[0102] Referring to FIG. 5, according to one embodiment, the electronic device (201) may provide the user with only a final output (502) that includes only text (e.g., text output by the generative model in response to the user utterance (501)) in response to the user utterance (501).
[0103] Referring to FIG. 6, according to one embodiment, the electronic device (201) may provide the user with only a final output (602) that includes only text (e.g., text output by the generative model in response to the user utterance (601)) in response to the user utterance (601). The electronic device (201) may receive an additional input (603) asking the user how the final output (602) was obtained. In response to the additional input (603), the electronic device (201) may provide the user with an intermediate output (604) that includes a training process based on hierarchically structured prompt-response pairs.
[0104] Referring to FIG. 7, according to one embodiment, the electronic device (201) may provide the user with an intermediate output (702) that includes a training process based on hierarchically structured prompt-response pairs in response to a user utterance (701). The electronic device (201) that has provided all of the intermediate outputs (702) may provide the user with a final output (703) that includes only text (e.g., text output by the generative model in response to the user utterance (701). Whether the intermediate output (702) is provided may vary depending on the settings. For example, the electronic device (201) may be configured to explicitly provide the intermediate output (702) from the beginning, or the electronic device (201) may be configured to provide it through a hidden menu upon the user's request.
[0105] According to one embodiment, the output provided by the electronic device (201) to the user may vary depending on the user's settings. The user's settings may be set based on the user's interaction with the user interface (e.g., touch, speech). The intermediate output (e.g., 604 of FIG. 6 and 702 of FIG. 7) may also be utilized as a speech guide for the user. By providing the intermediate output (e.g., 604 of FIG. 6 and 702 of FIG. 7), the electronic device (201) can reduce user complaints regarding malfunctions.
[0106]
[0107] FIG. 8 illustrates a flowchart of a method of operating an electronic device according to one embodiment.
[0108] Actions 810 to 860 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (810 to 860) may be changed, and at least two actions may be performed in parallel.
[0109] According to one embodiment, operations 810 to 860 may be understood to be performed in a processor (e.g., processor (220) of FIG. 3) of an electronic device (e.g., electronic device (201) of FIG. 2).
[0110] In operation 810, an electronic device according to one embodiment may receive a user utterance.
[0111] In operation 820, an electronic device according to one embodiment may determine whether training of a generative model is required to obtain a response corresponding to the utterance.
[0112] In operation 830, an electronic device according to one embodiment may generate hierarchically organized prompt-response pairs based on utterances when training of a generative model is required.
[0113] In operation 840, an electronic device according to one embodiment can train a generative model based on hierarchically organized prompt-response pairs.
[0114] In operation 850, an electronic device according to one embodiment may generate a prompt corresponding to the utterance.
[0115] In operation 860, an electronic device according to one embodiment may input a prompt to a trained generative model to obtain a response corresponding to the utterance.
[0116]
[0117] According to one embodiment, a method of operating an electronic device (e.g., electronic device 101 of FIG. 1, electronic device 201 of FIG. 2) may include an operation of receiving a user utterance. The method may include an operation of determining whether training of a generative model is required to obtain a response corresponding to the utterance. If training of the generative model is required, the method may include an operation of generating hierarchically organized prompt-response pairs based on the utterance. The method may include an operation of training the generative model based on the hierarchically organized prompt-response pairs. The method may include an operation of generating a prompt corresponding to the utterance. The method may include an operation of inputting the prompt to a trained generative model to obtain a response corresponding to the utterance.
[0118] In one embodiment, the hierarchically structured prompt-response pairs may be hierarchically structured so that the generative model can correctly respond to prompts of progressively higher difficulty. The generative model may perform sequential learning, starting from the lowest-level prompt-response pair among the hierarchically structured prompt-response pairs.
[0119] According to one embodiment, the hierarchically structured prompt-response pairs may be generated based on an actor-critic algorithm of hierarchical reinforcement learning, a template corresponding to the utterance, or a syntactic analysis of a sentence corresponding to the utterance.
[0120] In one embodiment, the actor of the hierarchical reinforcement learning generates prompt-response pairs for each layer, and the critic of the hierarchical reinforcement learning evaluates the prompt-response pairs for each layer generated by the actor, and the actor and the critic may be trained together.
[0121] In one embodiment, the operation of determining whether training of the generative model is required may be performed based on a parameter including at least one of the perplexity of the sentence corresponding to the utterance, the complexity of the sentence, or the uncertainty of the sentence. Depending on the parameter, there may be a difference in the composition of the hierarchically structured prompt-response pairs.
[0122] In one embodiment, the act of training the generative model may include performing curriculum learning on the generative model based on the hierarchically structured prompt-response pairs.
[0123] In one embodiment, the operation of training the generative model may include an operation of selecting a first prompt-first response pair, which is a lowest layer among the hierarchically structured prompt-response pairs. The operation of training the generative model may include an operation of inputting the first prompt to the generative model to obtain an output response. The operation of training the generative model may include an operation of verifying a correspondence between the output response and the first response.
[0124] In one embodiment, the operation of training the generative model may further include an operation of inputting a second prompt, which is one level higher than the first prompt, into the generative model if the output response and the first response correspond. The operation of training the generative model may further include an operation of debugging the first prompt if the output response and the first response do not correspond.
[0125] In one embodiment, the method may further include receiving a next utterance following the utterance. The method may further include generating new prompt-response pairs based on the hierarchically structured prompt-response pairs when generation of new prompt-response pairs is required for a response to the next utterance. The new prompt-response pairs may include at least one of the prompt-response pairs included in the hierarchically structured prompt-response pairs.
[0126] In one embodiment, the method may further include providing the user with an output corresponding to the response. The output may include at least one of: a final output comprising only text corresponding to the response; or an intermediate output comprising a training process based on the hierarchically structured prompt-response pairs.
[0127] An electronic device according to one embodiment (e.g., electronic device 101 of FIG. 1, electronic device 201 of FIG. 2) may include one or more processors (e.g., processor 120 of FIG. 1, processor 120 of FIG. 2). The electronic device may include a memory (e.g., memory 130 of FIG. 1, memory 130 of FIG. 2) that stores instructions. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to receive a user utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to determine whether training of a generative model is required to obtain a response corresponding to the utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to generate hierarchically organized prompt-response pairs based on the utterance, when training of the generative model is required. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to train the generative model based on the hierarchically organized prompt-response pairs. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to generate a prompt corresponding to the utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to input the prompt to a trained generative model to obtain a response corresponding to the utterance.
[0128] In one embodiment, the hierarchically structured prompt-response pairs may be hierarchically structured so that the generative model can correctly respond to prompts of progressively higher difficulty. The generative model may perform sequential learning, starting from the lowest-level prompt-response pair among the hierarchically structured prompt-response pairs.
[0129] According to one embodiment, the hierarchically structured prompt-response pairs may be generated based on an actor-critic algorithm of hierarchical reinforcement learning, a template corresponding to the utterance, or a syntactic analysis of a sentence corresponding to the utterance.
[0130] In one embodiment, the actor of the hierarchical reinforcement learning generates prompt-response pairs for each layer, and the critic of the hierarchical reinforcement learning evaluates the prompt-response pairs for each layer generated by the actor, and the actor and the critic may be trained together.
[0131] In one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to determine whether training of the generative model is required based on a parameter including at least one of perplexity of a sentence corresponding to the utterance, complexity of the sentence, or uncertainty of the sentence. Depending on the parameter, there may be a difference in the composition of the hierarchically structured prompt-response pairs.
[0132] In one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to perform curriculum learning on the generative model based on the hierarchically organized prompt-response pairs.
[0133] In one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to select a first prompt-first response pair, which is a lowest layer among the hierarchically organized prompt-response pairs. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to input the first prompt into the generative model to obtain an output response. In one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to verify a correspondence between the output response and the first response.
[0134] According to one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to: input a second prompt, which is one level higher than the first prompt, into the generative model if the output response and the first response correspond; or debug the first prompt if the output response and the first response do not correspond.
[0135] In one embodiment, the instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to receive a subsequent utterance following the utterance. The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to generate new prompt-response pairs based on the hierarchically organized prompt-response pairs when generation of new prompt-response pairs is required for a response to the subsequent utterance. The new prompt-response pairs may include at least one of the prompt-response pairs included in the hierarchically organized prompt-response pairs.
[0136]
[0137] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0138] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0139] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0140] Various embodiments of the present document may be implemented as software (e.g., a program) including one or more instructions stored in a storage medium (e.g., built-in memory or external memory) readable by a machine (e.g., an electronic device). For example, a processor (e.g., a processor) of the machine (e.g., an electronic device) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one instruction called. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' only means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0141] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0142] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In the operating method of an electronic device (101; 201), Action to receive user speech; An action for determining whether training of a generative model is required to obtain a response corresponding to the above utterance; When training of the generative model is required, an operation of generating hierarchically organized prompt-response pairs based on the utterance; An operation of training the generative model based on the hierarchically structured prompt-response pairs; An action to generate a prompt corresponding to the above utterance; and An action of inputting the above prompt to a trained generative model to obtain a response corresponding to the above utterance. A method of operation, comprising:
2. In paragraph 1, The above hierarchically structured prompt-response pairs are: The generative model is hierarchically structured so that it can correctly respond to prompts of progressively higher difficulty, or The actor-critic algorithm of hierarchical reinforcement learning, the template corresponding to the utterance, or the syntactic analysis of the sentence corresponding to the utterance, is generated. The above generative model is, Sequential learning is performed starting from the prompt-response pair of the lowest layer among the above hierarchically structured prompt-response pairs. How it works.
3. In either of paragraphs 1 and 2, The actor of the hierarchical reinforcement learning generates prompt-response pairs for each layer, and the critic of the hierarchical reinforcement learning evaluates the prompt-response pairs for each layer generated by the actor, and the actor and the critic are trained together. How it works.
4. In any one of paragraphs 1 to 3, The action for determining whether training of the above generative model is required is: It is performed based on a parameter including at least one of the perplexity of the sentence corresponding to the above utterance, the complexity of the sentence, or the uncertainty of the sentence, Depending on the above parameters, there is a difference in the composition of the hierarchically structured prompt-response pairs. How it works.
5. In any one of paragraphs 1 to 4, The operation of training the above generative model is as follows: An operation of performing curriculum learning on the generative model based on the above hierarchically structured prompt-response pairs. A method of operation, comprising:
6. In any one of paragraphs 1 to 5, The operation of training the above generative model is as follows: An action of selecting a first prompt-first response pair, which is the lowest layer among the above hierarchically structured prompt-response pairs; An operation of inputting the first prompt to the generative model to obtain an output response; and An operation for verifying the correspondence between the above output response and the above first response. A method of operation, comprising:
7. In any one of paragraphs 1 to 6, The operation of training the above generative model is as follows: If the output response and the first response correspond, an operation of inputting a second prompt, which is one level higher than the first prompt, into the generative model; or If the above output response and the above first response do not correspond, an action for debugging the above first prompt A method of operation, further comprising:
8. In any one of paragraphs 1 to 7, An action of receiving the next utterance following the above utterance; and When generation of new prompt-response pairs is required for the response to the above next utterance, an operation of generating the new prompt-response pairs based on the hierarchically organized prompt-response pairs. Including more, The above new prompt-response-pairs are, A method of operation, comprising at least one of the prompt-response pairs included in the hierarchically structured prompt-response pairs.
9. In any one of paragraphs 1 to 8, An action that provides the user with output corresponding to the above response. Including more The above output is, A final output containing only the text corresponding to the above response; or Intermediate output including a training process based on the above hierarchically structured prompt-response pairs. A method of operation, comprising at least one of:
10. In an electronic device (101; 201), One or more processors (120; 220); and Memory for storing instructions (130; 230) Including, The above instructions, when individually or collectively executed by one or more processors (120; 220), cause the electronic device (101; 201) to: Receive user speech, Determine whether training of a generative model is required to obtain a response corresponding to the above utterance, When training of the above generative model is required, hierarchically organized prompt-response pairs are generated based on the above utterances, Based on the above hierarchically structured prompt-response pairs, the generative model is trained, Generate a prompt corresponding to the above utterance, By inputting the above prompt to the trained generative model, it obtains a response corresponding to the above utterance. Electronic devices (101; 201).
11. In paragraph 10, The above hierarchically structured prompt-response pairs are: The generative model is hierarchically structured so that it can correctly respond to prompts of progressively higher difficulty, or The actor-critic algorithm of hierarchical reinforcement learning, the template corresponding to the utterance, or the syntactic analysis of the sentence corresponding to the utterance, is generated. The above generative model is, Sequential learning is performed starting from the prompt-response pair of the lowest layer among the above hierarchically structured prompt-response pairs. Electronic devices (101; 201).
12. In either of paragraphs 10 and 11, The actor of the hierarchical reinforcement learning generates prompt-response pairs for each layer, and the critic of the hierarchical reinforcement learning evaluates the prompt-response pairs for each layer generated by the actor, and the actor and the critic are trained together. Electronic devices (101; 201).
13. In any one of paragraphs 10 to 12, The above instructions, when individually or collectively executed by one or more processors (120; 220), cause the electronic device (101; 201) to: Determine whether training of the generative model is required based on a parameter including at least one of the perplexity of the sentence corresponding to the utterance, the complexity of the sentence, or the uncertainty of the sentence, Depending on the above parameters, there is a difference in the composition of the hierarchically structured prompt-response pairs. Electronic devices (101; 201).
14. In any one of paragraphs 10 to 13, The above instructions, when individually or collectively executed by one or more processors (120; 220), cause the electronic device (101; 201) to: Based on the above hierarchically structured prompt-response pairs, curriculum learning is performed for the generative model. Electronic devices (101; 201).
15. In any one of paragraphs 10 to 14, The above instructions, when individually or collectively executed by one or more processors (120; 220), cause the electronic device (101; 201) to: Select the first prompt-first response pair, which is the lowest level among the above hierarchically structured prompt-response pairs, Input the above first prompt into the generative model to obtain an output response, To verify the correspondence between the above output response and the above first response, Electronic devices (101; 201).
Citation Information
Patent Citations
Intelligent customer service application method and device, equipment and storage medium
CN117093685A
Training support apparatus of Raspberry Pi with mounting display device
KR102591187B1
Generating Language Models
US20150348541A1
Speech processing system and method
US20200320987A1
Identifying chat correction pairs for trainig model to automatically correct chat inputs
US20230177263A1