Electronic device and method of processing user utterance
Patent Information
- Application Number
- PCT/KR2025/002761
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-02-27
- Publication Date
- 2025-10-02
AI Technical Summary
Existing voice assistants struggle with effectively processing user utterances in languages other than English due to the dominance of English data in training large-scale language models, leading to inefficiencies in recognizing and responding to commands in minority languages.
An electronic device and method that rephrases user inputs in minority languages into a dominant language, utilizing a generative model to generate prompts that are better recognized by the language model, enhancing the recognition and processing of commands in voice assistants.
Improves the recognition and processing of user commands in minority languages by adapting the input to better align with the dominant language model, thereby enhancing the functionality and effectiveness of voice assistants.
Smart Images

Figure KR2025002761_02102025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for processing user speech
[0001] Embodiments of the present invention relate to an electronic device and a method for processing user speech.
[0002] Generative models (e.g., generative artificial intelligence) are artificial intelligence models that generate new data based on input data. Generative models can be utilized for user utterance processing systems, such as voice agents (e.g., voice assistants).
[0003] When training language models used in voice assistants (e.g., large-scale language models), a method utilizing a large amount of English data and a small amount of data from other languages is widely used. Large-scale language models utilize a large amount of training data (e.g., corpora) and can learn languages other than English with just a small amount of cross-reference data. The proportion of English data is increasing due to its advantage in various benchmarks.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0005] A method of operating an electronic device according to one embodiment may include an operation of obtaining a text indicating a command to be performed by the electronic device based on a user input. The method may include an operation of rephrasing the text. The method may include an operation of generating, based on the text, information about a candidate function translated into a language different from a language constituting the supported functions among the functions supported by the electronic device. The method may include an operation of generating a prompt based on the rephrased text and the information about the candidate function. The method may include an operation of inputting the prompt into a language model to obtain an output corresponding to the user input. The rephrased text may be obtained by a language model different from the language model.
[0006] An electronic device according to one embodiment may include at least one processor. The electronic device may include a memory that stores instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain text indicating a command to be performed by the electronic device based on a user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to rephrase the text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, based on the text, information about candidate functions supported by the electronic device that are translated into a language different from a language that constitutes the supported functions. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to generate a prompt based on the information about the rephrased text and the candidate function. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to input the prompt into a language model to obtain an output corresponding to the user input. The rephrased text may be obtained by a language model different from the language model.
[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0008] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.
[0009] FIG. 3 is a diagram for explaining a user speech processing system according to one embodiment.
[0010] FIG. 4 is a diagram for explaining information about functions supported by an electronic device according to one embodiment.
[0011] FIGS. 5 to 7 are drawings for explaining the operation of a user speech processing system according to one embodiment.
[0012] FIG. 8 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0013] FIG. 9 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0014] FIG. 10 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0015] FIG. 11 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0016] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.
[0017]
[0018] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0019] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0020] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134).
[0021] According to one embodiment, the processor (120) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (120) may include one or more processors. For example, the processor (120) may include a combination of one or more processors such as a CPU, a GPU, an MPU, an AP, and a CP.
[0022] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0023] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0024] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto.
[0025] According to one embodiment, the memory (130) may include one or more memories. Instructions stored in the memory (130) may be stored in a single memory. Instructions stored in the memory (130) may be divided and stored in multiple memories. Instructions stored in the memory (130) may be executed by a single processor (e.g., a main processor (121) or a secondary processor (123) such as a communication processor) or may be executed by multiple processors operating cooperatively (e.g., a main processor (121) and a secondary processor (123)).
[0026] According to one embodiment, the instructions stored in the memory (130) may be individually or collectively executed by the processor (120) to cause the electronic device (101) to perform and / or control the user speech processing method described with reference to FIGS. 2 to 11. The instructions stored in the memory (130) may be individually or collectively executed by a plurality of processors (e.g., the main processor (121) and / or the auxiliary processor (123)) to cause the electronic device (101) to perform and / or control the user speech processing method described with reference to FIGS. 2 to 11. According to one embodiment, the memory (130) may include a volatile memory (132) or a nonvolatile memory (134).
[0027] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0032] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0033] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0035] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0036] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0042] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0045]
[0046] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.
[0047] Referring to FIG. 2, according to one embodiment, the generative artificial intelligence system (200) may be a program (e.g., a software module) implemented on an electronic device (e.g., an electronic device (101) of FIG. 1) and / or a server (e.g., a server (108) of FIG. 1).
[0048] According to one embodiment, a User Query / Response Interface (210) can receive user input. The user input can be any type of input, such as natural language, image, audio, and / or video. Additionally, context information can be transmitted together with the user input. The context information can include various side information related to the time at which the user input is input into the artificial intelligence system (200). For example, the context information can include information about the application currently being used by the user or information about the user's location. Additionally, the user input can be a mixed type of input that includes natural language, image, audio, video, and / or context information as described above. Additionally, the user input can include non-natural language input, such as selecting a menu.
[0049] According to one embodiment, a user query / response interface (210) may provide output from a generative artificial intelligence system to a user. The output may include a natural language-based response and / or specific content. The output may also include an action requested by the user.
[0050] In one embodiment, an AI framework (220) may receive user input. Based on the user input (e.g., a user's query), the AI framework (220) may coordinate and / or control one or more components necessary to perform an action corresponding to the user's intent.
[0051] According to one embodiment, user input received from the user query / response interface (210) may be transmitted to a prompt design component (221). The prompt design component (221) may be used to generate a prompt suitable as input to a generative model (e.g., a large language model (LLM) and / or a large multimodal model (LMM)) based on the user input.
[0052] In one embodiment, the prompt design component (221) may be an AI component that utilizes a machine learning algorithm or a neural network. The prompt design component (221) may generate improved prompts over time through learning. The prompt design component (221) may access a knowledge repository (230) to generate prompts based on user input. The knowledge repository (230) may include user preference data, a prompt library, and / or prompt examples. The prompt design component (223) may provide the generated prompts to a generative model (e.g., an LLM and / or an LMM).
[0053] According to one embodiment, the APIs / Plugins management component (223) can communicate with an external information source based on a request for additional information when user input is transmitted to the generative model.
[0054] According to one embodiment, the APIs / Plugins management component (223) can establish a communication channel for communication with the outside of the system (200) via the API. The APIs / Plugins management component (223) can enable access to various data sources via the communication channel. The acquired information can be used to generate prompts by the prompt design component (221) together with user input, or can be used as input for the input of the generative model (250).
[0055] According to one embodiment, the APIs / Plugins management component (223) may request a final action via an API when the final action corresponding to user input, rather than an intermediate action, must be performed by an application or service.
[0056] In one embodiment, the refiner component (225) can fine-tune the output of the generative model (250). For example, the refiner component (225) can determine the relevance (e.g., a score) between the output (e.g., content) of the generative model and the user input. For example, the refiner component (225) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (225) can determine whether the output contains harmful information (e.g., violent content or profanity).
[0057] In one embodiment, the refinement component (225) may determine the degree of matching (e.g., a score) between the output of the generative model (250) and the user input (e.g., the intent of the user input). If the refinement component (225) determines that the output of the generative model (250) does not correspond to the user input, the refinement component (225) may modify the output to correspond to the user input.
[0058] In one embodiment, the refinement component (225) may provide hints to the user (e.g., hints for prompt generation) to enable the user to obtain information that matches the user's intent from the generative model (250).
[0059] According to one embodiment, a generative model (250) may refer to an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on user input (e.g., user utterance). The generative model (250) may include an image generation model and / or a language generation model.
[0060] In one embodiment, the image generation model may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). An example of an image generation model is a diffusion-based generative model having the structure of a VAE and a transformer.
[0061] In one embodiment, a language generation model (e.g., ChatGPT) may be a model trained to generate statistically most appropriate output based on input. The language generation model may include an LMM. The LMM can identify various types of input, such as text, images, audio (e.g., speech), and / or video, and generate new data corresponding to the input.
[0062]
[0063] FIG. 3 is a drawing for explaining a user speech processing system according to one embodiment, and FIG. 4 is a drawing for explaining information about a function supported by an electronic device according to one embodiment.
[0064] Referring to FIG. 3, according to one embodiment, the user speech processing system (300) may be a system for a voice agent (e.g., a voice assistant). The user speech processing system (300) may include a language model (350) (e.g., the generative model (250) of FIG. 2).
[0065] According to one embodiment, the user speech processing system (300) may generate a prompt suitable for a language model (e.g., language model (350)) based on training data composed of a first language and / or a second language. The amount of training data composed of the first language may be greater than the amount of training data composed of the second language. The first language may be a dominant language (e.g., English), and the second language may be a minority language (e.g., a language other than English, including Korean). The prompt generated by the user speech processing system (300) may be composed of the minority language (e.g., Korean). Instead of generating (or translating) the prompt in the dominant language, the user speech processing system (300) may rephrase the prompt composed of the minority language so that it can be matched with the dominant language. The prompt of the user speech processing system (300) may have a high recognition rate for the language model (e.g., a generative model). The prompt of the user speech processing system (300) may be generated taking into account the functions supported by the voice assistant.
[0066] According to one embodiment, the user speech processing system (300) may use an artificial intelligence model to generate a prompt.
[0067] According to one embodiment, the artificial intelligence model (or AI neural network) may include various foundation models such as a language model, a code model, an image model, and / or other artificial intelligence neural network models. The artificial intelligence model may include a large language model (LLM) and / or a large vision model (LVM). For convenience of explanation, the present disclosure will describe the large language model (LLM) and / or the large vision model (LVM) as examples.
[0068] The AI model that can be used in this disclosure may include an LLM, an AI neural network-based language model that has learned a large amount of text data through pre-training. The LLM may contain a relatively larger number of parameters (e.g., approximately 10 billion or more) than existing general language models. The LLM may utilize a transformer AI neural network structure based on an attention mechanism.
[0069] In one embodiment, the training of the LLM may include pre-training and / or fine-tuning. Pre-training may involve training the LLM to acquire general language knowledge using a large amount of text data. For example, pre-training may involve self-supervised learning, which predicts the next word in a text string using a previous word string. Fine-tuning may involve training the LLM to be suitable for a specific domain (e.g., chatbot, AI assistant, translation, summary generation, question answering) and / or task. Fine-tuning may involve further training (e.g., supervised learning, adaptive learning) the LLM using a dataset corresponding to the specific domain and / or task based on the pre-trained model. The LLM may perform a task based on text input containing natural language, referred to as a prompt.
[0070] In one embodiment, fine-tuning can be omitted in LLM learning. Users can control the prompts provided to the LLM to improve performance on a desired task. For example, users can control whether the prompts provide additional examples of tasks and / or guidance for performing the task, such as in-context learning, zero-shot learning, and / or few-shot learning. Publicly available LLMs include Bidirectional Encoder Representations from Transformer (BERT) and generative pre-trained transformer (GPT).
[0071] The term "LLM" can refer to the language neural network model itself, but it can also refer to the model of an LLM-based application (e.g., chatbot, AI assistant, translation, summary generation, text classification, sentence generation). For example, LLM-based chatbots like ChatGPT or LLM-based translators can also be referred to as "LLM."
[0072] "LLM" may include an inference engine utilizing the LLM neural network model. For example, "inputting an input prompt to the LLM" may mean "inputting the input prompt to an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may mean the output information of the last neural network layer of the LLM obtained when the input prompt is input to the LLM-based inference engine, and / or the output information modified through additional processing.
[0073] The attention mechanism is a technique that allows an AI model to focus (attention) on important parts of input data. The attention mechanism can be used to predict output data by predicting the extent to which a portion of time-series input data (e.g., time-series input data such as voice or video, or input data of some layers of a neural network) contributes to the output of the intermediate layers and / or the final output of the neural network. While a recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, may exhibit poor prediction performance when there is information dependence between long time-series distances, the attention mechanism can account for information dependence between long time-series distances by controlling the level of weight concentration (attention) within the entire and / or partial context of the input data. A transformer can be configured as an encoder-decoder structure. The encoder can process the input data and output compressed information (e.g., a contextual representation). The decoder can process compressed information and output data in token units. Each encoder and decoder can include an independent attention network, and may further include a cross-attention network connecting the encoder and decoder.
[0074] According to one embodiment, the user speech processing system (300) may include one or more components (310-360). The components (310-360) may be software modules and / or neural network models implemented on an electronic device (e.g., the electronic device (101) of FIG. 1) or a server (e.g., the server (108) of FIG. 1). The components (310-360) are illustrated as an example for describing the user speech processing system (300). Accordingly, the user speech processing system (300) may include various variations of the components (310-360) as long as the operations of the user speech processing system (300) described in the present disclosure can be implemented. For example, two or more components may be combined, or one or more components may be added or omitted. Alternatively, the user speech processing system (300) may further include one or more components (e.g., components (210 to 250) of FIG. 2) of a generative artificial intelligence system (e.g., generative artificial intelligence system (200) of FIG. 2).
[0075] According to one embodiment, the user speech processing system (300) may include an automatic speech recognition module (ASR) (310), a preprocessor (315), a prompt generation module (320), a candidate selection module (325), a database (330), a rephrasing module (340), a translation module (335), a language model (350) (e.g., the generative model (250) of FIG. 2), an alignment module (345), a postprocessor (355), and / or an application (360).
[0076] According to one embodiment, the ASR module (310) can convert a speech signal (e.g., an analog signal) (e.g., a user input (30)) into text data. When the user input is in text form, the operation of the ASR module (310) can be omitted.
[0077] In one embodiment, the preprocessor (315) may perform preprocessing on text data. For example, the preprocessor (315) may perform tokenization, which splits the text into tokens (e.g., words, phrases, and / or symbols). For example, the preprocessor (315) may remove unnecessary characters (e.g., characters such as '!' or '?') to generate a response (35) from the text data.
[0078] According to one embodiment, a prompt generation module (320) (e.g., a prompt design component (221) of FIG. 2) can generate a prompt (e.g., text indicating a command to be performed by a user utterance processing system (300) (e.g., an electronic device (101) of FIG. 1)). The prompt can be data (e.g., text, JSON (JavaScript object notation), image, audio, and / or video) that is provided to a reframing module (340), a translation module (335), and / or a language model (350) to generate new data (e.g., a response). The prompt generation module (320) can generate a prompt that is understandable by the reframing module (340), the translation module (335), and / or the language model (350) based on conditions associated with the response (35) (e.g., length or writing style) and text data corresponding to the user utterance.
[0079] In one embodiment, the rephrasing module (340) may be configured to rephrase a prompt generated by the prompt generation module (320). The rephrasing module (340) may perform a rephrasing operation. For example, the rephrasing module (340) may perform the rephrasing operation using a syntactic parser and / or rules (e.g., predefined rules) and / or artificial intelligence models (e.g., image generation models and / or language models). The rules may be predefined rules and may be updated through learning.
[0080] In one embodiment, the reframing module (340) may include an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on input (e.g., a prompt). The reframing module (340) may include an image generation model and / or a language model (e.g., a language generation model).
[0081] According to one embodiment, the reframing module (340) may be implemented as a generative model. The translation module (335) and the language model (350) may also be implemented as generative models. In this case, the reframing module (340), the translation module (335), and the language model (350) may be distinguished according to the training data. Each of the reframing module (340), the translation module (335), and the language model (350) may be trained with different training data. Each of the reframing module (340), the translation module (335), and the language model (350) may be trained with different data depending on the purpose.
[0082] According to one embodiment, the rephrasing module (340) may input a prompt (e.g., text indicating a command to be performed by the user speech processing system (300) (e.g., the electronic device (101) of FIG. 1)) and output a rephrased prompt (e.g., text). In relation to the rephrasing, the rephrasing module (340) may update the endings or particles of words included in the input prompt (e.g., text). In relation to the rephrasing, the rephrasing module (340) may update foreign language notation words or foreign words included in the input prompt (e.g., text). In relation to the rephrasing, the rephrasing module (340) may update verbs included in the input prompt (e.g., text). In relation to the rephrasing, the rephrasing module (340) may generally normalize the updated prompt (e.g., text). The rephrased prompt (e.g., text) may be rephrased to increase the recognizability of the language model (350). That is, the rewriting module (340) can transform an input prompt (e.g., text) according to certain rules to make it easier for the language model (350) to use. In this case, the certain rules may be based on the learning data (e.g., corpus) of the language model (350).
[0083] In one embodiment, the training data (e.g., corpus) of the language model (350) may be composed of literary language. Therefore, user utterances composed of colloquial language (e.g., user utterances with particles omitted) may be different from the training data of the language model (350). The rephrasing module (340) may update the word endings or particles of the input text (e.g., prompt) to correspond to the written language. For example, if text (e.g., prompt) containing "Add a vacation group" is input, the rephrasing module (340) may update it to "Add a vacation group." For example, if text containing "Delete my exercise reminder" is input, the rephrasing module (340) may update it to "Delete the reminder called exercise." Although examples of particle updates in Korean have been described, the rewriting module (340) can update the input text (e.g., prompt) to take into account particles or suffixes in other languages (e.g., gender suffixes in German).
[0084] In one embodiment, foreign language notation words or loanwords can be expressed using multiple words. For example, the English word "stopwatch" can be expressed (e.g., translated) as the Korean words "스톱왁싱," "스톱왁싱," or "초계" (chosigye). Accordingly, the rephrasing module (340) can update foreign language notation words or loanwords included in the input text (e.g., a prompt) by considering the word occupancy rates included in the training data (e.g., corpus) of the language model (350).
[0085] In one embodiment, users may include verbs in their utterances to avoid distinguishing between imperative, request, and declarative sentences. Accordingly, depending on the training data of the language model (350), the expression form of a verb (e.g., a verb included in a prompt input to the language model (350)) may result in differences in recognition rates. For example, the language model (350) may recognize "save" better than "save me." The reframing module (340) may update (e.g., normalize) verbs included in the input text.
[0086] In one embodiment, the rephrasing module (340) can normalize updated text (e.g., updating word endings or particles, updating foreign words or loanwords, and / or updating verbs). While individual updates to each element are reasonable, combining the updated elements may result in an inappropriate representation. The rephrasing module (340) can then normalize the updated text (e.g., prompts) as a whole.
[0087] In one embodiment, the text (e.g., prompt) rephrasing is described as being performed by the rephrasing module (340), but the text rephrasing may also be performed based on predefined rules or based on a syntactic parser. Furthermore, the rephrasing module (340) may be implemented as a combination of multiple software modules and / or multiple generative models rather than a single generative model. The rephrasing module (340) may be trained with a smaller amount of data compared to the language model (350).
[0088] According to one embodiment, the translation module (335) can translate information about a candidate feature in a dominant language (e.g., a second language) (e.g., English) into a weaker language (e.g., a first language) (e.g., Korean).
[0089] In one embodiment, the translation module (335) may include an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on input (e.g., a prompt). The translation module (335) may include an image generation model and / or a language model (e.g., a language generation model).
[0090] According to one embodiment, the translation module (335) may be implemented as a generative model. As described above, the reframing module (340) and the language model (350) may also be implemented as generative models, and the reframing module (340), the translation module (335), and the language model (350) may be distinguished according to the training data. The reframing module (340), the translation module (335), and the language model (350) may each be trained with different training data. The reframing module (340), the translation module (335), and the language model (350) may each be trained with different data depending on the purpose.
[0091] According to one embodiment, a function in the present disclosure may be a function supported by a user speech processing system (300) (e.g., an electronic device (101) of FIG. 1). That is, the function hereinafter may correspond to a function supported by a voice assistant (e.g., a function of performing a task or providing information through a voice command) (e.g., a weather information providing function, a schedule management function, an alarm setting function, a search and information providing function, a translation service function, and / or a memo writing function). Information about the functions supported by the user speech processing system (300) (e.g., an electronic device (101) of FIG. 1) may be stored in a database (330). Before describing the translation module (335), information about the functions supported by the user speech processing system (300) (e.g., an electronic device (101) of FIG. 1) will first be described.
[0092] Referring to FIG. 4, according to one embodiment, an example of information (40) about a function stored in a database (330) can be confirmed. Information (40) about a function supported by an electronic device (101) (e.g., a user speech processing system (300)) may include, in relation to a supportable function (41), a function_name (42), function_descriptions (43), function_instructions (44), and function_responses (45).
[0093] In one embodiment, the function_name (42) may be the name of the support function (41). The function_description (43) may include a general description of the function. The function_description (43) may include an overview of what functions are provided by the support function (41) and / or how the support function (41) operates. The function_instructions (44) may include instructions on how to use the support function (41) and / or step-by-step instructions. The function_instructions (44) may provide detailed and specific steps on how to utilize the support function (41). The function_instructions (44) may assist in the correct utilization of the function. The function_response (45) may include a method of responding to the support function (41). For example, in relation to a function that responds to a question about a new feature of an electronic device (101), the function_name (42) may include "answer_question_for_new_features", the function_description (43) may include "a function that responds to a query about a new feature of the device. For example, a function that responds to a query such as 'what is new in S24?'", and the function_instruction (44) may include "when a user asks a question related to a new feature of the device, such as 'what is new in Galaxy Fold 5', 'what is Samsung Wallet', call `support function{}`. / n params: / n - `user_saying(str)`: what the user said / n - `devices(str)`: devices such as `TV`, `mobile`, etc.", and the function_response (45) may include "result of `answer_question_for_new_features": generates a response by referencing the output field." For easy explanation, information (40) about the function stored in the database (330) is expressed in Korean, but information (40) about the function may be stored in a dominant language (e.g., second language) (e.g., English).
[0094] According to one embodiment, the candidate selection module (325) can select information about candidate functions from among information about functions stored in the database (330). The candidate selection module (325) can be implemented through an artificial neural network. For example, the candidate selection module (325) can be implemented through a question answering (QA) module (e.g., a module including a retriever, a reranker, and / or a reader) or a retrieval augmented generation (RAG) module. The candidate selection module (325) can be utilized to reduce the search space for the database (330).
[0095] According to one embodiment, the candidate selection module (325) may select information about a candidate feature based on the similarity between information about the feature and text (e.g., a prompt) (e.g., an output of the prompt generation module (320)). As described above, the information about the feature may be stored in a database (330) in a second language (e.g., a dominant language) (e.g., English), and the text (e.g., a prompt) may be stored in a first language (e.g., a minor language) (e.g., Korean). The candidate selection module (325) may include a multilingual model trained based on data composed of the first language and data composed of the second language, and the similarity between two pieces of information composed in different languages may be calculated through the multilingual model.
[0096] According to one embodiment, the translation module (335) can translate information about candidate features composed of a dominant language (e.g., a second language) (e.g., English) into a third language (e.g., a first language) (e.g., Korean). The translation module (335) can include an embedding identical to the embedding of the language model (350). An embedding can be associated with a process in natural language processing of converting a human's natural language into a machine-understandable vector. The embedding can be implemented as a neural network and included in the translation module (335). As described above, the training data of the language model (350) can be associated with a voice assistant, and the training data associated with the voice assistant can be different from the training data of a typical language model. By utilizing the translation module (335) including the same embedding as the language model (350) to translate text (e.g., a prompt) to be input to the language model (350), the recognition rate of the language model (350) can be improved.
[0097] In one embodiment, the translation module (335) may include a compressed form of the embedding compared to the language model (350) and may be trained with a smaller amount of data. It should be noted that the reframing module (340) and / or the translation module (335) are also updated as the language model (350) and / or the database (330) are updated.
[0098] According to one embodiment, the alignment module (345) can generate candidate prompts based on (e.g., by matching) the rephrased text (e.g., output of the rephrasing module (340)) and information about candidate features (e.g., output of the translation module (335)). The matching between the rephrased text (e.g., output of the rephrasing module (340)) and information about candidate features (e.g., output of the translation module (335)) can include machine learning-based matching, natural language grammar-based syntactic matching, meaning-based semantic matching, sentence similarity-based similarity matching, and / or index matching. Both the rephrased text (e.g., output of the rephrasing module (340)) and the information about the candidate features (e.g., output of the translation module (335)) may be composed of a first language (e.g., a thirteenth language), and the candidate prompts generated by the aforementioned matching may include priorities assigned based on linguistic features of the first language (e.g., a thirteenth language). The sorting module (345) may select prompts to be input to the language model (350) from among the candidate prompts based on the priorities.
[0099] In one embodiment, the language model (350) may receive a prompt and generate an output (e.g., output (58) of FIG. 5) corresponding to the user input (30). As described above, the output of the language model (350) may correspond to a function supported by a voice assistant (e.g., a user speech processing system (300)).
[0100] In one embodiment, the language model (350) may include an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on input (e.g., a prompt). The language model (350) may include an image generation model and / or a language model (e.g., a language generation model).
[0101] According to one embodiment, the language model (350) may be implemented as a generative model. As described above, the reframing module (340) and the translation module (335) may also be implemented as generative models, and the reframing module (340), the translation module (335), and the language model (350) may be distinguished according to training data. Each of the reframing module (340), the translation module (335), and the language model (350) may be trained with different training data. Each of the reframing module (340), the translation module (335), and the language model (350) may be trained with different data according to the purpose. The language model (350) may be for generating an output corresponding to the user input (30) (e.g., the output (58) of FIG. 5). The output of the language model (350) may be associated with a function supported by the user speech processing system (300) (e.g., the electronic device (101) of FIG. 1). That is, the language model (350) may be trained with data associated with functions supported by the voice assistant. The training data of the language model (350) may include training data composed of a dominant language (e.g., a second language) (e.g., English) and training data composed of a third language (e.g., a first language) (e.g., Korean). In the training data of the language model (350), the training data composed of the dominant language (e.g., a second language) (e.g., English) may be greater than the training data composed of the third language (e.g., a first language) (e.g., Korean). According to one embodiment, the post-processor (355) may perform post-processing on the output (e.g., a natural language-based response) generated by the language model (350). For example, the post-processor (355) may add characters (e.g., emoticons) or remove unnecessary words and / or unnecessary sentences.
[0102] In one embodiment, the application (360) can execute output (e.g., a command) generated by the language model (350). That is, the application (360) can perform an action corresponding to the output. After the application (360) executes, the user speech processing system (300) can provide a natural language-based response (35) to the user.
[0103]
[0104] FIGS. 5 to 7 are drawings for explaining the operation of a user speech processing system according to one embodiment.
[0105] Referring to FIG. 5, according to one embodiment, a user speech processing system (300) (e.g., an electronic device (101) of FIG. 1) can generate an output (a command Delete_reminder(exercise) to be performed in an application (360) and a natural language-based response "Successfully deleted the reminder called exercise") (58) corresponding to a user input ("Delete exercise reminder") (50). The user speech processing system (300) can perform an action corresponding to the output (58) through the application (360) and provide a response (59) corresponding to the output (58) to the user. The user input (“Delete the exercise reminder”) (50) and response (“Successfully deleted the exercise reminder”) (59) may be in a first language (e.g., a dominant language) (e.g., Korean), and the language model (350) for the voice assistant may be trained with data in a second language (e.g., a dominant language) (e.g., English).
[0106] According to one embodiment, the user speech processing system (300) can obtain text data ('delete exercise reminder') (51) from user input ('delete exercise reminder') (50) based on the ASR module (310).
[0107] According to one embodiment, the user speech processing system (300) can preprocess text data ('delete exercise reminder') (51) based on the preprocessor (315).
[0108] According to one embodiment, the user speech processing system (300) may generate a prompt (53) from text data ('delete exercise reminder') (52) (e.g., output of the preprocessor (315)) based on the prompt generation module (320). The prompt (53) may correspond to a text indicating a command to be performed by the user speech processing system (300) (e.g., electronic device (101) of FIG. 1).
[0109] According to one embodiment, the user speech processing system (300) may rephrase text (e.g., prompt) (53) based on the rephrasing module (340). The rephrased text ('delete the reminder called exercise') (54) may be rephrased to be similar to the training data (e.g., corpus) of the language model (350). Since the text rephrasing is described in detail in FIG. 3, a redundant description will be omitted. The rephrasing module (340) may be trained with a smaller amount of data compared to the language model (350).
[0110] According to one embodiment, the user speech processing system (300) can select information on candidate functions ('set alarm', 'add reminder', 'delete reminder' 쪋) (55) from among information stored in the database (330) (e.g., information on functions supported by the user speech processing system (300) (e.g., electronic device (101))) based on the candidate selection module (325).
[0111] According to one embodiment, the user speech processing system (300) can translate information about candidate functions ('set alarm', 'add reminder', 'delete reminder' 쪋) (55) composed of a second language (e.g., dominant language) (e.g., English) into a first language (e.g., minority language) (e.g., Korean) based on the translation module (335) (e.g., 'set alarm', 'add reminder', 'delete reminder' 쪋 (56)).
[0112] According to one embodiment, the user utterance processing system (300) can generate a prompt (delete 'exercise' reminder) (57) from the rephrased text ('delete 'exercise' reminder') (54) and information about candidate functions ('set alarm', 'add reminder', 'delete reminder, 쪋) (56) based on the alignment module (345).
[0113] According to one embodiment, the user speech processing system (300) can generate output (a command Delete_reminder(exercise) to be performed in the application (360) and a natural language-based response "Successfully deleted the reminder called exercise") (58) from a prompt (delete 'exercise' reminder) (57) based on the language model (350). The training data of the language model (350) may include more training data composed of a second language (e.g., dominant language) (e.g., English) than training data composed of a first language (e.g., minor language) (e.g., Korean).
[0114] In one embodiment, the user speech processing system (300) may post-process the output (e.g., natural language-based response) generated by the language model (350) based on a post-processor (355). For example, the post-processor (355) may add characters (e.g., emoticons) or remove unnecessary words and / or unnecessary sentences.
[0115] In one embodiment, the user speech processing system (300) can perform an action corresponding to the output (58) via the application (360). For example, the user speech processing system (300) can delete a reminder associated with 'exercise' in a schedule management application (e.g., a calendar application). After the application (360) performs the action, the user speech processing system (300) can provide a natural language-based response ("Successfully deleted the reminder called 'exercise'") (59) to the user.
[0116] Referring to FIG. 6, according to one embodiment, a user speech processing system (300) (e.g., an electronic device (101) of FIG. 1) can generate an output (a command to be performed in an application (360) called Stop_stopwatch and a natural language-based response "I stopped the stopwatch") (68) corresponding to a user input ("Stop the stopwatch") (60). The user speech processing system (300) can perform an action corresponding to the output (68) through the application (360) and provide a response (69) corresponding to the output (68) to the user.
[0117] The stopwatch included in the user input (“Stop the stopwatch”) (50) is a foreign word (e.g., based on a thirteenth language (e.g., Korean)) and can be expressed with multiple words. For example, the English word ‘stopwatch’ can be expressed as the Korean words ‘stopwatch’, ‘stopwatch’, or ‘chosikwa’. However, the word ‘stopwatch’ (e.g., the expression included in the user input (60)) may be included very little in the training data (e.g., corpus) of the language model (350). In addition, ‘stopwatch’ is translated into chosikwa by a typical translation model, but similarly, the word ‘chosikwa’ may be included very little in the training data (e.g., corpus) of the language model (350). The user utterance processing system (300) according to one embodiment may update foreign language notation words or foreign words included in the input text (e.g., prompt) by considering the occupancy rate of words included in the training data (e.g., corpus) of the language model (350). That is, the user speech processing system (300) can rephrase a prompt (e.g., text) composed of a dominant language (e.g., a first language) (e.g., Korean) so that it can be accurately matched with a dominant language (e.g., a second language) (e.g., English).
[0118] According to one embodiment, the user speech processing system (300) can obtain text data ('Stop the stopwatch') (61) from the user input ("Stop the stopwatch") (60) based on the ASR module (310).
[0119] According to one embodiment, the user speech processing system (300) can preprocess text data ('stop the stopwatch') (61) based on the preprocessor (315).
[0120] According to one embodiment, the user speech processing system (300) may generate a prompt (63) from text data ('Stop the stopwatch') (62) (e.g., output of the preprocessor (315)) based on the prompt generation module (320). The prompt (63) may correspond to a text indicating a command to be performed by the user speech processing system (300) (e.g., the electronic device (101) of FIG. 1).
[0121] According to one embodiment, the user speech processing system (300) may rephrase text (e.g., prompt) (63) based on the rephrasing module (340). The rephrased text ('stop the stopwatch') (64) may be rephrased to resemble training data (e.g., corpus) of the language model (350). Since the text rephrasing is described in detail in FIG. 3, a redundant description will be omitted. The rephrasing module (340) may be trained with a smaller amount of data compared to the language model (350).
[0122] According to one embodiment, the user speech processing system (300) can select information on candidate functions ('set alarm', 'turn off stopwatch', 'turn on stopwatch' 쪋) (65) from among information stored in the database (330) (e.g., information on functions supported by the user speech processing system (300) (e.g., electronic device (101))) based on the candidate selection module (325).
[0123] According to one embodiment, the user speech processing system (300) can translate information about candidate functions ('set alarm', 'turn off stopwatch', 'turn on stopwatch' 쪋) (65) composed of a second language (e.g., dominant language) (e.g., English) into a first language (e.g., dominant language) (e.g., Korean) (e.g., 'set alarm', 'stop stopwatch', 'run stopwatch' 쪋 (66)) based on the translation module (335).
[0124] According to one embodiment, the user utterance processing system (300) can generate a prompt ('Stop the stopwatch') (67) from the rephrased text ('Stop the stopwatch') (64) and information about candidate functions ('Set an alarm', 'Stop the stopwatch', 'Run the stopwatch', etc.) (66) based on the alignment module (345).
[0125] According to one embodiment, the user speech processing system (300) can generate output (a command to be performed in the application (360) Stop_stopwatch and a natural language-based response "Stopwatch stopped") (68) from a prompt ('Stop the stopwatch') (67) based on the language model (350). The training data of the language model (350) may include more training data composed of a second language (e.g., dominant language) (e.g., English) than training data composed of a first language (e.g., minor language) (e.g., Korean).
[0126] According to one embodiment, the user speech processing system (300) may perform post-processing on the output (e.g., natural language-based response) generated by the language model (350) based on the post-processor (355). For example, the post-processor (355) may add characters (e.g., emoticons) or remove unnecessary words and / or unnecessary sentences.
[0127] In one embodiment, the user speech processing system (300) can perform an action corresponding to the output (68) through the application (360). For example, in a stopwatch application, the user speech processing system (300) can stop a running stopwatch. After the application (360) performs the action, the user speech processing system (300) can provide a natural language-based response ("stopwatch stopped") (69) to the user.
[0128] Referring to FIG. 7, according to one embodiment, a user speech processing system (300) (e.g., an electronic device (101) of FIG. 1) can generate an output (a command to be performed in an application (360) called “get user nickname” and a natural language-based response “your nickname is Sam”) (78) corresponding to a user input (“Tell me your nickname”) (70). The user speech processing system (300) can perform an action corresponding to the output (78) through the application (360) and provide a response (79) corresponding to the output (78) to the user.
[0129] The nickname included in the user input (“Tell me your nickname”) (70) may be a word used to replace a nickname depending on the context. In Korean, a nickname may be an alternative expression for a person, a place, or an object. For example, the nickname “God of Soccer” may be given to a soccer player. On the other hand, a nickname may be an informal name used mainly within an individual or a group. A nickname may be a name or nickname used to refer to a specific person among family, friends, and / or colleagues. The training data (e.g., corpus) of the language model (350) may contain very few words called nicknames (e.g., expressions included in the user input (70)). The language model (350) may not provide a voice assistant function related to nicknames. The user speech processing system (300) according to one embodiment may update words included in the input text (e.g., prompt) by taking into account the function (e.g., voice assistant function) of the language model (350).
[0130] According to one embodiment, the user speech processing system (300) can obtain text data ('Tell me my nickname') (71) from user input ("Tell me my nickname") (70) based on the ASR module (310).
[0131] According to one embodiment, the user speech processing system (300) can preprocess text data ('Tell me my nickname') (71) based on the preprocessor (315).
[0132] According to one embodiment, the user speech processing system (300) may generate a prompt (73) from text data ('Tell me my nickname') (72) (e.g., output of the preprocessor (315)) based on the prompt generation module (320). The prompt (73) may correspond to text indicating a command to be performed by the user speech processing system (300) (e.g., electronic device (101) of FIG. 1).
[0133] According to one embodiment, the user speech processing system (300) may rephrase text (e.g., prompt) (73) based on the rephrasing module (340). The rephrased text ('what is my nickname') (74) may be rephrased to resemble training data (e.g., corpus) of the language model (350). Since the text rephrasing is described in detail in FIG. 3, a redundant description will be omitted. The rephrasing module (340) may be trained with a smaller amount of data compared to the language model (350).
[0134] According to one embodiment, the user speech processing system (300) can select information about a candidate function ('get user nickname') (75) from among information stored in a database (330) (e.g., information about functions supported by the user speech processing system (300) (e.g., electronic device (101))) based on the candidate selection module (325).
[0135] According to one embodiment, the user speech processing system (300) can translate information about a candidate feature ('get user nickname') (75) composed of a second language (e.g., dominant language) (e.g., English) into a first language (e.g., dominant language) (e.g., Korean) (e.g., 'what's my nickname') (76) based on the translation module (335).
[0136] According to one embodiment, the user speech processing system (300) can generate a prompt ('what's my nickname') (77) from the rephrased text ('what's my nickname') (74) and information about the candidate feature ('what's my nickname') (76), based on the alignment module (345).
[0137] According to one embodiment, the user speech processing system (300) can generate output (a command get user nickname to be performed in the application (360) and a natural language-based response "Your nickname is OO") (78) from a prompt ('What is my nickname?') (77) based on a language model (350). The training data of the language model (350) may include more training data composed of a second language (e.g., dominant language) (e.g., English) than training data composed of a first language (e.g., minor language) (e.g., Korean).
[0138] According to one embodiment, the user speech processing system (300) may perform post-processing on the output (e.g., natural language-based response) generated by the language model (350) based on the post-processor (355).
[0139] In one embodiment, the user speech processing system (300) can perform an action corresponding to the output (78) via the application (360). For example, the user speech processing system (300) can collect the user's nickname from a messaging application. The user speech processing system (300) can provide a natural language-based response ("Your nickname is Sam") (79) to the user.
[0140]
[0141] FIG. 8 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0142] Actions 810 to 850 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (810 to 850) may be changed, and at least two actions may be performed in parallel.
[0143] According to one embodiment, operations 810 to 850 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0144] In operation 810, an electronic device according to one embodiment may obtain text (e.g., a prompt) indicating a command to be performed by the electronic device based on a user input.
[0145] In operation 820, an electronic device according to one embodiment may rephrase text.
[0146] In operation 830, the electronic device according to one embodiment may generate information about a candidate function translated into a language that constitutes the user input among the functions supported by the electronic device based on the re-rendered text.
[0147] In operation 840, the electronic device according to one embodiment may generate a prompt based on information about the refreshed text and candidate features.
[0148] In operation 850, an electronic device according to one embodiment may input a prompt into a language model to obtain output corresponding to the user input.
[0149]
[0150] FIG. 9 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0151] Actions 910 to 950 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (910 to 950) may be changed, and at least two actions may be performed in parallel.
[0152] According to one embodiment, operations 910 to 950 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0153] In operation 910, an electronic device according to one embodiment may obtain text (e.g., a prompt) indicating a command to be performed by the electronic device based on a user input.
[0154] In operation 920, an electronic device according to one embodiment may rephrase text using a language model (e.g., language model (350)) and a different language model. The electronic device may input text into a rephrasing module, which is a language model different from the generative model, to obtain rephrased text.
[0155] In operation 930, the electronic device according to one embodiment may generate information about a candidate function translated into a language that constitutes the user input among the functions supported by the electronic device based on the re-rendered text.
[0156] In operation 940, the electronic device according to one embodiment may generate a prompt based on information about the refreshed text and candidate features.
[0157] In operation 950, an electronic device according to one embodiment may input a prompt into a language model (e.g., language model (350)) to obtain an output corresponding to the user input.
[0158]
[0159] FIG. 10 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0160] Actions 1010 to 1050 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (1010 to 1050) may be changed, and at least two actions may be performed in parallel.
[0161] According to one embodiment, operations 1010 to 1050 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0162] In operation 1010, an electronic device according to one embodiment may obtain text (e.g., a prompt) indicating a command to be performed by the electronic device based on a user input.
[0163] In operation 1020, an electronic device according to one embodiment may rephrase text based on rules. The rules may be predefined rules and may be updated through learning.
[0164] In operation 1030, the electronic device according to one embodiment may generate information about a candidate function translated into a language that constitutes the user input among the functions supported by the electronic device based on the re-rendered text.
[0165] In operation 1040, an electronic device according to one embodiment may generate a prompt based on information about the refreshed text and candidate features.
[0166] In operation 1050, an electronic device according to one embodiment may input a prompt into a language model to obtain output corresponding to the user input.
[0167]
[0168] FIG. 11 is a flowchart illustrating operations performed by an electronic device according to one embodiment.
[0169] Actions 1110 to 1150 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (1110 to 1150) may be changed, and at least two actions may be performed in parallel.
[0170] According to one embodiment, operations 1110 to 1150 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0171] In operation 1110, an electronic device according to one embodiment may obtain text (e.g., a prompt) indicating a command to be performed by the electronic device based on a user input.
[0172] In operation 1120, an electronic device according to one embodiment may rephrase text using a parser.
[0173] In operation 1130, the electronic device according to one embodiment may generate information about a candidate function translated into a language that constitutes the user input among the functions supported by the electronic device based on the re-rendered text.
[0174] In operation 1140, the electronic device according to one embodiment may generate a prompt based on information about the refreshed text and candidate features.
[0175] In operation 1150, an electronic device according to one embodiment may input a prompt into a language model to obtain output corresponding to the user input.
[0176]
[0177] According to one embodiment, a method of operating an electronic device (e.g., the electronic device 101 of FIG. 1) may include an operation of obtaining a text indicating a command to be performed by the electronic device based on a user input. The method may include an operation of rephrasing the text. The method may include an operation of generating, based on the text, information about a candidate function translated into a language different from a language in which the supported functions are configured among the functions supported by the electronic device. The method may include an operation of generating a prompt based on the rephrased text and the information about the candidate function. The method may include an operation of inputting the prompt into a language model to obtain an output corresponding to the user input. The rephrased text may be obtained by a language model different from the language model.
[0178] In one embodiment, the text, the rephrased text, the prompt, and the information about the candidate function may be configured in a first language that constitutes the user input. The language model is trained based on training data configured in the first language and training data configured in a second language different from the first language, and the training data configured in the second language may be greater than the training data configured in the first language.
[0179] In one embodiment, the act of rephrasing the text may include an act of updating the endings or particles of words included in the text. The act of rephrasing the text may include an act of updating foreign words or loanwords included in the text. The act of rephrasing the text may include an act of updating verbs included in the text. The act of rephrasing the text may include an act of normalizing the updated text.
[0180] According to one embodiment, if the language model is implemented as a generative model, the refreshed text may be obtained by inputting the text into a generative model other than the generative model.
[0181] In one embodiment, the operation of generating information about the candidate function may include an operation of selecting information about the candidate function based on similarity with the text among information about the function composed and stored in the second language. The operation of generating information about the candidate function may include an operation of translating the information about the candidate function composed in the second language into the first language based on a translation module.
[0182] In one embodiment, the similarity may be calculated using a multilingual model trained on data composed of the first language and data composed of the second language. The translation module may include an embedding identical to the embedding of the language model.
[0183] In one embodiment, the act of generating the prompt may include an act of generating candidate prompts based on information about the candidate function and the refreshed text. The act of generating the prompt may include an act of selecting the prompt from among the candidate prompts.
[0184] In one embodiment, the candidate prompts may include priorities assigned based on linguistic features of the first language.
[0185] According to one embodiment, the method of operating the electronic device may further include performing an operation corresponding to the output or providing a response corresponding to the output to the user.
[0186] An electronic device (e.g., electronic device 101 of FIG. 1) according to one embodiment may include at least one processor (e.g., processor 120 of FIG. 1). The electronic device may include a memory (e.g., memory 130 of FIG. 1) that stores instructions. The instructions, based on being individually or collectively executed by the at least one processor, may cause the electronic device to obtain text indicating a command to be performed by the electronic device based on a user input. The instructions, based on being individually or collectively executed by the at least one processor, may cause the electronic device to rephrase the text. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to generate, based on the text, information about a candidate function translated into a language different from a language that constitutes the supported functions among the functions supported by the electronic device. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to generate a prompt based on the re-rendered text and the information about the candidate function. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to input the prompt into a language model to obtain an output corresponding to the user input.The above-mentioned re-rendered text may be obtained by a language model other than the above-mentioned language model.
[0187] In one embodiment, the text, the refreshed text, the prompt, and the information about the candidate function may be in a first language that constitutes the user input. The language model may be trained based on training data in the first language and training data in a second language different from the first language, and the training data in the second language may be greater than the training data in the first language.
[0188] In one embodiment, the instructions, individually or collectively executed by the at least one processor, may cause the electronic device to update a suffix or particle of a word included in the text. The instructions, individually or collectively executed by the at least one processor, may cause the electronic device to update a foreign language word or foreign word included in the text. The instructions, individually or collectively executed by the at least one processor, may cause the electronic device to update a verb included in the text. The instructions, individually or collectively executed by the at least one processor, may cause the electronic device to normalize the updated text.
[0189] According to one embodiment, if the language model is implemented as a generative model, the refreshed text may be obtained by inputting the text into a generative model other than the generative model.
[0190] In one embodiment, the instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to select information about a candidate function based on similarity with the text among information about functions configured and stored in the second language. The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to translate, based on a translation module, the information about the candidate function configured in the second language into the first language.
[0191] In one embodiment, the similarity may be calculated using a multilingual model trained on data composed of the first language and data composed of the second language. The translation module may include an embedding identical to the embedding of the language model.
[0192] In one embodiment, the instructions, individually or collectively executed by the at least one processor, may cause the electronic device to generate candidate prompts based on information about the candidate function and the refreshed text. The instructions, individually or collectively executed by the at least one processor, may cause the electronic device to select the prompt from among the candidate prompts.
[0193] In one embodiment, the candidate prompts may include priorities assigned based on linguistic features of the first language.
[0194] According to one embodiment, the instructions, based on being individually or collectively executed by the at least one processor, may cause the electronic device to perform an action corresponding to the output or provide a response corresponding to the output to the user.
[0195]
[0196] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0197] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0198] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more operations. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0199] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one operation according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0200] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0201] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more operations of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In the operating method of an electronic device (101), An action of obtaining a text indicating a command to be performed by the electronic device (101) based on a user input; The action of rephrasing the above text; An operation of generating information about candidate functions translated into a language other than the language that constitutes the functions supported by the electronic device (101) based on the text; An operation of generating a prompt based on information about the above-mentioned refreshed text and the above-mentioned candidate function; and An action of inputting the above prompt into a language model to obtain output corresponding to the user input. Including, A method of operation, wherein the above-mentioned refreshed text is obtained by a language model different from the above-mentioned language model.
2. In paragraph 1, Information about the above text, the above refreshed text, the above prompt, and the above candidate function, It is composed of the first language that constitutes the above user input, The above language model is, It is learned based on learning data composed of the first language and learning data composed of a second language different from the first language, and the learning data composed of the second language is greater than the learning data composed of the first language. How it works.
3. In either of paragraphs 1 and 2, The action of re-raising the above text is: An action to update the endings or particles of words contained in the above text; An action to update foreign language words or loanwords contained in the above text; An action to update the verbs contained in the above text; and Action to normalize the updated text; A method of operation, comprising:
4. In any one of paragraphs 1 to 3, If the above language model is implemented as a generative model, The above refreshed text is, Obtained by inputting the text into a generative model other than the above generative model, How it works.
5. In any one of paragraphs 1 to 4, The action of generating information about the above candidate function is: An operation of selecting information about a candidate function based on similarity with the text among information about functions stored in the second language; and An operation of translating information about a candidate feature composed of the second language into the first language based on the translation module. A method of operation, comprising:
6. In any one of paragraphs 1 to 5, The above similarity is, It is calculated through a multilingual model learned based on data composed of the first language and data composed of the second language, The above translation module, Containing an embedding identical to the embedding of the above language model, How it works.
7. In any one of paragraphs 1 to 6, The action that generates the above prompt is: An operation of generating candidate prompts based on information about the candidate function and the refreshed text; and The action of selecting the above prompt from among the above candidate prompts A method of operation, comprising:
8. In any one of paragraphs 1 to 7, The above candidate prompts are: Including priorities assigned based on the linguistic characteristics of the first language, How it works.
9. In any one of paragraphs 1 to 8, An action that performs an action corresponding to the above output or provides a response corresponding to the above output to the user. A method of operation, further comprising:
10. A computer program stored on a computer-readable recording medium to execute any one of the methods of claims 1 to 9 in combination with hardware.
11. In the electronic device (101), At least one processor (120); and Memory for storing instructions (130) Including, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Based on user input, obtain a text indicating a command to be performed by the electronic device (101), Rephrase the above text, Based on the above text, information is generated about candidate functions translated into a language other than the language that constitutes the functions supported by the electronic device (101), Generate a prompt based on the information about the refreshed text and the candidate features, By inputting the above prompt into the language model, an output corresponding to the user input is obtained, The above-mentioned refreshed text is obtained by a language model different from the above-mentioned language model. Electronic device (101).
12. In paragraph 11, Information about the above text, the above refreshed text, the above prompt, and the above candidate function, It is composed of the first language that constitutes the above user input, The above language model is, It is learned based on learning data composed of the first language and learning data composed of a second language different from the first language, and the learning data composed of the second language is greater than the learning data composed of the first language. Electronic device (101).
13. In any one of paragraphs 11 and 12, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Update the endings or particles of words contained in the above text, Update foreign language words or loanwords included in the above text, Update the verbs contained in the above text, To normalize the updated text, Electronic device (101).
14. In any one of paragraphs 11 to 13, If the above language model is implemented as a generative model, The above refreshed text is, Obtained by inputting the text into a generative model other than the above generative model, Electronic device (101).
15. In any one of paragraphs 11 to 14, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Among the information about functions stored in the second language, information about candidate functions is selected based on the similarity with the text, Based on the translation module, information about the candidate function composed of the second language is translated into the first language. Electronic device (101).