Electronic device for processing user speech, operating method thereof, and computer-readable recording medium

The electronic device employs a dual generative model approach to process user speech, ensuring coherent and contextually relevant responses, addressing inconsistencies in existing systems by generating accurate natural language and action-based outputs.

WO2025183328A1PCT designated stage Publication Date: 2025-09-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/021477
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2024-12-30
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing user speech processing systems struggle to generate coherent and contextually relevant responses to user utterances, often leading to inconsistencies and inaccuracies in natural language-based interactions.

Method used

An electronic device equipped with a generative model that processes user utterances through multiple stages of prompt generation and refinement, utilizing both a first and second generative model to ensure accurate and contextually appropriate responses, including both natural language and action-based outputs.

Benefits of technology

The system provides consistent and contextually relevant responses by generating natural language-based outputs that align with user intent, while also performing actions requested by the user, thereby enhancing the overall user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021477_04092025_PF_FP_ABST
    Figure KR2024021477_04092025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device for processing a user utterance, an operating method thereof, and a computer-readable recording medium are disclosed. The operating method of the electronic device, according to one embodiment, may comprise an operation of obtaining a user utterance. The operating method may comprise the operations of: generating a first prompt on the basis of the user utterance; generating a second prompt on the basis of first data for invoking a function of an application related to the first prompt; and generating a natural language-based response to the user utterance by using a first generative model (555) to process the second prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for processing user speech, method of operation thereof, and recording medium

[0001] The disclosure below relates to an electronic device for processing user speech, a method of operating the same, and a computer-readable recording medium.

[0002] A generative model is an artificial intelligence model that can generate new data based on input data. Generative models can be used in user utterance processing systems, such as voice agents.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] An electronic device according to one embodiment may include at least one processor and a memory storing instructions. The instructions may be individually or collectively executed by the at least one processor, and may cause the electronic device to acquire a user utterance.

[0005] The instructions, which are individually or collectively executed by the at least one processor, may cause the electronic device to generate a first prompt based on the user utterance.

[0006] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second prompt based on the first data for invoking a function of an application associated with the first prompt.

[0007] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to process the second prompt using the first generative model to generate a natural language-based response to the user utterance.

[0008] The first data may be generated based on processing the first prompt using a second generative model.

[0009] A method of operating an electronic device (101) according to one embodiment may include an operation of obtaining a user utterance.

[0010] The above method of operation may include an operation of generating a first prompt based on the user utterance.

[0011] The above method of operation may include an operation of generating a second prompt based on first data for invoking a function of an application associated with the first prompt.

[0012] The above method of operation may include an operation of processing the second prompt using the first generative model (555) to generate a natural language-based response to the user utterance.

[0013] The above first data can be generated based on processing the first prompt using the second generative model (525).

[0014] According to one embodiment, a computer-readable recording medium storing one or more computer programs may include instructions for performing the method in a processor.

[0015] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0016] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0017] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0018] FIG. 3 is a diagram for explaining a user speech processing system according to one embodiment.

[0019] FIG. 4 is a diagram illustrating a final response provided by a user speech processing system according to one embodiment.

[0020] FIG. 5 is a diagram for explaining a user speech processing system according to one embodiment.

[0021] FIG. 6 is a diagram illustrating an encoder according to one embodiment.

[0022] Figure 7 is a drawing for explaining a prompt according to one embodiment.

[0023] Figure 8 is a drawing for explaining an action according to one embodiment.

[0024] FIGS. 9 and 10 are diagrams illustrating final responses provided by a user speech processing system according to one embodiment.

[0025] FIG. 11 and FIG. 12 are diagrams illustrating final responses provided by a user speech processing system according to one embodiment.

[0026] FIG. 13 is a diagram illustrating a final response provided by a user speech processing system according to one embodiment.

[0027] FIG. 14 is a flowchart illustrating operations performed by an electronic device according to one embodiment.

[0028] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0029]

[0030] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0031] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0032] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0033] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)) and input data or output data for commands related thereto. The memory (130) can include a volatile memory (132) or a non-volatile memory (134). According to one embodiment, instructions stored in the memory (130) can cause the electronic device (101) to perform one or more operations based on being individually or collectively executed by at least one processor (e.g., the main processor (121) and / or the auxiliary processor (123)). For example, instructions stored in the memory (130) may be executed by one processor (e.g., a main processor (121) or an auxiliary processor (123) such as a communication processor) or by multiple processors operating cooperatively (e.g., a main processor (121) and an auxiliary processor (123)).

[0034] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0035] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0036] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0037] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0038] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0039] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0040] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0041] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0042] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0043] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0044] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0045] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0046] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0047] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0048] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0049] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0050] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0051] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0052]

[0053] An electronic device according to an embodiment disclosed in this document may take various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to an embodiment of this document is not limited to the aforementioned devices.

[0054]

[0055] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0056] Referring to FIG. 2, according to one embodiment, the generative artificial intelligence system (200) may be a program (e.g., a software module) implemented on an electronic device (e.g., an electronic device (101) of FIG. 1) and / or a server (e.g., a server (108) of FIG. 1).

[0057] According to one embodiment, a user query / response interface (210) may receive user input. The user input may be an input of a type (or modality) such as natural language, image, audio, and / or video. Additionally, context information may also be transmitted when the user input is transmitted. The context information may include various side information related to the time when the user input is input to the artificial intelligence system (200). For example, there is side information such as application information currently being used by the user or location information of the user. Additionally, the user input may be an input of a mixed type of natural language, image, audio, video, and / or context information as described above. Additionally, the user input may include non-natural language input, such as selecting a menu.

[0058] According to one embodiment, a user query / response interface (210) may provide output from a generative artificial intelligence system to a user. The output may include a natural language-based response and / or specific content. The output may also include an action requested by the user.

[0059] In one embodiment, an AI framework (220) may receive user input. Based on the user input (e.g., a user's query), the AI ​​framework (220) may coordinate and / or control one or more components necessary to perform an action corresponding to the user's intent.

[0060] According to one embodiment, user input received from the user query / response interface (210) may be transmitted to a prompt design component (221). The prompt design component (221) may be used to generate a prompt suitable as input to a generative model (e.g., a large language model (LLM) and / or a large multimodal model (LMM)) based on the user input.

[0061] In one embodiment, the prompt design component (221) may be an AI component that utilizes a machine learning algorithm or a neural network. The prompt design component (221) may generate improved prompts over time through learning. The prompt design component (221) may access a knowledge repository (230) to generate prompts based on user input. The knowledge repository (230) may include user preference data, a prompt library, and / or prompt examples. The prompt design component (223) may provide the generated prompts to a generative model (e.g., an LLM and / or an LMM).

[0062] According to one embodiment, the APIs / Plugins management component (223) can communicate with an external information source based on a request for additional information when user input is transmitted to the generative model.

[0063] In one embodiment, the APIs / Plugins management component (223) can establish a communication channel for communication with the outside of the system (200) via the API. The APIs / Plugins management component (223) can enable access to various data sources via the communication channel. The acquired information can be used to generate prompts by the prompt design component (221) along with user input, or can be used as input to the generative model (250).

[0064] According to one embodiment, the APIs / Plugins management component (223) may request a final action via an API when the final action corresponding to user input, rather than an intermediate action, must be performed by an application or service.

[0065] In one embodiment, the refiner component (225) can fine-tune the output of the generative model (250). For example, the refiner component (225) can determine the relevance (e.g., a score) between the output (e.g., content) of the generative model and the user input. For example, the refiner component (225) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (225) can determine whether the output contains harmful information (e.g., violent content or profanity).

[0066] In one embodiment, the refinement component (225) may determine the degree of matching (e.g., a score) between the output of the generative model (250) and the user input (e.g., the intent of the user input). If the refinement component (225) determines that the output of the generative model (250) does not correspond to the user input, the refinement component (225) may modify the output to correspond to the user input.

[0067] In one embodiment, the refinement component (225) may provide hints to the user (e.g., hints for prompt generation) to enable the user to obtain information that matches the user's intent from the generative model (250).

[0068] According to one embodiment, a generative model (250) may refer to an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on user input (e.g., user utterance). The generative model (250) may include an image generation model and / or a language generation model.

[0069] In one embodiment, the image generation model may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). An example of an image generation model is a diffusion-based generative model having the structure of a VAE and a transformer.

[0070] In one embodiment, a language generation model (e.g., ChatGPT) may be a model trained to generate statistically most appropriate output based on input. The language generation model may include an LMM. The LMM can identify various types of input, such as text, images, audio (e.g., speech), and / or video, and generate new data corresponding to the input.

[0071]

[0072] FIG. 3 is a diagram for explaining a user speech processing system according to one embodiment.

[0073] Referring to FIG. 3, according to one embodiment, a user utterance processing system (300) may acquire a user utterance (30) (e.g., a speech utterance and / or a text utterance) and provide a final response (35) corresponding to the user utterance. For example, the user utterance processing system (300) may visually and / or audibly provide the user with a final response, "It is raining now," for the user utterance, "Tell me the current weather."

[0074] According to one embodiment, the user speech processing system (300) may include one or more components (310-345). The components (310-345) may be software modules implemented on an electronic device (e.g., the electronic device (101) of FIG. 1) or a server (e.g., the server (108) of FIG. 1). The components (310-345) are illustrated as an example for describing the user speech processing system (300). Accordingly, the user speech processing system (300) may include various variations of the components (310-345) as long as the operations of the user speech processing system (300) described in the present disclosure can be implemented. For example, two or more components may be combined, or one or more components may be added or omitted. Alternatively, the user speech processing system (300) may further include one or more components (e.g., components (210 to 250) of FIG. 2) of a generative artificial intelligence system (e.g., generative artificial intelligence system (200) of FIG. 2).

[0075] According to one embodiment, the user speech processing system (300) may include an automatic speech recognition module (ASR) (310), a preprocessor (315), a prompt generation module (320), a generative model (325) (e.g., the generative model (250) of FIG. 2), a first application (330), a postprocessor (335), a capsule (340), and a second application (345).

[0076] According to one embodiment, the ASR module (310) can convert a speech signal (e.g., an analog signal) into text data. When the user's speech is a text speech, the operation of the ASR module (310) can be omitted.

[0077] In one embodiment, the preprocessor (315) may perform preprocessing on text data. For example, the preprocessor (315) may perform tokenization, which splits the text into tokens (e.g., words, phrases, and / or symbols). For another example, the preprocessor (315) may remove unnecessary characters (e.g., characters such as '!' or '?') to generate a final response (35) from the text data.

[0078] In one embodiment, the prompt generation module (320) can generate a prompt. The prompt can be data (e.g., text, JavaScript object notation (JSON), image, audio, and / or video) provided to the generative model (325) to generate new data (e.g., a response). The prompt generation module (320) can generate a prompt that the generative model (325) can understand by using text data corresponding to the user utterance along with conditions (e.g., length or writing style) related to the final response (35).

[0079] In one embodiment, the generative model (325) can generate a natural language-based response (e.g., text such as "Today's weather is sunny") corresponding to a user utterance (e.g., "How's the weather today?"). To generate the natural language-based response, the generative model (325) can obtain necessary information from a first application (330) (e.g., a web page providing weather information).

[0080] In one embodiment, the post-processor (335) may perform post-processing on the natural language-based response generated by the generative model (325). For example, the post-processor (335) may add characters (e.g., emoticons) or remove unnecessary words or sentences.

[0081] According to one embodiment, the generative model (325) can generate data for performing an action corresponding to a user utterance. In the present disclosure, an action may refer to a specific task performed by an application (e.g., an application, a widget, a domain, or a web page) other than the generative model using a function (e.g., an application programming interface (API) or a deep link). For example, the generative model (325) can generate data for invoking a function of a second application (345) (e.g., a weather application) related to a user utterance (e.g., “How is the weather today?”). The user utterance processing system (300) can obtain data corresponding to the user utterance (e.g., image data such as a weather card) from the second application (345) using the data generated by the generative model (325).

[0082] According to one embodiment, the capsule (340) may include files and / or data necessary for interaction between the user speech processing system (300) and the second application (345).

[0083] According to one embodiment, the user speech processing system (300) may provide a final response (35) including data acquired through an action (e.g., image data such as a 'weather card' acquired through a second application (345)) and a natural language-based response.

[0084]

[0085] FIG. 4 is a diagram illustrating a final response provided by a user speech processing system according to one embodiment.

[0086] Referring to FIG. 4, according to one embodiment, the electronic device (101) may provide a final response (35) to a user utterance using the user utterance processing system (300). For example, when the electronic device (101) detects the user utterance “Tell me the weather today,” the electronic device (101) may generate a final response (35) using the user utterance processing system (300). The final response (35) may include an action-based response (42) obtained through an action and an output (44) (e.g., a natural language-based response) of a generative model (e.g., the generative model (325) of FIG. 3).

[0087] In one embodiment, the user utterance processing system (300) generates data for an action (e.g., data for interaction with the second application (345) of FIG. 3) and a natural language-based response in parallel using a single generative model (e.g., the generative model (325) of FIG. 3), so that the action-based response (42) may not match the natural language-based response (44). For example, the user's current location for the first application (330) may be set to a specific city in the United States, and the user's current location for the second application (345) may be set to a specific city in Korea. In this case, the action-based response (42) corresponding to the user utterance "Tell me the weather today" may provide weather information for a specific city in Korea (e.g., Gangnam-gu, Seoul), whereas the natural language-based response (44) may provide weather information for a specific city in the United States (e.g., Columbus, Ohio). The electronic device (101) can process user speech using a user speech processing system (e.g., the user speech processing system (500) of FIG. 5) that can prevent inconsistencies between information included in the final response (35).

[0088]

[0089] FIG. 5 is a diagram for explaining a user speech processing system according to one embodiment.

[0090] Referring to FIG. 5, according to one embodiment, a user utterance processing system (500) may acquire a user utterance (50) (e.g., a speech utterance and / or a text utterance) and provide a final response (55) corresponding to the user utterance. For example, the user utterance processing system (500) may visually and / or audibly provide the user with a final response, "It is raining now," to the user utterance, "Tell me the current weather."

[0091] According to one embodiment, the user speech processing system (500) may include one or more components (510-560). The components (510-560) may be software modules implemented on an electronic device (e.g., the electronic device (101) of FIG. 1) or a server (e.g., the server (108) of FIG. 1). The components (510-560) are illustrated as an example for describing the user speech processing system (500). Accordingly, the user speech processing system (500) may include various variations of the components (510-560) as long as the operations of the user speech processing system (500) described in the present disclosure can be implemented. For example, two or more components may be combined, or one or more components may be added or omitted. Alternatively, the user speech processing system (500) may further include one or more components (e.g., components (210 to 250) of FIG. 2) of a generative artificial intelligence system (e.g., generative artificial intelligence system (200) of FIG. 2).

[0092] According to one embodiment, the user speech processing system (500) may include an ASR module (510) (e.g., the ASR module (310) of FIG. 3), a preprocessor (515) (e.g., the preprocessor (315) of FIG. 3), a first prompt generation module (520), a first generative model (525) (e.g., the generative model (250) of FIG. 2), a postprocessor (530) (e.g., the postprocessor (335) of FIG. 3), a capsule (535) (e.g., the capsule (340) of FIG. 3), an encoder (545), a second prompt generation module (550), a second generative model (555) (e.g., the generative model (250) of FIG. 2), and a verifier (560).

[0093] According to one embodiment, the ASR module (510) can convert a speech signal (e.g., an analog signal) into text data. When the user's speech is a text speech, the operation of the ASR module (510) can be omitted.

[0094] In one embodiment, the preprocessor (515) may perform preprocessing on text data. For example, the preprocessor (515) may perform tokenization, which splits the text into tokens (e.g., words, phrases, and / or symbols). For example, the preprocessor (515) may remove unnecessary characters (e.g., characters such as '!' or '?') to generate a final response (55) from the text.

[0095] According to one embodiment, the first prompt generation module (520) can generate a prompt for the first generative model (525). The first prompt generation module (520) can generate the prompt for the first generative model (525) based on text corresponding to the user utterance (50) (e.g., 'Tell me the current weather') and information related to a function of the application (540) corresponding to the user utterance (50). The information related to the function of the application (540) can include a function identifier (e.g., a function name), a function description, and / or a function instruction. The function identifier can provide information for identifying the function. The function description can provide information about the purpose and / or behavior of the function so that the first generative model (525) can select a function necessary for processing the user utterance (50). Function instructions may provide information about when to call a function, how to call a function, and / or whether two or more functions are linked together.

[0096] According to one embodiment, the first prompt generation module (520) may send constraints on the format of text (e.g., 'Tell me the current weather'), a function identifier (e.g., a function name), and / or an output (e.g., a natural language-based response) of the second generative model (555) corresponding to the user utterance (50) to the second prompt generation module (550). The second prompt generation module (550) may generate a prompt of the second generative model (555) using information and / or data received from the first prompt generation module (520).

[0097] According to one embodiment, the first generative model (525) can generate data necessary to perform an action corresponding to a user utterance (50). For example, the first generative model (525) can generate data for calling a function of an application (540) related to a prompt generated by the first prompt generation module (520). For example, when the user utterance is an utterance for obtaining information about the weather, such as “Tell me the current weather,” the first generative model (525) can generate data for calling a function necessary to obtain weather information from a weather application.

[0098] In one embodiment, when multiple functions are required to perform an action related to a user utterance, the first generative model (525) can generate data to correlate the multiple functions.

[0099] According to one embodiment, the first generative model (525) can perform semantic inference based on information included in the prompt (e.g., a function description).

[0100] In one embodiment, the first generative model (525) may be a large generative model. The first generative model (525) may be a cloud-based model implemented on a server (e.g., server (108) of FIG. 1), but is not limited thereto. For example, the first generative model (525) may be an on-device model implemented on an electronic device (e.g., electronic device (101) of FIG. 1). The user utterance processing system (500) may use data (e.g., data for an action) generated by the first generative model (525) to obtain various types (or data modalities) of data (e.g., text data, image data, audio data, and / or video data) corresponding to the user utterance through the application (540). For example, the user speech processing system (500) can obtain image data, such as a weather card, corresponding to the user speech “Tell me the current weather” from a weather application.

[0101] According to one embodiment, the post-processor (530) can perform post-processing on data generated by the first generative model (525).

[0102] According to one embodiment, the capsule (535) may include files and / or data necessary for interaction between the user speech processing system (500) and the application (540). The capsule (535) may transmit data obtained through the application (540) to the encoder (545) and / or the verifier (560). For example, the capsule (535) may transmit a weather card obtained from a weather application to the encoder (545) and / or the verifier (560).

[0103] According to one embodiment, the encoder (545) can generate new data from data acquired through the application (540). For example, the encoder (545) can generate text data (e.g., output data) that describes or explains the image data (e.g., input data) based on features of the image data. The operation of the encoder (545) will be described in detail with reference to FIG. 6.

[0104] In one embodiment, the encoder (545) may be configured to generate new types of output data based on input data. For example, the encoder (545) may be a neural network-based encoder (e.g., a trained neural network-based encoder).

[0105] According to one embodiment, the second prompt generation module (550) can generate a prompt of the second generative model (555) based on data generated by the encoder (545) and data received from the first prompt generation module (520) (e.g., function name, text corresponding to the user utterance (50), and / or constraints on the format of the output of the second generative model (555)).

[0106] In one embodiment, the second generative model (555) can generate new data (e.g., text data, image data, audio data, and / or data) based on a prompt generated by the second prompt generation module (550). Since the second generative model (555) generates new data based on a prompt that includes information about data obtained from the application (540) (e.g., image data such as a weather card), the second generative model (555) can generate output (e.g., a natural language-based response) that matches the data obtained from the application (540).

[0107] In one embodiment, the second generative model (555) may be a generative model that is smaller than the first generative model (525), but is not limited thereto. For example, although FIG. 5 illustrates the first generative model (525) and the second generative model (555) as two different models, the second generative model (555) and the first generative model (525) may be implemented as a single generative model (not shown). For example, the first generative model (525) and the second generative model (555) may be different sub-models included in a single generative model (not shown). In this case, the first generative model (525) may be a sub-model that is larger than the second generative model (555). For another example, a generative model (not shown) may perform one or more of the operations of a first generative model (525) and the operations of a second generative model (555) based on input.

[0108] In one embodiment, the second generative model (555) may be, but is not limited to, an on-device model implemented on an electronic device (e.g., the electronic device (101) of FIG. 1 ). For example, the first generative model (525) may be a cloud-based model implemented on a server (e.g., the server (108) of FIG. 1 ).

[0109] In one embodiment, the verifier (560) can determine the relevance between the output of the second generative model (555) (e.g., a natural language-based response) and data acquired through the application (540). For example, the verifier (560) can calculate a score indicating the degree of matching based on the characteristics of each of the two data sets, and determine that the two data sets match when the calculated score satisfies a threshold value.

[0110] According to one embodiment, the verifier (560) can provide a final response (55) (e.g., the final response (1000) of FIG. 10, the final response (1200) of FIG. 12, or the final response (1300) of FIG. 13) to the user when the output (e.g., a natural language-based response) of the second generative model (555) and the data acquired through the application (540) match each other.

[0111] In one embodiment, the verifier (560) may provide feedback to the second generative model (555) notifying the second generative model (555) of a discrepancy between the output of the second generative model (555) (e.g., a natural language-based response) and data acquired through the application (540). In response to receiving feedback from the verifier (560), the second generative model (555) may regenerate data (e.g., a natural language-based response). To improve the performance of the second generative model (555), reinforcement learning such as reinforcement learning with human feedback (RLHF) or feedback techniques such as human-in-the-loop (HITL) may be used.

[0112] FIG. 6 is a diagram illustrating an encoder according to one embodiment.

[0113] Referring to FIG. 6, according to one embodiment, the encoder (545) can synthesize or generate new data based on data obtained through an application (e.g., application (540) of FIG. 5) related to a user utterance (e.g., user utterance (50) of FIG. 5).

[0114] According to one embodiment, the encoder (545) can generate a new type of output data containing information about a specific type of input data. For convenience of explanation, FIG. 6 illustrates an example of the encoder (545) generating text data from image data, but the scope of the present disclosure is not limited thereto.

[0115] According to one embodiment, the encoder (545) may generate text data (64) describing or explaining the image data (62) based on image data (62) (e.g., a weather card) obtained through a weather application related to the user utterance “Tell me the current weather.” The data (64) generated by the encoder (545) may be transmitted to a second prompt generation module (e.g., the second prompt generation module (550) of FIG. 5).

[0116]

[0117] Figure 7 is a drawing for explaining a prompt according to one embodiment.

[0118] Referring to FIG. 7, according to one embodiment, the first prompt generation module (520) may generate a prompt of the first generative model (e.g., the first generative model (525) of FIG. 5) based on a user utterance (e.g., the user utterance (50) of FIG. 5) or may transmit information to the second prompt generation module (550). For example, when the user utterance (50) is “Tell me the current weather” or “How strong is the wind today,” the first prompt generation module (520) may include information (77) about the utterance.

[0119] According to one embodiment, the first prompt generation module (520) may include a system prompt (71). The system prompt (71) may be set by a user. The system prompt (71) may provide information about basic constraints of a speech processing system (e.g., the speech processing system (500) of FIG. 5). For example, the system prompt (71) may include constraints on information (e.g., information on a specific topic), constraints on expression (e.g., profanity), and / or a persona of a generative model included in the speech processing system (500) (e.g., the first generative model (525) and / or the second generative model (555) of FIG. 5).

[0120] According to one embodiment, the first prompt generation module (520) may include information (73) related to a function of an application (e.g., application (540) of FIG. 5). For example, the information (73) may include a function identifier such as a function name, a function description, and / or function instructions. The function identifier, the function description, and the function instructions may be used to generate a prompt of the first generative model (525). The function identifier may be passed to a second prompt generation module (e.g., second prompt generation module (550) of FIG. 5).

[0121] In one embodiment, the first prompt generation module (520) may include information (75) (e.g., constraints) regarding the format of the output of the second generative model (555). For example, if the second generative model (555) generates a natural language-based response, the information (75) may include constraints that cause the second generative model (555) to generate a natural language-based response of two or fewer sentences.

[0122]

[0123] FIG. 8 is a diagram for explaining data acquired through an action according to one embodiment.

[0124] Referring to FIG. 8, according to one embodiment, the first prompt generation module (520) can generate a prompt of the first generative model (e.g., the generative model (525) of FIG. 5) based on a user utterance (e.g., the user utterance (50) of FIG. 5). For example, since the user utterance “Tell me the current weather” and the user utterance “How strong is the wind today” are both utterances for obtaining information about the weather, the first prompt generation module (520) can generate a prompt using information about a function of a weather application (e.g., get_Weather).

[0125] According to one embodiment, the first generative model (525) may generate data for calling a function of an application (e.g., application (540) of FIG. 5) based on a prompt. A user speech processing system (e.g., user speech processing system (500) of FIG. 5) may use the generated data to call a function and obtain data from the application (540). For example, the user speech processing system (500) may obtain image data (80), such as a weather card, from a weather application.

[0126] In Figure 8, the same weather card is shown to be obtained for two user utterances, “Tell me the current weather” and “How strong is the wind today”, but this is an example for illustrative purposes only.

[0127]

[0128] FIGS. 9 and 10 are diagrams illustrating a final response provided by a user speech processing system according to one embodiment. FIGS. 9 and 10 may be diagrams illustrating an operation of generating a final response (1100) corresponding to the user speech "Tell me the current weather."

[0129] Referring to FIGS. 9 and 10 , according to one embodiment, the electronic device (101) may generate a final response (1000) corresponding to a user utterance using a user utterance processing system (e.g., the user utterance processing system (500) of FIG. 5 ). The final response (1000) may include data (80) (e.g., image data such as a weather card) obtained through a weather application (e.g., the application (540) of FIG. 5) related to the user utterance and an output (82) (e.g., a natural language-based response) of a second generative model (555).

[0130] According to one embodiment, the second generative model (555) can generate a natural language-based response corresponding to a user utterance, such as "It's cloudy in New York. The temperature is 21 degrees, which is slightly hot," based on the prompt (90). The prompt (90) can be generated based on information received from the first prompt generation module (520) and data (80) obtained through an action (e.g., image data such as a weather card), as described above. Duplicate descriptions will be omitted.

[0131] According to one embodiment, the electronic device (101) can obtain data (80) from an application (540) using a first generative model (e.g., the first generative model (525) of FIG. 5) and generate a prompt (90) of a second generative model (555) using the data (80), thereby providing a final response (1000) containing consistent information to the user.

[0132]

[0133] FIGS. 11 and 12 are diagrams illustrating a final response provided by a user speech processing system according to one embodiment. FIGS. 11 and 12 may be diagrams illustrating an operation for generating a final response (1200) corresponding to the user speech "How strong is the wind today?"

[0134] Referring to FIGS. 11 and 12 , according to one embodiment, the electronic device (101) may generate a final response (1200) corresponding to a user utterance using a user utterance processing system (e.g., the user utterance processing system (500) of FIG. 5 ). The final response (1200) may include data (80) (e.g., image data such as a weather card) obtained through a weather application (e.g., the application (540) of FIG. 5) corresponding to the user utterance and an output (84) (e.g., a natural language-based response) of a second generative model (555).

[0135] According to one embodiment, the second generative model (555) can generate a natural language-based response corresponding to a user utterance based on a prompt (92), such as "The wind speed in New York is 9 km / h, so it is relatively windy." The prompt (92) can be generated based on information received from the first prompt generation module (520) and data (80) (e.g., a weather card) obtained through an action, as described above. Duplicate descriptions will be omitted.

[0136] According to one embodiment, the electronic device (101) can obtain data (80) from an application (540) using a first generative model (e.g., the first generative model (525) of FIG. 5) and generate a prompt (92) of a second generative model (555) using the data (80), thereby providing a final response (1200) containing consistent information to the user.

[0137]

[0138] FIG. 13 is a diagram illustrating a final response provided by a user utterance processing system according to one embodiment. FIG. 13 may be a diagram illustrating an operation for generating a final response (1300) corresponding to the user utterance "Tell me my schedule for today."

[0139] Referring to FIG. 13, according to one embodiment, the electronic device (101) may generate a final response (1300) corresponding to a user utterance by using a user utterance processing system (e.g., the user utterance processing system (500) of FIG. 5).

[0140] According to one embodiment, the electronic device (101) may generate a prompt of a second generative model (e.g., the second generative model (555) of FIG. 5) using data (e.g., image data or structured data including the user's actual schedule information) acquired through a calendar application (e.g., the application (540) of FIG. 4) corresponding to a user's utterance. The electronic device (101) may provide the user with a final response (1300) reflecting the user's actual schedule information through the second generative model (555).

[0141] According to one embodiment, the electronic device (101) can prevent hallucination by sequentially generating data using a first generative model (e.g., the first generative model (525) of FIG. 5) and a second generative model (555). Hallucination can refer to a phenomenon in which a generative model generates fictitious data that is not based on facts.

[0142]

[0143] FIG. 14 is a flowchart illustrating operations performed by an electronic device according to one embodiment.

[0144] Referring to FIG. 14, according to one embodiment, operations 1410 to 1440 may be performed sequentially, but are not limited thereto. For example, two or more operations may be performed in parallel. Operations 1410 to 1440 may be substantially the same as operations performed by the electronic device described with reference to FIGS. 1 and 5 to 13 (e.g., the electronic device (101) of FIGS. 1, 10, 12, and 13). Therefore, any redundant description will be omitted.

[0145] In operation 1410, the electronic device (101) may obtain a user utterance (e.g., user utterance (50) of FIG. 5).

[0146] At operation 1420, the electronic device (101) may generate a first prompt based on a user utterance (50).

[0147] In operation 1430, the electronic device (101) may generate a second prompt based on first data for calling a function of an application (e.g., application (540) of FIG. 5) associated with the first prompt. The first data may be generated based on processing the first prompt using a second generative model (e.g., first generative model (525) of FIG. 5).

[0148] In operation 1440, the electronic device (101) may generate a natural language-based response to the user utterance based on processing the second prompt using a second generative model (e.g., the second generative model (555) of FIG. 5).

[0149]

[0150] An electronic device (101) according to one embodiment may include at least one processor (120) and a memory (130) that stores instructions.

[0151] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain a user utterance.

[0152] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a first prompt based on the user utterance.

[0153] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a second prompt based on first data for invoking a function of an application (540) associated with the first prompt.

[0154] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to process the second prompt using the first generative model (555) to generate a response corresponding to the user utterance.

[0155] The above first data can be generated based on processing the first prompt using the second generative model (525).

[0156] The first prompt may include at least one of an identifier of the function and a purpose of the function.

[0157] The above first generation model (555) may be different from the above second generation model (525).

[0158] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain second data including information corresponding to the first prompt via the application (540) based on the first data.

[0159] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate the second prompt based on the second data.

[0160] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate third data of a data modality different from a data modality of the second data based on features of the second data.

[0161] The instructions, which are individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate the second prompt based on constraints on the format of the third data and the natural language-based response.

[0162] The second data may include at least one of image data and audio data.

[0163] The above third data may include text data.

[0164] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate the third data using a learned neural network-based encoder.

[0165] The instructions, which are individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to compare the output of the first generative model (555) generated based on the second data and the second prompt.

[0166] The above instructions, based on being individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to provide a final response including the natural language-based response and the second data to the user based on determining that the natural language-based response and the second data match each other.

[0167] A method of operating an electronic device (101) according to one embodiment may include an operation of obtaining a user utterance.

[0168] The above method of operation may include an operation of generating a first prompt based on the user utterance.

[0169] The above method of operation may include an operation of generating a second prompt based on first data for invoking a function of an application (540) related to the first prompt.

[0170] The above method of operation may include an operation of generating a natural language-based response to the user utterance by processing the second prompt using the first generative model (555).

[0171] The above first data can be generated based on processing the first prompt using the second generative model (525).

[0172] The first prompt may include at least one of an identifier of the function and a purpose of the function.

[0173] The above first generation model (555) may be different from the above second generation model (525).

[0174] The operation of generating the second prompt may include an operation of obtaining second data including information corresponding to the first prompt via the application (540) based on the first data.

[0175] The act of generating the second prompt may include an act of generating the second prompt based on the second data.

[0176] The operation of generating the second prompt based on the second data may include an operation of generating third data of a data modality different from a data modality of the second data based on features of the second data.

[0177] The operation of generating the second prompt based on the second data may include an operation of generating the second prompt based on constraints on the format of the third data and the output of the first generative model (555).

[0178] The second data may include at least one of image data and audio data.

[0179] The above third data may include text data.

[0180] The operation of generating the third data may include an operation of generating the third data using a neural network-based encoder.

[0181] The above operating method may further include an operation of comparing an output of the first generative model (555) generated based on the second data and the second prompt.

[0182] The above method of operation may include an operation of providing a final response including the natural language-based response and the second data to the user based on determining that the natural language-based response and the second data match each other.

[0183]

[0184] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from this document.

[0185] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish the corresponding component from other corresponding components and do not limit the corresponding components in any other respect (e.g., importance or order). When a component (e.g., a first) is referred to as "coupled" or "connected" to another (e.g., a second) component, with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0186] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0187] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0188] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0189] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101), At least one processor (120); and Memory for storing instructions (130) Including, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Obtain user utterance, Generate a first prompt based on the user utterance; Generate a second prompt based on first data for invoking a function of an application (540) related to the first prompt, By processing the second prompt using the first generative model (555), a natural language-based response to the user utterance is generated, The above first data is, An electronic device (101) generated based on processing the first prompt using a second generative model (525).

2. In paragraph 1, The first prompt above is, At least one of the identifier of the above function and the purpose of the above function An electronic device (101) comprising:

3. In any one of paragraphs 1 and 2, The above first generation model (555) is An electronic device (101) different from the above second generation model (525).

4. In any one of paragraphs 1 to 3, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Obtaining second data including information corresponding to the first prompt through the application (540) based on the first data, An electronic device (101) that generates the second prompt based on the second data.

5. In any one of paragraphs 1 to 4, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: Based on the features of the second data, third data having a data modality different from that of the second data is generated, An electronic device (101) that generates the second prompt based on constraints on the format of the third data and the natural language-based response.

6. In any one of paragraphs 1 to 5, The above second data is, At least one of image data and audio data Including, The above third data is, Text data An electronic device (101) comprising:

7. In any one of paragraphs 1 to 6, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: An electronic device (101) that generates the third data using a neural network-based encoder (545).

8. In any one of paragraphs 1 to 7, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: An electronic device (101) that allows comparison between the second data and the natural language-based response.

9. In any one of paragraphs 1 to 8, The above instructions are individually or collectively executed by the at least one processor (120), and cause the electronic device (101) to: An electronic device (101) that provides a final response including the natural language-based response and the second data to a user based on determining that the natural language-based response and the second data match each other.

10. In the operating method of the electronic device (101) according to the embodiment, The act of obtaining user utterance; An action to generate a first prompt based on the user utterance; An operation of generating a second prompt based on first data for invoking a function of an application (540) related to the first prompt; and An operation of processing the second prompt using the first generative model (555) to generate a natural language-based response to the user utterance. Including, The above first data is, A method generated based on processing the first prompt using a second generative model (525).

11. In paragraph 10, The first prompt above is, At least one of the identifier of the above function and the purpose of the above function A method comprising:

12. In any one of paragraphs 10 to 11, The action of generating the above second prompt is: An operation of obtaining second data including information corresponding to the first prompt through the application (540) based on the first data; An operation of generating third data of a data modality different from the data modality of the second data based on features of the second data; and An operation of generating the second prompt based on constraints on the format of the third data and the output of the first generative model (555). A method comprising:

13. In any one of paragraphs 10 to 12, The above second data is, At least one of image data and audio data Including, The above third data is, Text data A method comprising:

14. In any one of paragraphs 10 to 13, An operation of comparing the output of the first generative model (555) generated based on the second data and the second prompt. A method further comprising:

15. In any one of paragraphs 10 to 14, An operation of providing a final response including the natural language-based response and the second data to the user based on determining that the natural language-based response and the second data match each other. A method further comprising:

Citation Information

Patent Citations

  • Solution plan generation method and device, equipment and storage medium

    CN117407514A

  • Method, program and information processing device for collecting character information printed on printed matter

    JP7430437B1

  • Coolant system for vehicle

    KR1020250114985A

  • Systems, devices, methods and programs for controlling security using multiple artificial intelligence models

    KR102596977B1

  • System and method with entity type clarification for fine-grained factual knowledge retrieval

    US20230316001A1