Electronic device and method for providing voice service via electronic device

By integrating cameras, microphones, sensors, and AI applications, the electronic device addresses inefficiencies in AI-based voice services by accurately recognizing objects and managing conversation sessions, resulting in enhanced user interaction and efficient response generation.

WO2026005483A1PCT designated stage Publication Date: 2026-01-02SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008908
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing electronic devices struggle to effectively utilize large language models for providing AI-based voice services by accurately recognizing objects and managing conversation sessions based on user inputs, leading to inefficiencies in response generation and output.

Method used

The electronic device employs a camera, microphone, sensors, and processors to recognize external objects, create customized conversation sessions, determine audio paths, and generate responses using AI applications, including large and small language models, to provide spatial sound effects and manage multiple conversation sessions.

Benefits of technology

This approach enables precise object recognition, efficient conversation session management, and effective voice service delivery, enhancing user interaction through spatial audio and multimodal inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008908_02012026_PF_FP_ABST
    Figure KR2025008908_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment of the present disclosure comprises: a camera; a microphone; at least one sensor; a communication circuit; a memory that stores at least one AI application that provides an AI-based voice service; and at least one processor, wherein the memory may store instructions that, when executed by the at least one processor, cause the electronic device to: acquire, by using the camera, an image including at least one external object; receive a user input specifying a first object among the at least one external object; recognize information related to the first object by using at least one of the image, the at least one sensor, or the communication circuit; generate a first conversation session for the first object through the at least one AI application; determine an audio output path of the first conversation session on the basis of information related to the electronic device or the information related to the first object; receive a first voice input of a user related to the AI-based voice service through the microphone; generate a response to the first voice input on the basis of the first conversation session through the at least one AI application; and provide the response to the first voice input through the audio output path of the first conversation session. Various other embodiments understood through the specification are also possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and methods for providing voice services through electronic devices

[0001] The embodiments disclosed in this document relate to a technology for providing artificial intelligence-based voice services.

[0002] Electronic devices can perform actions in response to user voice commands through speech recognition and natural language processing. Recently, language models that process user speech (or text corresponding to user speech) based on artificial intelligence have become widespread. For example, large language models (LLMs) are language models comprised of artificial neural networks with numerous parameters, and various large language models are being developed. For example, AI-based voice services can utilize language models to generate responses corresponding to user voice input and provide them to the user.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] An electronic device according to an embodiment disclosed in the present document includes a camera, a microphone, at least one sensor, a communication circuit, a memory storing at least one AI application providing an artificial intelligence (AI)-based voice service, and at least one processor, wherein the memory, when executed by the at least one processor, causes the electronic device to obtain an image including at least one external object using the camera, receive a user input specifying a first object among the at least one external object, recognize information related to the first object using at least one of the image, the at least one sensor, or the communication circuit, create a first conversation session for the first object through the at least one AI application, determine an audio output path of the first conversation session based on information related to the electronic device or information related to the first object, receive a first voice input of a user related to the AI-based voice service through the microphone, create a response to the first voice input based on the first conversation session through the at least one AI application, and provide the response to the first voice input through the audio output path of the first conversation session. Instructions can be stored.

[0005] In addition, a method according to an embodiment disclosed in the present document may include an operation of acquiring an image including at least one external object using a camera of an electronic device, an operation of receiving a user input specifying a first object among the at least one external object, an operation of recognizing information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device, an operation of generating a first conversation session for the first object through at least one AI application that provides an artificial intelligence-based voice service included in the electronic device, an operation of determining an audio output path of the first conversation session based on information related to the electronic device or information related to the first object, an operation of receiving a first voice input of a user related to the AI-based voice service through the microphone, an operation of generating a response to the first voice input based on the first conversation session through the at least one AI application, and an operation of providing a response to the first voice input through the audio output path of the first conversation session.

[0006] In addition, a storage medium according to an embodiment disclosed in the present document may store instructions and / or a program that, when executed by at least one processor of an electronic device, causes the electronic device to obtain an image including at least one external object using a camera of the electronic device, receive a user input specifying a first object among the at least one external object, recognize information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device, create a first conversation session for the first object through at least one AI application that provides an artificial intelligence-based voice service included in the electronic device, determine an audio output path of the first conversation session based on information related to the electronic device or information related to the first object, receive a first voice input of a user related to the AI-based voice service through the microphone, create a response to the first voice input based on the first conversation session through the at least one AI application, and provide the response to the first voice input through the audio output path of the first conversation session.

[0007] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0008] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0009] FIG. 3 is a diagram illustrating an operation of recognizing an object and creating a conversation session for an object in an electronic device according to one embodiment.

[0010] FIG. 4 illustrates examples of an electronic device providing a voice service according to one embodiment.

[0011] FIGS. 5A and 5B are diagrams illustrating operations for managing a conversation session of an electronic device according to one embodiment.

[0012] Figure 6 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0013] Fig. 7 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0014] Fig. 8 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0015] Fig. 9 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0016] Fig. 10 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0017] FIG. 11 illustrates an electronic device within a network environment according to various embodiments.

[0018] FIG. 12 is a block diagram illustrating an integrated intelligence system according to one embodiment.

[0019] FIG. 13 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.

[0020] FIG. 14 is a diagram illustrating a user terminal displaying a screen for processing voice input received through an intelligent app, according to one embodiment.

[0021] FIG. 15 illustrates a generative artificial intelligence system according to one embodiment.

[0022] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0023] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0024] According to one embodiment, an electronic device (100) (e.g., an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIGS. 12 to 14, or an artificial intelligence system (1500) of FIG. 15) includes a camera (110) (e.g., a camera (210) of FIG. 2 or a camera module (1180) of FIG. 11), a microphone (120) (e.g., a microphone (220) of FIG. 2, an audio module (1170) of FIG. 11, or a microphone (1270) of FIG. 12), at least one sensor (130), a communication circuit (140) (e.g., a communication circuit (240) of FIG. 2, a communication module (1190) of FIG. 11, or a communication interface (1290) of FIG. 12), a memory (150) (e.g., an object storage module (233) of FIG. 2, a memory (1130) of FIG. 11, or a 12 memory (1230)), and at least one processor (160) (e.g., processor (1120) of FIG. 11 or processor (1220) of FIG. 12).

[0025] According to one embodiment, the camera (110) can acquire an image including at least one object external to the electronic device (100). For example, the at least one object can include an electronic device (e.g., a charger, a speaker, a vehicle, or a navigation device, etc.) and can include a non-electronic device (e.g., a doll, etc.).

[0026] In one embodiment, the microphone (120) can receive a user's voice input (e.g., a user's speech). The microphone (120) can generate voice data corresponding to the received voice input.

[0027] According to one embodiment, at least one sensor (130) can detect the distance, direction, position, or shape of an object, the position of the electronic device (100), the direction of the user's gaze, the direction of the user's head, or movement. The at least one sensor (130) can include, but is not limited to, a position sensor (e.g., GPS), a proximity sensor, an inertial sensor (e.g., an accelerometer sensor, a gyro sensor, a geomagnetic sensor), an image sensor, an infrared sensor, and / or an ultrasonic sensor.

[0028] According to one embodiment, the communication circuit (140) can transmit and receive information and / or data with at least one external electronic device (100) (for example, if the external object is the electronic device (100), the external object can be included). The communication circuit (140) can connect the electronic device (100) and the at least one external electronic device (100) through a designated communication method (for example, a communication protocol). The communication circuit (140) can detect at least one external electronic device (100) around the electronic device (100).

[0029] According to one embodiment, the memory (150) may store instructions that control the operation of the electronic device (100) when executed individually or collectively by at least one processor (160). According to one embodiment, the instructions may be stored in one memory (150) or in a plurality of memories (150). The memory (150) may at least temporarily store information and / or data related to operations of the electronic device (100). For example, the memory (150) may at least temporarily store images acquired through the camera (110), data measured through the sensor (130), information related to recognized objects, information related to the electronic device (100), information related to an external electronic device (100), information related to user input, properties of a conversation session (hereinafter, referred to as a 'customized conversation session' or a 'customized AI conversation session' in the present disclosure), audio input / output paths of the conversation session, and / or input and response contents for each conversation session. According to various embodiments, the information and / or data stored in the memory (150) is not limited to those listed above.

[0030] According to one embodiment, the memory (150) may include at least one artificial intelligence (AI) application (151) (e.g., AI application (251) of FIG. 2, application (1146) of FIG. 11, or apps (1235ㅁ, 1235b) of FIG. 12). For example, the AI ​​application (151) may support an AI-based voice service. For example, the AI-based voice service may include a service that provides a response to a user's voice input (e.g., a prompt) by using at least one language model (e.g., a large language model (LLM) and / or a small large language model (sLLM)). The at least one AI application (151) may include an application (151) that supports an AI-based voice service. For example, the at least one AI application (151) may include an application (151) that provides a response to a user's voice input (e.g., a user utterance) by using a language model. For example, the language model may be included in an AI application (151), an electronic device (100), or an external server (e.g., an AI server).

[0031] According to one embodiment, the processor (160) can control the operation of the electronic device (100) by individually or collectively executing instructions stored in the memory (150). For example, operations described as being performed by the 'processor (160)' in the present disclosure can be understood as being performed by at least one processor (160) individually or collectively. For example, at least one processor (160) can independently or collectively control each of the operations of the electronic device (100) described below. According to one embodiment, at least one processor (160) can include a circuit such as a central processing unit (CPU), a microprocessor unit (MPU), an application processor (AP), a communication processor (CP), a system on chip (SoC), and / or an integrated circuit (IC).

[0032] According to one embodiment, the processor (160) can obtain an image including at least one external object using the camera (110). According to one embodiment, the processor (160) can recognize at least one external object included in the image.

[0033] According to one embodiment, the processor (160) may receive a user input specifying a first object among at least one external object. For example, the processor (160) may receive a user input selecting a first object among at least one external object included in an image. For example, the user input selecting the first object may include a touch input, a gesture input, and / or a voice input. According to one embodiment, the processor (160) may specify one object (e.g., a first object) or multiple objects (e.g., a first object and a second object) based on the user input.

[0034] In one embodiment, the processor (160) may receive user input specifying a name of a first object. In one embodiment, the processor (160) may register the name of the first object as a call word for a first conversation session for the first object.

[0035] According to one embodiment, the processor (160) may recognize information related to the first object using at least one of an image, at least one sensor (130), or a communication circuit (140). According to one embodiment, the information related to the first object may include at least one of the type of the first object, identification information of the first object, the location of the first object, the shape of the first object, or the name of the first object, but is not limited thereto. For example, the processor (160) may recognize the type of the first object, the shape of the first object, and / or other objects existing around the first object through image analysis. For example, the processor (160) may recognize the location of the first object, the distance between the first object and the electronic device (100), and / or the shape of the first object through at least one sensor (130). For example, when the first object is an external electronic device (100), the processor (160) may receive information related to the first object from the first object (e.g., the name of the first object, identification information of the first object, and / or resource (e.g., component) information of the first object). According to one embodiment, when a plurality of objects are specified, the processor (160) may recognize information related to each of the specified objects.

[0036] According to one embodiment, the processor (160) may generate a first conversation session for a first object through at least one AI application (151). According to one embodiment, the first conversation session for the first object may be a customized conversation session corresponding to the first object. According to one embodiment, when a plurality of objects are specified based on a user input, the processor (160) may generate a conversation session for each of the specified objects through at least one AI application (151).

[0037] According to one embodiment, the processor (160) may determine an audio output path of the first conversation session based on information related to the electronic device (100) or information related to the first object. For example, the processor (160) may determine a path for outputting a response of the first conversation session for the first object. According to one embodiment, the information related to the electronic device (100) may include, but is not limited to, at least one of a status of a function or operation being performed by the electronic device (100), a communication connection status of the electronic device (100), and / or information about a component included in the electronic device (100) (or a resource of the electronic device (100). For example, the audio output path may include a speaker of the electronic device (100), a speaker of the first object (if the first object is an external electronic device (100) including a speaker), and / or a speaker of an external electronic device (100) connected to the electronic device (100) (e.g., including an external electronic device (100) excluding the first object). For example, the processor (160) may determine the same or different audio output paths for each conversation session for the object. According to one embodiment, if there are multiple objects specified based on a user input, the processor (160) may determine an audio output path for each conversation session for each of the specified objects based on information related to the electronic device (100) or information related to each of the multiple specified objects.

[0038] According to one embodiment, the processor (160) may determine the same or different input paths for each conversation session for each specified object. For example, the input path may be a path for receiving a user's voice input, and may include a microphone (120) of the electronic device (100), a microphone (120) of the first object (if the first object is an external electronic device (100) including a microphone (120), and / or a microphone (120) of an external electronic device (100) connected to the electronic device (100) (e.g., including an external electronic device (100) excluding the first object).

[0039] According to one embodiment, the processor (160) may determine attributes of a response (which may be referred to herein as a 'pre-prompt') to be provided based on a conversation session (e.g., the first conversation session) for a specified object based on at least one of information associated with the specified object (e.g., the first object) or a state of the electronic device (100). For example, attributes of the response may include, but are not limited to, a character of the response (e.g., a character of an AI assistant set in the first conversation session), a voice, a speech pattern, a tone, an age (e.g., an age of an AI assistant set in the first conversation session), and / or a level of knowledge (e.g., a type and / or depth of information used to generate the response).

[0040] In one embodiment, the processor (160) may receive a first voice input (e.g., a first user utterance) from a user related to an AI-based voice service via the microphone (120). For example, the processor (160) may provide the first voice input as input (e.g., a prompt) for a first conversation session. For example, the processor (160) may provide the first voice input to at least one AI application (151).

[0041] According to one embodiment, when a conversation session for a plurality of objects exists, the processor (160) may determine which object among the plurality of objects the first voice input is for based on at least one of the direction of the first voice input (e.g., direction of speech), the direction of the user's gaze, the direction of the head, the movement of the user, the position of the electronic device (100), the position of each of the plurality of objects, or the state of the electronic device (100). For example, the processor (160) may recognize the object corresponding to the first voice input.

[0042] According to various embodiments, the user input obtained by the processor (160) is not limited to voice input, and may include multimodal input (e.g., at least one of text, image, touch, gesture, or voice input).

[0043] According to one embodiment, the processor (160) may generate a response to the first voice input based on the first conversation session through at least one AI application (151). For example, the at least one AI application (151) may generate a response to the user's first voice input on its own (e.g., using a language model included in the electronic device (100)) or may generate a response to the user's first voice input through an external AI server (e.g., using a language model included in the external AI server). According to one embodiment, the at least one AI application (151) may generate a response to the first voice input using a language model customized for the first conversation session (e.g., a large language model (LLM) and / or a small large language model (sLLM)). For example, the customized language model may include a language model learned or specialized based on properties of the first conversation session (e.g., a pre-prompt). According to various embodiments, the model used by at least one AI application (151) is not limited to a language model, and may include various artificial intelligence-based models (e.g., a multimodal model (e.g., a large multimodal model (LMM)).

[0044] According to one embodiment, the processor (160) may generate a response based on properties of a conversation session (e.g., a first conversation session) for an object (e.g., a first object) corresponding to a first voice input among a plurality of objects.

[0045] In one embodiment, the processor (160) may provide a response to the first voice input through an output path of the first conversation session. In one embodiment, the processor (160) may provide a response through an output path of the conversation session for an object corresponding to the first voice input among a plurality of objects.

[0046] According to one embodiment, the processor (160) can recognize the speech direction of the first voice input through the microphone (120) and / or at least one sensor (130). The processor (160) can recognize the relative position of the first object with respect to the user based on at least one of the speech direction, the position of the electronic device (100), information related to the first object, and the user's gaze direction or the user's head direction detected through the at least one sensor (130). The processor (160) can control at least one of the output size or direction of a sound corresponding to the response based on the recognized relative position. For example, the processor (160) can control the output of the sound so that the user can perceive the sound corresponding to the response as being heard from the direction of the first object, thereby providing a spatial sound effect.

[0047] According to one embodiment, the processor (160) can independently store and / or manage the contents of inputs and responses corresponding to multiple conversation sessions for multiple objects. According to one embodiment, when generating a response to a third voice input following a second voice input based on a session for an object corresponding to the second voice input, the processor (160) can generate a response to the third voice input by reflecting the second voice input and the response to the second voice input. When generating a response to the third voice input based on a conversation session for an object that does not correspond to the second voice input, the processor (160) can generate a response to the third voice input without reflecting the second voice input and the response to the second voice input.

[0048] According to one embodiment, the processor (160) may determine, based on a user input, whether to share at least some of the input and response corresponding to a first conversation session for a first object to a second conversation session for a second object. The processor (160) may receive a third voice input of the user related to the second object through the microphone (120). If at least some of the input and response corresponding to the first conversation session is shared to the second conversation session, the processor (160) may generate a response to the third voice input by reflecting at least some of the shared input and response based on the second conversation session.

[0049] According to one embodiment, the processor (160) may call the first conversation session after generating a conversation session (e.g., the first conversation session) for at least one object (e.g., the first object) (e.g., after at least one of operations 640 to 680). For example, the processor (160) may call the first conversation session through a user input (e.g., a user utterance) that includes a call word for the registered first object. According to one embodiment, the processor (160) may recognize at least one of a user's gaze direction, a user's head direction, a user's movement, or a position of the first object through at least one sensor (130). The processor (160) may call the first conversation session based on at least one of a state of the electronic device (100), the user's gaze direction, the user's head direction, the user's movement, or the position of the first object. For example, the processor (160) may invoke the first conversation session at least partially based on recognizing that the user's gaze, head direction, and / or movement is directed toward the first object. For example, the processor (160) may invoke the first conversation session at least partially based on a location of the electronic device (100), a location of the first object, a relative location of the electronic device (100) and the first object, and / or a relationship (e.g., a relative positional relationship) between at least one external object and the first object. For example, the processor (160) may invoke the first conversation session at least partially based on a communication connection and / or signal transmission and reception with the first object (e.g., when the first object is an external electronic device (100).

[0050] According to various embodiments, the electronic device (100) may recognize an object and create a conversation session for the recognized object in a manner other than a method of creating a conversation session for the object through object recognition using an image. According to one embodiment, the processor (160) may recognize context information including at least one of an operating state of the electronic device (100) or information related to an external electronic device (100) connected to the electronic device (100). The processor (160) may create a third conversation session for providing an AI-based voice service corresponding to the context information through at least one AI application (151). The processor (160) may determine a response attribute of the AI ​​service provided through the third conversation session based on the context information.

[0051] According to various embodiments, the configuration of the electronic device (100) is not limited to that illustrated in FIG. 1, and at least some components may be omitted or at least one component (e.g., at least one of the components of FIG. 2 and FIGS. 11 to 15) may be added.

[0052]

[0053] Fig. 2 is a block diagram of an electronic device according to one embodiment. Hereinafter, descriptions overlapping with those in Fig. 1 are briefly described or omitted.

[0054] According to one embodiment, the electronic device (200) (e.g., the electronic device (100) of FIG. 1, the electronic device (1101) of FIG. 11, or the user terminal (1201) of FIG. 12) includes a camera (210) (e.g., the camera (110) of FIG. 1 or the camera module (1180) of FIG. 11), a microphone (220) (e.g., the microphone (120) of FIG. 1, the audio module (1170) of FIG. 11, or the microphone (1270) of FIG. 12), an object judgment module (231), an object storage module (233), an audio processing module (235), a communication circuit (240) (e.g., the communication circuit (130) of FIG. 1, the communication module (1190) of FIG. 11, or the communication interface (1290) of FIG. 12), an AI application (251) (e.g., the AI ​​application (151) of FIG. 1, the application (1146) of FIG. 11, or the 12 apps (1235a, 1235b)), and an on-device AI module (260) (e.g., the generative AI system (1500) of FIG. 15).

[0055] According to one embodiment, the camera (210) can capture an image that includes at least one external object.

[0056] According to one embodiment, the microphone (220) can receive a user's voice input.

[0057] According to one embodiment, the object judgment module (231) can recognize at least one object included in an image. The object judgment module (231) can recognize one or more objects (e.g., the first object (201) and / or the second object (203)) among the recognized at least one object for generating or invoking a customized AI conversation session.

[0058] According to one embodiment, the object storage module (233) may map and store information related to an object and information of a customized AI conversation session for the object. For example, the object storage module (233) may store information related to an object for a customized AI conversation session, information related to an electronic device (200), specified call conditions, information of an audio input / output path of the customized AI conversation session, information of a response attribute of the customized AI conversation session, and / or a call word. According to one embodiment, the object storage module (233) may independently store input and response contents of each customized AI conversation session (e.g., a first customized AI conversation session for the first object (201) or a second customized AI conversation session for the second object (203)) for each object (e.g., a first customized AI conversation session for the first object (201) or a second customized AI conversation session for the second object (203)).

[0059] According to one embodiment, the audio processing module (235) may transmit a user's voice input received through the microphone (220) to a customized AI conversation session having a corresponding audio input path, and / or transmit a response generated based on the customized AI conversation session to an audio output path of the customized AI conversation session.

[0060] According to one embodiment, the communication circuit (240) may communicate with an external electronic device (280) (e.g., may include an external object if the external object is an external electronic device (280)) and / or an external server (290) (e.g., an AI server (290)), or may transmit and receive information and / or data with the external electronic device (280) and / or the external server (290).

[0061] According to one embodiment, the AI ​​application (251) may include an application (251) that supports an artificial intelligence-based voice service. For example, the AI ​​application (251) may include an application (251) initially stored in the electronic device (200) and / or an application (251) provided by a third party. For example, the AI ​​application (251) may include an application (251) that provides a response to a user input based on AI (e.g., a language model and / or a multimodal model).

[0062] According to one embodiment, the on-device AI module (260) may include a voice-to-text processing module (261), an object-specific text processing module (265), and a text post-processing module (263). For example, the on-device AI module (260) may include a generative AI module (e.g., the generative artificial intelligence system (1500) of FIG. 15).

[0063] According to one embodiment, a voice-to-text processing module (261) can convert a user's voice input into text. The voice-to-text processing module (261) can transmit the converted text to a text post-processing module (263).

[0064] According to one embodiment, the text post-processing module (263) may determine at least one customized AI conversation session to which the text is to be delivered, and determine whether to reflect (reference) or not to reflect (not reference) the text and the response to the text in the customized AI conversation session.

[0065] According to one embodiment, the object-specific text processing module (265) may transmit each text corresponding to a voice input to a customized AI dialogue session corresponding to the voice input. For example, the object-specific text processing module (265) may determine an object corresponding to the voice input based on the voice input, the state of the electronic device (200), the user's gaze direction, head direction, movement, and / or the position of the electronic device (200), and transmit the text corresponding to the voice input to a customized AI dialogue session for the determined object.

[0066] According to one embodiment, the electronic device (200) may perform at least some operations for providing an AI-based voice service in conjunction with an external electronic device (280). For example, the electronic device (200) may be connected to the external electronic device (280) through a communication circuit (240). The electronic device (200) may recognize an external object (e.g., a first object (201) and / or a second object (203)) through the external electronic device (280), receive a user input (e.g., a voice input) through the external electronic device (280), and / or provide a response to a user input generated based on a customized AI conversation session through the external electronic device (280).

[0067] According to one embodiment, the AI ​​server (290) (e.g., the server (1108) of FIG. 11 or the intelligent server (1300) or service server (1400) of FIG. 12) may store at least one external AI application (251). According to one embodiment, the AI ​​server (290) may provide an AI-based voice service according to a request of the AI ​​application (251). For example, the AI ​​server (290) may include at least one response model. The AI ​​server (290) may receive information related to a user's voice input from the AI ​​application (251), and may generate a response to the voice input using at least one response model and provide the response to the AI ​​application (251).

[0068] According to various embodiments, at least some of the components of FIG. 2 may be omitted or at least one component (e.g., at least one of the components of FIG. 1, FIG. 11 to FIG. 15) may be added. According to various embodiments, at least some of the components of FIG. 2 may be implemented as a single integrated component (module). For example, the object judgment module (231), the object storage module (233), the audio processing module (235), or the on-device AI (260) module may be implemented as at least one processor (e.g., the processor (160) of FIG. 1, the processor (1120) of FIG. 11, or the processor (1220) of FIG. 12).

[0069]

[0070] FIG. 3 is a diagram illustrating an operation of recognizing an object and creating a conversation session for an object in an electronic device according to one embodiment.

[0071] According to one embodiment, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIG. 12, or a system (1500) of FIG. 15) may execute an AI application (e.g., an AI application (151) of FIG. 1, an AI application (251) of FIG. 2, or an application (1146) of FIG. 11, or an app (1235a, 1235b) of FIG. 12). For example, the electronic device may activate an AI-based voice service of the electronic device. For example, the electronic device may activate the AI-based voice service based on receiving a user input that calls an AI assistant.

[0072] According to one embodiment, an electronic device may acquire an image (e.g., a still image or a moving image, and may include a preview image) including at least one external object using a camera (e.g., a camera (110) of FIG. 1, a camera (210) of FIG. 2, or a camera module (1180) of FIG. 11). According to one embodiment, the electronic device may analyze the image to recognize at least one object included in the image.

[0073] According to one embodiment, the electronic device may receive a user input for selecting a first object from at least one image. For example, the user input may include, but is not limited to, a touch input, a gesture input, and / or a voice input for selecting the first object included in the image. For example, the electronic device may select the first object by controlling a camera (or a capture function of the camera) so that only the first object is included in an image acquired through the camera. According to one embodiment, the electronic device may display an indication indicating the selected first object.

[0074] According to one embodiment, an electronic device may generate a customized AI conversation session for a selected first object. The electronic device may set an audio output path (and / or an audio input path) of the customized AI conversation session for the first object based on a user input, information related to the first object, and / or information related to the electronic device. The electronic device may determine attributes of a response of the customized AI conversation session for the first object based on the user input, information related to the first object, and / or information related to the electronic device. According to one embodiment, when generating a customized AI conversation session for the first object, the electronic device may store information regarding at least one of whether the first object is an electronic device, a characteristic of the first object, a location of the first object, information about objects surrounding the first object, and a locational relationship between the first object and surrounding objects, in association with the customized AI conversation session for the first object. The information stored in association with the customized AI conversation session for the first object may be used to specify the customized AI conversation session for the first object when the customized AI conversation session is subsequently called.

[0075] In one embodiment, an electronic device may receive a user input specifying the name of a selected first object. The electronic device may register the name of the first object as a trigger word for a customized AI conversation session for the first object. Upon receiving a user input including the trigger word, the electronic device may activate a customized AI conversation session corresponding to the trigger word.

[0076] For example, referring to 300a, the electronic device may receive a user input specifying a name of a selected first object (e.g., a rabbit doll) as “Bunny.” The electronic device may register “Bunny” as a trigger word for a customized AI conversation session for the first object. Referring to 300b, the electronic device may receive a user input specifying a name of a selected first object (e.g., a puppy doll) as “Puppy.” The electronic device may register “Puppy” as a trigger word for a customized AI conversation session for the first object.

[0077]

[0078] FIG. 4 illustrates examples of an electronic device providing a voice service according to one embodiment.

[0079] Referring to FIG. 400a, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIG. 12, and a system (1500) of FIG. 15) can recognize an external object (e.g., a teddy bear) through a camera. The electronic device can create a customized conversation session for the recognized object in order to provide an AI-based voice service. After creating the customized conversation session for the recognized object, the electronic device can call the customized conversation session for the object when receiving a user input including the name (call word) of the object. When the electronic device executes an application that provides an AI-based voice service (e.g., a user calls an AI assistant), the electronic device can call the customized conversation session for the object recognized in the image based on image recognition through the camera. When an electronic device detects a user's speech direction, gaze direction, head direction, and / or movement, and executes an application that provides an AI-based voice service (e.g., when a user calls an AI assistant), the electronic device may recognize an external object based on the user's speech direction, gaze direction, head direction, and / or movement, and call a customized conversation session for the recognized object. For example, the electronic device may reconnect to a previously created and / or previously connected customized conversation session, and continue a conversation with the customized AI assistant that was previously conducted through the customized conversation session. According to one embodiment, when information related to surrounding objects of an object corresponding to a customized conversation session (e.g., type of surrounding object, relative position) has been stored in advance, the electronic device may recognize an object corresponding to the customized conversation session by using the information related to the surrounding objects among objects included in an image, and call a customized conversation session for the recognized object.

[0080] Referring to FIG. 400b, an electronic device may provide an AI-based voice service by interfacing with an external electronic device (e.g., a vision see-through (VST)). For example, the electronic device may recognize a user's speech direction, gaze direction, head direction, and / or movement through the external electronic device. The electronic device may recognize an external object based on information about the user's speech direction, gaze direction, head direction, and / or movement received from the external electronic device. When the electronic device executes an application that provides an AI-based voice service (e.g., calls an AI assistant), the electronic device may invoke a customized conversation session for the recognized object.

[0081] Referring to FIG. 400c, if the external object is an external electronic device (e.g., a charger or an electronic device capable of communicating with an electronic device), the electronic device may be connected to the external object through a communication circuit, or, if it obtains information about the external object, create a customized conversation session for the recognized external object based on the information about the external object. For example, the electronic device may create an AI assistant (e.g., a bear-shaped virtual AI assistant) corresponding to the customized conversation session based on the characteristics (e.g., identification information, shape, and / or function) of the recognized external object. The electronic device may recognize the external object through the communication circuit, or recognize the external object based on information received from the external object. The electronic device may invoke a customized conversation session for the recognized external object. For example, if the electronic device executes an application that provides an AI-based voice service (e.g., invokes an AI assistant), the electronic device may identify the external object (external electronic device) with which the communication is connected, and invoke a customized conversation session for the connected external object.

[0082]

[0083] FIGS. 5A and 5B are diagrams illustrating operations for managing a conversation session of an electronic device according to one embodiment.

[0084] According to one embodiment, an electronic device (e.g., the electronic device 100 of FIG. 1, the electronic device 200 of FIG. 2, the electronic device 1101 of FIG. 11, the user terminal 1201 of FIG. 12, and the system 1500 of FIG. 15) may generate a plurality of customized conversation sessions for a plurality of objects in order to provide an AI-based voice service. The electronic device may independently store and / or manage the contents of inputs and responses corresponding to each of the plurality of customized conversation sessions. For example, the electronic device may individually store and / or manage the conversation contents of each of the plurality of customized conversation sessions. In FIGS. 5A and 5B , it is assumed that the plurality of customized conversation sessions include a first conversation session and a second conversation session, but the number of customized conversation sessions is not limited thereto. For example, FIG. 5A illustrates a flow of providing a user's voice input and response during provision of an AI-based voice service, and FIG. 5B illustrates a case in which the input and response contents of the conversation session corresponding to FIG. 5A are stored. For example, each input and response content of each conversation session may include information regarding whether the electronic device will refer to or not refer to that content in each conversation session. For example, the electronic device may generate a response to a subsequent user input by reflecting the content of the input and response determined to be referred to in each conversation session. The electronic device may generate a response to a subsequent user input without reflecting the content of the input and response determined not to be referred to (i.e., not referenced) in each conversation session.

[0085] In one embodiment, in operation 501, the electronic device may initiate an AI-based voice service. For example, the electronic device may execute at least one AI application to provide the AI-based voice service.

[0086] According to one embodiment, in operation 503, the electronic device may create a first conversation session for a first object. After creating the first conversation session for the first object, the electronic device may determine a first AI assistant corresponding to the first conversation session. For example, the electronic device may determine attributes of the first AI assistant (e.g., response attributes of the first conversation session) based on information related to the first object and / or information related to the electronic device. According to one embodiment, the electronic device may designate a name of the first AI assistant. For example, when creating the first conversation session, the electronic device may designate a name (e.g., X) of the first AI assistant corresponding to the first object. The electronic device may register the name as a call word of the first conversation session (e.g., the first AI assistant). According to one embodiment, the electronic device may determine an input path (e.g., an audio input path) and an output path (e.g., an audio output path) of the first conversation session. In the following, it is assumed that the first object is a doll, the input path of the first conversation session is the terminal's microphone, and the output path is the terminal's speaker.

[0087] According to one embodiment, in operation 505, the electronic device may create a second conversation session for a second object. After creating the second conversation session for the second object, the electronic device may determine a second AI assistant corresponding to the second conversation session. For example, the electronic device may determine attributes of the second AI assistant (e.g., response attributes of the second conversation session) based on information related to the second object and / or information related to the electronic device. According to one embodiment, the electronic device may designate a name of the second AI assistant. For example, when creating a first conversation session, the electronic device may designate a name (e.g., Y) of the first AI assistant corresponding to the first object. The electronic device may register the name as a call word for the second conversation session (e.g., the second AI assistant). According to one embodiment, the electronic device may determine an input path (e.g., an audio input path) and an output path (e.g., an audio output path) of the second conversation session. In the following, it is assumed that the second object is a speaker device, the input path of the second conversation session is the microphone of the earphone connected to the terminal, and the output path is the speaker of the second object.

[0088] Referring to 511, the electronic device may receive the user's first voice input (“How are you?”) in the first conversation session. For example, the electronic device may receive the user's first voice input through the microphone of the terminal, which is an input path of the first conversation session. According to one embodiment, the user's first voice input may be received through an input path (e.g., a microphone of earphones) of the second conversation session. The electronic device may determine that the first voice input is an input for the first object (i.e., the first AI assistant (X)) based on at least one of the first voice input, the user's gaze direction, the user's head direction, the user's movement, the electronic device (user), the first object, and / or the position of the second object. In this case, the electronic device may determine not to reference the first voice input in the second conversation session.

[0089] Referring to 521, the electronic device may obtain the first response (“Yes, I was waiting for you at home”) from X based on the first conversation session. For example, the electronic device may output the first response through an output path (e.g., a speaker of the terminal) of the first conversation session. For example, the first response may be received through an input path (e.g., earphones) of the second conversation session. The electronic device may decide not to reference the first response in the second conversation session.

[0090] Referring to 531, the electronic device may receive the user's second voice input (“How are you?”) in the second conversation session. For example, the electronic device may receive the user's second voice input through the microphone of the earphone, which is an input path of the second conversation session. According to one embodiment, the user's second voice input may be received through the input path of the first conversation session (e.g., the microphone of the terminal). The electronic device may determine that the second voice input is an input for the second object (i.e., the second AI assistant (Y)) based on at least one of the second voice input, the user's gaze direction, the user's head direction, the user's movement, the electronic device (the user), the first object, and / or the position of the second object. In this case, the electronic device may determine not to reference the second voice input in the first conversation session.

[0091] Referring to 541, the electronic device may obtain a second response (“Today was a nice day, so it was pleasant”) from Y based on the second conversation session. For example, the electronic device may output the second response through an output path of the second conversation session (e.g., a speaker of the second object). For example, the second response may be received through an input path of the first conversation session (e.g., a microphone of the terminal). The electronic device may decide not to reference the second response in the first conversation session.

[0092] Referring to 551, an electronic device can conduct simultaneous conversations in a first conversation session and a second conversation session. For example, the electronic device can receive a third voice input (e.g., “Shall we talk about X, Y together?”) from the user to conduct simultaneous conversations in multiple conversation sessions through an input path of the first conversation session and / or an input path of the second conversation session. Based on the third voice input, the electronic device can simultaneously activate the first conversation session and the second conversation session. For example, the electronic device can provide the user with an AI-based voice service, such as a conversation with a first object (e.g., X) and a second object (e.g., Y), based on the first conversation session and the second conversation session. For example, the electronic device can decide to share or cross-reference the contents of the user’s voice input and response in each of the multiple conversation sessions that are simultaneously activated.

[0093] Referring to 561, the electronic device may receive the user's fourth voice input (“Where would be a good place to go for fun tomorrow”) in the second conversation session. For example, the electronic device may receive the user's fourth voice input through the microphone of the earphone, which is an input path of the second conversation session. According to one embodiment, the user's fourth voice input may be received through the input path of the first conversation session (e.g., the microphone of the terminal) and / or the input path of the second conversation session. For example, if the fourth voice input is received through the input path of the second conversation session (e.g., the microphone of the earphone), the electronic device may obtain information about the fourth voice input received through the input path of the second conversation session and provide the information about the fourth voice input to the first conversation session (e.g., X). According to one embodiment, the electronic device may determine to refer to the fourth voice input in the first conversation session.

[0094] Referring to 571, the electronic device can obtain a third response (“It’s going to rain tomorrow, so go home”) from Y based on the second conversation session. For example, the electronic device can output the third response through an output path of the second conversation session. For example, the third response can be received through an input path of the first conversation session (e.g., a microphone of the terminal) and / or an input path of the second conversation session (e.g., a microphone of the earphone). For example, if the third response is received through an input path of the second conversation session (e.g., a microphone of the earphone), the electronic device can obtain information about the third response received through the input path of the second conversation session and provide the information about the third response to the first conversation session (e.g., X). The electronic device can decide to refer to the third response in the first conversation session.

[0095] Referring to 581, the electronic device may receive a fifth voice input (“What do you think?”) from a user in a first conversation session. For example, the electronic device may receive the fifth voice input through an input path (e.g., a microphone of the terminal) of the first conversation session and / or an input path (e.g., a microphone of the earphones) of the second conversation session. For example, if the fifth voice input is received through an input path (e.g., a microphone of the earphones) of the second conversation session, the electronic device may obtain information about the fifth voice input received through the input path of the second conversation session and provide the information about the fifth voice input to the second conversation session (e.g., Y). According to one embodiment, the electronic device may determine to refer to the fifth voice input in the second conversation session.

[0096] Referring to 591, the electronic device can obtain a fourth response (“Me too, home, let’s watch a movie at home together”) from X based on the first conversation session. For example, the electronic device can output the fourth response through an output path of the first conversation session. For example, the fourth response can be received through an input path of the first conversation session (e.g., a microphone of the terminal) and / or an input path of the second conversation session (e.g., a microphone of the earphone). For example, if the fourth response is received through an input path of the first conversation session (e.g., a microphone of the earphone), the electronic device can obtain information about the fourth response received through the input path of the first conversation session and provide the information about the fourth response to the second conversation session (e.g., Y). The electronic device can decide to refer to the fourth response in the second conversation session.

[0097] According to one embodiment, the electronic device may provide the user with the contents of inputs and responses stored for each conversation session. The electronic device may also provide information on whether each of the contents of inputs and responses stored for each conversation session is referenced or not. Referring to FIG. 5B, a first area relates to a first voice input of 511, a second area relates to a first response of 521, a third area relates to a second voice input of 531, a fourth area relates to a second response of 541, a fifth area relates to a third voice input of 551, a sixth area relates to a fourth voice input of 561, a seventh area relates to a third response of 571, an eighth area relates to a fifth voice input of 581, and a ninth area relates to a fourth response of 591. For example, although the contents of the first and second conversation sessions are shown together in FIG. 5B, the electronic device may provide the contents of the first and second conversation sessions together or separately. For example,

[0098] According to one embodiment, the electronic device may change whether each input and response of each conversation session is referenced or non-referenced based on a user input. For example, the electronic device may change at least one voice input and / or at least one response that was set as non-referenced in a first conversation session (and / or a second conversation session) to be referenced in the first conversation session (and / or the second conversation session) based on a user input. Conversely, the electronic device may change at least one voice input and / or at least one response that was set as referenced in the first conversation session (and / or the second conversation session) to be non-referenced in the first conversation session (and / or the second conversation session) based on a user input.

[0099]

[0100] An electronic device according to one embodiment of the present disclosure may include a camera, a microphone, at least one sensor, a communication circuit, a memory storing at least one AI application that provides an artificial intelligence (AI)-based voice service, and at least one processor.

[0101] The memory may store instructions that, when executed by the at least one processor, cause the electronic device to obtain an image including at least one external object using the camera.

[0102] The instructions, when executed by the at least one processor, may cause the electronic device to receive a user input specifying a first object of the at least one external object.

[0103] The instructions, when executed by the at least one processor, may cause the electronic device to recognize information related to the first object using at least one of the image, the at least one sensor, or the communication circuit.

[0104] The instructions, when executed by the at least one processor, may cause the electronic device to create a first conversation session for the first object through the at least one AI application.

[0105] The instructions, when executed by the at least one processor, may cause the electronic device to determine an audio output path of the first conversation session based on information associated with the electronic device or information associated with the first object.

[0106] The above instructions, when executed by the at least one processor, may cause the electronic device to receive a first voice input of a user related to the AI-based voice service through the microphone.

[0107] The instructions, when executed by the at least one processor, may cause the electronic device to generate a response to the first voice input based on the first conversation session through the at least one AI application.

[0108] The instructions, when executed by the at least one processor, may cause the electronic device to provide a response to the first voice input via an audio output path of the first conversation session.

[0109] The instructions, when executed by the at least one processor, may cause the electronic device to receive a user input specifying a name of the first object.

[0110] The above instructions, when executed by the at least one processor, may cause the electronic device to register the name of the first object as a call word of the first conversation session.

[0111] Information related to the first object may include at least one of the type of the first object, identification information of the first object, the location of the first object, the shape of the first object, or the name of the first object.

[0112] Information related to the electronic device may include at least one of the status of a function or operation being performed by the electronic device, the status of a communication connection of the electronic device, or information about a component included in the electronic device.

[0113] The instructions, when executed by the at least one processor, may cause the electronic device to determine attributes of a response to be provided based on the first conversation session, based on at least one of information related to the first object or a state of the electronic device.

[0114] The instructions, when executed by the at least one processor, may cause the electronic device to recognize at least one of a direction of the user's gaze, a direction of the user's head, a movement of the user, or a position of the first object through the at least one sensor.

[0115] The instructions, when executed by the at least one processor, may cause the electronic device to invoke the first conversation session based on at least one of a state of the electronic device, a gaze direction of the user, a head direction of the user, a movement of the user, or a position of the first object.

[0116] The above instructions, when executed by the at least one processor, may cause the electronic device to recognize a speech direction corresponding to the first voice input through the microphone.

[0117] The instructions, when executed by the at least one processor, may cause the electronic device to recognize a relative position of the first object to the user based on at least one of the direction of speech, the position of the electronic device, information related to the first object, and the direction of the user's gaze or the direction of the user's head detected by the at least one sensor.

[0118] The instructions, when executed by the at least one processor, may cause the electronic device to control at least one of an output size or direction of a sound corresponding to the response based on the relative position.

[0119] The above instructions, when executed by the at least one processor, may cause the electronic device to create a second conversation session with a second object through the at least one AI application.

[0120] The instructions, when executed by the at least one processor, may cause the electronic device to determine an audio output path of the second conversation session based on at least one of information associated with the electronic device and information associated with the second object.

[0121] The above instructions, when executed by the at least one processor, may cause the electronic device to receive a second voice input of a user related to the AI-based voice service through the microphone.

[0122] The above instructions, when executed by the at least one processor, may cause the electronic device to recognize relative positions of the first object and the second object with respect to a user.

[0123] The instructions, when executed by the at least one processor, may cause the electronic device to determine an object corresponding to the second voice input among the first object and the second object based on at least one of a speech direction corresponding to the second voice input, a user's gaze direction, a user's head direction, or the relative position.

[0124] The instructions, when executed by the at least one processor, may cause the electronic device to generate a response to the second voice input based on a conversation session for an object corresponding to the second voice input among the first conversation session and the second conversation session.

[0125] The instructions, when executed by the at least one processor, may cause the electronic device to provide a response to the second voice input via an audio output path of a conversation session for the determined object.

[0126] The above instructions, when executed by the at least one processor, may cause the electronic device to independently store the contents of the input and response corresponding to the first conversation session and the contents of the input and response corresponding to the second conversation session.

[0127] The instructions, when executed by the at least one processor, may cause the electronic device to determine, based on a user input, whether to share at least some of the input and response corresponding to the first conversation session with the second conversation session.

[0128] The instructions, when executed by the at least one processor, may cause the electronic device to receive a third voice input of a user related to the second object through the microphone.

[0129] The instructions, when executed by the at least one processor, may cause the electronic device to generate a response to the third voice input by reflecting at least some of the shared input and response based on the second conversation session, if at least some of the input and response corresponding to the first conversation session is shared with the second conversation session.

[0130] The instructions, when executed by the at least one processor, may cause the electronic device to generate a response to a third voice input following the second voice input based on a session for an object corresponding to the second voice input, by reflecting the second voice input and the response to the second voice input.

[0131] The instructions, when executed by the at least one processor, may cause the electronic device to generate a response to the third voice input based on a conversation session for an object that does not correspond to the second voice input, without reflecting the second voice input and the response to the second voice input.

[0132] The instructions, when executed by the at least one processor, may cause the electronic device to recognize context information including at least one of an operating state of the electronic device or information related to an external electronic device connected to the electronic device.

[0133] The above instructions, when executed by the at least one processor, may cause the electronic device to generate a third conversation session for providing the AI-based voice service corresponding to the context information through the at least one AI application.

[0134] The above instructions, when executed by the at least one processor, may cause the electronic device to determine a response attribute of an AI-based voice service provided through the third conversation session based on the context information.

[0135] According to various embodiments of the present disclosure, an electronic device can provide AI voice services based on physical objects, thereby providing diverse and intuitive interactions between a user and an AI assistant, and can improve user experience and increase usability of the AI ​​voice services. According to various embodiments of the present disclosure, an electronic device can provide a variety of personalized and / or customized voice services to a user by generating and providing customized AI conversation sessions (customized AI assistants) for each of a plurality of physical objects (and / or surrounding situations, or states of the electronic device). According to various embodiments of the present disclosure, an electronic device can provide a user experience similar to interacting with an actual physical object to a user by setting an audio path (e.g., an audio input path and / or an audio output path) for providing an AI voice service (e.g., an AI conversation session) based on the location, type, and / or state of an object, the location of the user, and / or the location and / or state of the electronic device.

[0136]

[0137] Figure 6 is a flowchart of a method for providing voice service of an electronic device according to one embodiment.

[0138] According to one embodiment, in operation 610, an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (1101) of FIG. 11, the user terminal (1201) of FIG. 12, and the system (1500) of FIG. 15) may acquire an image including at least one external object using a camera. For example, the external object may include an electronic device and / or a non-electronic device. According to one embodiment, the electronic device may recognize at least one external object included in the image.

[0139] According to one embodiment, in operation 620, the electronic device may receive a user input specifying a first object from among at least one external object. For example, the electronic device may receive a user input selecting a first object from among at least one external object included in an image. For example, the user input selecting the first object may include a touch input, a gesture input, and / or a voice input. According to one embodiment, the electronic device may specify one object (e.g., the first object) or multiple objects (e.g., the first object and the second object) based on the user input.

[0140] In one embodiment, the electronic device may receive a user input specifying a name of a first object. In one embodiment, the electronic device may register the name of the first object as a call word for a first conversation session for the first object.

[0141] According to one embodiment, in operation 630, the electronic device may recognize information related to the first object using at least one of an image, at least one sensor, or a communication circuit. According to one embodiment, the information related to the first object may include at least one of the type of the first object, identification information of the first object, the location of the first object, the shape of the first object, or the name of the first object, but is not limited thereto. For example, the electronic device may recognize the type of the first object, the shape of the first object, and / or other objects existing around the first object through image analysis. For example, the electronic device may recognize the location of the first object, the distance between the first object and the electronic device, and / or the shape of the first object through at least one sensor. For example, when the first object is an external electronic device, the electronic device may receive information related to the first object (e.g., the name of the first object, identification information of the first object, and / or resource (e.g., component) information of the first object) from the first object. According to one embodiment, when a plurality of objects are specified, the electronic device can recognize information associated with each of the specified objects.

[0142] According to one embodiment, in operation 640, the electronic device may create a first conversation session for a first object through at least one AI application. The at least one AI application may include an application that supports AI-based voice services. For example, the at least one AI application may include an application that provides a response to a user's voice input (e.g., a user utterance) using a language model. For example, the language model may be included in the AI ​​application, included in the electronic device, or included in an external server (e.g., an AI server). According to one embodiment, the first conversation session for the first object may be a customized conversation session corresponding to the first object. According to one embodiment, when a plurality of objects are specified based on a user input, the electronic device may create a conversation session for each of the specified objects through at least one AI application.

[0143] According to one embodiment, in operation 650, the electronic device may determine an audio output path of the first conversation session based on information related to the electronic device or information related to the first object. For example, the electronic device may determine a path for outputting a response of the first conversation session for the first object. According to one embodiment, the information related to the electronic device may include, but is not limited to, at least one of a status of a function or operation being performed by the electronic device, a communication connection status of the electronic device, and / or information about a component included in the electronic device (or a resource of the electronic device). For example, the audio output path may include a speaker of the electronic device, a speaker of the first object (if the first object is an external electronic device including a speaker), and / or a speaker of an external electronic device connected to the electronic device (e.g., including an external electronic device excluding the first object). For example, the electronic device may determine the same or different audio output paths for each conversation session for the object. According to one embodiment, when there are multiple objects specified based on user input, the electronic device can determine an audio output path for each of the conversation sessions for each of the specified objects based on information associated with the electronic device or information associated with each of the multiple specified objects.

[0144] According to one embodiment, the electronic device may determine the same or different input paths for each conversation session for each specific object. For example, the input path may include a microphone of the electronic device, a microphone of a first object (if the first object is an external electronic device including a microphone), and / or a microphone of an external electronic device connected to the electronic device (e.g., including an external electronic device other than the first object).

[0145] According to one embodiment, the electronic device may determine attributes of a response (which may be referred to herein as a 'pre-prompt') to be provided based on a conversation session (e.g., the first conversation session) for a specified object based on at least one of information associated with the specified object (e.g., the first object) or a state of the electronic device. For example, attributes of the response may include, but are not limited to, a character of the response (e.g., a character of an AI assistant set in the first conversation session), a voice, a speech pattern, a tone, an age (e.g., an age of an AI assistant set in the first conversation session), and / or a level of knowledge (e.g., a type and / or depth of information used to generate the response).

[0146] In one embodiment, at step 660, the electronic device may receive a first user voice input (e.g., a first user utterance) related to an AI-based voice service via a microphone. For example, the electronic device may provide the first voice input as input (e.g., a prompt) for a first conversation session. For example, the electronic device may provide the first voice input to at least one AI application.

[0147] According to one embodiment, when a conversation session for a plurality of objects exists, the electronic device may determine which object among the plurality of objects the first voice input is for based on at least one of the direction of the first voice input (e.g., direction of speech), the direction of the user's gaze, the direction of the head, the movement of the user, the position of the electronic device, the position of each of the plurality of objects, or the state of the electronic device. For example, the electronic device may recognize the object corresponding to the first voice input.

[0148] According to various embodiments, the user input received by the electronic device is not limited to voice input, but may include multimodal input (e.g., at least one of text, image, touch, gesture, or voice input).

[0149] According to one embodiment, in operation 670, the electronic device may generate a response to the first voice input based on the first conversation session through at least one AI application. For example, the at least one AI application may generate a response to the user's first voice input on its own (e.g., using a language model included in the electronic device) or may generate a response to the user's first voice input through an external AI server (e.g., using a language model included in the external AI server). According to one embodiment, the at least one AI application may generate a response to the first voice input using a language model customized for the first conversation session (e.g., a large language model (LLM) and / or a small large language model (sLLM)). For example, the customized language model may include a language model learned or specialized based on properties of the first conversation session (e.g., a pre-prompt).

[0150] According to one embodiment, the electronic device may generate a response based on properties of a conversation session (e.g., a first conversation session) for an object (e.g., a first object) corresponding to a first voice input among a plurality of objects.

[0151] According to various embodiments, the model utilized by at least one AI application is not limited to a language model, but may include various artificial intelligence-based models (e.g., multimodal models (e.g., large multimodal models (LMM)).

[0152] In one embodiment, at operation 680, the electronic device may provide a response to the first voice input through an output path of the first conversation session. In one embodiment, the electronic device may provide a response through an output path of the conversation session for an object corresponding to the first voice input among a plurality of objects.

[0153] According to one embodiment, the electronic device can recognize the speech direction of the first voice input through a microphone and / or at least one sensor. The electronic device can recognize the relative position of the first object with respect to the user based on at least one of the speech direction, the position of the electronic device, information related to the first object, and the user's gaze direction or the user's head direction detected through the at least one sensor. The electronic device can control at least one of the output size or direction of a sound corresponding to the response based on the recognized relative position. For example, the electronic device can provide a spatial sound effect by controlling the output of the sound so that the user can perceive the sound corresponding to the response as being heard from the direction of the first object.

[0154] According to one embodiment, the electronic device can independently store and / or manage the contents of inputs and responses corresponding to multiple conversation sessions for multiple objects. According to one embodiment, when generating a response to a third voice input following a second voice input based on a session for an object corresponding to the second voice input, the electronic device can generate a response to the third voice input by reflecting the second voice input and the response to the second voice input. When generating a response to the third voice input based on a conversation session for an object that does not correspond to the second voice input, the electronic device can generate a response to the third voice input without reflecting the second voice input and the response to the second voice input.

[0155] According to one embodiment, the electronic device may determine, based on a user input, whether to share at least a portion of an input and response corresponding to a first conversation session for a first object with a second conversation session for a second object. The electronic device may receive a third voice input of the user related to the second object through a microphone. If at least a portion of the input and response corresponding to the first conversation session is shared with the second conversation session, the electronic device may generate a response to the third voice input by reflecting at least a portion of the shared input and response based on the second conversation session.

[0156] According to one embodiment, the electronic device may invoke the first conversation session after generating a conversation session (e.g., the first conversation session) for at least one object (e.g., the first object) (e.g., after at least one of operations 640 to 680). For example, the electronic device may invoke the first conversation session through a user input (e.g., a user utterance) that includes a wake word for the registered first object. According to one embodiment, the electronic device may recognize at least one of a user's gaze direction, a user's head direction, a user's movement, or a position of the first object through at least one sensor. The electronic device may invoke the first conversation session based on at least one of a state of the electronic device, the user's gaze direction, the user's head direction, the user's movement, or the position of the first object. For example, the electronic device may invoke the first conversation session at least in part based on recognizing that the user's gaze, head direction, and / or movement is directed toward the first object. For example, the electronic device may invoke the first conversation session based at least in part on a location of the electronic device, a location of the first object, a relative location of the electronic device and the first object, and / or a relationship (e.g., a relative positional relationship) between at least one external object and the first object. For example, the electronic device may invoke the first conversation session based at least in part on a communication connection and / or signal transmission and reception with the first object (e.g., if the first object is an external electronic device).

[0157] According to various embodiments, although FIG. 6 describes a case where a customized conversation session is created based on image acquisition via a camera, the method for creating a customized conversation session is not limited thereto. For example, an electronic device may recognize context information including at least one of an operating state of the electronic device or information related to an external electronic device connected to the electronic device. The electronic device may create a third conversation session to provide an AI-based voice service corresponding to the context information through at least one AI application. Based on the context information, the electronic device may determine a response attribute of the AI ​​service provided through the third conversation session.

[0158] According to various embodiments, the order of at least some of the operations of FIG. 6 may be changed, some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIGS. 7 to 10) may be added. For example, at least some of the operations of FIGS. 6 to 10 may be performed independently or in conjunction with one another.

[0159]

[0160] Fig. 7 is a flowchart of a method for providing voice services in an electronic device according to one embodiment. Below, any descriptions that overlap with those in Fig. 6 will be briefly explained or omitted.

[0161] According to one embodiment, in operation 710, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIG. 12, or a system (1500) of FIG. 15) may acquire an external image through a camera. For example, the external image may include at least one external object.

[0162] According to one embodiment, in operation 720, the electronic device may recognize at least one external object included in an external image. For example, the electronic device may select at least one of the at least one external object. For example, the electronic device may receive a user input specifying a name of the at least one selected object. The electronic device may register the name of the selected object as a call word for a conversation session regarding the object.

[0163] According to one embodiment, in operation 730, the electronic device may obtain information related to the selected object. The electronic device may obtain information related to the object using an image containing the selected object, at least one sensor, and a communication circuit. For example, the information related to the object may include the name (call word) of the object specified based on a user input in operation 720.

[0164] For example, an electronic device can recognize the positional relationship of a selected object with respect to a user (electronic device) (e.g., in front, behind, left, or right of the electronic device). For example, an electronic device can recognize the relationship between a selected object and surrounding objects (surrounding objects). For example, an electronic device can recognize the types of surrounding objects located around the selected object and the relative positions of the selected object with respect to the surrounding objects (e.g., to the left, right, in front, or behind the surrounding objects). The electronic device can store information related to surrounding objects (e.g., types and / or positions of surrounding objects) together with information related to the selected object, and can then use the information related to the surrounding objects to identify an object corresponding to a user's voice input. For example, an electronic device can recognize the type of a selected object (e.g., whether the selected object is an external electronic device or a non-electronic device).

[0165] According to one embodiment, at step 740, the electronic device may create a conversation session for an AI-based voice service. For example, the electronic device may create an AI-based customized conversation session for a selected object through an AI application. For example, the customized conversation session may be created independently for each object.

[0166] According to one embodiment, at operation 750, the electronic device may generate a pre-prompt (which may be referred to herein as an “attribute of the conversation session” or an “attribute of the response”) based on information associated with the object. For example, the pre-prompt may include at least one of a personality, voice, tone, speech, age, and / or knowledge level of the AI ​​assistant used to generate the response of the conversation session (e.g., difficulty of description included in the response and / or scope of information used to generate the response and / or depth of information).

[0167] According to one embodiment, at step 760, the electronic device may perform an AI-based voice conversation based on a pre-prompt. For example, the electronic device may receive a user input (e.g., a user voice input (e.g., a user utterance)) and provide the voice input as an input (e.g., a prompt) for a customized conversation session for an object. The electronic device may generate a response to the user's voice input based on the pre-prompt through an AI application. The electronic device may generate the response to the user input using an AI-based response model (e.g., an LLM, an sLLM, and / or an LMM). For example, the electronic device may generate the response based on a generative AI module (or a generative AI model) including a response model (e.g., the generative AI system (1500) of FIG. 15). The electronic device may provide the generated response.

[0168] According to various embodiments, the order of at least some of the operations of FIG. 7 may be changed, some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 6 or FIGS. 8 to 10) may be added.

[0169]

[0170] Fig. 8 is a flowchart of a method for providing voice services in an electronic device according to one embodiment. Below, descriptions that overlap with those in Figs. 6 and 7 are briefly described or omitted.

[0171] According to one embodiment, in operation 805, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIG. 12, or a system (1500) of FIG. 15) may execute an AI application. For example, the AI ​​application may include an application that provides an AI-based voice service (e.g., an AI assistant or an AI (e.g., a language model)-based voice conversation).

[0172] According to one embodiment, at operation 810, the electronic device may determine whether a wake word for a customized AI conversation session has been received. For example, the electronic device may determine whether a user's voice input (e.g., a user utterance) including the wake word has been received. For example, the wake word may be registered for each customized AI conversation session. For example, the wake word may include the name of an object corresponding to the customized AI conversation session. According to one embodiment, the electronic device may perform operation 815 when the wake word for the customized AI conversation session has been received, and may perform operation 820 when the wake word has not been received.

[0173] In one embodiment, at operation 815, the electronic device may invoke a customized AI conversation session corresponding to the wake word.

[0174] According to one embodiment, at operation 820, the electronic device may determine whether the camera is activated. According to one embodiment, the electronic device may perform operation 825 if the camera is activated, and may perform operation 835 if the camera is deactivated.

[0175] According to one embodiment, in operation 825, the electronic device may acquire an image including at least one external object using a camera. The electronic device may recognize the object included in the image through image analysis. The electronic device may acquire information related to the object. For example, the electronic device may extract information related to the object from the image. The electronic device may acquire information related to the object using at least one sensor and / or communication circuit.

[0176] In one embodiment, at operation 830, the electronic device may invoke a customized AI conversation session corresponding to information associated with the object.

[0177] According to one embodiment, in operation 835, the electronic device may determine whether a specified condition is satisfied. For example, the specified condition may include, but is not limited to, whether the electronic device is in a specified operating state, whether the electronic device is performing a specified function, whether the electronic device is in a specified location, and / or whether an external electronic device connected to the electronic device is a specified device. For example, the specified condition may be a condition mapped to a pre-generated customized AI conversation session. According to one embodiment, the electronic device may perform operation 840 if the specified condition is satisfied, and may perform operation 855 if the specified condition is not satisfied.

[0178] According to one embodiment, in operation 840, the electronic device may recognize information of a connected external electronic device and / or an operational state of the electronic device. For example, the electronic device may obtain information related to a specified condition (e.g., an operational state of the electronic device, information of an external electronic device connected to the electronic device, a function being performed by the electronic device, a location of the electronic device).

[0179] In one embodiment, at operation 845, the electronic device may invoke a customized AI conversation session corresponding to the recognized information (e.g., information related to a specified condition). For example, the electronic device may invoke a customized AI conversation session mapped to the specified condition.

[0180] In one embodiment, at step 850, the electronic device may initiate a conversation in the invoked personalized AI conversation session. For example, the electronic device may receive a user's voice input based on the invoked personalized AI conversation session and provide a response to the user's voice input.

[0181] According to one embodiment, in operation 855, the electronic device may initiate a conversation in a basic AI conversation session. For example, the electronic device may invoke the basic AI conversation session, receive a user's voice input based on the basic AI conversation session, and provide a response to the user's voice input. For example, the basic AI conversation session may be an AI conversation session excluding a customized AI conversation session. For example, the basic AI conversation session may refer to an AI conversation session that is not mapped to a specific object and / or a specific condition. For example, the basic AI conversation session may refer to a conversation session that has default response properties (or initial response properties) of an AI conversation session created through an AI application without setting properties of a response corresponding to a specific object and / or a specific condition.

[0182] According to various embodiments, the order of at least some of the operations of FIG. 8 may be changed, some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIGS. 6, 7, 9, or 10) may be added.

[0183]

[0184] Figure 9 is a flowchart of a method for providing voice services in an electronic device according to one embodiment. Below, any descriptions that overlap with those in Figures 6 to 8 will be briefly explained or omitted.

[0185] According to one embodiment, in operation 910, an electronic device (e.g., the electronic device 100 of FIG. 1, the electronic device 200 of FIG. 2, the electronic device 1101 of FIG. 11, the user terminal 1201 of FIG. 12, or the system 1500 of FIG. 15) may initiate at least one customized AI conversation session. For example, the customized AI conversation session may mean a conversation session having customized properties for at least one object. For example, the electronic device may recognize at least one external object and generate a customized AI conversation session for each of the at least one external object based on information related to the recognized at least one external object. The electronic device may then call (or execute) the customized AI conversation session for the at least one previously generated external object.

[0186] In one embodiment, at operation 920, the electronic device may recognize registered object information. For example, the electronic device may recognize information related to an object corresponding to each customized AI conversation session.

[0187] According to one embodiment, at operation 930, the electronic device may recognize resource information of the electronic device. For example, the electronic device may recognize components (e.g., input devices and / or output devices) included in the electronic device. For example, if the object corresponding to the customized AI conversation session is an external electronic device, the electronic device may recognize the resource of the object.

[0188] According to one embodiment, in operation 940, the electronic device may determine an input path and / or an output path of a customized AI conversation session based on resource information (e.g., resource information of the electronic device and / or resource information of an object (external electronic device)). For example, the electronic device may determine the same or different input paths and / or output paths for each customized AI conversation session. For example, assume that the electronic device includes a microphone and a speaker, a first object is a non-electronic device (e.g., a doll), a second object is a Bluetooth speaker, and a third object is a smart speaker including a microphone and a speaker. For example, the electronic device may determine an input path of a first customized AI conversation session for the first object as the microphone of the electronic device, and an output path of the first customized AI conversation session as the speaker of the electronic device. For example, the electronic device may determine an input path of a second customized AI conversation session for the second object as the microphone of the electronic device, and an output path of the second customized AI conversation session as the speaker of the second object. For example, the electronic device may determine the input path of the third customized AI conversation session for the third object as the microphone of the third object, and determine the output path of the third customized AI conversation session as the speaker of the third object. For example, when a voice input is received through the microphone of the electronic device, the electronic device may provide the voice input to the first customized AI conversation session and / or the second customized AI conversation session based on the speech direction of the voice input, the user's gaze direction, head direction, and / or movement direction, and when the voice input is received through the microphone of the third object, the electronic device may provide the voice input to the third customized AI conversation session. For example, when a voice input is received through multiple input paths, the electronic device may share the voice input to multiple customized AI conversation sessions corresponding to the multiple input paths (i.e., provide the voice input to each customized AI conversation session).For example, when an electronic device generates a response to a voice input based on a first customized AI conversation session, the response may be output through a speaker of the electronic device; when a response to a voice input based on a second customized AI conversation session, the response may be output through a speaker of a second object; and when a response to a voice input based on a third customized AI conversation session, the response may be output through a speaker of a third object.

[0189] In one embodiment, at operation 950, the electronic device may display information about the determined input path and / or output path on the display. In one embodiment, the electronic device may change the input path and / or output path for each determined customized AI conversation session based on the user input.

[0190] According to various embodiments, the order of at least some of the operations of FIG. 9 may be changed, some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIGS. 6 to 8 or 10) may be added.

[0191]

[0192] Fig. 10 is a flowchart of a method for providing voice services in an electronic device according to one embodiment. Below, any descriptions that overlap with those in Figs. 6 to 9 will be briefly explained or omitted.

[0193] According to one embodiment, in operation 1010, an electronic device (e.g., electronic device (100) of FIG. 1, electronic device (200) of FIG. 2, electronic device (1101) of FIG. 11, user terminal (1201) of FIG. 12, system (1500) of FIG. 15) may initiate multiple customized AI conversation sessions.

[0194] According to one embodiment, in operation 1020, the electronic device may recognize the locations and resources of registered objects. For example, the electronic device may recognize the locations and / or resources of objects corresponding to each of a plurality of customized AI conversation sessions. For example, if the object is a non-electronic device, the electronic device may recognize that the object does not have resources.

[0195] According to one embodiment, at operation 1030, the electronic device may recognize a resource of the electronic device.

[0196] According to one embodiment, in operation 1040, the electronic device may recognize a relative position between a user (e.g., the electronic device) and an object. For example, the electronic device may recognize the relative position between the electronic device and each of the objects through image analysis acquired via at least one sensor or camera.

[0197] In one embodiment, at operation 1050, the electronic device may receive a user's voice input. For example, the user's voice input may include input (e.g., a prompt) for at least one customized AI conversation session.

[0198] According to one embodiment, in operation 1060, the electronic device may select at least one object from among objects corresponding to multiple customized AI conversation sessions based on at least one of a voice input, a speech direction, a gaze direction, information related to an object (e.g., location information and / or resource information), or information related to the electronic device (e.g., location information and / or resource information). For example, the electronic device may select at least one object based on a wake word (e.g., a name of an object) included in the voice input. The electronic device may recognize a speech direction of the voice input using multiple microphones included in the electronic device, and select at least one object based on the speech direction. The electronic device may select at least one object recognized from an image acquired through a camera. The electronic device may recognize the locations of objects using a sensor, and select at least one object located near the electronic device (user). When the electronic device receives a voice input through a specific audio input path (e.g., a microphone), the electronic device may select an object of a customized AI conversation session having the corresponding audio input path. The electronic device may select at least one object based on information related to the state of the electronic device (e.g., charging or in charging dock mode), the posture of the electronic device (e.g., whether a foldable terminal is folded), and an external object (external electronic device) connected to the electronic device.

[0199] In one embodiment, the electronic device may activate a customized AI conversation session corresponding to a selected object. For example, activating a customized AI conversation session may indicate a state for transmitting voice input to the customized AI conversation session and receiving a response to the voice input. For example, the electronic device may not transmit voice input to a disabled customized AI conversation session or may not obtain a response from a disabled customized AI conversation session.

[0200] According to one embodiment, in operation 1070, the electronic device may control the output sound of a response of an activated personalized AI conversation session based on the position of a selected object and / or the movement of a user (e.g., the electronic device). For example, assume that a first object (e.g., a doll) is located on the left side of the electronic device, a second object (e.g., a Bluetooth speaker) is located in front of the electronic device, and a third object (e.g., a smart speaker) is located on the right side of the electronic device. When outputting a response generated based on a personalized AI conversation session for the selected object, the electronic device may apply a spatial sound effect to the output response. For example, when the electronic device provides a response generated based on a first personalized AI conversation session for the first object, the electronic device may control the direction and / or volume of the output sound so that the user may perceive that the output sound of the response is heard from the left. For example, when the electronic device provides a response generated based on a second personalized AI conversation session for the second object, the electronic device may control the direction and / or volume of the output sound so that the user may perceive that the output sound of the response is heard from the front. For example, if an electronic device provides a response generated based on a third personalized AI conversation session for a third object, the direction and / or volume of the output sound may be controlled so that the user perceives the output sound of the response as coming from the right. For example, the electronic device may recognize the relative position of the selected object with respect to the user's gaze direction and / or head direction, and control the output sound of the response so that the response generated based on the personalized AI conversation session for the selected object is perceived as coming from the relative position of the selected object.

[0201] According to various embodiments, the order of at least some of the operations of FIG. 10 may be changed, some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIGS. 6 to 9) may be added.

[0202]

[0203] According to one embodiment of the present disclosure, a method for providing a voice service of an electronic device may include an operation of obtaining an image including at least one external object using a camera of the electronic device.

[0204] The method may include receiving a user input specifying a first object among the at least one external object.

[0205] The method may include an operation of recognizing information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device.

[0206] The method may include an operation of generating a first conversation session for the first object through at least one AI application that provides an artificial intelligence-based voice service included in the electronic device.

[0207] The method may include an action of determining an audio output path of the first conversation session based on information related to the electronic device or information related to the first object.

[0208] The method may include an operation of receiving a first voice input of a user related to the AI-based voice service through the microphone.

[0209] The method may include generating a response to the first voice input based on the first conversation session through the at least one AI application.

[0210] The method may include providing a response to the first voice input via an audio output path of the first conversation session.

[0211] The method may include receiving user input specifying a name for the first object.

[0212] The method may include an action of registering the name of the first object as a call word of the first conversation session.

[0213] The method may include an operation of determining an attribute of a response provided based on the first conversation session based on at least one of information related to the first object or a state of the electronic device.

[0214] The method may include an operation of recognizing at least one of a direction of the user's gaze, a direction of the user's head, a movement of the user, or a position of the first object through the at least one sensor.

[0215] The method may include an action of invoking the first conversation session based on at least one of a state of the electronic device, a gaze direction of the user, a head direction of the user, a movement of the user, or a position of the first object.

[0216] The method may include an operation of recognizing a speech direction corresponding to the first voice input through the microphone.

[0217] The method may include an operation of recognizing a relative position of the first object to the user based on at least one of the direction of speech, the position of the electronic device, information related to the first object, a direction of the user's gaze detected through the at least one sensor, or a direction of the user's head.

[0218] The method may include an operation of controlling at least one of an output size or direction of a sound corresponding to the response based on the relative position.

[0219] The method may include an action of generating a second conversation session for a second object via the at least one AI application.

[0220] The method may include an operation of determining an audio output path of the second conversation session based on at least one of information related to the electronic device or information related to the second object.

[0221] The method may include an operation of receiving a second voice input of a user related to the AI-based voice service through the microphone.

[0222] The method may include an operation of recognizing relative positions of the first object and the second object with respect to the user.

[0223] The method may include an operation of determining an object corresponding to the second voice input among the first object and the second object based on at least one of a speech direction corresponding to the second voice input, a direction of the user's gaze, a direction of the user's head, or the relative position.

[0224] The method may include an operation of generating a response to the second voice input based on a conversation session for an object corresponding to the second voice input among the first conversation session and the second conversation session.

[0225] The method may include providing a response to the second voice input via an audio output path of a conversation session for the determined object.

[0226] The method may include an operation of independently storing the contents of input and response corresponding to the first conversation session and the contents of input and response corresponding to the second conversation session.

[0227] The method may include an operation of generating a response to a third voice input following the second voice input based on a session for an object corresponding to the second voice input, by reflecting the second voice input and the response to the second voice input.

[0228] The method may include generating a response to the third voice input based on a conversation session for an object that does not correspond to the second voice input, without reflecting the second voice input and the response to the second voice input.

[0229] The method may include an operation of determining, based on user input, whether to share at least some of the input and response corresponding to the first conversation session with the second conversation session.

[0230] The method may include receiving a third voice input of a user related to the second object through the microphone.

[0231] The method may include an operation of generating a response to the third voice input by reflecting at least some of the shared input and response based on the second conversation session, if at least some of the input and response corresponding to the first conversation session is shared with the second conversation session.

[0232] According to one embodiment of the present disclosure, a storage medium may store instructions and / or a program that, when executed by at least one processor of an electronic device, causes the electronic device to obtain an image including at least one external object using a camera of the electronic device, receive a user input specifying a first object among the at least one external object, recognize information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device, create a first conversation session for the first object through at least one AI application that provides an artificial intelligence-based voice service included in the electronic device, determine an audio output path of the first conversation session based on information related to the electronic device or information related to the first object, receive a first voice input of a user related to the AI-based voice service through the microphone, create a response to the first voice input based on the first conversation session through the at least one AI application, and provide the response to the first voice input through the audio output path of the first conversation session.

[0233] According to various embodiments of the present disclosure, by providing an AI voice service based on a physical object, it is possible to provide diverse and intuitive interactions between a user and an AI assistant, thereby improving the user experience and increasing the usability of the AI ​​voice service. According to various embodiments of the present disclosure, by creating and providing customized AI conversation sessions (customized AI assistants) for each of a plurality of physical objects (and / or surrounding situations, or states of electronic devices), it is possible to provide a variety of personalized and / or customized voice services to the user. According to various embodiments of the present disclosure, by setting an audio path (e.g., an audio input path and / or an audio output path) for providing an AI voice service (e.g., an AI conversation session) based on the location, type, and / or state of an object, the location of the user, and / or the location and / or state of the electronic device, it is possible to provide the user with a user experience similar to interacting with an actual physical object.

[0234]

[0235] FIG. 11 is a block diagram of an electronic device (1101) within a network environment (1100) according to various embodiments. Referring to FIG. 11, in the network environment (1100), the electronic device (1101) may communicate with the electronic device (1102) via a first network (1198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1104) or the server (1108) via a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) via the server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), a memory (1130), an input module (1150), an audio output module (1155), a display module (1160), an audio module (1170), a sensor module (1176), an interface (1177), a connection terminal (1178), a haptic module (1179), a camera module (1180), a power management module (1188), a battery (1189), a communication module (1190), a subscriber identification module (1196), or an antenna module (1197). In some embodiments, the electronic device (1101) may omit at least one of these components (e.g., the connection terminal (1178)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).

[0236] The processor (1120) may, for example, execute software (e.g., a program (1140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1120) may store commands or data received from other components (e.g., a sensor module (1176) or a communication module (1190)) in a volatile memory (1132), process the commands or data stored in the volatile memory (1132), and store result data in a non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1121). For example, when the electronic device (1101) includes the main processor (1121) and the auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a given function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as a part thereof.

[0237] The auxiliary processor (1123) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1160), the sensor module (1176), or the communication module (1190)) of the electronic device (1101), for example, on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1180) or a communication module (1190)). In one embodiment, the auxiliary processor (1123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0238] The memory (1130) can store various data used by at least one component (e.g., the processor (1120) or the sensor module (1176)) of the electronic device (1101). The data can include, for example, software (e.g., the program (1140)) and input data or output data for commands related thereto. The memory (1130) can include a volatile memory (1132) or a non-volatile memory (1134).

[0239] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).

[0240] The input module (1150) can receive commands or data to be used in a component of the electronic device (1101) (e.g., a processor (1120)) from an external source (e.g., a user) of the electronic device (1101). The input module (1150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0241] The audio output module (1155) can output audio signals to the outside of the electronic device (1101). The audio output module (1155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0242] The display module (1160) can visually provide information to an external party (e.g., a user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0243] The audio module (1170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150), output sound through the sound output module (1155), or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1101).

[0244] The sensor module (1176) can detect the operating status (e.g., power or temperature) of the electronic device (1101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0245] The interface (1177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1101) with an external electronic device (e.g., the electronic device (1102)). In one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0246] The connection terminal (1178) may include a connector through which the electronic device (1101) may be physically connected to an external electronic device (e.g., the electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0247] The haptic module (1179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0248] The camera module (1180) can capture still images and videos. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0249] The power management module (1188) can manage power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0250] A battery (1189) may power at least one component of the electronic device (1101). In one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0251] The communication module (1190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may operate independently from the processor (1120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, a corresponding communication module can communicate with an external electronic device (1104) via a first network (1198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can verify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1196).

[0252] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1192) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1192) may support various requirements specified in the electronic device (1101), an external electronic device (e.g., the electronic device (1104)), or a network system (e.g., the second network (1199)). According to one embodiment, the wireless communication module (1192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0253] The antenna module (1197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1198) or the second network (1199), may be selected from the plurality of antennas by, for example, the communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1197).

[0254] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0255] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0256] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) via a server (1108) connected to a second network (1199). Each of the external electronic devices (1102 or 104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations executed in the electronic device (1101) may be executed in one or more of the external electronic devices (1102, 1104, or 1108). For example, when the electronic device (1101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server utilizing machine learning and / or a neural network.In one embodiment, an external electronic device (1104) or server (1108) may be included in the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0257]

[0258] FIG. 12 is a block diagram illustrating an integrated intelligence system according to one embodiment.

[0259] Referring to FIG. 12, an integrated intelligence system of one embodiment may include a user terminal (1201), an intelligent server (1300), and a service server (1400).

[0260] A user terminal (1201) of one embodiment (e.g., electronic device (1101) of FIG. 11) may be a terminal device (or electronic device) that can connect to the Internet, and may be, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a television, white goods, a wearable device, an HMD (head mounted device), or a smart speaker.

[0261] According to the illustrated embodiment, the user terminal (1201) may include a communication interface (1290), a microphone (1270), a speaker (1255), a display (1260), a memory (1230), and / or a processor (1220). The above-listed components may be operatively or electrically connected to each other.

[0262] A communication interface (1290) (e.g., a communication module (1190) of FIG. 11) may be configured to be connected to an external device and transmit and receive data. A microphone (1270) (e.g., an audio module (1170) of FIG. 11) may receive sound (e.g., a user's speech) and convert it into an electrical signal. A speaker (1255) (e.g., an audio output module (1155) of FIG. 11) may output an electrical signal as sound (e.g., a voice). A display (1260) (e.g., a display module (1160) of FIG. 11) may be configured to display an image or video. In one embodiment, the display (1260) may also display a graphical user interface (GUI) of an app (or application program) being executed.

[0263] The memory (1230) of one embodiment (e.g., the memory (1130) of FIG. 11) may store a client module (1231), a software development kit (SDK) (1233), and multiple applications. The client module (1231) and the SDK (1233) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (1231) or the SDK (1233) may constitute a framework for processing voice input.

[0264] The above-described multiple applications (e.g., 1235a, 1235b) may be programs for performing a designated function. According to one embodiment, the multiple applications may include a first app (1235a) and / or a second app (1235b). According to one embodiment, each of the multiple applications may include a plurality of operations for performing a designated function. For example, the applications may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the multiple applications may be executed by the processor (1220) to sequentially execute at least some of the multiple operations.

[0265] The processor (1220) of one embodiment can control the overall operation of the user terminal (1201). For example, the processor (1220) can be electrically connected to a communication interface (1290), a microphone (1270), a speaker (1255), and a display (1260) to perform a specified operation. For example, the processor (1220) can include at least one processor.

[0266] The processor (1220) of one embodiment may also execute a program stored in the memory (1230) to perform a designated function. For example, the processor (1220) may execute at least one of the client module (1231) or the SDK (1233) to perform the following operations for processing voice input. The processor (1220) may control the operations of multiple applications, for example, through the SDK (1233). The following operations described as operations of the client module (1231) or the SDK (1233) may be operations performed by the execution of the processor (1220).

[0267] The client module (1231) of one embodiment can receive a voice input. For example, the client module (1231) can receive a voice signal corresponding to a user utterance detected through a microphone (1270). The client module (1231) can transmit the received voice input (e.g., a voice signal) to the intelligent server (1300). The client module (1231) can transmit status information of the user terminal (1201) to the intelligent server (1300) together with the received voice input. The status information can be, for example, execution status information of an app.

[0268] The client module (1231) of one embodiment can receive a result corresponding to the received voice input from the intelligent server (1300). For example, the client module (1231) can receive a result corresponding to the received voice input if the intelligent server (1300) can produce a result corresponding to the received voice input. The client module (1231) can display the received result on the display (1260).

[0269] In one embodiment, the client module (1231) can receive a plan corresponding to the received voice input. The client module (1231) can display the results of executing multiple operations of the app according to the plan on the display (1260). For example, the client module (1231) can sequentially display the results of executing multiple operations on the display. The user terminal (1201) can, for example, display only some of the results of executing multiple operations (e.g., the result of the last operation) on the display.

[0270] According to one embodiment, the client module (1231) may receive a request from the intelligent server (1300) to obtain information necessary to produce a result corresponding to a voice input. According to one embodiment, the client module (1231) may transmit the necessary information to the intelligent server (1300) in response to the request.

[0271] The client module (1231) of one embodiment can transmit result information of executing multiple operations according to a plan to the intelligent server (1300). The intelligent server (1300) can use the result information to confirm that the received voice input has been processed correctly.

[0272] The client module (1231) of one embodiment may include a voice recognition module. According to one embodiment, the client module (1231) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (1231) may execute an intelligent app for processing voice inputs by performing organic actions in response to a specified voice input (e.g., "Wake up!").

[0273] An intelligent server (1300) of one embodiment may receive information related to a user voice input from a user terminal (1201) via a network (1299) (e.g., the first network (1198) and / or the second network (1199) of FIG. 11). According to one embodiment, the intelligent server (1300) may convert data related to the received voice input into text data. According to one embodiment, the intelligent server (1300) may generate at least one plan for performing a task corresponding to the user voice input based on the text data.

[0274] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI ​​system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) and / or a recurrent neural network (RNN)), or a combination of the above or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI ​​system may select at least one plan from a plurality of predefined plans.

[0275] An intelligent server (1300) of one embodiment may transmit results according to a generated plan to a user terminal (1201), or transmit the generated plan to the user terminal (1201). According to one embodiment, the user terminal (1201) may display results according to the plan on a display. According to one embodiment, the user terminal (1201) may display results of executing an operation according to the plan on a display.

[0276] An intelligent server (1300) of one embodiment may include a front end (1310), a natural language platform (1320), a capsule database (1330), an execution engine (1340), an end user interface (1350), a management platform (1360), a big data platform (1370), or an analytic platform (1380).

[0277] The front end (1310) of one embodiment can receive a voice input received by the user terminal (1201) from the user terminal (1201). The front end (1310) can transmit a response corresponding to the voice input to the user terminal (1201).

[0278] According to one embodiment, the natural language platform (1320) may include an automatic speech recognition module (ASR module) (1321), a natural language understanding module (NLU module) (1323), a planner module (1325), a natural language generator module (NLG module) (1327), and / or a text to speech module (TTS module) (1329).

[0279] An automatic speech recognition module (1321) of one embodiment can convert a voice input received from a user terminal (1201) into text data. A natural language understanding module (1323) of one embodiment can use the text data of the voice input to determine the user's intention. For example, the natural language understanding module (1323) can perform syntactic analysis and / or semantic analysis to determine the user's intention. The natural language understanding module (1323) of one embodiment can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the voice input, and can match the meaning of the determined words to the intention to determine the user's intention.

[0280] In one embodiment, the planner module (1325) can generate a plan using the intent and parameters determined by the natural language understanding module (1323). According to one embodiment, the planner module (1325) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (1325) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (1325) can determine parameters necessary to execute the determined plurality of operations or result values ​​output by the execution of the plurality of operations. The parameters and the result values ​​can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and / or a plurality of concepts determined by the user's intent. The planner module (1325) can determine the relationship between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (1325) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (1325) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (1325) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (1325) can generate the plan using information stored in a capsule database (1330) in which a set of relationships between concepts and actions is stored.

[0281] The natural language generation module (1327) of one embodiment can convert specified information into text format. The information converted into text format may be in the form of natural language speech. The text-to-speech conversion module (1329) of one embodiment can convert information in text format into information in speech format.

[0282] According to one embodiment, some or all of the functions of the natural language platform (1320) may also be implemented in the user terminal (1201). For example, the user terminal (1201) may include an automatic speech recognition module and / or a natural language understanding module. After the user terminal (1201) recognizes a user voice command, it may transmit text information corresponding to the recognized voice command to the intelligent server (1300). For example, the user terminal (1201) may include a text-to-speech conversion module. The user terminal (1201) may receive text information from the intelligent server (1300) and output the received text information as voice.

[0283] The capsule database (1330) may store information on the relationships between multiple concepts and actions corresponding to multiple domains. According to one embodiment, a capsule may include multiple action objects (or action information) and / or concept objects (or concept information) included in a plan. According to one embodiment, the capsule database (1330) may store multiple capsules in the form of a concept action network (CAN). According to one embodiment, the multiple capsules may be stored in a function registry included in the capsule database (1330).

[0284] The capsule database (1330) may include a strategy registry that stores strategy information required when determining a plan corresponding to a voice input. The strategy information may include reference information for determining one plan when there are multiple plans corresponding to a voice input. According to one embodiment, the capsule database (1330) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database (1330) may include a layout registry that stores layout information of information output through the user terminal (1201). According to one embodiment, the capsule database (1330) may include a vocabulary registry that stores vocabulary information included in capsule information. According to one embodiment, the capsule database (1330) may include a dialog registry that stores information on dialogue (or interaction) with a user. The capsule database (1330) may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans.The developer tool may include a dialog editor that creates a dialogue with the user. The developer tool may also include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on the currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database (1330) may also be implemented within the user terminal (1201).

[0285] The execution engine (1340) of one embodiment can produce a result using the generated plan. The end user interface (1350) can transmit the produced result to the user terminal (1201). Accordingly, the user terminal (1201) can receive the result and provide the received result to the user. The management platform (1360) of one embodiment can manage information used in the intelligent server (1300). The big data platform (1370) of one embodiment can collect user data. The analysis platform (1380) of one embodiment can manage the quality of service (QoS) of the intelligent server (1300). For example, the analysis platform (1380) can manage the components and processing speed (or efficiency) of the intelligent server (1300).

[0286] In one embodiment, a service server (1400) may provide a designated service (e.g., food ordering or hotel reservation) to a user terminal (1201). According to one embodiment, the service server (1400) may be a server operated by a third party. In one embodiment, the service server (1400) may provide information for generating a plan corresponding to a received voice input to an intelligent server (1300). The provided information may be stored in a capsule database (1330). In addition, the service server (1400) may provide result information according to the plan to the intelligent server (1300). The service server (1400) may communicate with the intelligent server (1300) and / or the user terminal (1201) via a network (1299). The service server (1400) may communicate with the intelligent server (1300) via a separate connection. Although the service server (1400) is depicted as a single server in FIG. 12, the embodiments of this document are not limited thereto. At least one of the services (1401, 1402, and 1403) of the service server (1400) may be implemented as a separate server.

[0287] In the integrated intelligence system described above, the user terminal (1201) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.

[0288] In one embodiment, the user terminal (1201) may provide a voice recognition service through an intelligent app (or voice recognition app) stored internally. In this case, for example, the user terminal (1201) may recognize a user utterance or voice input received through the microphone and provide the user with a service corresponding to the recognized voice input.

[0289] In one embodiment, the user terminal (1201) may perform a designated action based on the received voice input, either alone or in conjunction with the intelligent server and / or service server. For example, the user terminal (1201) may execute an app corresponding to the received voice input and perform a designated action through the executed app.

[0290] In one embodiment, when a user terminal (1201) provides a service together with an intelligent server (1300) and / or a service server, the user terminal may detect user speech using the microphone (1270) and generate a signal (or voice data) corresponding to the detected user speech. The user terminal may transmit the voice data to the intelligent server (1300) using a communication interface (1290).

[0291] According to one embodiment, an intelligent server (1300) may generate a plan for performing a task corresponding to a voice input received from a user terminal (1201), or a result of performing an operation according to the plan, in response to the voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to the user's voice input and / or a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations or result values ​​output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and / or the plurality of concepts.

[0292] In one embodiment, the user terminal (1201) can receive the response using the communication interface (1290). The user terminal (1201) can output a voice signal generated within the user terminal (1201) to the outside using the speaker (1255), or can output an image generated within the user terminal (1201) to the outside using the display (1260).

[0293] FIG. 13 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.

[0294] The capsule database (e.g., capsule database (1330)) of the above intelligent server (1300) can store capsules in the form of a CAN (concept action network). The capsule database can store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network).

[0295] The capsule database may store a plurality of capsules (e.g., Capsule A (1331), Capsule B (1334)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., Capsule A (1331)) may correspond to one domain (e.g., location (geo), application). In addition, one capsule may correspond to at least one service provider's capsule (e.g., CP 1 (1332), CP 2 (1333), CP3 (1335), and / or CP4 (1336)) for performing a function for a domain related to the capsule. According to one embodiment, one capsule may include at least one operation (1330a) and at least one concept (1330b) for performing a specified function.

[0296] The natural language platform (1320) can generate a plan for performing a task corresponding to a received voice input using capsules stored in the capsule database (1330). For example, the planner module (1325) of the natural language platform can generate a plan using capsules stored in the capsule database. For example, a plan (1337) can be generated using the actions (1331a, 1332a) and concepts (1331b, 1332b) of capsule A (1330) and the action (1334a) and concept (1334b) of capsule B (1334).

[0297]

[0298] FIG. 14 is a diagram showing a screen in which a user terminal processes voice input received through an intelligent app according to one embodiment.

[0299] The user terminal (1201) can execute an intelligent app to process user input through an intelligent server (1300).

[0300]

[0301] According to one embodiment, in the first screen (1410), when the user terminal (1201) recognizes a designated voice input (e.g., wake up!) or receives an input via a hardware key (e.g., a dedicated hardware key), the user terminal (1201) may execute an intelligent app for processing the voice input. For example, the user terminal (1201) may execute the intelligent app while the schedule app is running. According to one embodiment, the user terminal (1201) may display an object (e.g., an icon) (1411) corresponding to the intelligent app on the display (1260). According to one embodiment, the user terminal (1201) may receive a voice input by a user's speech. For example, the user terminal (1201) may receive a voice input such as "Tell me my schedule for this week!" According to one embodiment, the user terminal (1201) can display a user interface (UI) (1213) (e.g., an input window) of an intelligent app on which text data of a received voice input is displayed.

[0302] According to one embodiment, in the second screen (1415), the user terminal (1201) may display a result corresponding to the received voice input on the display. For example, the user terminal (1201) may receive a plan corresponding to the received user input and display "This Week's Schedule" on the display according to the plan.

[0303]

[0304] FIG. 15 illustrates a generative artificial intelligence system according to one embodiment.

[0305] Referring to FIG. 15, in a generative artificial intelligence system (1500), a user question / response interface (1510) can receive user input. The user input may be in the form of natural language, images, and / or videos. Additionally, contextual information may also be transmitted when the user input is transmitted. The contextual information may include various additional information at the time of the user input. For example, information on the application currently being used by the user or information on the user's location. Furthermore, the user input may also be in a form that combines the aforementioned natural language, images, sounds, and contextual information. Furthermore, the user input may also be in a form other than natural language, such as selecting a menu.

[0306] According to one embodiment, the user question / response interface (1510) may output the results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user question / response interface (1510) may output the results of a generative artificial intelligence system (1500) to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user.

[0307] According to one embodiment, the artificial intelligence framework (1520) can receive user input and coordinate and control each component or module necessary to perform the user's intention based on the user's query.

[0308] In one embodiment, user input received from the user question / response interface (1510) may be transmitted to a prompt design module (1521). The prompt design module (1521) may be used to generate prompts suitable for inputting the user input into a large language model (LMM) or a large multi-modal model (LMM). The prompt design module (1521) may be an artificial intelligence component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design module (1521) may access a database (1530) containing user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and may transmit the generated prompts to the LLM or LMM.

[0309] According to one embodiment, the API / plug-in management module (1522) may communicate with external information when there is a request for additional information when passing user input as input to the generative artificial intelligence model (1550). The API / plug-in management module (1522) may establish a channel for communicating with the outside of the artificial intelligence interface through an application programming interface (API) and may enable access to various data sources through the established channel. In addition, the API / plug-in management module (1522) may request an action through the API that ultimately performs the user input, rather than an intermediate result, if the application / service module (1540) needs to perform the action. Information obtained from the outside may be used to generate a prompt in the prompt design module (1521) together with the user input or may be passed as an input to the generative artificial intelligence model (1550).

[0310] In one embodiment, the transformation module (1523) can fine-tune the output from the generative artificial intelligence model (1550). For example, the transformation module (1523) can verify whether the content generated through the LLM and / or LMM is irrelevant, biased, or harmful. In addition, the transformation module (1523) can determine to what extent the content matches the user's desired result and, if necessary, perform additional processing. The transformation module (1523) can additionally configure and provide the user with hints to avoid undesired output.

[0311] According to one embodiment, a generative artificial intelligence model (1550) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative artificial intelligence model (1550) may include a model that generates images and / or a model that generates language. Representative models that generate images include a generative adversarial network (GAN) and a variational autoencoder (VAE), and examples include a diffusion-based generative model that uses a VAE and a transformer structure. A model that generates language is a model that is trained to statistically output the most appropriate output value based on an input value, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM that can recognize various types of data input such as text, images, and voice and generate new data corresponding thereto.

[0312]

[0313] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0314] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0315] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0316] Various embodiments of the present document may be implemented as software (e.g., a program (1140)) including one or more instructions stored in a storage medium (e.g., an internal memory (1136) or an external memory (1138)) readable by a machine (e.g., an electronic device (1101)). For example, a processor (e.g., a processor (1120)) of the machine (e.g., an electronic device (1101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0317] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0318] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

In electronic devices, camera; mike; At least one sensor; communication circuit; A memory storing at least one AI application that provides artificial intelligence (AI)-based voice services; and Contains at least one processor, The above memory, when executed by the at least one processor, causes the electronic device to: Obtaining an image including at least one external object using the above camera, Receiving a user input specifying a first object among the at least one external object, Recognize information related to the first object using at least one of the image, the at least one sensor, or the communication circuit, Generating a first conversation session for the first object through at least one AI application; Determine the audio output path of the first conversation session based on information related to the electronic device or information related to the first object, Receive a first voice input of a user related to the AI-based voice service through the microphone; Generating a response to the first voice input based on the first conversation session through the at least one AI application; An electronic device storing instructions for providing a response to said first voice input through an audio output path of said first conversation session. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Receiving user input specifying a name for the first object; An electronic device that registers the name of the first object as a call word of the first conversation session. In claim 2, Information related to the first object includes at least one of the type of the first object, identification information of the first object, the location of the first object, the shape of the first object, or the name of the first object. An electronic device, wherein information related to the electronic device includes at least one of the status of a function or operation being performed by the electronic device, the status of a communication connection of the electronic device, or information about a component included in the electronic device. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device that determines properties of a response to be provided based on the first conversation session based on at least one of information related to the first object or a state of the electronic device. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Recognize at least one of the user's gaze direction, the user's head direction, the user's movement, or the position of the first object through at least one sensor, An electronic device that calls the first conversation session based on at least one of a state of the electronic device, a direction of the user's gaze, a direction of the user's head, a movement of the user, or a position of the first object. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Recognize the speech direction corresponding to the first voice input through the microphone, Recognize the relative position of the first object to the user based on at least one of the direction of the speech, the position of the electronic device, information related to the first object, the direction of the user's gaze detected through the at least one sensor, or the direction of the user's head, An electronic device that controls at least one of the output size or direction of a sound corresponding to the response based on the relative position. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Generating a second conversation session for a second object through at least one AI application; Determine the audio output path of the second conversation session based on at least one of information related to the electronic device or information related to the second object; Receive a second voice input from a user related to the AI-based voice service through the microphone; Recognize the relative positions of the first object and the second object with respect to the user, Based on at least one of the speech direction corresponding to the second voice input, the user's gaze direction, the user's head direction, or the relative position, an object corresponding to the second voice input among the first object and the second object is determined, Generate a response to the second voice input based on a conversation session for an object corresponding to the second voice input among the first conversation session and the second conversation session, An electronic device that provides a response to said second voice input through an audio output path of a conversation session for said determined object. In claim 7, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device that independently stores the contents of input and response corresponding to the first conversation session and the contents of input and response corresponding to the second conversation session. In claim 8, The above instructions, when executed by the at least one processor, cause the electronic device to: Based on the user input, determine whether to share at least some of the input and response corresponding to the first conversation session to the second conversation session; Receive a third voice input of the user related to the second object through the microphone, An electronic device that generates a response to the third voice input by reflecting at least some of the shared input and response based on the second conversation session, if at least some of the input and response corresponding to the first conversation session is shared in the second conversation session. In claim 8, The above instructions, when executed by the at least one processor, cause the electronic device to: When generating a response to a third voice input following the second voice input based on a session for an object corresponding to the second voice input, a response to the third voice input is generated by reflecting the second voice input and the response to the second voice input, An electronic device that generates a response to the third voice input based on a conversation session for an object that does not correspond to the second voice input, without reflecting the second voice input and the response to the second voice input. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Recognize context information including at least one of the operating status of the electronic device or information related to an external electronic device connected to the electronic device, Generating a third conversation session to provide the AI-based voice service corresponding to the context information through the at least one AI application; An electronic device that determines response properties of an AI-based voice service provided through the third conversation session based on the above context information. In a method for providing voice service of an electronic device, An action of obtaining an image including at least one external object using a camera of the electronic device; An action of receiving a user input specifying a first object among the at least one external object; An operation of recognizing information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device; An action of creating a first conversation session for the first object through at least one AI application providing an artificial intelligence-based voice service included in the electronic device; An action of determining an audio output path of the first conversation session based on information related to the electronic device or information related to the first object; An action of receiving a first voice input of a user related to the AI-based voice service through the microphone; An operation of generating a response to the first voice input based on the first conversation session through at least one AI application; and A method comprising providing a response to said first voice input through an audio output path of said first conversation session. In claim 12, An action of receiving user input specifying a name of the first object; and A method comprising an action of registering the name of the first object as a call word of the first conversation session. In claim 12, A method comprising an action of determining properties of a response provided based on the first conversation session based on at least one of information related to the first object or a state of the electronic device. In a non-transitory storage medium that stores instructions, The above instructions, when executed by at least one processor of an electronic device, cause the electronic device to: Obtaining an image including at least one external object using a camera of the electronic device, Receiving a user input specifying a first object among the at least one external object, Recognizing information related to the first object using at least one of the image, at least one sensor of the electronic device, or a communication circuit of the electronic device, Generating a first conversation session for the first object through at least one AI application providing an artificial intelligence-based voice service included in the electronic device; Determine the audio output path of the first conversation session based on information related to the electronic device or information related to the first object, Receive a first voice input of a user related to the AI-based voice service through the microphone; Generating a response to the first voice input based on the first conversation session through the at least one AI application; A storage medium that provides a response to the first voice input through an audio output path of the first conversation session.

Citation Information

Patent Citations

  • Electronic apparatus

    JP2012029107A

  • Airbag cushion Protector

    KR1020240157415A

  • Systems, methods, and devices for image-responsive automated assistants

    KR102297392B1

  • A method and a system for providing a service to conduct a conversation with a virtual person simulating the deceased

    KR102407132B1

  • Data processing method and apparatus, device, and readable storage medium

    US20230362333A1