Electronic device for providing response to user command and control method thereof
The electronic device uses neural network models to simulate human-like conversation by outputting sounds during processing, addressing the mechanical response issue and enhancing user interaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-13
- Publication Date
- 2026-05-15
AI Technical Summary
Electronic devices providing responses to user utterances through neural network models lack the ability to simulate human-like conversation, resulting in a mechanical and unsatisfying user experience.
An electronic device equipped with processors and a speaker that utilize multiple neural network models to identify and output sounds during the neural network operation, simulating a conversational response by outputting sounds before and after the neural network model provides an answer, thereby reducing perceived wait times and enhancing user interaction.
The solution provides a more natural and engaging user experience by minimizing wait times between user input and response, creating a conversational feel with the electronic device.
Smart Images

Figure KR2025016058_15052026_PF_FP_ABST
Abstract
Description
Electronic device for providing a response to a user command and a method for controlling the same
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more specifically, to an electronic device and a method for controlling the same that provides a response to a user command.
[0002] Thanks to advancements in electronic technology, various types of electronic devices are being developed. In particular, user convenience is improving as electronic devices that provide responses to user utterances through neural network models are becoming more widespread.
[0003] However, electronic devices that provide responses to user utterances through neural network models merely provide mechanical answers, so they could not provide the user with the feeling of conversing with a person.
[0004] According to one embodiment of the present disclosure for achieving the above objectives, an electronic device comprises one or more processors including a memory for storing a plurality of sounds and instructions, a speaker, and a processing circuitry, wherein when the instructions are executed individually or collectively by the one or more processors, when a user command is obtained, at least one sound among the plurality of sounds is identified, and a first neural network model is requested to perform an operation on the user command, and while the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and when an answer based on the operation of the first neural network model is obtained, a sound corresponding to the answer is output through the speaker.
[0005] Additionally, the system further includes a communication interface, wherein the instructions control the communication interface to transmit the user command to a server performing operations of the first neural network model when executed individually or collectively by the one or more processors, and output the at least one sound through the speaker after identifying the at least one sound until the sound corresponding to the answer is output, and when the answer is received from the server through the communication interface, the sound corresponding to the answer can be output through the speaker.
[0006] And, the memory further stores the first neural network model, and when the instructions are executed individually or collectively by the one or more processors, the first neural network model performs operations on the user command, outputs at least one sound through the speaker while the operations of the first neural network model are being performed, and when the answer is obtained, outputs a sound corresponding to the answer through the speaker.
[0007] Additionally, the memory further stores a second neural network model, and when the instructions are executed individually or collectively by the one or more processors, the second neural network model performs operations on the user command to identify the estimated operation time until the answer based on the first neural network model is provided, and can identify at least one sound among the plurality of sounds based on the estimated operation time.
[0008] And, the plurality of sounds include sounds with different playback times, and when the instructions are executed individually or collectively by the one or more processors, the sound having the playback time closest to the expected operation time among the plurality of sounds can be identified.
[0009] Additionally, when the above instructions are executed individually or collectively by the one or more processors, if the user command is obtained, a sound corresponding to a part of the user command can be identified as the at least one sound.
[0010] And, when the above instructions are executed individually or collectively by the one or more processors, the at least one sound is output through the speaker while the operation of the first neural network model is performed, and when the answer is obtained, the sound corresponding to the remainder of the answer excluding part of the user command can be output through the speaker.
[0011] Additionally, when the above instructions are executed individually or collectively by the one or more processors, a prompt is given to output a part of the user command as sound and to request the execution of an operation of the first neural network model on the user command, and while the operation of the first neural network model is being performed, at least one sound is output through the speaker, and when the answer is obtained, a sound corresponding to the answer is output through the speaker, and the answer may be an answer from which a part of the user command is excluded.
[0012] And, further comprising a microphone, when the instructions are executed individually or collectively by the one or more processors, if the user utterance is received through the microphone, the user command is obtained based on the user utterance, and the at least one sound is identified based on at least one of the content or tone of the user command.
[0013] Additionally, when the above instructions are executed individually or collectively by the one or more processors, the plurality of sounds can be updated based on the tone of the user command.
[0014] Meanwhile, according to one embodiment of the present disclosure, a control method for an electronic device may include the steps of: identifying at least one sound among a plurality of sounds when a user command is obtained; requesting the execution of a first neural network model operation for the user command; outputting the at least one sound through a speaker of the electronic device while the operation of the first neural network model is being performed; and, when an answer based on the operation of the first neural network model is obtained, outputting a sound corresponding to the answer through the speaker.
[0015] Additionally, the requesting step transmits the user command to a server that performs operations of the first neural network model, and the step of outputting at least one sound through the speaker outputs the at least one sound through the speaker after identifying the at least one sound and before outputting the sound corresponding to the answer, and the step of outputting the sound corresponding to the answer through the speaker can output the sound corresponding to the answer through the speaker when the answer is received from the server.
[0016] And, the requesting step can perform operations of the first neural network model for the user command.
[0017] Additionally, the identifying step may perform an operation of a second neural network model on the user command to identify the estimated operation time until the answer based on the operation of the first neural network model is provided, and identify at least one sound among the plurality of sounds based on the estimated operation time.
[0018] And, the plurality of sounds include sounds with different playback times, and the identifying step can identify the sound among the plurality of sounds that has a playback time closest to the expected calculation time.
[0019] Additionally, when the user command is obtained, the identifying step can identify a sound corresponding to a part of the user command as the at least one sound.
[0020] And, in the step of outputting a sound corresponding to the above answer through the speaker, when the above answer is obtained, a sound corresponding to the remainder of the above answer, excluding a part of the user command, can be output through the speaker.
[0021] Additionally, the requesting step requests a prompt to output a part of the user command as sound and requests the execution of an operation of the first neural network model on the user command, and the step of outputting a sound corresponding to the answer through the speaker outputs a sound corresponding to the answer through the speaker when the answer is obtained, and the answer may be an answer from which a part of the user command has been excluded.
[0022] And, the identification step can obtain the user command based on the user utterance when the user utterance is received through a microphone included in the electronic device, and identify the at least one sound based on at least one of the content or tone of the user command.
[0023] Additionally, the method may further include a step of updating the plurality of sounds based on the tone of the user command.
[0024] FIG. 1 is a block diagram showing an electronic system according to one embodiment of the present disclosure.
[0025] FIG. 2 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.
[0026] FIG. 3 is a block diagram showing the detailed configuration of an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 4 is a drawing for explaining a generative artificial intelligence system (400) according to one embodiment of the present disclosure.
[0028] FIG. 5 is a flowchart illustrating an operation to identify at least one sound using an estimated computation time according to one embodiment of the present disclosure.
[0029] FIG. 6 is a flowchart illustrating an operation for identifying a sound corresponding to a part of a user command as at least one sound according to one embodiment of the present disclosure.
[0030] FIG. 7 is a drawing for explaining a second neural network model according to one embodiment of the present disclosure.
[0031] FIG. 8 is a drawing for explaining a plurality of sounds according to one embodiment of the present disclosure.
[0032] FIG. 9 is a drawing for explaining the operation of outputting a sound corresponding to a part of a user command as at least one sound according to one embodiment of the present disclosure.
[0033] FIG. 10 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0034] The purpose of the present disclosure is to provide an electronic device and a method for controlling the same for providing a user with the sensation of conversing with a person in the process of providing a response to a user's utterance through a neural network model.
[0035] The present disclosure will be described in detail below with reference to the attached drawings.
[0036] The terms used in the embodiments of this disclosure have been selected to be as widely used as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant explanatory section of this disclosure. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.
[0037] In this specification, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, operations, or components such as parts) and do not exclude the presence of additional features.
[0038] The expression "at least one of A or / and B" should be understood as representing either "A" or "B" or "A and B".
[0039] Expressions such as "first," "second," "first," or "second" used in this specification may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0040] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "consisting of" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0041] In this specification, the term "user" may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).
[0042] Various embodiments of the present disclosure will be described in more detail below with reference to the attached drawings.
[0043] FIG. 1 is a block diagram showing an electronic system (1000) according to one embodiment of the present disclosure. As shown in FIG. 1, the electronic system (1000) may include an electronic device (100) and a server (200).
[0044] The electronic device (100) is a device that provides a response to a user command and can be implemented as a device such as a smartphone, tablet PC, smart glasses, smart watch, speaker, soundbar, desktop PC, laptop, TV, projector, etc. However, it is not limited thereto, and the electronic device (100) may be any device that can provide a response to a user command.
[0045] The electronic device (100) can acquire a user command, transmit the user command to a server, receive a response to the user command from the server, and provide the received response to the user.
[0046] The electronic device (100) can output at least one sound until a user command is obtained and a response is provided, and can output a sound corresponding to the response when a response is obtained.
[0047] The server (200) may be a device that obtains a response to a user command. For example, the server (200) may receive a user command from an electronic device (100), input the user command into a first neural network model to obtain a response, and provide the obtained response to the electronic device (100).
[0048] FIG. 2 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.
[0049] According to FIG. 2, the electronic device (100) includes a memory (110), a speaker (120), and a processor (130).
[0050] Memory (110) may refer to hardware that stores information, such as data, in an electrical or magnetic form so that a processor (130), etc., can access it. To this end, memory (110) may be implemented as at least one piece of hardware among non-volatile memory, volatile memory, flash memory, hard disk drive (HDD) or solid-state drive (SSD), RAM, ROM, etc.
[0051] At least one instruction required for the operation of an electronic device (100) or a processor (130) may be stored in the memory (110). Here, the instruction is a unit of code that directs the operation of the electronic device (100) or the processor (130), and may be written in machine language, which is a language that a computer can understand. Alternatively, a plurality of instructions that perform a specific task of the electronic device (100) or the processor (130) may be stored in the memory (110) as an instruction set.
[0052] Data, which is information in bit or byte units capable of representing characters, numbers, images, etc., can be stored in the memory (110). For example, a plurality of sounds, neural network models, etc., can be stored in the memory (110). The plurality of sounds may include at least some of the sounds corresponding to user commands.
[0053] Here, the neural network model may include at least one of a first neural network model and a second neural network model. The first neural network model may be a model trained to output an answer to a user command. The second neural network model may be a model trained to output an estimated computation time until an answer is provided based on the computation of the first neural network model for the user command. The first neural network model may have a larger capacity and slower processing speed compared to the second neural network model. Accordingly, significant delays may occur while computation is performed by the first neural network model, and the user may need to wait until an answer is provided even after a user command has been entered.
[0054] The memory (110) is accessed by the processor (130), and the processor (130) may perform read / write / modify / delete / update, etc. on instructions, instruction sets, or data.
[0055] The speaker (120) is a component that outputs various audio data processed by the processor (130), as well as various notification sounds or voice messages.
[0056] The processor (130) controls the overall operation of the electronic device (100). Specifically, the processor (130) can control the overall operation of the electronic device (100) by being connected to each component of the electronic device (100). For example, the processor (130) can control the operation of the electronic device (100) by being connected to components such as memory (110), speaker (120), etc.
[0057] One or more processors (130) may include one or more of a CPU, a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core), a NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. One or more processors (130) may control one or any combination of other components of the electronic device (100) and may perform operations or data processing related to communication. One or more processors (130) may execute one or more programs or instructions stored in memory (110). For example, one or more processors (130) may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory (110).
[0058] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an artificial intelligence dedicated processor).
[0059] One or more processors (130) may be implemented as a single-core processor including one core, or as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors (130) are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0060] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one of the plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in a multi-core processor, or the first operation and the second operation may be performed by a first core included in a multi-core processor and the third operation may be performed by a second core included in a multi-core processor.
[0061] In the embodiments of the present disclosure, one or more processors (130) may refer to a system-on-chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, NPU, hardware accelerator, or machine learning accelerator, but the embodiments of the present disclosure are not limited thereto. However, for convenience of explanation, the operation of the electronic device (100) is described below using the expression "processor (130)."
[0062] The processor (130) can obtain user commands. For example, the processor (130) can receive user speech through a microphone included in the electronic device (100) and obtain user commands from the user speech. Alternatively, the processor (130) may obtain user commands from another electronic device through a communication interface. Alternatively, the processor (130) may obtain user commands through a user interface such as a keyboard. However, it is not limited to these methods, and there may be various ways for the processor (130) to obtain user commands.
[0063] When a user command is obtained, the processor (130) can identify at least one sound among a plurality of sounds. For example, the processor (130) can identify the estimated computation time until an answer based on the computation of a first neural network model for the user command is provided, and can identify at least one sound among a plurality of sounds based on the estimated computation time. For instance, the plurality of sounds may include sounds with different playback times, and the processor (130) can identify the sound among the plurality of sounds that has a playback time closest to the estimated computation time. In one embodiment, if the estimated computation time is 4.5 seconds and the plurality of sounds are each 3 seconds and 4 seconds, the processor (130) can identify the sound of 4 seconds. Or, if the estimated computation time is 4.5 seconds and the plurality of sounds are each 3 seconds, 4 seconds, and 5 seconds, the processor (130) can identify the sound of 5 seconds. Since a 4-second sound is shorter than the expected computation time, a user waiting time may occur, and if there are multiple sounds with a playback time closest to the expected computation time, the processor (130) can identify the sound with a playback time longer than the expected computation time. However, it is not limited to this, and the processor (130) may identify a 4-second sound and output the 4-second sound for 4.5 seconds. That is, the processor (130) may identify the sound with a playback time closest to the expected computation time among multiple sounds and change the playback speed of the identified sound to output it for the expected computation time.
[0064] The processor (130) can perform operations on a second neural network model for a user command to identify the estimated operation time until an answer based on operations on a first neural network model for a user command is provided. However, it is not limited thereto, and the processor (130) may also identify the estimated operation time until an answer based on operations on a first neural network model for a user command is provided based on rules. For example, the processor (130) may identify the estimated operation time until an answer based on operations on a first neural network model for a user command is provided based on the capacity of the user command.
[0065] The processor (130) may request the execution of a first neural network model operation for a user command and output at least one sound through the speaker (120) while the first neural network model operation is being performed. However, it is not limited thereto, and the processor (130) may request the execution of a first neural network model operation for a user command and perform a second neural network model operation for a user command to identify the estimated operation time until an answer based on the first neural network model operation for a user command is provided. The second neural network model may have a smaller capacity and faster processing speed compared to the first neural network model. Accordingly, even if the processor (130) requests the execution of a first neural network model operation for a user command and performs a second neural network model operation for a user command, the time to obtain the estimated operation time may not be significantly delayed.
[0066] When the processor (130) obtains an answer based on the operation of the first neural network model, it can output a sound corresponding to the answer through the speaker (120). Through this operation, the user can have a natural question-and-answer session with the electronic device (100) because at least one sound is output between the user's speech and the output of the sound corresponding to the answer. For example, when the processor (130) receives the speech "What is the first market in our country?", it can output at least one sound such as "Hmm" before providing an answer to this, and provide an answer such as "It is the AA market." The user can hear answers such as "Hmm" and "It is the AA market" without a waiting time (e.g., 3 seconds) after the speech "What is the first market in our country?", and can feel as if they are having a conversation with a person. According to the prior art, when a waiting time (e.g., 3 seconds) elapses after an utterance such as "What is the first market in our country?", an answer such as "It is Market AA." is provided, which may make the user feel awkward; however, according to the present disclosure, this feeling is not experienced, thereby improving user convenience.
[0067] The processor (130) controls a communication interface to transmit a user command to a server (200) that performs operations on a first neural network model, and outputs at least one sound through a speaker (120) after identifying at least one sound and before outputting a sound corresponding to an answer, and when an answer is received from the server (200) through the communication interface, outputs a sound corresponding to the answer through the speaker (120).
[0068] However, it is not limited thereto, and the electronic device (100) may store the first neural network model. For example, the memory (110) may further store the first neural network model, and the processor (130) may perform operations on the first neural network model for a user command, output at least one sound through the speaker (120) while the operations on the first neural network model are being performed, and when an answer is obtained, output a sound corresponding to the answer through the speaker (120).
[0069] Alternatively, the processor (130) may determine a device to perform operations based on the first neural network model based on a user command. For example, the processor (130) may identify the estimated operation time until an answer based on the operation of the first neural network model for the user command is provided, and if the estimated operation time is less than a preset time, use the first neural network model stored in the electronic device (100), and if the estimated operation time is greater than or equal to the preset time, use the first neural network model stored in the server (200).
[0070] However, it is not limited thereto, and the server (200) may store a teacher model, and the electronic device (100) may store a student model obtained from the teacher model through a knowledge distillation technique, and the processor (130) may use the student model stored in the electronic device (100) if the expected computation time is less than a preset time, and may use the teacher model stored in the server (200) if the expected computation time is longer than the preset time.
[0071] Alternatively, the server (200) may store a first neural network model, the electronic device (100) may store a third neural network model with reduced capacity and lower performance from the first neural network model, and the processor (130) may determine a device to perform operations based on the neural network model based on a user command. For example, the processor (130) may identify the estimated operation time until an answer based on the operation of the first neural network model for the user command is provided, and if the estimated operation time is less than a preset time, use the third neural network model stored in the electronic device (100), and if the estimated operation time is greater than or equal to the preset time, use the first neural network model stored in the server (200).
[0072] In the above description, the processor (130) is described as identifying at least one sound based on the expected computation time, but it is not limited thereto. For example, when a user command is obtained, the processor (130) may identify at least one sound corresponding to a part of the user command. The processor (130) may output at least one sound through the speaker (120) while the computation of the first neural network model is being performed, and when an answer is obtained, output a sound corresponding to the remainder of the answer excluding a part of the user command through the speaker (120).
[0073] For example, when the processor (130) receives an utterance such as “Where is the oldest cave in our country?”, it identifies “the oldest cave in our country” as at least one sound and outputs “the oldest cave in our country” while the operation of the first neural network model for the user command corresponding to the utterance is performed, and then when an answer such as “the oldest cave in our country is AA Cave.” is obtained, it outputs “AA Cave.” excluding “the oldest cave in our country” which has already been output as at least one sound.
[0074] Alternatively, the processor (130) may request a first neural network model to perform an operation on the user command and a prompt to output part of the user command as sound, and while the operation of the first neural network model is being performed, output at least one sound through the speaker (120), and when an answer is obtained, output a sound corresponding to the answer through the speaker (120). Here, the answer may be an answer from which part of the user command is excluded. That is, as the processor (130) provides the first neural network model with not only the user command but also a prompt to output part of the user command as sound, the first neural network model may provide an answer from which part of the user command is excluded.
[0075] The electronic device (100) further includes a microphone, and when a user utterance is received through the microphone, the processor (130) may obtain a user command based on the user utterance and identify at least one sound based on at least one of the content or tone of the user command. For example, when a user utterance is received through the microphone, the processor (130) may obtain a user command based on the user utterance and identify at least one sound among a plurality of sounds that is most similar to the tone of the user command. Alternatively, when a user utterance is received through the microphone, the processor (130) may obtain a user command based on the user utterance, identify at least one sound among a plurality of sounds, modulate at least one sound based on at least one of the content or tone of the user command, and output at least one modulated sound. Alternatively, the processor (130) may update a plurality of sounds based on the tone of the user command.
[0076] Meanwhile, although the second neural network model has been described above as being stored in the electronic device (100), it is not limited thereto. For example, the second neural network model may also be stored in the server (200). In this case, the processor (130) transmits a user command to the server (200), and the server (200) can compute the user command through the first neural network model and the second neural network model. When the server (200) obtains an estimated computation time from the second neural network model, it transmits the estimated computation time to the electronic device (100), and subsequently, when an answer is obtained from the first neural network model, it may transmit the answer to the electronic device (100).
[0077] Additionally, although it has been described above that at least one sound and an answer are provided through a speaker, this is not limited thereto. For example, the processor (130) may display information corresponding to at least one sound through a display, and then, when an answer is obtained, display information corresponding to the answer through a display. Alternatively, the processor (130) may transmit information corresponding to at least one sound to a user electronic device, and then, when an answer is obtained, transmit information corresponding to the answer to the user electronic device.
[0078] Meanwhile, the artificial intelligence-related functions according to the present disclosure can be operated through the processor (130) and memory (110).
[0079] The processor (130) may be composed of one or more processors. In this case, the one or more processors may be a general-purpose processor such as a CPU, AP, DSP, etc., a graphics-dedicated processor such as a GPU, VPU (Vision Processing Unit), or an artificial intelligence-dedicated processor such as an NPU.
[0080] One or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory (110). Alternatively, if one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model. The predefined operation rules or artificial intelligence models are characterized by being created through learning.
[0081] Here, "created through learning" means that a basic artificial intelligence model is trained using multiple learning data by a learning algorithm, thereby creating a predefined rule of operation or an artificial intelligence model configured to perform a desired characteristic (or objective). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0082] An artificial intelligence model can be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through calculations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights can be updated during the learning process so that the loss or cost values obtained by the artificial intelligence model are reduced or minimized.
[0083] Artificial neural networks may include deep neural networks (DNNs), such as, but are not limited to, Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), Generative Adversarial Networks (GANs), or Deep Q-Networks.
[0084] Meanwhile, although the electronic device (100) in FIG. 2 is described as including a speaker, it is not limited thereto. For example, the electronic device (100) may include a communication interface and may transmit at least one sound and a response to an external speaker through the communication interface. For instance, the electronic device (100) may transmit at least one sound and a response to a Bluetooth speaker or wired earphones through the communication interface, and the Bluetooth speaker or wired earphones may output at least one sound and a response.
[0085] FIG. 3 is a block diagram showing the detailed configuration of an electronic device (100) according to one embodiment of the present disclosure. The electronic device (100) may include a memory (110), a speaker (120), and a processor (130). Additionally, according to FIG. 3, the electronic device (100) may further include a communication interface (140), a microphone (150), a display (160), a user interface (170), and a camera (180). Detailed descriptions of parts of the components shown in FIG. 3 that overlap with the components shown in FIG. 2 are omitted.
[0086] The communication interface (140) is a configuration that performs communication with various types of external devices according to various types of communication methods. For example, an electronic device (100) can perform communication with a server through the communication interface (140).
[0087] The communication interface (140) may include a Wi-Fi module, a Bluetooth module, an infrared communication module, and a wireless communication module. Here, each communication module may be implemented in the form of at least one hardware chip.
[0088] The Wi-Fi module and Bluetooth module perform communication using the Wi-Fi and Bluetooth methods, respectively. When using the Wi-Fi or Bluetooth module, various connection information, such as the SSID and session key, is transmitted and received first; after establishing a communication connection using this information, various data can be transmitted and received. The infrared communication module performs communication according to infrared communication (IrDA, Infrared Data Association) technology, which wirelessly transmits data over short distances using infrared rays that lie between visible light and millimeter waves.
[0089] In addition to the communication method described above, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), and 5G (5th Generation).
[0090] Alternatively, the communication interface (140) may include a wired communication interface such as HDMI, DP, Thunderbolt, USB, RGB, D-SUB, DVI, etc.
[0091] In addition, the communication interface (140) may include at least one of a LAN (Local Area Network) module, an Ethernet module, or a wired communication module that performs communication using a pair cable, a coaxial cable, or a fiber optic cable.
[0092] The microphone (150) is configured to receive sound input and convert it into an audio signal. The microphone (150) is electrically connected to the processor (120) and can receive sound under the control of the processor (120).
[0093] For example, the microphone (150) may be formed as an integrated unit on the upper side, front side, or side side of the electronic device (100). Alternatively, the microphone (150) may be provided in a remote control or the like, separate from the electronic device (100). In this case, the remote control may receive sound through the microphone (150) and provide the received sound to the electronic device (100).
[0094] The microphone (150) may include various configurations such as a microphone that collects analog sound, an amplifier circuit that amplifies the collected sound, an A / D conversion circuit that samples the amplified sound and converts it into a digital signal, and a filter circuit that removes noise components from the converted digital signal.
[0095] Meanwhile, the microphone (150) may be implemented in the form of a sound sensor, and any configuration capable of collecting sound is acceptable.
[0096] The display (160) is configured to display an image and can be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, and a PDP (Plasma Display Panel). The display (160) may also include a driving circuit, a backlight unit, etc., which can be implemented in forms such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, and an OTFT (organic TFT). Meanwhile, the display (160) can be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, etc.
[0097] The user interface (170) may be implemented as a button, touchpad, mouse, and keyboard, or as a touch screen capable of performing display functions and operation input functions. Here, the button may be a various type of button, such as a mechanical button, touchpad, or wheel, formed in any area of the exterior of the main body of the electronic device (100), such as the front, side, or back.
[0098] The camera (180) is configured to capture still images or video. The camera (180) can capture a still image at a specific point in time, but can also capture a series of still images.
[0099] The camera (180) includes a lens, a shutter, an aperture, a solid-state image sensor, an AFE (Analog Front End), and a TG (Timing Generator). The shutter controls the time when light reflected from a subject enters the camera (180), and the aperture controls the amount of light incident on the lens by mechanically increasing or decreasing the size of the opening through which light enters. When light reflected from a subject accumulates as photocharge, the solid-state image sensor outputs an image based on the photocharge as an electrical signal. The TG outputs a timing signal for reading out pixel data from the solid-state image sensor, and the AFE samples and digitizes the electrical signal output from the solid-state image sensor.
[0100] As described above, the electronic device (100) can provide the user with the feeling of having a conversation with a person by outputting at least one sound before the answer is provided by the neural network operation.
[0101] The operation of the electronic device (100) will be described in more detail below through FIGS. 4 to 9. FIGS. 4 to 9 describes individual embodiments for the convenience of explanation. However, the individual embodiments of FIGS. 4 to 9 may be implemented in any combination.
[0102] FIG. 4 is a drawing for explaining a generative artificial intelligence system (400) according to one embodiment of the present disclosure.
[0103] The User Query / Response Interface (410) can receive user input. User input may be in the form of natural language, images, and / or videos. Additionally, context information may be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, information about the application currently being used by the user or the user's location information. Additionally, user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, user input may be in a non-natural language form, such as selecting a menu. The User Query / Response Interface (410) can output results from the generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of actions requested by the user.
[0104] The AI framework (420) receives user input and can coordinate and control each component necessary to perform the user's intent based on the user's query.
[0105] User input received from the User Query / Response Interface (410) can be sent to the Prompt design component (421). The Prompt design component (421) can be used to generate a prompt suitable for inputting user input into an LLM or LMM. The Prompt design component (421) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. Based on user input, the Prompt design component (421) can generate a prompt by accessing a knowledge component (430) containing user preference data, a prompt library, and prompt examples, and can send the generated prompt to the LLM or LMM.
[0106] The API / Plug-in management component (422) can perform the role of communicating with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (422) establishes a channel to communicate with the outside of the AI Interface via the API, and can enable access to various data sources through the established channel. Additionally, the API / Plug-in management component (422) can request an action via the API if the application or service needs to perform an action that executes the user input as a final step rather than an intermediate result. Information obtained from the outside (e.g., the Applications / service component (440)) can be used to generate a prompt in the Prompt design component (421) along with user input, or it can be passed as input to the generative model.
[0107] The Refiner component (423) can fine-tune the output from the generative model. For example, the Refiner component (423) can verify whether the content generated through LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the Refiner component (423) can determine how well the output matches the user's desired result and, if additional processing is required, proceed with that process. The Refiner component (423) can also configure and provide hints to the user to avoid unwanted output.
[0108] A Generative AI Model (450) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A Generative AI Model (450) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include Diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs (large multimodal models) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0109] FIG. 5 is a flowchart illustrating an operation to identify at least one sound using an estimated computation time according to one embodiment of the present disclosure.
[0110] The processor (130) can obtain a user command (S510) and identify the estimated computation time until an answer based on the computation of the first neural network model for the user command is provided (S520).
[0111] The processor (130) can identify at least one sound based on the estimated computation time (S530). For example, a plurality of sounds may be stored in the memory (110). For instance, the plurality of sounds may include a first sound of 1 second length, a second sound of 2 seconds length, and a third sound of 3 seconds length, and the processor (130) can identify the third sound as at least one sound if the estimated computation time is 3 seconds.
[0112] However, it is not limited to this, and the processor (130) may identify a sound shorter than the expected computation time as at least one sound. For example, if the expected computation time is 3 seconds, the processor (130) may identify a first sound of 1 second length as at least one sound by creating a margin before and after the time interval in which at least one sound is output, such as by creating a margin of 1 second before the time interval in which at least one sound is output and 1 second after the time interval in which at least one sound is output. In this case, after the processor (130) identifies the first sound as at least one sound, it may output the first sound after 1 second has elapsed. If an answer is obtained that matches the expected computation time 1 second after the output of the first sound is finished, it may immediately output a sound corresponding to the answer. Through this operation, a buffer time is secured between the at least one sound and the sound corresponding to the answer, so that the answer can be provided in a more natural manner.
[0113] The processor (130) can output at least one sound while the operation of the first neural network model for the user command is being performed (S540). In the example described above, the processor (130) outputs a third sound, and when the operation of the first neural network model is completed and an answer is obtained, it can output a sound corresponding to the answer.
[0114] FIG. 6 is a flowchart illustrating an operation for identifying a sound corresponding to a part of a user command as at least one sound according to one embodiment of the present disclosure.
[0115] The processor (130) acquires a user command (S610) and can identify at least one sound corresponding to a part of the user command (S620). The processor (130) can output at least one sound while the operation of the first neural network model for the user command is performed (S630).
[0116] For example, when the processor (130) receives an utterance such as “Where is the oldest cave in our country?”, it identifies “the oldest cave in our country” as at least one sound and outputs “the oldest cave in our country” while the operation of the first neural network model for the user command corresponding to the utterance is performed, and then when an answer such as “the oldest cave in our country is AA Cave.” is obtained, it outputs “AA Cave.” excluding “the oldest cave in our country” which has already been output as at least one sound.
[0117] FIGS. 5 and 6 describe an operation of identifying and outputting at least one sound before outputting a sound corresponding to an answer. However, an answer may be obtained before outputting at least one sound, or an answer may not be obtained even after outputting at least one sound. For example, if the processor (130) is provided with an answer before completing the output of at least one sound, it may output a sound corresponding to the answer after completing the output of at least one sound. Alternatively, if the processor (130) is provided with an answer before completing the output of at least one sound, it may output a sound corresponding to the answer after completing the output of at least one sound and after a preset time (e.g., 1 second) has elapsed. In this case, the processor (130) may update the second neural network model using the user command and the actual time taken until the answer is provided.
[0118] Alternatively, if the processor (130) does not receive an answer even after completing the output of at least one sound, it may output a first preset sound among multiple sounds. For example, when the processor (130) receives an utterance such as "Where is the oldest cave in our country?", it may output "The oldest cave in our country" as at least one sound. If the processor (130) does not receive an answer even after the output is completed, it may output a first preset sound such as "sound".
[0119] The processor (130) outputs a first preset sound at least once, and if no answer is obtained even after the output is completed, it may output a second preset sound among a plurality of sounds. For example, if the processor (130) outputs a first preset sound such as "Um" three times and no answer is obtained, it may provide a second preset sound such as "I don't know," "There is a delay. I will let you know again when an answer is obtained."
[0120] By providing a first preset sound, natural question-and-answer interaction is possible even when obtaining a response is delayed. Additionally, by providing a second preset sound, the problem of the user waiting indefinitely can be resolved.
[0121] FIG. 7 is a drawing for explaining a second neural network model according to one embodiment of the present disclosure.
[0122] The first neural network model may be a model trained to output an answer to a user command, and since it may be large in size and slow in processing speed, the user may feel bored during neural network computations by the first neural network model.
[0123] The processor (130) can identify the estimated computation time until the answer based on the computation of the first neural network model is provided. For example, the processor (130) can identify the estimated computation time until the answer based on the computation of the first neural network model is provided by performing a computation of the second neural network model on a user command.
[0124] The second neural network model may be a model obtained by learning multiple sample user commands and multiple sample operation times. For example, the top of FIG. 7 represents the actual operation time, and the bottom of FIG. 7 represents information regarding the expected operation time, and the processor (130) may obtain the expected operation time from the information regarding the expected operation time of the second neural network model. For instance, 710 in FIG. 7 represents the actual operation time of 0.1 seconds, and 720 in FIG. 7 represents information regarding the expected operation time, with a 20% probability of 0.0 seconds and an 80% probability of 0.1 seconds, and the processor (130) may obtain 0.08 seconds as the expected operation time by weighting the arithmetic mean of the information regarding the expected operation time. Additionally, 730 in FIG. 7 represents an actual operation time of 1.6 seconds, and 740 in FIG. 7 represents information regarding the expected operation time, with a probability of 1.6 seconds of 30%, a probability of 1.7 seconds of 20%, and a probability of 1.8 seconds of 50%, and the processor (130) can obtain an expected operation time of 1.72 seconds by weighting the arithmetic mean of the information regarding the expected operation time.
[0125] However, this is not limited to this, and the second neural network model may be trained to output the estimated computation time rather than information about the estimated computation time.
[0126] FIG. 8 is a drawing for explaining a plurality of sounds according to one embodiment of the present disclosure.
[0127] The processor (130) can identify at least one sound among a plurality of sounds based on the estimated computation time.
[0128] For example, as illustrated in FIG. 8, the processor (130) can identify "sound" as at least one sound when the budget calculation time is 1 second, and identify "about that" as at least one sound when the budget calculation time is 3 seconds.
[0129] Alternatively, the processor (130) may identify "he" as at least one sound if the budget calculation time is 1 second based on the user's usual speech pattern, and identify "so" as at least one sound if the budget calculation time is 3 seconds.
[0130] In FIG. 8, the general sounds may be additional sounds used by ordinary people, which are sounds stored in the electronic device (100) at the time of manufacture. In FIG. 8, the speech pattern used by the user may be a speech pattern stored by the user of the electronic device (100) through calls, recordings, video recordings, etc., or may be designated as a speech pattern preferred by the user.
[0131] The processor (130) may provide a sound in a general tone or a tone of voice used by the user based on a user command. Alternatively, the processor (130) may provide a sound that is a mixture of a general tone and a tone of voice used by the user. For example, if the processor (130) has an estimated processing time of 5 seconds, it may provide the sound “That is… so…”.
[0132] FIG. 9 is a drawing for explaining the operation of outputting a sound corresponding to a part of a user command as at least one sound according to one embodiment of the present disclosure.
[0133] For example, as illustrated in FIG. 9, when a user utterance such as “Where are the oldest cave and temple in our country?” is received, the device of the conventional (910) can take 3 seconds to search for an answer and then provide an answer such as “The name of the oldest cave in our country is A, and the oldest temple is Temple B in Uiseong.”
[0134] In this regard, according to the present disclosure (920), when a user utterance such as “Where are the oldest cave and temple in our country?” is received, the processor (130) may provide “The name of the oldest cave in our country” while searching for an answer, and when an answer is obtained, provide an answer such as “A, and the oldest temple is Temple B in Uiseong.”
[0135] That is, according to the present disclosure, it is possible to provide the user with the feeling that the interaction is maintained even during the 3 seconds of searching for an answer.
[0136] FIG. 10 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0137] First, when a user command is obtained, at least one of a plurality of sounds is identified (S1010). Then, a request is made to perform an operation of the first neural network model for the user command (S1020). Then, while the operation of the first neural network model is being performed, at least one sound is output through the speaker of the electronic device (S1030). Then, when an answer based on the operation of the first neural network model is obtained, a sound corresponding to the answer is output through the speaker (S1040).
[0138] Additionally, the requesting step (S1020) transmits a user command to a server that performs operations on the first neural network model, and the step of outputting at least one sound through a speaker (S1030) outputs at least one sound through a speaker after identifying at least one sound and before outputting a sound corresponding to the answer, and the step of outputting a sound corresponding to the answer through a speaker (S1040) can output a sound corresponding to the answer through a speaker when an answer is received from the server.
[0139] And, the requesting step (S1020) can perform operations of the first neural network model for the user command.
[0140] Additionally, the identifying step (S1010) performs an operation of the second neural network model on the user command to identify the estimated operation time until an answer based on the operation of the first neural network model is provided, and can identify at least one sound among a plurality of sounds based on the estimated operation time.
[0141] And, the plurality of sounds includes sounds with different playback times, and the identification step (S1010) can identify the sound among the plurality of sounds that has a playback time closest to the expected operation time. In this case, the playback speed of the sound that has a playback time closest to the expected operation time may be changed. For example, if the expected operation time is 4.5 seconds and the identified sound is 4 seconds, the playback speed of the identified sound may be changed so that the sound is output for 4.5 seconds. That is, the playback time of the identified sound may be increased or decreased so that the sound is output during the expected operation time.
[0142] Additionally, the identifying step (S1010) can identify a sound corresponding to a part of the user command as at least one sound when a user command is obtained.
[0143] And, in the step (S1040) of outputting a sound corresponding to the answer through a speaker, when the answer is obtained, a sound corresponding to the remainder of the answer, excluding part of the user command, can be output through a speaker.
[0144] Additionally, the requesting step (S1020) requests a prompt to output a part of the user command as sound and requests the execution of a first neural network model operation on the user command, and the step of outputting a sound corresponding to the answer through a speaker (S1040) outputs a sound corresponding to the answer through a speaker when the answer is obtained, and the answer may be an answer from which a part of the user command has been excluded.
[0145] And, the identification step (S1010) can obtain a user command based on the user command when a user utterance is received through a microphone included in an electronic device, and can identify at least one sound based on at least one of the content or tone of the user command.
[0146] Additionally, it may include a step of updating multiple sounds based on the tone of the user command.
[0147] An electronic device according to one embodiment as described above includes one or more processors including a memory for storing a plurality of sounds and instructions, a speaker, and a processing circuitry, wherein when the instructions are executed individually or collectively by the one or more processors, when a user command is obtained, at least one sound among the plurality of sounds is identified, and a first neural network model is requested to perform an operation on the user command, and while the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and when an answer based on the operation of the first neural network model is obtained, a sound corresponding to the answer is output through the speaker.
[0148] According to one example, the system further includes a communication interface, wherein the instructions control the communication interface to transmit the user command to a server performing operations of the first neural network model when executed individually or collectively by the one or more processors, and output the at least one sound through the speaker after identifying the at least one sound until the sound corresponding to the answer is output, and when the answer is received from the server through the communication interface, the sound corresponding to the answer can be output through the speaker.
[0149] According to one example, the memory further stores the first neural network model, and when the instructions are executed individually or collectively by the one or more processors, the first neural network model performs operations on the user command, outputs at least one sound through the speaker while the operations of the first neural network model are being performed, and when the answer is obtained, outputs a sound corresponding to the answer through the speaker.
[0150] According to one example, the memory further stores a second neural network model, and when the instructions are executed individually or collectively by the one or more processors, the second neural network model performs operations on the user command to identify the estimated operation time until the answer based on the first neural network model is provided, and can identify at least one sound among the plurality of sounds based on the estimated operation time.
[0151] According to one example, the plurality of sounds include sounds with different playback times, and when the instructions are executed individually or collectively by the one or more processors, the sound having the playback time closest to the expected operation time among the plurality of sounds can be identified.
[0152] According to one example, when the instructions are executed individually or collectively by the one or more processors, if the user command is obtained, a sound corresponding to a part of the user command can be identified as the at least one sound.
[0153] According to one example, when the instructions are executed individually or collectively by the one or more processors, the at least one sound is output through the speaker while the operation of the first neural network model is being performed, and when the answer is obtained, the sound corresponding to the remainder of the answer, excluding part of the user command, can be output through the speaker.
[0154] According to one example, when the instructions are executed individually or collectively by the one or more processors, a prompt is given to output a part of the user command as a sound and to request the execution of an operation of the first neural network model on the user command, and while the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and when the answer is obtained, a sound corresponding to the answer is output through the speaker, and the answer may be an answer from which a part of the user command has been excluded.
[0155] According to one example, the system further includes a microphone, and when the instructions are executed individually or collectively by one or more processors, if the user utterance is received through the microphone, the system can obtain the user command based on the user utterance and identify the at least one sound based on at least one of the content or tone of the user command.
[0156] According to one example, when the instructions are executed individually or collectively by one or more processors, the plurality of sounds can be updated based on the tone of the user command.
[0157] A control method for an electronic device according to one embodiment may include the steps of: identifying at least one sound among a plurality of sounds when a user command is obtained; requesting the execution of a first neural network model operation for the user command; outputting the at least one sound through a speaker of the electronic device while the operation of the first neural network model is being performed; and, when an answer based on the operation of the first neural network model is obtained, outputting a sound corresponding to the answer through the speaker.
[0158] According to one example, the requesting step transmits the user command to a server that performs operations of the first neural network model, and the step of outputting at least one sound through the speaker outputs the at least one sound through the speaker after identifying the at least one sound and before outputting the sound corresponding to the answer, and the step of outputting the sound corresponding to the answer through the speaker can output the sound corresponding to the answer through the speaker when the answer is received from the server.
[0159] According to one example, the requesting step can perform operations of the first neural network model on the user command.
[0160]
[0161] According to one example, the identifying step may perform an operation of a second neural network model on the user command to identify the estimated operation time until the answer based on the operation of the first neural network model is provided, and identify at least one sound among the plurality of sounds based on the estimated operation time.
[0162] According to one example, the plurality of sounds includes sounds with different playback times, and the identifying step can identify the sound among the plurality of sounds that has a playback time closest to the expected calculation time.
[0163] According to one example, when the user command is obtained, the identifying step may identify a sound corresponding to a part of the user command as the at least one sound.
[0164] According to one example, the step of outputting a sound corresponding to the above answer through the speaker may, when the above answer is obtained, output a sound corresponding to the remainder of the above answer, excluding a part of the user command, through the speaker.
[0165] According to one example, the requesting step requests a prompt to output a part of the user command as sound and requests the execution of an operation of the first neural network model on the user command, and the step of outputting a sound corresponding to the answer through the speaker outputs a sound corresponding to the answer through the speaker when the answer is obtained, and the answer may be an answer from which a part of the user command has been excluded.
[0166] According to one example, the identifying step may, when the user utterance is received through a microphone included in the electronic device, obtain the user command based on the user utterance and identify the at least one sound based on at least one of the content or tone of the user command.
[0167] According to one example, the method may further include the step of updating the plurality of sounds based on the tone of the user command.
[0168] According to various embodiments of the present disclosure as described above, the electronic device can provide the user with the feeling of conversing with a person by outputting at least one sound before an answer is provided by neural network operation.
[0169] The electronic device according to one or more embodiments disclosed in this disclosure may be a device of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this disclosure is not limited to the devices described above.
[0170] One or more embodiments of the present disclosure and the terms used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In the present disclosure, each of phrases such as “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B or C”, “at least one of A, B and C”, and “at least one of A, B, or C” may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationally,” it means that said component may be connected to said other component directly (e.g., wired), wirelessly, or through a third component.
[0171] The term “module” as used in one or more embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0172] One or more embodiments of the present disclosure may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0173] According to one embodiment, the method according to one or more embodiments disclosed in this disclosure may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0174] According to one or more embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one or more embodiments, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one or more embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, Memory for storing multiple sounds and instructions; Speaker; and One or more processors including processing circuitry; and When the above instructions are executed individually or collectively by the one or more processors, When a user command is obtained, at least one sound among the plurality of sounds is identified, and Requesting the execution of a first neural network model operation for the above user command, and While the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and An electronic device that outputs a sound corresponding to the answer through the speaker when an answer based on the operation of the first neural network model is obtained.
2. In Paragraph 1, It further includes a communication interface; and When the above instructions are executed individually or collectively by the one or more processors, Control the communication interface to transmit the user command to a server that performs operations on the first neural network model, and After identifying at least one sound, output the at least one sound through the speaker until outputting a sound corresponding to the answer, and An electronic device that outputs a sound corresponding to the answer through the speaker when the answer is received from the server through the communication interface.
3. In Paragraph 1, The above memory is, The above-mentioned first neural network model is further stored, When the above instructions are executed individually or collectively by the one or more processors, Performing operations of the first neural network model for the above user command, While the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and An electronic device that outputs a sound corresponding to the above answer through the speaker when the above answer is obtained.
4. In Paragraph 1, The above memory is, Save the second neural network model, When the above instructions are executed individually or collectively by the one or more processors, Performing operations of the second neural network model on the above user command to identify the estimated operation time until the answer based on the first neural network model's operations is provided, and An electronic device that identifies at least one sound among the plurality of sounds based on the above-mentioned expected operation time.
5. In Paragraph 4, The above plurality of sounds are, It includes sounds with different playback times, When the above instructions are executed individually or collectively by the one or more processors, An electronic device that identifies the sound having the playback time closest to the estimated computation time among the plurality of sounds.
6. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, An electronic device that, when the above user command is obtained, identifies a sound corresponding to a part of the above user command as the at least one sound.
7. In Paragraph 6, When the above instructions are executed individually or collectively by the one or more processors, While the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and An electronic device that, when the above answer is obtained, outputs a sound through the speaker corresponding to the remainder of the above answer, excluding a part of the above user command.
8. In Paragraph 6, When the above instructions are executed individually or collectively by the one or more processors, A prompt to output a part of the above user command as sound and a request to perform an operation of the above first neural network model on the above user command, and While the operation of the first neural network model is being performed, the at least one sound is output through the speaker, and When the above answer is obtained, a sound corresponding to the above answer is output through the speaker, and The above answer is, An electronic device that is an answer with part of the above user command excluded.
9. In Paragraph 1, Includes a microphone; and When the above instructions are executed individually or collectively by the one or more processors, When the user utterance is received through the microphone, the user command is obtained based on the user utterance, and An electronic device that identifies at least one sound based on at least one of the content or tone of the above user command.
10. In Paragraph 9, When the above instructions are executed individually or collectively by the one or more processors, An electronic device that updates the plurality of sounds based on the tone of the above user command.
11. In a method for controlling an electronic device, When a user command is obtained, a step of identifying at least one sound among a plurality of sounds; A step of requesting the execution of a first neural network model operation for the above user command; A step of outputting at least one sound through a speaker of the electronic device while the operation of the first neural network model is being performed; and A control method comprising the step of outputting a sound corresponding to the answer through the speaker when an answer based on the operation of the first neural network model is obtained.
12. In Paragraph 11, The above requested step is, The user command is transmitted to a server that performs operations of the first neural network model, and The step of outputting at least one sound through the speaker is, After identifying at least one sound, output the at least one sound through the speaker until outputting a sound corresponding to the answer, and The step of outputting a sound corresponding to the above answer through the speaker is, A control method that outputs a sound corresponding to the answer through the speaker when the answer is received from the server.
13. In Paragraph 11, The above requested step is, A control method for performing operations of the first neural network model on the above user command.
14. In Paragraph 11, The above identification step is, Performing operations of a second neural network model on the above user command to identify the estimated operation time until the answer based on the operation of the first neural network model is provided, and A control method for identifying at least one sound among a plurality of sounds based on the above-mentioned expected operation time.
15. In Paragraph 14, The above plurality of sounds are, It includes sounds with different playback times, The above identification step is, A control method for identifying the sound having the playback time closest to the expected calculation time among the plurality of sounds.