Electronic device and method for providing continuous command function of virtual assistant, and non-transitory computer-readable storage medium
The electronic device enhances virtual assistant systems by allowing seamless command execution through identification of registered and temporarily permitted users, addressing the unnatural user experience caused by repeated wake words.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-05
AI Technical Summary
Existing virtual assistant systems require repeated wake words or user inputs for each command, leading to an unnatural and disruptive user experience, especially in environments with multiple speakers.
An electronic device that recognizes and executes continuous commands without additional wake words by using identification information to differentiate between registered and temporarily permitted users, allowing seamless command execution.
Enhances user experience by enabling natural and smooth conversations with virtual assistants, accurately processing commands from both registered and temporarily authorized users without disrupting the continuous command function.
Smart Images

Figure KR2025013022_05032026_PF_FP_ABST
Abstract
Description
Electronic device, method and non-transitory computer-readable storage medium for providing continuous command function of virtual assistant
[0001] The present disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for providing a continuous command function of a virtual assistant.
[0002] An electronic device can acquire a sound signal from an external source via a microphone. For example, the sound signal may include a speech signal uttered by a speaker. For example, the electronic device can provide a virtual assistant (or virtual assistant function) based on the speech signal.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] Aspects will be partly explained in the following description, and partly will be made clear in the description or can be learned through practice of the examples presented.
[0005] According to an aspect of the present disclosure, an electronic device comprises an input interface configured to receive sound data; a memory storing instructions, the memory including one or more storage media; and at least one processor including a processing circuit, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: receive, through the input interface, a first voice signal including a first command based on activation of a virtual assistant; generate, based on the first voice signal, first identification information corresponding to a first speaker of the first voice signal; execute a function corresponding to the first command of the first voice signal; receive, through the input interface, a second voice signal including a second command following the first voice signal; generate, based on the second voice signal, second identification information corresponding to a second speaker of the second voice signal; and execute a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information being greater than or equal to a reference value.
[0006] According to an aspect of the present disclosure, a method performed by an electronic device having an input interface configured to receive sound data may include: receiving, through the input interface, a first voice signal including a first command based on activation of a virtual assistant; generating, based on the first voice signal, first identification information corresponding to a first speaker of the first voice signal; executing a function corresponding to the first command of the first voice signal; receiving, through the input interface, a second voice signal including a second command following the first voice signal; generating, based on the second voice signal, second identification information corresponding to a second speaker of the second voice signal; and executing a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information being greater than or equal to a reference value.
[0007] According to an aspect of the present disclosure, a non-transitory computer-readable storage medium may store one or more programs, which, when individually or collectively executed by at least one processor of an electronic device having an input interface configured to receive sound data, cause the electronic device to receive, through the input interface, a first voice signal including a first command, based on activation of a virtual assistant; generate, based on the first voice signal, first identification information corresponding to a first speaker of the first voice signal; execute a function corresponding to the first command of the first voice signal; receive, through the input interface, a second voice signal including a second command following the first voice signal; generate, based on the second voice signal, second identification information corresponding to a second speaker of the second voice signal; and execute a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information being greater than or equal to a reference value.
[0008] According to an aspect of the present disclosure, an electronic device comprises an input interface configured to receive sound data; a memory storing instructions, the memory including one or more storage media; and at least one processor including a processing circuit, the instructions causing the electronic device, when individually or collectively executed by the at least one processor: receiving, through the input interface, a first voice signal including a first command based on activation of a virtual assistant; generating, based on the first voice signal, identification information corresponding to a first speaker of the first voice signal; executing a function corresponding to the first command of the first voice signal; receiving, through the input interface, a second voice signal including a second command following the first voice signal; identifying, based on the identification information and the second voice signal, whether the first speaker corresponds to a second speaker of the second voice signal; executing, based on the first speaker corresponding to the second speaker, a function corresponding to the second command of the second voice signal; And based on the first speaker not corresponding to the second speaker, it can cause the function corresponding to the second command of the second voice signal to be refrained from being executed.
[0009] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0010] FIGS. 1A and 1B illustrate examples of a continuous command function of a virtual assistant according to various embodiments.
[0011] FIG. 2 illustrates an example of a continuous command function of a virtual assistant based on temporary allow, according to an embodiment.
[0012] Figure 3 is a schematic view of an exemplary electronic device.
[0013] FIGS. 4A, 4B, 4C, 4D, 4E, and 4F illustrate examples of trained models included within an electronic device, according to an embodiment.
[0014] FIG. 5 illustrates an example of a flow of operations for training a model for a virtual assistant for a specific user, according to an embodiment.
[0015] FIG. 6A illustrates an example of an operational flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a registered user with respect to a virtual assistant, according to an embodiment.
[0016] FIG. 6b illustrates an example of an operational flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a registered user with respect to a virtual assistant in a continuous command function of a virtual assistant according to an embodiment.
[0017] FIG. 7 illustrates an example of an operational flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a temporarily permitted user in a continuous command function of a virtual assistant according to an embodiment.
[0018] FIG. 8 illustrates an example of an operation flow for a method of verifying a second user who uttered a second voice signal and executing a function corresponding to a command of the second voice signal, based on whether there is a correspondence between a first user who uttered a first voice signal and a registered user in a continuous command function of a virtual assistant according to an embodiment.
[0019] FIG. 9 illustrates an example of a method for recognizing a voice signal in a continuous command function of a virtual assistant based on temporary permission, according to an embodiment.
[0020] FIG. 10 illustrates an example of a user interface for settings related to a continuous command function of a virtual assistant, according to an embodiment.
[0021] FIG. 11 is a block diagram of an electronic device within a network environment according to various embodiments.
[0022] FIG. 12 is a block diagram illustrating an integrated intelligence system according to various embodiments.
[0023] Figure 13 is a diagram showing the form in which relationship information between concepts and operations according to various embodiments is stored in a database.
[0024] FIG. 14 is a diagram illustrating a screen for processing voice input received through an intelligent app by an electronic device according to various embodiments.
[0025] The terms used in this disclosure are used merely to describe specific embodiments and may not be intended to limit the scope of the present disclosure. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning as the meaning they have in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.
[0026] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.
[0027] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled. However, this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." Conditions described as "more than" may be replaced with "more than," conditions described as "less than," and conditions described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of the elements from A (including A) to B (including B).
[0028] FIGS. 1A and 1B illustrate examples of a continuous command function of a virtual assistant according to various embodiments.
[0029] FIGS. 1A and 1B illustrate examples of providing a continuous command function while an electronic device (101) activates a virtual assistant. For example, the virtual assistant may be a software application (or software agent) that processes tasks requested by a user of the electronic device (101) and provides services. For example, the virtual assistant may be referred to as a voice assistant, a digital assistant, an intelligent automated assistant, an automatic digital assistant, an artificial intelligence assistant, an intelligent assistant, a personal assistant, a mobile assistant, an intelligent agent, and / or equivalent technical terms. As a non-limiting example, the virtual assistant may include Bixby. However, the present disclosure is not limited thereto.
[0030] For example, the electronic device (101) may receive a voice signal through an input interface (e.g., the input interface (320) of FIG. 3). For example, the voice signal (or speech signal) may include a voice. For example, the voice signal may include the voice, noise, or background sound. For example, the voice may be uttered or spoken by a user of the electronic device (101) (e.g., the user (100) of FIGS. 1A and 1B) or another user (e.g., the user (109) of FIG. 1B). For example, the voice signal may be referred to as a sound signal, a signal, a speech signal, or a user signal.
[0031] For example, the voice signal may include a wake-up word. For example, the wake-up word may include a word or sentence designated to activate the virtual assistant. For example, activating the virtual assistant may include calling the virtual assistant (or executing a voice recognition function by the virtual assistant) while an application for the virtual assistant (or a virtual assistant application) is running within the electronic device (101) (or while running within the background of the electronic device (101). For example, the electronic device (101) may activate the virtual assistant upon identifying that the received voice signal includes the wake-up word. In other words, the virtual assistant may be activated in response to the wake-up word. For example, the wake word may be referred to as a wake-up voice input, a lexical trigger, a hot-phrase, a hot-word, a trigger word, a trigger phrase, a trigger expression, and / or equivalent technical terms.
[0032] In the above example, the virtual assistant is activated upon receiving a voice signal including the call word, but the present disclosure is not limited thereto. For example, the electronic device (101) may also activate the virtual assistant in response to an input to the electronic device (101). For example, the input may include an input to a physical button of the electronic device (101) or a visual object (or icon) displayed through the display of the electronic device (101). For example, the input may be referred to as a user input.
[0033] For example, the voice signal may include a command word. For example, the command word may include a word or phrase that causes the activated virtual assistant to perform a specific function. For example, when the electronic device (101) identifies that the received voice signal includes the command word while the virtual assistant is activated, the electronic device (101) may execute a function corresponding to the command word. For example, the execution of the function corresponding to the command word may be triggered by the virtual assistant. As a non-limiting example, if the command word is "play music," the function word may include playing music through a software application that provides music. In other words, the function word may be executed in response to the command word. For example, the command word may be referred to as a command voice input word.
[0034] For example, the continuous command function may be a function of the activated virtual assistant. For example, the continuous command function may include executing a function corresponding to a received command without recognizing an additional wake word (or without additional user input) while the virtual assistant is activated. As a non-limiting example, the electronic device (101) may execute a function corresponding to a first command received after the virtual assistant of the electronic device (101) is activated while the continuous command function of the virtual assistant is executed, and then execute a function corresponding to a second command received without a wake word or user input for activating the virtual assistant. For example, the continuous command function may be referred to as a continuous conversation function, a continuous command mode, a smart follow up mode, a multiple command function, and / or similar technical terms.
[0035] FIG. 1A illustrates examples (110, 120, 130) in which an electronic device (101) activates a virtual assistant upon receiving a voice signal spoken by a user (100) of the electronic device (101) and executes a function corresponding to a command while a continuous command function of the virtual assistant is being executed. For example, the user (100) of the electronic device (101) may be a registered user (or speaker) with respect to the virtual assistant of the electronic device (101) before the virtual assistant of the electronic device (101) is activated. For example, the registered user (or speaker) may represent a user (or speaker) who can perform (or authorize) the activation of the virtual assistant and the execution of a function based on the virtual assistant within an application (or virtual assistant application) providing the virtual assistant. In other words, the registered user may be an allowed (or authorized) user with respect to the virtual assistant.
[0036] Referring to example (110), the user (100) may utter a wake word (115) to activate the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the wake word (115) uttered by the user (100) through a microphone. In example (110), for convenience of explanation, the case where the user (100) utters the wake word (115) is exemplified, but the present disclosure is not limited thereto. For example, the user (100) may utter a voice signal including the wake word (115). As a non-limiting example, the voice signal uttered by the user (100) may include the wake word (115), the wake word (115) and a command, or a word (or sentence) other than the wake word (115) and a command. For example, the electronic device (101) can recognize (or identify) a wake word (115). For example, upon recognizing the wake word (115), the electronic device (101) can execute a function corresponding to the wake word (115). As a non-limiting example, the function corresponding to the wake word (115) may include playing (or outputting) a message (117). As a non-limiting example, the message (117) may include “Hello.” In the present disclosure, the function of playing the identified text within the electronic device (101) as a voice, such as playing the message (117), may be referred to as TTS (text to speech) (or TTS playback, playback of TTS synthesized sound). Furthermore, as a non-limiting example, the function corresponding to the wake word (115) may include displaying a screen (119). As a non-limiting example, the screen (119) may include a dialogue window that includes a visual object representing a message (117).
[0037] Referring to example (120), the user (100) may utter a command (125) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (125) uttered by the user (100) through a microphone. As a non-limiting example, the command (125) may include, “Tell me the weather today.” As described above, in example (120), the electronic device (101) of example (110) is illustrated as receiving a command (125) that is distinct from the wake word (115) after receiving the wake word (115), but the present disclosure is not limited thereto. For example, the electronic device (101) may recognize (or identify) the command (125). For example, the electronic device (101) may, upon recognizing the command (125), execute a function corresponding to the command (125). As a non-limiting example, the function corresponding to the command (125) may include playing (or outputting) a message (127). As a non-limiting example, the message (127) may include "Today's weather is clear." Furthermore, as a non-limiting example, the function corresponding to the command (125) may include displaying a screen (129). As a non-limiting example, the screen (129) may include a screen including a visual object representing today's weather. As a non-limiting example, the virtual assistant application may cause the weather providing application to display a screen including the visual object representing today's weather, or may obtain and display a screen including the visual object representing today's weather from the weather providing application.
[0038] In examples (110) and (120), the electronic device (101) can continuously receive (or acquire) a command (125) in response to a call word (115) through a microphone. For example, the command (125) and the call word (115) can be continuously spoken by the user (100). In this case, the electronic device (101) can skip playing (or outputting) the message (117) upon recognizing the call word (115) and execute the function corresponding to the command (125).
[0039] Referring to example (130), the user (100) may utter a command (135) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (135) uttered by the user (100) through a microphone. As a non-limiting example, the command (135) may include “play music.” The command (135) of example (130) may be received after the command (125) of example (120) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (135) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (125) and executing the function corresponding to the command (125), the electronic device (101) may receive an additional command (e.g., command (135)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (135). For example, upon recognizing the command (135), the electronic device (101) may execute the function corresponding to the command (135). As a non-limiting example, the function corresponding to the command (135) may include playing (or outputting) a message (137). As a non-limiting example, the message (137) may include, “Yes, play music.” Also, as a non-limiting example, the function corresponding to the command (135) may include displaying a screen (139). As a non-limiting example, the screen (139) may include a screen that includes a visual object representing music being played.As a non-limiting example, the virtual assistant application may cause an application that provides music (or plays music) to display a screen including a visual object representing the music being played, or may obtain and display a screen including a visual object representing the music being played from the application that provides music.
[0040] Referring to FIG. 1A, the electronic device (101) can receive commands received from a user (100) while the continuous command function of the virtual assistant is being executed, and execute functions corresponding to the commands, without receiving an additional wake word (or obtaining a user input for activating the virtual assistant). In examples (110, 120, 130) of FIG. 1A, a case is illustrated where commands (125, 135) spoken by a user (100) registered with respect to the virtual assistant are received.
[0041] In the following FIG. 1b, an example is provided in which an electronic device (101) executes functions corresponding to commands spoken by a registered user (100) and commands spoken by other users while the continuous command function of the virtual assistant is being executed.
[0042] FIG. 1B illustrates examples (140, 150, 160, 170) in which an electronic device (101) activates a virtual assistant upon receiving a voice signal uttered by a user (100) of the electronic device (101), and executes a function corresponding to a command while a continuous command function of the virtual assistant is executed upon receiving voice signals uttered by the user (100) and the user (109). For example, the user (100) of the electronic device (101) may be a registered user (or speaker) with respect to the virtual assistant before the virtual assistant of the electronic device (101) is activated. For example, the user (109) may not be a registered user with respect to the virtual assistant.
[0043] Referring to example (140), the user (100) may utter a wake word (145) to activate the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the wake word (145) uttered by the user (100) through a microphone. For example, the electronic device (101) may recognize (or identify) the wake word (145). For example, upon recognizing the wake word (145), the electronic device (101) may execute a function corresponding to the wake word (145). As a non-limiting example, the function corresponding to the wake word (145) may include playing (or outputting) a message (147). As a non-limiting example, the message (147) may include "Hello." Additionally, as a non-limiting example, the function corresponding to the call word (145) may include displaying a screen (149). As a non-limiting example, the screen (149) may include a dialogue window including a visual object representing a message (147).
[0044] Referring to example (150), the user (109) may utter a command (155) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). As described above, the user (109) may not be a registered user with respect to the virtual assistant of the electronic device (101). In this case, the electronic device (101) may recognize a situation intended by the user (100) when receiving the command (155) uttered by the user (109) after the virtual assistant is activated by the wake word (145) uttered by the user (100). In this case, the command (155) may be the first command received after the virtual assistant is activated. In other words, even if the command (155) is uttered by another user, that is, a user (109), after the virtual assistant is activated by the user (100), the electronic device (101) can execute a function corresponding to the command (155). For example, the electronic device (101) can receive (or acquire) the command (155) uttered by the user (109) through the microphone. As a non-limiting example, the command (155) can include “Tell me today’s weather.” For example, the electronic device (101) can recognize (or identify) the command (155). For example, upon recognizing the command (155), the electronic device (101) can execute a function corresponding to the command (155). As a non-limiting example, the function corresponding to the command (155) can include playing (or outputting) a message (157). As a non-limiting example, the message (157) may include, "Today's weather is clear." Furthermore, as a non-limiting example, the function corresponding to the command (155) may include displaying a screen (159). As a non-limiting example, the screen (159) may include a screen that includes a visual object representing today's weather.
[0045] Referring to example (160), the user (100) may utter a command (165) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (165) uttered by the user (100) through a microphone. As a non-limiting example, the command (165) may include “play music.” The command (165) of example (160) may be received after the command (155) of example (150) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (165) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (155) and executing the function corresponding to the command (155), the electronic device (101) may receive an additional command (e.g., command (165)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (165). For example, upon recognizing the command (165), the electronic device (101) may execute the function corresponding to the command (165). As a non-limiting example, the function corresponding to the command (165) may include playing (or outputting) a message (167). As a non-limiting example, the message (167) may include, “Yes, play music.” Also, as a non-limiting example, the function corresponding to the command (165) may include displaying a screen (169). As a non-limiting example, the screen (169) may include a screen that includes a visual object representing music being played.
[0046] Referring to example (160), when the continuous command function of the virtual secretary is running, after the virtual secretary activated by the call word (145) executes the function corresponding to the command (155), a function corresponding to the command (165) spoken by the user (100) registered with respect to the virtual secretary can be executed without receiving an additional call word (or user input).
[0047] In an embodiment, referring to example (170), the user (109) may utter a command (175) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (175) uttered by the user (109) through a microphone. As a non-limiting example, the command (175) may include “play music.” The command (175) of example (170) may be received after the command (155) of example (150) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (175) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (155) and executing the function corresponding to the command (155), the electronic device (101) may perform reception of an additional command (e.g., command (175)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (175). For example, even if the electronic device (101) recognizes the command (175), it may refrain from executing (or, cease, stop, skip, bypass) the function corresponding to the command (175) (or, not execute the function corresponding to the command (175)). For example, the electronic device (101) may not recognize the command (175). Since the user (109) is not a registered user of the virtual assistant, like the user (100), the electronic device (101) may not recognize the command (175) uttered by the user (109), or even if it recognizes the command (175), it may refrain from executing a function corresponding to the command (175).As a non-limiting example, the electronic device (101) may refrain from playing a message by refraining from executing the function corresponding to the command (175). Furthermore, the electronic device (101) may display the screen (159) by refraining from executing the function corresponding to the command (175). For example, the screen (159) may be a screen displayed by executing the function corresponding to the command (155) in example (150). In other words, the electronic device (101) may maintain the display of the screen (159) displayed by the function corresponding to the previously received command (155), rather than displaying the screen (159) as the function corresponding to the command (175).
[0048] Referring to FIG. 1B, the electronic device (101) may receive commands received from the user (100) while the continuous command function of the virtual assistant is being executed, without receiving an additional wake word (or obtaining a user input for activating the virtual assistant), and execute functions corresponding to the commands. In addition, the electronic device (101) may execute a function corresponding to a command received from the user (109) immediately after being activated according to the wake word received from the user (100) while the continuous command function of the virtual assistant is being executed. However, even if the electronic device (101) receives additional commands without an additional wake word after executing the function corresponding to the command received from the user (109), it may not execute the function corresponding to the additionally received command.
[0049] According to an embodiment, the electronic device (101) may receive one wake word when receiving one command from a user (100) who is a registered user of the virtual assistant of the electronic device (101) in order to provide a continuous command function of the virtual assistant. Referring to the above, the electronic device (101) may receive a wake word for each command in order to provide a continuous command function of the virtual assistant, or may only allow commands by users registered in relation to the virtual assistant. The virtual assistant may be a function of the electronic device (101) that performs a task on behalf of the user. When users execute the continuous command function of the virtual assistant, they may desire to have a more natural conversation with the virtual assistant and to have continuous commands requested in the conversation processed (or, performed). However, since a call word is required for each command or only commands by registered users are allowed, registered users (e.g., user (100)) and unregistered users (e.g., user (109)) using the continuous command function of the virtual assistant may perceive the conversation as unnatural (or, not smooth).
[0050] Hereinafter, the present disclosure can receive users' commands and execute functions corresponding to the commands based on temporary permission when the continuous command function of the virtual assistant is executed without receiving additional wake words (or user input). For example, the temporary permission can indicate that other users, in addition to registered users, are granted the same privileges as temporarily registered users. For example, the temporary permission can be referred to as on-the-fly registration. Accordingly, the virtual assistant according to the present disclosure can provide a more natural and smooth user experience. In other words, the present disclosure can process commands uttered by multiple users in a continuous conversational manner by extending the privileges provided only to registered users (or speakers) to temporarily granted users (or speakers). In addition, the present disclosure can more accurately receive and process voice signals uttered by temporarily granted users (or speakers) other than registered users (or speakers). Accordingly, in an environment where multiple speakers exist, the present disclosure can accurately recognize commands uttered by both registered users and temporarily authorized users and provide corresponding functions. An example of the continuous command function of a virtual assistant based on temporary authorization is illustrated and described below with reference to Figure 2.
[0051] Figure 2 illustrates an example of a virtual assistant's continuous command function based on temporary allow.
[0052] FIG. 2 illustrates examples (210, 220, 230, 240) in which an electronic device (101) activates a virtual assistant upon receiving a voice signal uttered by a user (100) of the electronic device (101), and executes a function corresponding to a command while a continuous command function of the virtual assistant is executed upon receiving voice signals uttered by a user (100) who is a registered user of the virtual assistant and a user (109) who is a temporarily permitted user. For example, the user (100) of the electronic device (101) may be a registered user (or speaker) of the virtual assistant before the virtual assistant of the electronic device (101) is activated. For example, the user (109) may not be a registered user of the virtual assistant, but may be a temporarily permitted user. For example, the temporarily permitted user may be a user who utters a command that is first received after the virtual assistant is activated. As a non-limiting example, the temporarily permitted user may correspond to some of the registered users. For example, the correspondence of the temporarily permitted user to the registered user may indicate that the temporarily permitted user is included in the registered users or matches (or is identical to) some of the registered users.
[0053] The user (100) may utter a wake word to activate the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the wake word uttered by the user (100) through a microphone. For example, the electronic device (101) may recognize (or identify) the wake word. For example, the electronic device (101) may activate the virtual assistant upon recognizing the wake word. For specific details related thereto, reference may be made to example (110) of FIG. 1A or example (140) of FIG. 1B. Hereinafter, redundant descriptions are omitted. In addition, the electronic device (101) may receive a user input for activating the virtual assistant instead of receiving the wake word. For example, the electronic device (101) may activate the virtual assistant upon receiving the user input.
[0054] Referring to example (210), the user (109) may utter a command (215) to execute a function of the electronic device (101) using the virtual assistant of the electronic device (101). As described above, the user (109) may not be a registered user with respect to the virtual assistant of the electronic device (101). In this case, the electronic device (101) may receive the command (215) uttered by the user (109) after the virtual assistant is activated. For example, the command (215) may be the first command received after the virtual assistant is activated. For example, the electronic device (101) may receive (or acquire) the command (215) uttered by the user (109) through the microphone. As a non-limiting example, the command (215) may include "Tell me today's weather." For example, the electronic device (101) can recognize (or identify) a command (215). For example, upon recognizing the command (215), the electronic device (101) can execute a function corresponding to the command (215). As a non-limiting example, the function corresponding to the command (215) may include playing (or outputting) a message (217). As a non-limiting example, the message (217) may include “Today’s weather is clear.” Also, as a non-limiting example, the function corresponding to the command (215) may include displaying a screen (219). As a non-limiting example, the screen (219) may include a screen including a visual object representing today’s weather.
[0055] Referring to example (220), the user (100) may utter a command (225) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (225) uttered by the user (100) through a microphone. As a non-limiting example, the command (225) may include “play music.” The command (225) of example (220) may be received after the command (215) of example (210) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (225) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (215) and executing the function corresponding to the command (215), the electronic device (101) may receive an additional command (e.g., command (225)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (225). For example, upon recognizing the command (225), the electronic device (101) may execute the function corresponding to the command (225). As a non-limiting example, the function corresponding to the command (225) may include playing (or outputting) a message (227). As a non-limiting example, the message (227) may include, “Yes, play music.” Also, as a non-limiting example, the function corresponding to the command (225) may include displaying a screen (229). As a non-limiting example, the screen (229) may include a screen that includes a visual object representing music being played.
[0056] Referring to example (230), the user (109) may utter a command (235) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (235) uttered by the user (109) through a microphone. As a non-limiting example, the command (235) may include “play music.” The command (235) of example (230) may be received after the command (215) of example (210) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (235) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (215) and executing the function corresponding to the command (215), the electronic device (101) may receive an additional command (e.g., command (235)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (235). For example, upon recognizing the command (235), the electronic device (101) may execute the function corresponding to the command (235). As a non-limiting example, the function corresponding to the command (235) may include playing (or outputting) a message (237). As a non-limiting example, the message (237) may include, “Yes, play music.” Also, as a non-limiting example, the function corresponding to the command (235) may include displaying a screen (239). As a non-limiting example, the screen (239) may include a screen that includes a visual object representing music being played.
[0057] Unlike FIG. 1B, referring to FIG. 2, since the user (109) is the user who first uttered the command (215) received after the virtual assistant was activated, the electronic device (101) may temporarily allow the user (109) to interact with the virtual assistant. As a non-limiting example, the electronic device (101) may generate identification information indicating the user (109) who uttered the command (215). For example, the identification information may include a vector value indicating the user (109). For example, the electronic device (101) may use the identification information indicating the user (109) to identify whether additionally received commands (e.g., command (235)) are commands uttered by the user (109). Accordingly, the electronic device (101) can execute functions corresponding to commands (e.g., commands (235)) that are spoken by the user (109) and additionally received.
[0058] In an embodiment, referring to example (240), the user (209) may utter a command (245) to execute a function of the electronic device (101) by using the virtual assistant of the electronic device (101). For example, the electronic device (101) may receive (or acquire) the command (245) uttered by the user (109) via a microphone. As a non-limiting example, the command (245) may include “play music.” The command (245) of example (240) may be received after the command (215) of example (210) is received. When the continuous command function of the virtual assistant of the electronic device (101) is executing, the electronic device (101) may receive the command (245) without receiving an additional wake word. In other words, rather than deactivating the virtual assistant after receiving the command (155) and executing the function corresponding to the command (155), the electronic device (101) may perform reception of an additional command (e.g., command (245)) via the microphone. For example, the electronic device (101) may recognize (or identify) the command (245). For example, even if the electronic device (101) recognizes the command (245), it may refrain from executing (or, cease, stop, skip, bypass) the function corresponding to the command (245) (or, not execute the function corresponding to the command (245)). For example, the electronic device (101) may not recognize the command (245). Since the user (209) is not a designated user for the virtual assistant, the electronic device (101) may not recognize the command (245) uttered by the user (209), or even if it recognizes the command (245), it may refrain from executing a function corresponding to the command (245).In the present disclosure, the designated user may include a user registered with the virtual assistant, such as user (100), and a user temporarily permitted with respect to the virtual assistant, such as user (109). As a non-limiting example, the electronic device (101) may refrain from playing a message by refraining from executing the function corresponding to the command (245). Furthermore, the electronic device (101) may display the screen (219) by refraining from executing the function corresponding to the command (245). For example, the screen (219) may be a screen displayed when the function corresponding to the command (215) in example (210) is executed. In other words, the electronic device (101) may maintain the display of the screen (219) displayed according to the function corresponding to the previously received command (215), rather than displaying the screen (219) as the function corresponding to the command (245).
[0059] Referring to FIG. 2, the electronic device (101) can receive commands received from the user (100) and the user (109) while the continuous command function of the virtual assistant is being executed, and execute functions corresponding to the commands, without receiving an additional wake word (or obtaining a user input for activating the virtual assistant). In other words, the electronic device (101) can execute functions corresponding to commands received from not only registered users but also temporarily permitted users with respect to the virtual assistant while the continuous command function of the virtual assistant is being executed.
[0060] Figure 3 is a schematic view of an exemplary electronic device.
[0061] FIG. 3 illustrates an example of an electronic device (101). The electronic device (101) of FIG. 3 may be an example of the electronic device (1101) of FIG. 11. For example, the electronic device (101) may include at least a portion of, or correspond to at least a portion of, the electronic device (1101).
[0062] For example, the electronic device (101) may be implemented in various form factors. For example, the electronic device (101) may include not only an electronic device including the display of the bar type, but also an electronic device including the display that is a flexible display. For example, the flexible display may include an electronic device including a foldable display, an electronic device including a multi-foldable display, or an electronic device including a rollable display. In addition, for example, the electronic device (101) may include a tablet PC. In addition, for example, the electronic device (101) may be implemented as a wearable device. For example, the wearable device may include a head mounted display (HMD) or a watch-shaped device. However, the present disclosure is not limited thereto.
[0063] Referring to FIG. 3, according to one embodiment, an electronic device (101) may include at least one processor (310), an input interface (320), a speaker (330), and a memory (340). However, the embodiments of the present disclosure are not limited thereto. For example, at least one processor (310), an input interface (320), a speaker (330), and a memory (340) may be electrically and / or operably coupled with each other by a communication bus. Hereinafter, operably coupled hardware components may mean that a direct connection or an indirect connection is established between the hardware components, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components. Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the hardware components illustrated in FIG. 3 (e.g., at least one processor (310) and at least a portion of the memory (340)) may be included in a single integrated circuit such as a system on a chip (SoC) or a system in package (SIP). The type and / or number of hardware components included in the electronic device (101) is not limited to those illustrated in FIG. 3. For example, the electronic device (101) may include only some of the hardware components illustrated in FIG. 3. For example, the electronic device (101) may not include a speaker (330).
[0064] According to one embodiment, at least one processor (310) of the electronic device (101) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). As an example, the hardware component for processing data may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a microcontroller (MCU), and / or a neural processing unit (NPU). The number of at least one processor (310) may be one or more. For example, at least one processor (310) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core. At least one processor (310) of FIG. 3 may be substantially identically applied to the processor (1120) of FIG. 11.
[0065] For example, at least one processor (310) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which may be configured to individually and / or collectively perform the various functions described below in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms encompass, for example, and without limitation, situations where one processor performs some of the recited functions and other processor(s) perform other parts of the recited functions, and also situations where one processor may perform all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. At least one processor may execute program instructions to achieve or perform the various functions.
[0066] According to one embodiment, the electronic device (101) may include an input interface (320) configured to receive sound data. For example, the input interface may include a microphone for acquiring sounds (e.g., voice, noise, audio) input from the outside of the electronic device (101). Depending on the embodiment, the microphone may be implemented as a component of the electronic device (101) or as an external microphone connected to the electronic device (101) by wire or wirelessly. As a non-limiting example, the input interface (320) may include at least one microphone. The microphone may be, but is not limited to, a digital microphone, an electronic condenser microphone (ECM), or a micro electro mechanical system (MEMS). Specific details of the input interface (320) of FIG. 3 may be substantially identical to those of the input module (1150) of FIG. 11.
[0067] According to one embodiment, the electronic device (101) may include a speaker (330) for outputting audio information (e.g., sound or audio data). As a non-limiting example, the audio information output through the speaker (330) may include TTS output when playing TTS synthesized sound based on the execution of the virtual assistant application (350) of the electronic device (101). However, the present disclosure is not limited thereto. Specific details regarding the speaker (330) of FIG. 3 may be substantially identically applied to the audio output module (1155) (and / or audio module (1170)) of FIG. 11.
[0068] According to one embodiment, the electronic device (101) may include a memory (340). The memory (340) may include a hardware component for storing data and / or instructions input to and / or output from at least one processor (310). The memory (340) may include, for example, a volatile memory such as a random-access memory (RAM), and / or a non-volatile memory such as a read-only memory (ROM). The volatile memory may include, for example, at least one of a dynamic RAM (DRAM), a static RAM (SRAM), a cache RAM, and a pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disc, and an embedded multimedia card (eMMC). The specific details of the memory (340) of FIG. 3 can be applied substantially identically to the details of the memory (1130) of FIG. 11.
[0069] According to one embodiment, one or more instructions (or commands) representing operations and / or actions to be performed on data by at least one processor (310) of the electronic device (101) may be stored in the memory (340) of the electronic device (101). A set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or an application. Hereinafter, when an application is installed in the electronic device (e.g., the electronic device (101)), it may mean that one or more instructions provided in the form of an application are stored in the memory (340), and that the one or more applications are stored in a format executable by the processor of the electronic device (e.g., a file having an extension designated by the operating system of the electronic device (101)). According to one embodiment, the electronic device (101) may execute one or more instructions stored in the memory (340) to perform the operations of FIGS. 5 to 8. For example, the one or more instructions, when executed by at least one processor (310), may cause the electronic device (101) to perform at least some of the operations of FIGS. 5 to 8.
[0070] The electronic device (101) may include a display. For example, the display may be used to display a screen to be displayed based on the execution of the virtual assistant application (350) of the electronic device (101). For example, specific details regarding the display may be substantially identical to those regarding the display module (1160) of FIG. 11 . In addition, the electronic device (101) may include a module for power supply. For example, the electronic device (101) may include a battery. Specific details regarding the battery may be substantially identical to those regarding the battery (1189) of FIG. 11 .
[0071] For example, the memory (340) may include (or store) a virtual assistant application (350). For example, the virtual assistant application (350) may be referred to as an application for providing a virtual assistant. The electronic device (101) may receive a voice signal through the input interface (320) periodically (or aperiodically) by executing the virtual assistant application (350) (or executing the virtual assistant application (350) in the background). For example, the electronic device (101) may activate the virtual assistant upon recognizing that a wake word is included in the voice signal received through the input interface (320). In other words, the electronic device (101) may activate the virtual assistant, which was deactivated while executing the virtual assistant application (350), upon recognizing the wake word.
[0072] For example, the memory (340) may include (or store) a call word recognition module (355). For example, the call word recognition module (355) may be used to determine (or identify, recognize) whether a call word is included in a voice signal received through the input interface (320). As a non-limiting example, the call word recognition module (355) may be used only to determine whether a call word is included in the voice signal. As a non-limiting example, the call word recognition module (355) may be used to determine whether a call word is included in the voice signal and whether the call word was spoken by a designated user. The designated user may include a registered user with respect to the virtual assistant and a temporarily permitted user with respect to the virtual assistant. For example, the call word recognition module (355) that determines whether a call word is included in a voice signal and whether the call word was spoken by the designated user may be referred to as text-dependent speaker verification (TDSV). For example, the TDSV may be a model trained to recognize a user who has uttered a specific word (e.g., a call word). For example, the description of the TDSV may be substantially identical to the description of the text-independent speaker verification (TISV) (375) below.
[0073] For example, the memory (340) may include (or store) an automatic speech recognition (ASR) module (360). For example, the ASR module (360) may be used to convert a voice signal received through the input interface (320) into text. For example, the text may include characters representing a voice portion within the voice signal. For example, the text may include a wake word, or a wake word and a command. The electronic device (101) may analyze the intention of the user (or speaker) who uttered the voice signal from the text through natural language understanding (NLU). For example, the electronic device (101) may recognize the intention of the user by analyzing the text converted from the voice signal based on NLU.
[0074] For example, the memory (340) may include (or store) a speaker feature extraction model (370). For example, the speaker feature extraction model (370) may be referred to as a model trained for speaker feature extraction. For example, the speaker feature extraction model (370) may be used to generate identification information indicating a user who uttered a voice signal received through the input interface (320). As a non-limiting example, the identification information may include at least one vector value indicating the user. For example, the vector value indicating the user may be referred to as an embedding, an embedding vector, a speaker feature vector, a speaker feature value, or a speaker model. In FIG. 3, the speaker feature extraction model (370) may also be included (or implemented) in the TISV (375) (or the TDSV). For specific details on the speaker feature extraction model (370), reference may be made to FIGS. 4a and 4b below.
[0075] For example, the memory (340) may include a TISV (375). The TISV (375) may be referred to as a model trained for speaker verification. For example, the TISV (375) may be used to generate identification information indicating a user who uttered the voice signal based on a voice signal received through the input interface (320), and to determine (or identify, determine, verify) whether the user indicated by the identification information corresponds to (or is identical to, or matches) the user indicated by the referenced identification information. The referenced identification information may be referred to as reference identification information. As a non-limiting example, the reference identification information may include a vector value indicating the user. For specific details regarding the TISV (375), reference may be made to FIG. 4C below.
[0076] For example, the memory (340) may include (or store) a voice filter (VF) (380). The VF (380) may be referred to as a trained model for voice filtering. For example, the VF (380) may be used to filter (or identify, detect) a speech portion within a speech signal received via the input interface (320). As a non-limiting example, the VF (380) may be used to filter a speech portion uttered by a user indicated by identification information used as a reference within the speech signal. Filtering the speech portion within the speech signal may include suppressing (or removing, reducing) a portion (e.g., noise, speech portions of other users) other than the speech portion within the speech signal. For specific details regarding the VF (380), reference may be made to FIG. 4D below.
[0077] For example, the memory (340) may include (or store) a voice activity detector (VAD) (385). For example, the VAD (385) may be referred to as a trained model for voice activity detection. As a non-limiting example, the VAD (385) may be used to detect a voice portion within a voice signal received via the input interface (320). As a non-limiting example, the VAD (385) may output a voice portion within the voice signal as 1 and a non-voice portion within the voice signal as 0. For specific details regarding the independent VAD (385) for a specific user, reference may be made to FIG. 4e below. For specific details regarding the dependent VAD (385) for a specific user, reference may be made to FIG. 4f below. The dependent VAD (385) for a specific user may replace or complement the TISV (375) in detecting the voice portion of a specific user. VAD (385) can perform EPD (end-point detection). For example, the EPD can be used to detect the end point (or point in time) of a speech portion within the speech signal.
[0078] For example, the memory (340) may include (or store) a set of identification information (390). For example, the set of identification information (390) may include first identification information (391), second identification information (392), and up to n-th identification information (393). Although FIG. 3 illustrates an identification information set (390) including three or more pieces of identification information, the present disclosure is not limited thereto. For example, the set of identification information (390) may include two or fewer pieces of identification information. For example, the identification information included in the set of identification information (390) may indicate a designated user. As a non-limiting example, the first identification information (391) may indicate a registered user with respect to the virtual assistant. As a non-limiting example, the second identification information (392) may indicate a temporarily permitted user with respect to the virtual assistant.
[0079] In the above example, if the second identification information (392) indicates a temporarily permitted user, the second identification information (392) may be deleted from the identification information set (390) (or, memory (340), electronic device (101)) as a specific event is identified. As a non-limiting example, the specific event may include at least one of: deactivation of the activated virtual assistant, deactivation of a continuous command function of the virtual assistant, or identification that the user indicated by the second identification information (392) matches a registered user with respect to the virtual assistant (e.g., the user indicated by the first identification information (391). In the above example, the second identification information (392) indicating the temporarily permitted user is depicted as being stored in a storage space of the memory (340) (e.g., the identification information set (390) within the memory (340), but the present disclosure is not limited thereto. For example, the second identification information (392) indicating the temporarily permitted user may be stored in another storage space of the memory (340). As a non-limiting example, the other storage space may be referred to as a storage space for temporary storage (or caching).
[0080] Although various aspects of the memory (340) are illustrated individually, it should be understood that two or more such elements, modules, or units may be combined into a single element, module, or unit that performs all of the operations or functions of the combined two or more elements, modules, or units. Furthermore, at least some of the functions of at least one of such elements, modules, or units may be performed by another element, module, or unit.
[0081] Hereinafter, in the present disclosure, providing a signal to a model (e.g., speaker feature extraction model (370), TISV (375), VF (380), VAD (385)) may refer to inputting, using as input, or feeding an input signal (or input data) to the model. Furthermore, in the present disclosure, generating a signal from models (e.g., speaker feature extraction model (370), TISV (375), VF (380), VAD (385)) may refer to outputting, using as output, or obtaining an output signal (or output data) from the model.
[0082] Figures 4a to 4f illustrate examples of trained models included within an electronic device.
[0083] FIGS. 4A and 4B illustrate examples of a speaker feature extraction model (370) included in the electronic device (101) (or memory (340)) of FIG. 3. Referring to FIG. 4A, an example (401) of a method for the speaker feature extraction model (370) to generate a speaker feature vector (419) based on a received speech signal (410) is illustrated. Referring to FIG. 4B, an example (402) of a method for the speaker feature extraction model (370) to generate identification information based on a plurality of speech signals is illustrated.
[0084] Referring to example (401), the electronic device (101) can provide (or feed, input) a voice signal (410) to the speaker feature extraction model (370). For example, the voice signal (410) can be received through the input interface (320) of the electronic device (101).
[0085] For example, the speaker feature extraction model (370) may perform preprocessing (411) on the provided speech signal (410). As a non-limiting example, the preprocessing (411) may include amplification of the magnitude of the speech signal (410). As a non-limiting example, the preprocessing (411) may include normalization of the speech signal (410). As a non-limiting example, the preprocessing (411) may include noise removal of the speech signal (410).
[0086] For example, the speaker feature extraction model (370) may perform feature extraction (413) on a speech signal (410) on which preprocessing (411) has been performed (hereinafter, referred to as a preprocessed speech signal (410)). For example, the feature extraction (413) may convert the preprocessed speech signal (410) into a vector. For example, the vector converted from the preprocessed speech signal (410) may be data representing the voice of the speech signal (410). According to the feature extraction (413), the amount of data of the preprocessed speech signal (410) may be reduced. As a non-limiting example, the feature extraction (413) may include mel-frequency cepstral coefficients (MFCC). As a non-limiting example, the feature extraction (413) may include a log-mel filter bank. As a non-limiting example, feature extraction (413) may include linear predictive coding (LPC). As a non-limiting example, feature extraction (413) may include a spectrogram.
[0087] For example, the speaker feature extraction model (370) can generate (or output) a speaker feature vector (419) using the speaker vector extraction model (415) for the speech signal (410) on which feature extraction (413) has been performed (hereinafter, the extracted speech signal (410)). For example, the speaker vector extraction model (415) can be used to generate a speaker feature vector (419) for indicating a user (or speaker) who uttered the speech signal (410) for input data (e.g., a vector or an extracted speech signal (410)). For example, the speaker vector extraction model (415) can use modeling. As a non-limiting example, the speaker vector extraction model (415) can use a GMM (Gaussian mixture model) supervector. As a non-limiting example, the speaker vector extraction model (415) can use an i-vector. Alternatively, for example, the speaker vector extraction model (415) may utilize deep learning.
[0088] As a non-limiting example, the speaker feature extraction model (370) may be included in the TISV (375) or the TDSV. The speaker feature extraction model (370) included in the TISV (375) may generate a speaker feature vector (419) indicating a user who uttered the speech signal (410), regardless of the words or sentences included in the speech signal (410). The speaker feature extraction model (370) included in the TDSV may recognize specific words or specific sentences included in the speech signal (410) and generate a speaker feature vector (419) indicating a user who uttered the speech signal (410).
[0089] As described above, the speaker feature extraction model (370) can generate identification information based on at least one voice signal of a specific speaker, thereby storing (or registering) the identification information for the specific speaker in the memory (340). For specific details related thereto, reference may be made to FIG. 4B.
[0090] Referring to example (402), the electronic device (101) can provide (or feed, input) at least one voice signal to the speaker feature extraction model (370). In example (402), the at least one voice signal 410 can include a first voice signal (421-1), a second voice signal (422-1), and a third voice signal (423-1). In FIG. 4B, for convenience of explanation, three voice signals are illustrated as being provided to the speaker feature extraction model (370), but the present disclosure is not limited thereto.
[0091] For example, the speaker feature extraction model (370) can generate (or output) at least one speaker feature vector based on the at least one speech signal. For example, the at least one speaker feature vector 419 can include a first speaker feature vector (421-2), a second speaker feature vector (422-2), and a third speaker feature vector (423-2). Referring to FIG. 4A, the speaker feature extraction model (370) generates the speaker feature vector based on the speech signal.
[0092] For example, the identification information (425) can be identified based on the first speaker feature vector (421-2), the second speaker feature vector (422-2), and the third speaker feature vector (423-2). For example, the identification information (425) can be identified as a representative value (e.g., mean value, median value) of the first speaker feature vector (421-2), the second speaker feature vector (422-2), and the third speaker feature vector (423-2).
[0093] In the example of Fig. 4b, three speaker feature vectors are illustrated as constituting identification information (425), but the present disclosure is not limited thereto. For example, the identification information (425) may be composed of one speaker feature vector. In addition, in Fig. 4b, one speech signal is illustrated as being output as one speaker feature vector, but the present disclosure is not limited thereto. For example, one speech signal may be used to output multiple speaker feature vectors, or one speech signal may be distinguished into multiple speech signals, and the distinguished multiple speech signals may be used to output multiple speaker feature vectors. For example, when one speech signal is distinguished into multiple speech signals, one speech signal may be distinguished by word within one speech signal, or the time length constituting one speech signal may be distinguished by a specific time length (or frame).
[0094] In FIG. 4b, an example (402) of generating identification information (425) based on at least one voice signal using a speaker feature extraction model (370) is illustrated. Accordingly, the electronic device (101) can generate identification information (425) indicating a user who uttered the at least one voice signal. At this time, generating (or storing) the identification information (425) can be referred to as registering the user with respect to a virtual assistant application (350). After registering the user with respect to the virtual assistant application (350), the electronic device (101) can analyze a voice signal received from the user while the virtual assistant is activated and recognize that the voice signal is uttered by the user using the identification information (425). For specific details related thereto, reference may be made to FIG. 4c below.
[0095] FIG. 4C illustrates an example of a TISV (375) included in the electronic device (101) (or, memory (340)) of FIG. 3. Referring to FIG. 4C, an example (403) of a method for determining the similarity between a user who uttered a voice signal (431) and a user indicated by reference identification information (435-1) based on the received voice signal (431) is illustrated. In example (403), for convenience of explanation, a method for determining the similarity between a user who uttered a voice signal (431) and a user indicated by reference identification information (435-1) is illustrated using the TISV (375), but the present disclosure is not limited thereto. The contents of the TISV (375) can be substantially equally applied to the TDSV.
[0096] Referring to example (403), TISV (375) may include a speaker feature extraction model (370). For example, the electronic device (101) may provide a voice signal (431) received through the input interface (320) to the speaker feature extraction model (370). For example, the speaker feature extraction model (370) of TISV (375) may generate a speaker feature vector (433) indicating a user who uttered the voice signal (431).
[0097] For example, the TISV (375) may perform similarity value identification (435) based on the speaker feature vector (433) and the reference identification information (435-1). For example, the similarity value identification (435) may refer to identifying a similarity value between the speaker feature vector (433) generated based on the speech signal (431) and the reference identification information (435-1) used as a reference. As a non-limiting example, the TISV (375) may include a model trained for identifying similarity values or a module for an operation for identifying similarity values. For example, the module may be referred to as an algorithm for the operation rather than using a trained model. In other words, the similarity value identification (435) may be performed by an operation of at least one processor (310) without the TISV (375). For example, the reference identification information (435-1) may be identification information within a set of identification information (390) stored within the memory (340). For example, the reference identification information (435-1) may indicate a designated user (e.g., a registered user or a temporarily permitted user) with respect to the virtual assistant. As a non-limiting example, the reference identification information (435-1) may include multiple identification information corresponding to multiple users. In one example, the reference identification information (435-1) may include identification information corresponding to a first registered user, identification information corresponding to a second registered user, and identification information corresponding to a temporarily permitted user.
[0098] For example, TISV (375) can perform a comparison (437) between the similarity value output according to the similarity value identification (435) and a reference value. For example, the reference value can be a value for determining whether the user who uttered the voice signal (431) corresponds to (or matches, is identical to) the user indicated by the reference identification information (435-1). For example, TISV (375) can output a first value (439-1) if the similarity value is greater than or equal to the reference value according to the comparison (437). Alternatively, TISV (375) can output a second value (439-2) if the similarity value is less than the reference value according to the comparison (437). For example, the first value (439-1) may indicate that the user who uttered the voice signal (431) corresponds to the user indicated by the reference identification information (435-1). For example, the second value (439-2) may indicate that the user who uttered the voice signal (431) does not correspond to (or is different from) the user indicated by the reference identification information (435-1). As a non-limiting example, the first value (439-1) may be 'True'. As a non-limiting example, the second value (439-2) may be 'False'. As a non-limiting example, the comparison (437) and the output of the result according to the comparison (437) may be performed by the operation of at least one processor (310) without the TISV (375).
[0099] Referring to the above, the electronic device (101) can determine (or judge, identify, verify) whether the user who uttered the voice signal (431) corresponds to (or matches, or is identical to) the user indicated by the reference identification information (435-1) by providing the voice signal (431) to the TISV (375) (and the speaker feature extraction model (370)), using the output result (e.g., the first value (439-1) or the second value (439-2)).
[0100] FIG. 4D illustrates an example of a VF (380) included in the electronic device (101) (or memory (340)) of FIG. 3. Referring to FIG. 4D, an example (404) of a method for extracting a speech portion uttered by a specific user from a speech signal (441) based on the received speech signal (441) is illustrated. For convenience of explanation, it is assumed that the speech signal (441) includes speech portions uttered by multiple users. However, the present disclosure is not limited thereto.
[0101] Referring to example (404), the electronic device (101) can provide (or feed, input) a voice signal (441) received through the input interface (320) to the VF (380). For example, the VF (380) can perform preprocessing (443) on the provided voice signal (441). For example, the specific details of the preprocessing (443) of FIG. 4D can be substantially identically applied to the details of the preprocessing (411) of FIG. 4A. However, the present disclosure is not limited thereto. For example, the preprocessing (443) may include processes other than the preprocessing (411). According to the preprocessing (443), the voice signal (441) can be converted into a frequency domain according to a fast Fourier transform (FFT), and magnitudes in the frequency domain can be extracted.
[0102] For example, the VF (380) may provide a speech signal (441) on which preprocessing (443) has been performed (hereinafter, referred to as a preprocessed speech signal (441)) to a speech enhancement model (445). For example, the speech enhancement model (445) may be implemented as a neural network (NN). For example, the speech enhancement model (445) may be used to maintain frequency energy (or magnitude in the frequency domain) corresponding to a portion of the user's speech indicated by the reference identification information (445-1) within the preprocessed speech signal (441) and reduce energy corresponding to the remaining portions of the user's speech. For example, the reference identification information (445-1) may be identification information within an identification information set (390) stored within a memory (340). For example, the reference identification information (445-1) may indicate a designated user (e.g., a registered user or a temporarily authorized user) with respect to the virtual assistant.
[0103] For example, the VF (380) can perform an IFFT (inverse-FFT) (447) on the speech signal (441) output from the speech enhancement model (445). For example, the VF (380) can convert (or restore) the speech signal (441) converted to the frequency domain back into a speech signal (or time domain) by performing the IFFT (447). Accordingly, the VF (380) can output a filtered speech signal (449). The filtered speech signal (449) can be a speech signal in which, compared to the speech signal (441), the part of the user's speech indicated by the reference identification information (445-1) is maintained, and the remaining parts of the user's speech are reduced. In other words, the speech signal (441) can be filtered based on the VF (380) so that the part of the user's speech indicated by the reference identification information (445-1) is left prominent.
[0104] FIG. 4E illustrates an example of a VAD (385) included in the electronic device (101) (or memory (340)) of FIG. 3. Referring to FIG. 4E, an example (405) of a method for detecting a speech portion from a speech signal (451) based on the received speech signal (451) is illustrated. For convenience of explanation, it is assumed that the speech signal (451) includes speech portions spoken by multiple users. However, the present disclosure is not limited thereto.
[0105] Referring to example (405), the electronic device (101) may provide (or feed, input) a voice signal (451) received through the input interface (320) to the VAD (385). For example, the VAD (385) may perform preprocessing (453) on the provided voice signal (451). For example, the specific details of the preprocessing (453) of FIG. 4E may be substantially identical to the details of the preprocessing (411) of FIG. 4A. However, the present disclosure is not limited thereto. For example, the preprocessing (453) may include processes other than the preprocessing (411).
[0106] For example, the VAD (385) can provide a speech signal (451) on which preprocessing (453) has been performed (hereinafter, referred to as a preprocessed speech signal (451)) to the VAD model (455). For example, the VAD model (455) can be implemented with statistical signal processing or a neural network (NN). For example, the VAD model (455) can be used to detect a speech portion within the preprocessed speech signal (451). For example, the detected speech portion can represent a speech spoken by a person, regardless of the users. For example, the VAD model (455) can detect whether a speech portion exists at a specific time interval (e.g., 20 ms (milliseconds)) (or frame) for the preprocessed speech signal (451).
[0107] For example, the VAD (385) can perform post-processing (457) on the speech signal (451) output from the VAD model (455). For example, the post-processing (457) may include compensating for distortion resulting from the pre-processing (453) or outputting a value indicating a speech portion detected by the VAD model (455). However, the present disclosure is not limited thereto. For example, the VAD (385) may skip the post-processing (457). Accordingly, the VAD (385) can generate (or output) the detected speech signal (459) without performing the post-processing (457).
[0108] For example, the VAD (385) may generate (or output) a detected speech signal (459) based on the speech signal (451). For example, the detected speech signal (459) may indicate '1' for a specific time interval if a speech portion exists within the specific time interval, and may indicate '0' for the specific time interval if a speech portion does not exist within the specific time interval. In the example (405), the detected speech signal (459) may indicate '000111111111000'. For example, the detected speech signal (459) may include a speech segment and a non-speech segment within the speech signal (451). For example, the speech segment may be '111111111'. For example, the non-speech segment may be '000', '000'. In other words, VAD (385) can detect the speech section of the speech part.
[0109] In some embodiments, the VAD (385) may be used not only to detect a speech portion within a speech signal (451), but also to identify a location (or a point in time) where the speech portion ends. For example, identifying a location where the speech portion ends may be referred to as EPD (end-point detection). In example (405), the location (or point in time) where the speech portion ends within the speech signal (451) may be a location (or a point in time) indicated by the last 1 of '000111111111000' of the detected speech signal (459) (or a location (or a point in time) where the last 1 changes to a consecutive 0). In the above example, identifying a location where the speech portion ends is illustrated, but the present disclosure is not limited thereto. For example, a location where a speech portion begins within the speech signal (451) may also be identified.
[0110] FIG. 4F illustrates an example of a VAD (385) included in the electronic device (101) (or memory (340)) of FIG. 3. Referring to FIG. 4F, an example (406) is illustrated of a method for detecting a speech portion uttered by a specific user from a speech signal (461) based on the received speech signal (461). In other words, unlike the VAD (385) of example (405), the VAD (385) of example (406) can be used to detect a speech portion of a specific user, rather than all speech portions within the speech signal. Accordingly, the VAD (385) of example (406) can be referred to as a personalized VAD. For convenience of explanation, it is assumed that the speech signal (461) includes speech portions uttered by multiple users. However, the present disclosure is not limited thereto.
[0111] Referring to example (406), the electronic device (101) can provide (or feed, input) a voice signal (461) received through the input interface (320) to the VAD (385). For example, the VAD (385) can perform preprocessing (463) on the provided voice signal (461). For example, the specific details of the preprocessing (463) of FIG. 4F can be substantially identically applied to the details of the preprocessing (411) of FIG. 4A. However, the present disclosure is not limited thereto. For example, the preprocessing (463) may include processes other than the preprocessing (411).
[0112] For example, the VAD (385) may provide a speech signal (461) on which preprocessing (463) has been performed (hereinafter, referred to as a preprocessed speech signal (461)) to a personalized VAD model (465). For example, the personalized VAD model (465) may be implemented as a neural network (NN). For example, the personalized VAD model (465) may be used to detect a speech portion of a user indicated by the reference identification information (465-1) within the preprocessed speech signal (451) using the reference identification information (465-1). At this time, the user indicated by the reference identification information (465-1) may be referred to as a target user (or target speaker). For example, the detected speech portion may represent a speech uttered by the user indicated by the reference identification information (465-1). For example, the personalized VAD model (465) can detect whether a voice portion exists at a specific time interval (e.g., 20 ms (milliseconds)) (or frame) for the preprocessed voice signal (461). For example, the reference identification information (465-1) can be identification information within a set of identification information (390) stored in the memory (340). For example, the reference identification information (465-1) can indicate a designated user (e.g., a registered user or a temporarily permitted user) with respect to the virtual assistant.
[0113] For example, the VAD (385) may perform post-processing (467) on a voice signal (461) output from a personalized VAD model (465). For example, the post-processing (467) may include compensating for distortion resulting from the pre-processing (463) or outputting a value indicating a voice portion detected by the personalized VAD model (465).
[0114] For example, the VAD (385) may generate (or output) a detected speech signal (469) based on the speech signal (461). For example, the detected speech signal (469) may indicate '1' for a specific time interval if a speech portion exists within the specific time interval, and may indicate '0' for the specific time interval if a speech portion does not exist within the specific time interval. In example (406), the detected speech signal (469) may indicate '000000111111000'. For example, the detected speech signal (469) may include a speech segment and a non-speech segment within the speech signal (461). For example, the speech segment may be '111111'. For example, the non-speech segment may be '000000', '000'. In other words, VAD (385) can detect the speech section of the speech part.
[0115] If the voice signal (451) of example (405) and the voice signal (461) of example (406) are the same, compared to the detected voice signal (459) of example (405), the detected voice signal (469) of example (406) can indicate '111111', which is the user's voice part indicated by the reference identification information (465-1).
[0116] In some embodiments, the VAD (385) may be used not only to detect speech segments within a speech signal (461), but also to identify the location (or point in time) where the speech segment ends. For example, identifying the location where the speech segment ends may be referred to as end-point detection (EPD). For specific details related thereto, refer to FIG. 4E.
[0117] As mentioned above, the VAD (385) of FIG. 4f can be used to replace or supplement the TISV (375) (or TDSV) because it can detect a portion of speech uttered by a specific user within a speech signal. In other words, the examples exemplified below in which the TISV (375) is used can be substantially equally applied to the VAD (385) instead of the TISV (375).
[0118] As a non-limiting example, the reference identification information (465-1) used in the VAD (385) of example (406), the reference identification information (445-1) used in the VF (380), and the reference identification information (435-1) used in the TISV (375) may be different from each other. For example, the reference identification information (465-1) indicating the specific user, the reference identification information (445-1) indicating the specific user, and the reference identification information (435-1) indicating the specific user may be different from each other. In other words, in models available for recognizing a user (e.g., TISV (375), VF (380), VAD (385) of example (406)), the reference identification information indicating the same user may be different from each other.
[0119] At least one of the components, elements, modules, or units represented by the blocks illustrated in FIGS. 4A-4F may be implemented by various combinations of hardware, software, and / or firmware structures that perform the respective functions described above, according to exemplary embodiments. For example, at least one of these components, elements, modules, or units may utilize direct circuit structures such as memory, processing, logic, lookup tables, etc., which may perform the respective functions under the control of one or more microprocessors or other control devices. Furthermore, at least one of these components, elements, modules, or units may be specifically implemented as part of a module, program, or code that includes one or more executable instructions for performing a specified logic function and is executed by one or more microprocessors or other control devices. Furthermore, at least one of these components, elements, modules, or units may further include a processor, such as a central processing unit (CPU), microprocessor, or the like, that performs the respective functions. Two or more of these components, elements, modules, or units may be combined into a single component, element, module, or unit that may perform all operations or functions of the two or more combined components, elements, modules, or units. Additionally, at least a portion of the functions of at least one of these components, elements, modules, or units may be performed by other components, elements, modules, or units. Furthermore, although a bus is not depicted in the block diagram, communication between the components, elements, modules, or units may be performed via a bus. The functional aspects of the exemplary embodiments may be implemented as algorithms executed on one or more processors. Furthermore, the components, elements, modules, or units represented as blocks or processing steps may utilize various related technologies for electronic configuration, signal processing and / or control, data processing, etc.
[0120] Figure 5 illustrates an example of a workflow for training a model for a virtual assistant for a specific user.
[0121] At least some of the methods of FIG. 5 may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0122] According to an embodiment, the electronic device (101) may execute a virtual assistant application (350). As a non-limiting example, the electronic device (101) may execute the virtual assistant application (350) in the foreground. The screen of the virtual assistant application (350) executed in the foreground may be displayed through the display of the electronic device (101). For example, the screen may include a menu for configuring the virtual assistant application (350). For example, an example of a menu for configuring the virtual assistant application (350) may be illustrated in FIG. 10.
[0123] Referring to FIG. 5, in operation (500), according to an embodiment, the electronic device (101) may display a learning text. For example, the electronic device (101) may display the learning text within the screen displayed through the display of the electronic device (101). For example, the learning text may be text used to register a user with a virtual assistant of the virtual assistant application (350) (or to train models for the user). The learning text may be referred to as spoken text. As a non-limiting example, the learning text may include a wake word, a command, a specific sentence, or a specific word.
[0124] In operation (500), a case where the learning text is displayed is illustrated, but the present disclosure is not limited thereto. For example, the electronic device (101) may output TTS synthesized sound corresponding to the learning text as audio information through the speaker (330).
[0125] In operation (505), according to one embodiment, the electronic device (101) may receive a voice signal. For example, the electronic device (101) may wait to receive the voice signal by activating the input interface (320). For example, activating the input interface (320) may indicate switching to a state in which a sound signal (or voice signal) can be received from outside the electronic device (101) through the input interface (320). In the above example, the input interface (320) is described as being activated to receive the voice signal, but the present disclosure is not limited thereto. For example, the electronic device (101) may also activate the input interface (320) regardless of the display of the learning text. For example, the electronic device (101) may receive the voice signal through the activated input interface (320).
[0126] In operation (510), according to one embodiment, the electronic device (101) may determine whether the voice signal corresponds to the training text based on speech verification. For example, the electronic device (101) may perform speech verification. For example, the electronic device (101) may determine (or verify, identify, or determine) whether the voice signal corresponds to the training text displayed in operation (500) by converting the voice signal received through the input interface (320) into text. For example, converting the voice signal into text may be performed using the ASR module (360) of the electronic device (101).
[0127] In operation (510), if the speech signal corresponds to the training text according to the speech verification, the electronic device (101) may perform operation (515). If the speech signal corresponds to the training text, the speech verification may be considered successful. In operation (515), if the speech signal does not correspond to the training text according to the speech verification, the electronic device (101) may perform operation (500) again. If the speech signal does not correspond to the training text, the speech verification may be considered failed.
[0128] In operation (515), according to one embodiment, the electronic device (101) may determine whether the utterance verification has been performed a reference number of times. For example, the electronic device (101) may determine whether the number of times the utterance verification has been successfully performed is the reference number of times. In operation (515), if the utterance verification has not been performed a reference number of times, the electronic device (101) may perform operation (500) again. In operation (515), if the utterance verification has been performed a reference number of times, the electronic device (101) may perform operation (520).
[0129] When operation (500) is performed again, the learning text may be changed to a different learning text. However, the present disclosure is not limited thereto. For example, the learning text may be maintained.
[0130] In operation (520), according to one embodiment, the electronic device (101) may perform training of the model. For example, the electronic device (101) may perform training of the model based on the voice signal received in operation (505). The model for which training is performed may include the speaker feature extraction model (370) of FIGS. 4A and 4B , the TISV (375) (or TDSV) of FIG. 4C , the VF (380) of FIG. 4D , or the VAD (385) of FIG. 4F . As a non-limiting example, the electronic device (101) may perform training of the model based on the voice signal including the call word. Accordingly, when the electronic device (101) receives a voice signal including the call word, the electronic device (101) may recognize the user who has spoken the call word using the trained model.
[0131] The electronic device (101) can generate identification information indicating a specific user who uttered a voice signal, verify a specific user who uttered a voice signal, filter a voice portion of a specific user within a voice signal, or detect a voice portion uttered by a specific user within a voice signal, based on a model trained according to the method of FIG. 5, as described later in FIGS. 6A to 8.
[0132] FIG. 6A illustrates an example of an operational flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a registered user with respect to a virtual assistant.
[0133] At least some of the methods of FIG. 6A may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0134] In some embodiments, prior to performing operation (600), the electronic device (101) may register (or store) a user with respect to the virtual assistant. For example, the user may be referred to as a registered user. Additionally, the electronic device (101) may execute the virtual assistant application (350) in the background.
[0135] In operation (600), according to one embodiment, the electronic device (101) may receive a voice signal. For example, the electronic device (101) may receive the voice signal through an activated input interface (320).
[0136] In operation (605), according to one embodiment, the electronic device (101) may determine whether a wake word is recognized. For example, the electronic device (101) may recognize a wake word within a received voice signal by converting the received voice signal into text.
[0137] In operation (605), if a call word is recognized in the received voice signal, the electronic device (101) may perform operation (610). In operation (605), if a call word is not recognized in the received voice signal, the electronic device (101) may perform operation (600) again.
[0138] For example, the electronic device (101) may activate the virtual assistant when a wake word is recognized in the received voice signal. In FIG. 6A, the virtual assistant is illustrated as being activated depending on whether the wake word is recognized, but the present disclosure is not limited thereto. For example, the electronic device (101) may also activate the virtual assistant in response to an input (or user input) to the electronic device (101). For example, the input may include an input to a physical button of the electronic device (101) or a visual object (or icon) displayed through the display of the electronic device (101).
[0139] As a non-limiting example, the electronic device (101) may identify the user who uttered the call word based on the recognized call word. For example, the electronic device (101) may recognize that the user who uttered the call word is the registered user.
[0140] In operation (610), according to one embodiment, the electronic device (101) may receive a voice signal including a command. For example, the electronic device (101) may receive the voice signal including the command through the input interface (320). In FIG. 6A, for convenience of explanation, the voice signal received in operation (600) is illustrated as including a wake word and the voice signal received in operation (610) as including a command. However, the present disclosure is not limited thereto. For example, a single voice signal may include both a wake word and a command.
[0141] In operation (615), according to one embodiment, the electronic device (101) may perform a voice filtering (VF) for a registered user. For example, the electronic device (101) may output a filtered voice signal through the VF (380) for the voice signal including the command received in operation (610). For example, the electronic device (101) may provide the voice signal including the command to the VF (380). For example, the electronic device (101) may filter a voice portion spoken by the registered user within the voice signal including the command through the VF (380) using reference identification information indicating the registered user. For example, the registered user may correspond to the user recognized in operation (605) of recognizing a wake word. For example, the electronic device (101) may perform VF (and / or TISV of operation (620) to be described later) using identification information corresponding to the user who uttered the call word while recognizing the call word.
[0142] In one example, if the voice signal including the command is uttered by the registered user, the filtered voice signal may have a relatively high frequency energy in the voice portion uttered by the registered user. Conversely, if the voice signal including the command is uttered by a user different from the registered user, the filtered voice signal may have a relatively low frequency energy throughout the voice signal. This may be because the frequency energy of users other than the registered user is reduced by the VF (380).
[0143] In operation (620), according to one embodiment, the electronic device (101) may perform TISV for the registered user. For example, the electronic device (101) may output a result indicating whether the user who uttered the voice signal including the command corresponds to the registered user through the TISV (375) for the voice signal filtered in operation (615). For example, the electronic device (101) may provide the filtered voice signal to the TISV (375). For example, the electronic device (101) may generate a speaker feature vector (or identification information) through the TISV (375) (or the speaker feature extraction model (370)) based on the filtered voice signal. For example, the electronic device (101) may identify a similarity value between the speaker feature vector (or identification information) and the reference identification information indicating the registered user. For example, the electronic device (101) can output a result (e.g., the first value (439-1) or the second value (439-2) of FIG. 4c) indicating whether the user who uttered the voice signal including the command is the registered user by performing a comparison between the similarity value and the reference value through the TISV (375).
[0144] In operation (625), according to one embodiment, the electronic device (101) may determine whether the user corresponds to a registered user. For example, the electronic device (101) may determine whether the user who uttered the voice signal including the command received in operation (610) corresponds to (or is identical to, or matches) the registered user.
[0145] In operation (625), if the user who uttered the voice signal including the command received in operation (610) corresponds to the registered user, operation (630) may be performed. In operation (625), if the user who uttered the voice signal including the command received in operation (610) does not correspond to the registered user, operation (600) may be performed again.
[0146] In operation (630), according to one embodiment, the electronic device (101) may execute a function corresponding to the command. For example, the electronic device (101) may execute the function corresponding to the received command based on the activated virtual assistant.
[0147] In FIG. 6A, TISV is described as being performed in operation (620) on the output result (e.g., a filtered voice signal) after performing operation (615), but the present disclosure is not limited thereto. For example, the electronic device (101) may perform VF and TISV on the voice signal received in operation (610), respectively. This may be to compensate for distortion that may occur as VF is performed. Depending on the embodiment, for example, the electronic device (101) may omit (or not perform, skip, refrain from performing, bypass) operation (615) and perform operation (620). Depending on the embodiment, for example, the electronic device (101) may omit (or not perform, skip, refrain from performing, bypass) at least a part of operation (620). For example, at least a portion of operation (620) may include identifying a similarity value, performing a comparison between the similarity value and a reference value, and outputting a result indicating whether the user is a registered user. In other words, in operation (620), the electronic device (101) may only generate identification information based on the voice signal received in operation (610).
[0148] Also, for example, the electronic device (101) can utilize the personalized VAD (385) of FIG. 4F. For example, instead of determining whether the user who uttered the voice signal received through the TISV (375) corresponds to the registered user in operation (620), the electronic device (101) can determine whether the user who uttered the voice signal received through the personalized VAD (385) corresponds to the registered user. According to an embodiment, before performing operation (620), the electronic device (101) can detect a voice portion of the registered user within a voice signal including a command through the personalized VAD (385) and provide the detected voice signal to the TISV (375). At this time, the reference identification information used as a reference in the personalized VAD (385) can indicate the registered user. As the personalized VAD (385) is further utilized, a voice signal with the parts spoken by other users and noise parts excluded is utilized in the TISV (375), so that more accurate user verification can be performed.
[0149] Although FIG. 6A illustrates an operation of receiving a voice signal containing a single command and performing a function corresponding to the command, the present disclosure is not limited thereto. For example, the electronic device (101) may also execute functions corresponding to consecutive voice signals as it executes the continuous command function of the activated virtual assistant. For specific details related thereto, reference may be made to FIG. 6B below.
[0150] FIG. 6b illustrates an example of an operation flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a registered user regarding a virtual assistant in a continuous command function of a virtual assistant.
[0151] At least some of the methods of FIG. 6B may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (101). In an embodiment, the operations may be performed sequentially, but are not necessarily required to be performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0152] In some embodiments, prior to performing operation (650), the electronic device (101) may register (or store) a user with respect to the virtual assistant. For example, the user may be referred to as a registered user. Additionally, the electronic device (101) may execute the virtual assistant application (350) in the background.
[0153] In operation (650), according to one embodiment, the electronic device (101) may activate a virtual assistant. For example, the electronic device (101) may receive a voice signal including a wake word and activate the virtual assistant by recognizing the wake word within the received voice signal. In an embodiment, the electronic device (101) may also activate the virtual assistant in response to an input to the electronic device (101).
[0154] In an embodiment, the electronic device (101) may execute a continuous command function of the virtual assistant. For example, the electronic device (101) may execute the continuous command function upon receiving a voice signal including a command for executing the continuous command function of the virtual assistant. For example, the electronic device (101) may execute the continuous command function in response to an input for executing the continuous command function of the virtual assistant. As a non-limiting example, the input for executing the continuous command function may include a touch input, a double-touch input, a press input (e.g., a long press), or an input to an external electronic device (e.g., a true wired stereo (TWS), a smart watch, or a smart ring) connected to the electronic device (101). For example, the external electronic device may transmit a signal to the electronic device (101) notifying that an input to the external electronic device has been received. The electronic device (101) may execute the continuous command function upon receiving the signal.
[0155] In operation (655), according to one embodiment, the electronic device (101) may determine whether a preset time period has expired. For example, in operation (650), the electronic device (101) may start a timer for the preset time period from the time the virtual assistant is activated. As a non-limiting example, the preset time period may be 7 seconds. However, the present disclosure is not limited thereto. For example, the electronic device (101) may adjust the length of the preset time period in a menu for setting the virtual assistant application (350). In an embodiment, the electronic device (101) may increase the length of the preset time period or reset the timer when the user uses or looks at the electronic device (101). For example, the electronic device (101) may identify whether a command (or a voice signal including a command) is received by activating the input interface (320) during the preset time period. For example, during the preset time period, the electronic device (101) may be in a standby state to receive a command (or a voice signal including a command) through the virtual assistant. The preset time period may be referred to as a standby time.
[0156] In operation (655), if the preset time period has expired, the electronic device (101) may perform operation (660). In operation (655), if the preset time period has not expired, the electronic device (101) may perform operation (665).
[0157] In operation (660), according to one embodiment, the electronic device (101) may deactivate the virtual assistant. For example, the electronic device (101) may deactivate the activated virtual assistant when the preset time period has expired. For example, the preset time period may be used as a trigger for deactivating the activated virtual assistant again.
[0158] In operation (665), according to one embodiment, the electronic device (101) may receive a voice signal including a command. For example, the electronic device (101) may receive the voice signal including the command through the input interface (320) before the preset time period expires.
[0159] In operation (670), according to one embodiment, the electronic device (101) may perform a VF for the registered user. For example, the electronic device (101) may output a filtered voice signal through the VF (380) with respect to the voice signal including the command received in operation (665). For example, the electronic device (101) may provide the voice signal including the command to the VF (380). For example, the electronic device (101) may filter a voice portion spoken by the registered user within the voice signal including the command through the VF (380) using reference identification information indicating the registered user.
[0160] In operation (675), according to one embodiment, the electronic device (101) may perform TISV for the registered user. For example, the electronic device (101) may output a result indicating whether the user who uttered the voice signal including the command corresponds to the registered user through the TISV (375) for the voice signal filtered in operation (670). For example, the electronic device (101) may provide the filtered voice signal to the TISV (375). For example, the electronic device (101) may generate a speaker feature vector (or identification information) through the TISV (375) (or the speaker feature extraction model (370)) based on the filtered voice signal. For example, the electronic device (101) may identify a similarity value between the speaker feature vector (or identification information) and the reference identification information indicating the registered user. For example, the electronic device (101) can output a result (e.g., the first value (439-1) or the second value (439-2) of FIG. 4c) indicating whether the user who uttered the voice signal including the command is the registered user by performing a comparison between the similarity value and the reference value through the TISV (375).
[0161] In operation (680), according to one embodiment, the electronic device (101) may determine whether the user corresponds to a registered user. For example, the electronic device (101) may determine whether the user who uttered the voice signal including the command received in operation (665) corresponds to (or is identical to, or matches) the registered user.
[0162] In operation (680), if the user who uttered the voice signal including the command received in operation (665) corresponds to the registered user, operation (685) may be performed. According to an embodiment, in operation (680), if the user who uttered the voice signal including the command received in operation (665) does not correspond to the registered user, operation (655) may be performed again. When operation (655) is performed again, the electronic device (101) may start (or initialize) a timer for the preset time period upon determining that the user who uttered the voice signal including the command received does not correspond to the registered user.
[0163] In operation (685), according to one embodiment, the electronic device (101) may execute a function corresponding to the command. For example, the electronic device (101) may execute the function corresponding to the received command based on the activated virtual assistant.
[0164] In operation (690), according to one embodiment, the electronic device (101) may determine whether another voice signal including a command is received. For example, the electronic device (101) may determine whether the other voice signal including the command is received while executing the function according to operation (685). For example, the function may include playing a TTS synthesized voice. In other words, the electronic device (101) may determine whether the other voice signal including the command is further received while executing the function.
[0165] In operation (690), if the electronic device (101) further receives another voice signal including a command while executing the function, the electronic device (101) may perform operation (685) again. If the function performed in operation (685) before operation (690) is the reproduction of TTS synthesized voice, the electronic device (101), when further receiving the other voice signal while executing the function, may remove a portion corresponding to the reproduction of the TTS synthesized voice from the other voice signal. For example, the electronic device (101) may remove a portion corresponding to the reproduction of the TTS synthesized voice from the other voice signal through an adaptive echo canceller (AEC). In one example, if the electronic device (101) further receives the other voice signal while reproducing the TTS synthesized voice, the electronic device (101) may pause the reproduction of the TTS synthesized voice or reduce the volume for reproducing the TTS synthesized voice.
[0166] In operation (690), if the electronic device (101) does not receive another voice signal containing a command while executing the function, the electronic device (101) may perform operation (655) again. For example, the electronic device (101) may start a timer for the preset time period as it performs operation (655) again.
[0167] Thereafter, in the re-performed operation (655), the electronic device (101) can determine whether a voice signal including a command is received within the preset time period, and, if received, can perform the operation (665) to the operation (690) again.
[0168] In the above example of FIG. 6B, it is illustrated that while executing the function corresponding to the command in operation (685), it is determined whether another voice signal including the command is further received in operation (690), but the present disclosure is not limited thereto. For example, the electronic device (101) may determine, between operation (665) and operation (685), whether another voice signal including the command is further received in operation (690), such as operation (690). Accordingly, the electronic device (101) may further receive another voice signal through the input interface (320) while processing (e.g., ASR, NLU, VR, TISV) the received voice signal. As a non-limiting example, the electronic device (101) may simultaneously process the command of the voice signal and the command of the other voice signal. In one example, when the electronic device (101) utilizes a large language model (LLM) in a virtual assistant application (350), the electronic device (101) may simultaneously provide the voice signal and the other voice signal as prompts for the LLM. As a non-limiting example, the electronic device (101) may process a command of the other voice signal after processing a command of the voice signal.
[0169] LLM refers to an artificial neural network-based language model that has learned from a large amount of text data through pre-training. Compared to typical language models, LLMs can contain a relatively large number of parameters (e.g., over 10 billion). LLMs are a type of machine learning model used in natural language processing. They can be trained on large amounts of text data and used to make predictions about new text data. LLMs can be utilized for tasks such as natural language understanding, sentence generation, translation, grammatical error correction, and summarization.
[0170] FIG. 6A illustrates an example of receiving a voice signal including a command through the virtual assistant when the continuous command function is not executed, and performing VF and / or TISV using reference identification information indicating a registered user for the received voice signal. In addition, FIG. 6B illustrates an example of receiving a voice signal including a command through the virtual assistant when the continuous command function is executed, and performing VF and / or TISV using reference identification information indicating a registered user for the received voice signal. In the following FIG. 7, an example of a method for an electronic device (101) to execute functions corresponding to commands based on temporary permission while the continuous command function of the virtual assistant is executed is described.
[0171] Figure 7 illustrates an example of an operational flow for a method of executing a function corresponding to a command of a voice signal based on a voice signal spoken by a temporarily permitted user in a continuous command function of a virtual assistant.
[0172] At least some of the methods of FIG. 7 may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0173] According to an embodiment, before performing operation (700), the electronic device (101) may register (or store) a user with respect to the virtual assistant. For example, the user may be referred to as a registered user. Additionally, the electronic device (101) may execute the virtual assistant application (350) in the background.
[0174] In operation (700), according to one embodiment, the electronic device (101) may receive a voice signal. For example, the electronic device (101) may receive the voice signal through an activated input interface (320).
[0175] In operation (705), according to one embodiment, the electronic device (101) may determine whether a wake word is recognized. For example, the electronic device (101) may recognize a wake word within a received voice signal by converting the received voice signal into text.
[0176] In operation (705), if a call word is recognized in the received voice signal, the electronic device (101) may perform operation (710). In operation (705), if a call word is not recognized in the received voice signal, the electronic device (101) may perform operation (700) again.
[0177] For example, the electronic device (101) may activate the virtual assistant when a wake word is recognized in the received voice signal. In FIG. 7, the virtual assistant is activated depending on whether the wake word is recognized, but the present disclosure is not limited thereto. For example, the electronic device (101) may also activate the virtual assistant in response to an input (or user input) to the electronic device (101). For example, the input may include an input to a physical button of the electronic device (101) or a visual object (or icon) displayed through the display of the electronic device (101).
[0178] According to an embodiment, the electronic device (101) may execute a continuous command function of the virtual assistant. For example, the electronic device (101) may execute the continuous command function upon receiving a voice signal including a command for executing the continuous command function of the virtual assistant. For example, the electronic device (101) may execute the continuous command function in response to an input for executing the continuous command function of the virtual assistant. As a non-limiting example, the input for executing the continuous command function may include a touch input, a double touch input, a press input (e.g., a long press), or an input to an external electronic device (e.g., a true wired stereo (TWS), a smart watch, or a smart ring) connected to the electronic device (101). For example, the external electronic device may transmit a signal to the electronic device (101) notifying that an input to the external electronic device has been received. The electronic device (101) may execute the continuous command function upon receiving the signal.
[0179] In operation (710), according to one embodiment, the electronic device (101) may receive a voice signal including a command. For example, the electronic device (101) may receive the voice signal including the command through the input interface (320). In FIG. 7, for convenience of explanation, the voice signal received in operation (700) is illustrated as including a wake word and the voice signal received in operation (710) as including a command, but the present disclosure is not limited thereto. For example, a single voice signal may include both a wake word and a command.
[0180] In operation (715), according to one embodiment, the electronic device (101) may determine whether a voice signal including a command is the first voice signal. For example, the electronic device (101) may determine whether a voice signal including a command received after the virtual assistant is activated is the first voice signal. The first voice signal may be a voice signal including a command that is first received since the time the electronic device (101) activates the virtual assistant.
[0181] In operation (715), the electronic device (101) may perform operation (720) if the voice signal including the received command is the first voice signal. According to an embodiment, in operation (715), the electronic device (101) may perform operation (730) if the voice signal including the received command is not the first voice signal.
[0182] In the following, for convenience of explanation, the voice signal received in the first performed operation (710) is assumed to be the first voice signal.
[0183] In operation (720), according to one embodiment, the electronic device (101) may generate identification information indicating a user of the voice signal. For example, the electronic device (101) may generate first identification information indicating a first user who uttered the first voice signal including a command. For example, the first identification information may be generated through the speaker feature extraction model (370) of FIG. 4A. As a non-limiting example, the first identification information may be generated based on the first voice signal. In one example, the first identification information may include a speaker feature vector (or vector value) generated based on the first voice signal. For example, the first identification information may be used to temporarily allow the first user who uttered the first voice signal, which is the first voice signal received after the virtual assistant is activated. In other words, the first user may be a temporarily allowed user. As a non-limiting example, the electronic device (101) may store (or cache, temporarily store) the first identification information in the memory (340).
[0184] In operation (725), according to one embodiment, the electronic device (101) may execute a function corresponding to the command. For example, the electronic device (101) may execute a function corresponding to the command of the first voice signal. After performing operation (725), the electronic device (101) may perform operation (710) again. In the example of FIG. 7, operation (710) is illustrated as being performed again after performing operation (725), but the present disclosure is not limited thereto. As described above in FIG. 6B, the electronic device (101) may perform operation (710) during or even before performing operation (725).
[0185] In the re-performed operation (710), the electronic device (101) can receive a voice signal including a command. Hereinafter, for convenience of explanation, the voice signal including the command received in the re-performed operation (710) is assumed to be a second voice signal.
[0186] In operation (715), the electronic device (101) may determine whether the voice signal including the command is the first voice signal. For example, the electronic device (101) may determine whether the second voice signal is the first voice signal. If the second voice signal is not the first voice signal, the electronic device (101) may perform operation (730).
[0187] In operation (730), according to one embodiment, the electronic device (101) can identify the user who uttered the voice signal based on the trained model. For example, the trained model can include at least one of the TISV (375), the VF (380), or the personalized VAD (385). For example, the electronic device (101) can identify the second user who uttered the second voice signal through the trained model using the first identification information as a reference. Identifying the second user can include determining whether the second user who uttered the second voice signal is the same user as the first user or a different user.
[0188] As a non-limiting example, the electronic device (101) may provide the second voice signal to the VF (380) using the first identification information as reference identification information. For example, the VF (380) may maintain the portion of the second voice signal uttered by the first user indicated by the first identification information, and suppress (or reduce) the portion uttered by another user.
[0189] As a non-limiting example, the electronic device (101) may provide the second voice signal to the TISV (375) using the first identification information as reference identification information. As a non-limiting example, the second voice signal provided to the TISV (375) may be a second voice signal filtered through the VF (380) as described above. As a non-limiting example, the second voice signal provided to the TISV (375) may be a second voice signal received through the input interface (320) (or a second voice signal not filtered through the VF (380). For example, the TISV (375) may generate second identification information indicating the second user who uttered the second voice signal based on the second voice signal. For example, the TISV (375) may identify a similarity value between the second identification information indicating the second user who uttered the second voice signal and the first identification information. For example, TISV (375) can perform a comparison between the similarity value and the reference value. For example, TISV (375) can output a result indicating whether the second user who uttered the second voice signal corresponds to the first user who uttered the first voice signal. For example, the result can include one of a first value indicating that the first user and the second user correspond (e.g., the first value (439-1) of FIG. 4C) and a second value indicating that the first user and the second user do not correspond (e.g., the second value (439-2) of FIG. 4C).
[0190] In the above example, TISV (375) is illustrated as utilizing one reference identification information (e.g., the first identification information), but the present disclosure is not limited thereto. For example, TISV (375) may utilize reference identification information indicating a temporarily permitted user (e.g., the first identification information) and reference identification information indicating a registered user as references. Accordingly, TISV (375) may be utilized to determine whether the second user who uttered the second voice signal corresponds to a temporarily permitted user or a registered user. In this case, the second voice signal provided to TISV (375) may be a second voice signal received via the input interface (320) (or a second voice signal not filtered via VF (380)). Specific details related thereto are exemplified and described below with reference to FIG. 8.
[0191] As a non-limiting example, the electronic device (101) may provide the second voice signal to a personalized VAD (385) that uses the first identification information as reference identification information. As a non-limiting example, the second voice signal provided to the personalized VAD (385) may be a second voice signal filtered through the VF (380) as described above. Alternatively, as a non-limiting example, the second voice signal provided to the personalized VAD (385) may be a second voice signal received through the input interface (320) (or a second voice signal not filtered through the VF (380). For example, the personalized VAD (385) may detect, based on the second voice signal, a second voice signal indicating a portion of speech spoken by the first user within the second voice signal. The electronic device (101) can determine whether the second user who uttered the second voice signal corresponds to the first user, based on whether the detected second voice signal output from the personalized VAD (385) includes a voice portion uttered by the first user.
[0192] In operation (735), according to one embodiment, the electronic device (101) may determine whether the identified user corresponds to a designated user. For example, the electronic device (101) may determine whether the second user identified in operation (730) is the designated user. For example, the designated user may include a user registered with the virtual assistant and a user temporarily permitted with respect to the virtual assistant (e.g., the first user).
[0193] In operation (735), the electronic device (101) may perform operation (725) if the second user is the designated user. According to an embodiment, in operation (735), the electronic device (101) may perform operation (710) if the second user is not the designated user.
[0194] In operation (725), the electronic device (101) may execute a function corresponding to the command. For example, if the second user corresponds to the first user who is a temporarily permitted user or corresponds to a registered user, the electronic device (101) may execute a function corresponding to the command of the second voice signal. After performing operation (725), the electronic device (101) may perform operation (710) again.
[0195] In FIG. 7, for convenience of explanation, an operation of determining whether a preset time period has expired in relation to the reception of a voice signal including a command is omitted, but the present disclosure is not limited thereto. For example, the electronic device (101) may start a timer for the preset time period from the time the virtual assistant is activated or the time when execution of a function corresponding to the command is completed. Accordingly, if a voice signal including a command is received before the timer expires, the timer may be initialized. According to an embodiment, when the timer expires, the electronic device (101) may deactivate the virtual assistant. For example, the electronic device (101) may adjust the length of the preset time period in a menu for setting the virtual assistant application (350). According to an embodiment, the electronic device (101) may increase the length of the preset time period or initialize the timer when the user uses the electronic device (101) or looks at the electronic device (101).
[0196] In the above example, for convenience of explanation, it is assumed that the first user who uttered the first voice signal, which is the initial voice signal, is not a registered user, but a temporarily permitted user. However, the present disclosure is not limited thereto. For example, the first user may be a registered user. Even if the first user is a registered user, by using the temporarily generated first identification information as a reference for the trained model, the electronic device (101) can determine whether the second user who uttered the second voice signal corresponds to the first user.
[0197] In one example, the electronic device (101) may determine whether to use the first identification information as a reference for the trained model by further determining whether the first user is a registered user. For example, in operation (720), the electronic device (101) may generate the first identification information indicating the first user who uttered the first voice signal. After this, the electronic device (101) can perform similarity value identification (e.g., similarity value identification (435) of FIG. 4C) and comparison (437) with a reference value through the TISV (375) based on the first identification information. At this time, the reference identification information used as a reference for the similarity value identification of the TISV (375) may be identification information indicating a registered user. At this time, the identification information indicating the registered user may be identification information stored in the memory (340) (or, identification information set (390)). The electronic device (101) may refrain from using (or not use) the first identification information as a reference for a trained model when the similarity value between the first identification information and the identification information indicating the registered user is greater than or equal to the reference value. For example, the electronic device (101) may remove the first identification information. The electronic device (101) may remove the first identification information between the first identification information and the identification information indicating the registered user. If the similarity value is less than the reference value, the first identification information may be stored (or cached, temporarily stored) in the memory (340) to be used as a reference for the trained model.
[0198] In the example of FIG. 7, the electronic device (101) may remove the first identification information indicating the temporarily permitted first user. For example, the electronic device (101) may remove the stored first identification information upon identifying a specific event. For example, the specific event may include at least one of: deactivation of the activated virtual assistant, deactivation of the continuous command function of the virtual assistant, or identification that the first user indicated by the first identification information matches a user registered with the virtual assistant.
[0199] As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant when the screen (or execution screen) displayed when the virtual assistant application (350) is executed in the foreground is displayed in a minimized state and is displayed in the minimized state for a specified period of time. As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant when another application is executed in response to the execution of a function corresponding to a command received based on the activated virtual assistant and the other application uses the input interface (320). For example, the other application may include a phone application, a video call application, or a voice recording application. As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant when media content (e.g., music or video) is played in response to the execution of a function corresponding to a command received based on the activated virtual assistant. As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant application (350) when the virtual assistant application (350) running in the foreground is run in the background. For example, the background execution of the virtual assistant application (350) may be performed in response to an input for entering a home screen application of the electronic device (101) (e.g., a home key) or an input for entering a previously running screen (or a previously running application) (e.g., a back key). As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant when the display of the electronic device (101) is in an off state (or a low-power consumption state). For example, an always-on display (AOD) screen may be displayed in the low-power consumption state.As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant upon receiving a command for deactivating the virtual assistant. For example, the command for deactivation may include "goodbye," "end," or "cancel." As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant upon expiration of a timer for a preset time period. As a non-limiting example, the electronic device (101) may deactivate the activated virtual assistant when the connected communication network is disconnected (or when airplane mode is activated). However, in one example, if the virtual assistant application can provide services without a communication network, the virtual assistant may remain activated.
[0200] As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant when a timer for a preset time period expires. In other words, the electronic device (101) may terminate the continuous command function of the virtual assistant when the virtual assistant is deactivated. As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant when another application is executed in response to the execution of a function corresponding to a command received based on the activated virtual assistant, and the other application uses the input interface (320). As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant when, during the reproduction of a TTS synthesized voice in response to the execution of a function corresponding to a command received based on the activated virtual assistant, the reproduction of the TTS synthesized voice is interrupted. As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant upon identifying an input to the execution screen (or dialogue window) of the virtual assistant application (350) (e.g., screen (119) of FIG. 1A). As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant upon receiving a command for terminating the continuous command function of the virtual assistant. As a non-limiting example, the electronic device (101) may terminate the continuous command function of the virtual assistant when the connected communication network is disconnected (or when airplane mode is activated).
[0201] FIG. 8 illustrates an example of an operational flow for a method of verifying a second user who uttered a second voice signal and a method of executing a function corresponding to a command of the second voice signal, based on whether there is a correspondence between a first user who uttered a first voice signal and a registered user in a continuous command function of a virtual assistant.
[0202] At least some of the methods of FIG. 8 may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be configured to be performed (or controlled) by at least one processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0203] According to an embodiment, before performing operation (800), the electronic device (101) may register (or store) a user with respect to the virtual assistant. For example, the user may be referred to as a registered user. Additionally, the electronic device (101) may execute the virtual assistant application (350) in the background.
[0204] In operation (800), according to one embodiment, the electronic device (101) may activate a virtual assistant. For example, the electronic device (101) may activate the virtual assistant when a wake word is recognized in a received voice signal. For example, the electronic device (101) may also activate the virtual assistant in response to an input (or user input) to the electronic device (101). For example, the input may include an input to a physical button of the electronic device (101) or a visual object (or icon) displayed through a display of the electronic device (101). For example, the electronic device (101) may execute a continuous command function of the activated virtual assistant.
[0205] In operation (805), according to one embodiment, the electronic device (101) may receive a first voice signal. For example, the electronic device (101) may receive the first voice signal through the input interface (320).
[0206] In operation (810), according to one embodiment, the electronic device (101) may generate first identification information indicating a first user of the first voice signal. For example, the first identification information may be generated through the speaker feature extraction model (370) of FIG. 4A. As a non-limiting example, the first identification information may be generated based on the first voice signal. In one example, the first identification information may include a speaker feature vector (or vector value) generated based on the first voice signal. For example, the first identification information may be used to temporarily allow the first user who uttered the first voice signal, which is the first voice signal received after the virtual assistant is activated.
[0207] In operation (815), according to one embodiment, the electronic device (101) may determine whether the first user corresponds to a registered user. For example, the electronic device (101) can perform similarity value identification (e.g., similarity value identification (435) of FIG. 4C) and comparison (437) with a reference value using the first identification information through the TISV (375). At this time, the reference identification information used as a reference for similarity value identification of the TISV (375) may be identification information indicating a registered user. At this time, the identification information indicating the registered user may be identification information stored in the memory (340) (or, the identification information set (390)). The electronic device (101) can identify that the first user is the registered user when the similarity value between the first identification information and the identification information indicating the registered user is greater than or equal to the reference value. According to an embodiment, the electronic device (101) can identify that the first user is different from the registered user when the similarity value between the first identification information and the identification information indicating the registered user is less than the reference value.
[0208] As a non-limiting example, if the virtual assistant is activated via a call word in operation (800), the registered user may be the user who uttered the call word, identified upon recognizing the call word. In other words, the electronic device (101) recognizing the call word may include the electronic device (101) identifying (or recognizing) the user who uttered the call word.
[0209] In operation (815), the electronic device (101) may perform operation (820) if the first user is a registered user. According to an embodiment, in operation (815), the electronic device (101) may perform operation (855) if the first user is not a registered user.
[0210] In operation (820), according to one embodiment, the electronic device (101) may execute a function corresponding to the first command included in the first voice signal and generate instruction information indicating that the first user is the registered user. For example, the electronic device (101) may execute a function corresponding to the first command included in the first voice signal. In addition, the electronic device (101) may generate instruction information (e.g., the first value (439-1) of FIG. 4C) indicating that the first user is the registered user as a result output through the TISV (375). As a non-limiting example, the electronic device (101) may remove the first identification information when the first user is the registered user. However, the present disclosure is not limited thereto. For example, even when the first user is the registered user, the electronic device (101) may use the first identification information as a reference for a trained model.
[0211] As a non-limiting example, the electronic device (101) may provide a visual effect through a screen using the generated instruction information. For example, the electronic device (101) may display a screen through the display of the electronic device (101) in response to a voice signal received from the first user. For example, the screen may include text indicating the first user (e.g., texts 1021 and 1022 of FIG. 10) or text indicating a command included in the voice signal (e.g., the first command). For example, when the electronic device (101) generates the instruction information indicating that the first user is the registered user, the electronic device (101) may apply a first visual effect to the screen based on the generated instruction information. In one example, the first visual effect may include at least one of an indicator (or text) indicating that the user is a registered user, a color of the text, a font of the text, or a size of the text.
[0212] In operation (825), according to one embodiment, the electronic device (101) may receive a second voice signal. For example, the electronic device (101) may receive the second voice signal following the first voice signal while the continuous command function of the activated virtual assistant is being executed.
[0213] In operation (830), according to one embodiment, the electronic device (101) may generate second identification information indicating a second user of the second voice signal. For example, the second identification information may be generated through the speaker feature extraction model (370) of FIG. 4A. As a non-limiting example, the second identification information may be generated based on the second voice signal. In one example, the second identification information may include a speaker feature vector (or vector value) generated based on the second voice signal.
[0214] In operation (835), according to one embodiment, the electronic device (101) may perform VF and TISV using the first identification information. For example, the electronic device (101) may provide the second voice signal to the VF (380) that uses the first identification information as reference identification information. For example, the electronic device (101) may provide the second voice signal to the TISV (375) that uses the first identification information as reference identification information.
[0215] In operation (840), according to one embodiment, the electronic device (101) may determine whether the second user corresponds to a registered user. For example, the electronic device (101) may perform similarity value identification (e.g., similarity value identification (435) of FIG. 4C) and comparison (437) with a reference value through the TISV (375) based on the second identification information. At this time, the reference identification information used as a reference for the similarity value identification of the TISV (375) may be identification information indicating a registered user (or the first identification information). The electronic device (101) may identify that the second user is the registered user when the similarity value between the second identification information and the first identification information is greater than or equal to the reference value. According to an embodiment, the electronic device (101) may identify that the second user is different from the registered user when the similarity value between the second identification information and the first identification information is less than the reference value.
[0216] In operation (840), the electronic device (101) may perform operation (845) if the second user is the registered user. According to an embodiment, in operation (840), the electronic device (101) may perform operation (850) if the second user is different from the registered user.
[0217] In operation (845), according to one embodiment, the electronic device (101) may execute a function corresponding to the second command. For example, the electronic device (101) may execute a function corresponding to the second command of the second voice signal when the second user is the same as the registered user (or the first user).
[0218] In operation (850), according to one embodiment, the electronic device (101) may refrain from executing a function corresponding to the second command. For example, the electronic device (101) may refrain from executing a function corresponding to the second command of the second voice signal if the second user is different from the registered user (or the first user).
[0219] In operation (855), according to one embodiment, the electronic device (101) may execute a function corresponding to the first command included in the first voice signal and generate instruction information indicating that the user is not a registered user. For example, the electronic device (101) may execute a function corresponding to the first command included in the first voice signal. In addition, the electronic device (101) may generate instruction information (e.g., the second value (439-2) of FIG. 4C) indicating that the first user is not the registered user as a result output through the TISV (375). As a non-limiting example, the electronic device (101) may store (or cache, temporarily store) the first identification information in the memory (340) when the first user is not the registered user.
[0220] As a non-limiting example, the electronic device (101) may provide a visual effect through a screen using the generated instruction information. For example, the electronic device (101) may display a screen through the display of the electronic device (101) in response to a voice signal received from the first user. For example, the screen may include text indicating the first user (e.g., texts 1021 and 1022 of FIG. 10) or text indicating a command included in the voice signal (e.g., the first command). For example, if the electronic device (101) generates instruction information indicating that the first user is not the registered user, the electronic device (101) may apply a second visual effect to the screen based on the generated instruction information. In one example, the second visual effect may include at least one of an indicator (or text) indicating that the user is an unregistered user, a color of the text, a font of the text, or a size of the text. In the above example, the second visual effect may be different from the first visual effect.
[0221] In operation (860), according to one embodiment, the electronic device (101) may receive a second voice signal. For example, the electronic device (101) may receive the second voice signal following the first voice signal while the continuous command function of the activated virtual assistant is being executed.
[0222] In operation (865), according to one embodiment, the electronic device (101) may generate second identification information indicating a second user of the second voice signal. For example, the second identification information may be generated through the speaker feature extraction model (370) of FIG. 4A. As a non-limiting example, the second identification information may be generated based on the second voice signal. In one example, the second identification information may include a speaker feature vector (or vector value) generated based on the second voice signal.
[0223] In operation (870), according to one embodiment, the electronic device (101) may perform TISV using the first identification information. For example, the electronic device (101) may provide the second voice signal to the TISV (375) using the first identification information as reference identification information. Unlike operation (835), in operation (870), since the first user is different from the registered user, the first identification information indicating the temporarily permitted user and the identification information indicating the registered user may be used as reference identification information, and thus the VF (380) may be omitted. In other words, unlike the TISV (375) that performs verification of whether a plurality of users (or speakers) are the same user (or speaker), the VF (380) may not perform filtering on a plurality of users (or speakers).
[0224] In one example, the electronic device (101) can learn (or generate) TISV (375) as reference identification information based on the first identification information. For example, the electronic device (101) can perform TISV (375) on the second voice signal using the learned reference identification information.
[0225] In operation (875), according to one embodiment, the electronic device (101) may determine whether the second user corresponds to a designated user. For example, the designated user may include a registered user and a temporarily permitted user. For example, the electronic device (101) may perform similarity value identification (e.g., similarity value identification (435) of FIG. 4C) and comparison (437) with a reference value through the TISV (375) based on the second identification information. At this time, the reference identification information used as a reference for the similarity value identification of the TISV (375) may be identification information (or the first identification information) indicating a registered user. The electronic device (101) may identify that the second user is the designated user when the similarity value between the second identification information and the first identification information is greater than or equal to the reference value. According to an embodiment, the electronic device (101) may identify that the second user is different from the designated user when the similarity value between the second identification information and the first identification information is greater than or equal to the reference value and less than or equal to the reference value.
[0226] In operation (875), the electronic device (101) may perform operation (880) if the second user is the designated user. According to an embodiment, in operation (875), the electronic device (101) may perform operation (885) if the second user is different from the designated user.
[0227] In operation (880), according to one embodiment, the electronic device (101) may execute a function corresponding to the second command. For example, the electronic device (101) may execute a function corresponding to the second command of the second voice signal if the second user is the same as the designated user (e.g., the first user or the registered user).
[0228] In operation (885), according to one embodiment, the electronic device (101) may refrain from executing a function corresponding to the second command. For example, the electronic device (101) may refrain from executing a function corresponding to the second command of the second voice signal if the second user is different from the designated user (e.g., the first user or the registered user).
[0229] Figure 9 illustrates an example of a method for recognizing a voice signal in a continuous command function of a virtual assistant based on temporary permission.
[0230] Referring to FIG. 9, an example (900) of a method for recognizing a voice signal in a continuous command function of a virtual assistant based on temporary permission is illustrated in the electronic device (101). In the example (900), a user (901) may be a registered user with respect to the virtual assistant. In the example (900), a user (902) and a user (903) may be unregistered users with respect to the virtual assistant.
[0231] Referring to example (900), the electronic device (101) can activate a virtual assistant. For example, the electronic device (101) can activate the virtual assistant when a wake word is recognized in a received voice signal. For example, the electronic device (101) can also activate the virtual assistant in response to an input (or user input) to the electronic device (101). For example, the input can include an input to a physical button of the electronic device (101) or a visual object (or icon) displayed through the display of the electronic device (101). For example, the electronic device (101) can execute a continuous command function of the activated virtual assistant.
[0232] For example, the electronic device (101) may receive a plurality of voice signals (910, 920, 930, 940, 950) through the input interface (320) while the continuous command function is being executed. For example, the voice signals (910, 920, 930, 940, 950) may be received through the input interface (320) in the following order: voice signal (910), voice signal (920), voice signal (930), voice signal (940), and voice signal (950). For example, the voice signal (910) may be “Pororo, Pororo, Pororo!” spoken by the user (902). For example, the voice signal (920) may be “Play Pororo on Youtube.” spoken by the user (901). For example, the voice signal (930) may be “Play the next one.” spoken by the user (902). For example, the voice signal (940) may be “Honey, dinner is ready, come around the table.” spoken by the user (903). For example, the voice signal (950) may be “Please turn on the volume.” spoken by the user (902).
[0233] For example, the electronic device (101) may generate identification information indicating the user (902) who first uttered the received voice signal (910) after the virtual assistant is activated. For example, the electronic device (101) may recognize the user (902) as a temporarily authorized user.
[0234] For example, the electronic device (101) can recognize the voice signals (910, 930, 950) of a user (902) who is a temporarily permitted user and the voice signal (920) of a user (901) who is a registered user among the voice signals (910, 920, 930, 940, 950), and reject recognition of the voice signal (940) of the user (903). For example, the electronic device (101) can recognize the voice signals (910, 930, 950) of a user (902) who is a temporarily permitted user and the voice signal (920) of a user (901) who is a registered user among the voice signals (910, 920, 930, 940, 950) through the VF (380) and / or the TISV (375). For example, the electronic device (101) can process recognized voice signals (910, 920, 930, 950) through the ASR module (360). The electronic device (101) can provide the voice signals (910, 920, 930, 950) processed through the ASR module (360) to the virtual assistant application (350). Accordingly, the virtual assistant application (350) can recognize the voice signals (915, 925, 935, 955) among the voice signals (915, 925, 935, 945, 955). At this time, the voice signal (945) may not be provided to the virtual assistant application (350) as it is rejected through the VF (380) and / or the TISV (375).
[0235] The virtual assistant application (350) can identify the command of each of the voice signals (915, 925, 935, 955) and execute a function corresponding to the identified command.
[0236] As a non-limiting example, the electronic device (101) may further include a module for discarding utterances. For example, the module for discarding utterances may be stored in the memory (340) of the electronic device (101). For example, the module for discarding utterances may analyze text(s) from voice signal(s) received through the input interface (320) through the ASR module (360) while the continuous command function is being executed. For example, the module for discarding utterances may identify texts (hereinafter, discarded texts) that do not require execution of a function corresponding to a command based on the virtual assistant based on the analysis of the texts. According to an embodiment, an example of the discarded text may be a voice signal (940) spoken by the user (903) in example (900). In other words, the discarded text may be a voice signal that is not a command for causing the virtual assistant to execute a specific function. As a non-limiting example, the module for discarding utterances may be implemented as (or utilize) an LLM to identify (or determine) the discarded text. For example, if the electronic device (101) determines the discarded text through the module for discarding utterances, it may consider that a voice signal corresponding to the discarded text has not been received. In one example, the electronic device (101) may refrain from initializing (or not perform initialization of) the timer even if a voice signal corresponding to the discarded text is received before the timer for a preset time period has expired after starting.
[0237] Figure 10 illustrates an example of a user interface for settings related to the continuous command function of a virtual assistant.
[0238] FIG. 10 illustrates an example of a user interface (1000) for setting up a virtual assistant application (350) of an electronic device (101). For example, the user interface (1000) may be referred to as a screen (or execution screen).
[0239] For example, the electronic device (101) may display a user interface (1000) through the display of the electronic device (101). For example, the user interface (1000) may be a user interface for settings related to a continuous command function of a virtual assistant. As a non-limiting example, the user interface (1000) may include a first settings menu (1010), a second settings menu (1020), and a third settings menu (1030).
[0240] For example, the first settings menu (1010) may be a menu for activating and deactivating the continuous command function. For example, the first settings menu (1010) may include a toggle (1015) for activating and deactivating the continuous command function. For example, the toggle (1015) may be referenced as an executable object or icon.
[0241] For example, the second settings menu (1020) may be a menu for registering and removing users (or speakers) who will allow continuous commands while the continuous command function is running. For example, users registered (or added) within the second settings menu (1020) may be registered users for the virtual assistant. For example, the second settings menu (1020) may include texts (1021, 1022) indicating currently registered users. For example, the second settings menu (1020) may include a button (1023) for newly adding a user who will allow continuous commands. For example, the button (1023) may be referenced as an executable object or an icon. Text indicating a user added based on at least one input to the button (1023) may be displayed within the second settings menu (1020), like the texts (1021, 1022) of the second settings menu (1020).
[0242] For example, the third setting menu (1030) may be a menu for setting the number of users (or speakers) to temporarily allow continuous commands while the continuous command function is running. For example, the number of users set in the third setting menu (1030) may be temporarily allowed users for the virtual assistant. For example, the third setting menu (1030) may include a button (1035) for changing (or adjusting) the number of users to temporarily allow. For example, the button (1035) may be referenced as an executable object or an icon. For example, the button (1035) may include text indicating the current number of temporarily allowed users (e.g., 2).
[0243] The user interface (1000) illustrated in FIG. 10 is merely an example for convenience of explanation, and the present disclosure is not limited thereto. For example, the user interface (1000) may further include other setting menus, or may omit at least some of the setting menus (1010, 1020, 1030) illustrated in FIG. 10.
[0244] FIG. 11 is a block diagram of an electronic device within a network environment according to various embodiments.
[0245] Referring to FIG. 11, in a network environment (1100), an electronic device (1101) may communicate with an electronic device (1102) via a first network (1198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1104) or a server (1108) via a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) via the server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), a memory (1130), an input module (1150), an audio output module (1155), a display module (1160), an audio module (1170), a sensor module (1176), an interface (1177), a connection terminal (1178), a haptic module (1179), a camera module (1180), a power management module (1188), a battery (1189), a communication module (1190), a subscriber identification module (1196), or an antenna module (1197). In some embodiments, the electronic device (1101) may omit at least one of these components (e.g., the connection terminal (1178)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).
[0246] The processor (1120) may, for example, execute software (e.g., a program (1140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1120) may store commands or data received from other components (e.g., a sensor module (1176) or a communication module (1190)) in a volatile memory (1132), process the commands or data stored in the volatile memory (1132), and store result data in a non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1121). For example, when the electronic device (1101) includes the main processor (1121) and the auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a given function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as a part thereof.
[0247] The auxiliary processor (1123) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1160), the sensor module (1176), or the communication module (1190)) of the electronic device (1101), for example, on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1180) or a communication module (1190)). In one embodiment, the auxiliary processor (1123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0248] The memory (1130) can store various data used by at least one component (e.g., the processor (1120) or the sensor module (1176)) of the electronic device (1101). The data can include, for example, software (e.g., the program (1140)) and input data or output data for commands related thereto. The memory (1130) can include a volatile memory (1132) or a non-volatile memory (1134).
[0249] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).
[0250] The input module (1150) can receive commands or data to be used in a component of the electronic device (1101) (e.g., a processor (1120)) from an external source (e.g., a user) of the electronic device (1101). The input module (1150) may also be referred to as an input interface. The input module (1150) may include, for example, a microphone, a mouse, a keyboard, keys (e.g., buttons), or a digital pen (e.g., a stylus pen).
[0251] The audio output module (1155) can output audio signals to the outside of the electronic device (1101). The audio output module (1155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0252] The display module (1160) can visually provide information to an external party (e.g., a user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0253] The audio module (1170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150), output sound through the sound output module (1155), or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1101).
[0254] The sensor module (1176) can detect the operating status (e.g., power or temperature) of the electronic device (1101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0255] The interface (1177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1101) with an external electronic device (e.g., the electronic device (1102)). In one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0256] The connection terminal (1178) may include a connector through which the electronic device (1101) may be physically connected to an external electronic device (e.g., the electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0257] The haptic module (1179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0258] The camera module (1180) can capture still images and videos. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0259] The power management module (1188) can manage power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0260] A battery (1189) may power at least one component of the electronic device (1101). In one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0261] The communication module (1190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may operate independently from the processor (1120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, a corresponding communication module can communicate with an external electronic device (1104) via a first network (1198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can verify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1196).
[0262] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1192) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1192) may support various requirements specified in the electronic device (1101), an external electronic device (e.g., the electronic device (1104)), or a network system (e.g., the second network (1199)). According to one embodiment, the wireless communication module (1192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0263] The antenna module (1197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1198) or the second network (1199), may be selected from the plurality of antennas by, for example, the communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1197).
[0264] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0265] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0266] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) via a server (1108) connected to a second network (1199). Each of the external electronic devices (1102 or 1104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations executed in the electronic device (1101) may be executed in one or more of the external electronic devices (1102, 1104, or 1108). For example, when the electronic device (1101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server utilizing machine learning and / or a neural network.In one embodiment, an external electronic device (1104) or server (1108) may be included in the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0267] For example, an external electronic device (1102) can render content data executed in an application and transmit it to an electronic device (1101), and the electronic device (1101) that receives the data can output the content data to a display module. If the electronic device (1101) detects user movement through an IMU sensor, the processor of the electronic device (1101) can correct the rendering data received from the external electronic device (1102) based on the movement information and output it to the display module. Alternatively, the processor can transmit the movement information to the external electronic device (1102) and request rendering so that the screen data is updated accordingly. According to various embodiments, the external electronic device (1102) can be various types of devices, such as a smartphone or a case device that can store and charge the electronic device (1101).
[0268] FIG. 12 is a block diagram illustrating an integrated intelligence system according to various embodiments.
[0269] Referring to FIG. 12, an integrated intelligent system of one embodiment may include an electronic device (101), an intelligent server (1200) (e.g., server (1108) of FIG. 11), and a service server (1290) (e.g., server (1108) of FIG. 11).
[0270] An electronic device (1101) of one embodiment may be a terminal device (or electronic device) that can connect to the Internet, and may be, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a TV, white goods, a wearable device, an HMD, or a smart speaker.
[0271] According to an embodiment, the electronic device (1101) may include an interface (1177), an input module (1150), an audio output module (1155), a display module (1160), a memory (1130), or a processor (1120). The above-listed components may be operatively or electrically connected to each other.
[0272] An interface (1177) of one embodiment may be configured to connect to an external device and transmit and receive data. An input module (1150) of one embodiment may receive sound (e.g., user speech) and convert it into an electrical signal. An audio output module (1155) of one embodiment may output the electrical signal as sound (e.g., voice).
[0273] The display module (1160) of one embodiment may be configured to display an image or video. The display module (1160) of one embodiment may also display a graphical user interface (GUI) of a running app (or application program). The display module (1160) of one embodiment may receive touch input via a touch sensor. For example, the display module (1160) may receive text input via a touch sensor in an on-screen keyboard area displayed within the display module (1160).
[0274] In one embodiment, the memory (1130) may store a client module (1151), a software development kit (SDK) (1153), and a plurality of apps (1146). The client module (1151) and the SDK (1153) may constitute a framework (or solution program) for performing general functions. In addition, the client module (1151) or the SDK (1153) may constitute a framework for processing user input (e.g., voice input, text input, touch input).
[0275] The plurality of apps (1146) stored in the memory (1130) of one embodiment may be programs for performing a specified function. According to one embodiment, the plurality of apps (1146) may include a first app (1146-1) and a second app (1146-2). According to one embodiment, each of the plurality of apps (1146) may include a plurality of operations for performing a specified function. For example, the apps may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the plurality of apps (1146) may be executed by the processor (1120) to sequentially execute at least some of the plurality of operations.
[0276] The processor (1120) of one embodiment can control the overall operation of the electronic device (1101). For example, the processor (1120) can be electrically connected to an interface (1177), an input module (1150), an audio output module (1155), and a display module (1160) to perform a specified operation.
[0277] The processor (1120) of one embodiment may also execute a program stored in the memory (1130) to perform a designated function. For example, the processor (1120) may execute at least one of the client module (1151) or the SDK (1153) to perform the following operations for processing user input. The processor (1120) may control the operations of multiple apps (1146), for example, through the SDK (1153). The following operations described as operations of the client module (1151) or the SDK (1153) may be operations executed by the processor (1120).
[0278] The client module (1151) of one embodiment can receive user input. For example, the client module (1151) can receive a voice signal corresponding to a user utterance detected through the input module (1150). Alternatively, the client module (1151) can receive a touch input detected through the display module (1160). Alternatively, the client module (1151) can receive a text input detected through a keyboard or a visual keyboard. In addition, the client module (1151) can receive various forms of user input detected through an input module included in the electronic device (1101) or an input module connected to the electronic device (1101). The client module (1151) can transmit the received user input to the intelligent server (1200). The client module (1151) can transmit status information of the electronic device (1101) to the intelligent server (1200) together with the received user input. The above status information may be, for example, the execution status information of the app.
[0279] The client module (1151) of one embodiment can receive a result corresponding to the received user input. For example, the client module (1151) can receive a result corresponding to the received voice input when the intelligent server (1200) can produce a result corresponding to the received user input. The client module (1151) can display the received result on the display module (1160). In addition, the client module (1151) can output the received result as audio through the audio output module (1155).
[0280] In one embodiment, the client module (1151) can receive a plan corresponding to the received user input. The client module (1151) can display the results of executing multiple operations of the app according to the plan on the display module (1160). For example, the client module (1151) can sequentially display the results of executing multiple operations on the display and output audio through the audio output module (1155). The electronic device (1101) can, for example, display only some results of executing multiple operations (e.g., the result of the last operation) on the display module (1160) and output audio through the audio output module (1155).
[0281] According to one embodiment, the client module (1151) may receive a request from the intelligent server (1200) to obtain information necessary to produce a result corresponding to a user input. According to one embodiment, the client module (1151) may transmit the necessary information to the intelligent server (1200) in response to the request.
[0282] The client module (1151) of one embodiment can transmit result information of executing multiple operations according to a plan to the intelligent server (1200). The intelligent server (1200) can use the result information to confirm that the received user input has been processed correctly.
[0283] The client module (1151) of one embodiment may include a voice recognition module. According to one embodiment, the client module (1151) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (1151) may execute an intelligent app to process voice inputs to perform organic actions through designated inputs (e.g., "Wake up!").
[0284] An intelligent server (1200) of one embodiment can receive information related to a user voice input from an electronic device (101) via a communication network. According to one embodiment, the intelligent server (1200) can convert data related to the received voice input into text data. According to one embodiment, the intelligent server (1200) can generate a plan for performing a task corresponding to the user voice input based on the text data.
[0285] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI system may be a rule-based system or a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)). The AI system may be a combination of the above or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI system may select at least one plan from a plurality of predefined plans.
[0286] An intelligent server (1200) of one embodiment may transmit a result according to a generated plan to an electronic device (1101), or transmit the generated plan to the electronic device (1101). According to one embodiment, the electronic device (1101) may display a result according to the plan on a display module (1160). According to one embodiment, the electronic device (1101) may display a result of executing an operation according to the plan on a display module (1160).
[0287] An intelligent server (1200) of one embodiment may include a front end (1210), a natural language platform (1220), a capsule DB (1230), an execution engine (1240), an end user interface (1250), a management platform (1260), a big data platform (1270), or an analytic platform (1280).
[0288] A front end (1210) of one embodiment can receive user input from an electronic device (101). The front end (1210) can transmit a response corresponding to the user input.
[0289] According to one embodiment, the natural language platform (1220) may include an automatic speech recognition module (ASR module) (1221), a natural language understanding module (NLU module) (1223), a planner module (1225), a natural language generator module (NLG module) (1227), or a text to speech module (TTS module) (1229).
[0290] An automatic speech recognition module (1221) of one embodiment can convert voice input received from an electronic device (1101) into text data. A natural language understanding module (1223) of one embodiment can use text data of the voice input to determine a user's intent. For example, the natural language understanding module (1223) can perform syntactic analysis or semantic analysis on user input in the form of text data to determine a user's intent. The natural language understanding module (1223) of one embodiment can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from user input, and can match the meaning of the determined words to the intent to determine the user's intent. The natural language understanding module (1223) can obtain intent information corresponding to a user's utterance. The intent information can be information indicating the user's intent determined by interpreting text data. The intent information can include information indicating an action or function that the user intends to execute using the device.
[0291] In one embodiment, the planner module (1225) can generate a plan using the intent and parameters determined by the natural language understanding module (1223). According to one embodiment, the planner module (1225) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (1225) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (1225) can determine parameters necessary to execute the determined plurality of operations or result values output by the execution of the plurality of operations. The parameters and the result values can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (1225) can determine the relationships between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (1225) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (1225) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (1225) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (1225) can generate the plan using information stored in a capsule database (1230) in which a set of relationships between concepts and actions is stored.
[0292] The natural language generation module (1227) of one embodiment can convert specified information into text format. The information converted into text format may be in the form of natural language speech. The text-to-speech conversion module (1229) of one embodiment can convert information in text format into information in speech format.
[0293] According to one embodiment, some or all of the functions of the natural language platform (1220) may also be implemented in the electronic device (1101).
[0294] The capsule database (1230) may store information on the relationships between a plurality of concepts and actions corresponding to a plurality of domains. According to one embodiment, a capsule may include a plurality of action objects (or action information) and concept objects (or concept information) included in a plan. According to one embodiment, the capsule database (1230) may store a plurality of capsules in the form of a concept action network (CAN). According to one embodiment, the plurality of capsules may be stored in a function registry included in the capsule database (1230).
[0295] The capsule database (1230) may include a strategy registry that stores strategy information required when determining a plan corresponding to a voice input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to a user input. According to one embodiment, the capsule database (1230) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database (1230) may include a layout registry that stores layout information of information output through the electronic device (1101). According to one embodiment, the capsule database (1230) may include a vocabulary registry that stores vocabulary information included in capsule information. According to one embodiment, the capsule database (1230) may include a dialog registry that stores information on dialogue (or interaction) with the user. The capsule database (1230) may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans.The developer tool may include a dialog editor that creates a dialogue with the user. The developer tool may also include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on the currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database (1230) may also be implemented within the electronic device (1101).
[0296] The execution engine (1240) of one embodiment can produce a result using the generated plan. The end user interface (1250) can transmit the produced result to the electronic device (1101). Accordingly, the electronic device (1101) can receive the result and provide the received result to the user. The management platform (1260) of one embodiment can manage information used in the intelligent server (1200). The big data platform (1270) of one embodiment can collect user data. The analysis platform (1280) of one embodiment can manage the quality of service (QoS) of the intelligent server (1200). For example, the analysis platform (1280) can manage the components and processing speed (or efficiency) of the intelligent server (1200).
[0297] In one embodiment, a service server (1290) may include CP service A (1291), CP service B (1292), and CP service C (1293). The service server (1290) in one embodiment may provide a service (e.g., food ordering or hotel reservation) specified for an electronic device (1101). According to one embodiment, the service server (1290) may be a server operated by a third party. The service server (1290) in one embodiment may provide information for generating a plan corresponding to a received user input to an intelligent server (1200). The provided information may be stored in a capsule database (1230). In addition, the service server (1290) may provide result information according to the plan to the intelligent server (1200).
[0298] In the integrated intelligence system of FIG. 12, the electronic device (1101) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.
[0299] In one embodiment, the electronic device (1101) may provide a voice recognition service through an intelligent app (or voice recognition app) stored internally. In this case, for example, the electronic device (1101) may recognize a user utterance or voice input received through the microphone and provide the user with a service corresponding to the recognized voice input.
[0300] In one embodiment, the electronic device (1101) may perform a designated operation, either alone or in conjunction with the intelligent server and / or service server, based on the received voice input. For example, the electronic device (1101) may execute an app corresponding to the received voice input and perform a designated operation through the executed app.
[0301] In one embodiment, when an electronic device (1101) provides a service together with an intelligent server (1200) and / or a service server (1290), the electronic device (1101) can detect a user utterance using an input module (1150) and generate a signal (or voice data) corresponding to the detected user utterance. The electronic device (1101) can transmit the voice data to the intelligent server (1200) using an interface (1177).
[0302] According to one embodiment, an intelligent server (1200) may generate a plan for performing a task corresponding to a voice input received from an electronic device (1101), or a result of performing an operation according to the plan. The plan may include, for example, a plurality of operations for performing a task corresponding to a user's voice input, and a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations, or result values output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.
[0303] An electronic device (1101) of one embodiment can receive the response using an interface (1177). The electronic device (1101) can output a voice signal generated within the electronic device (1101) to the outside using an audio output module (1155), or can output an image generated within the electronic device (1101) to the outside using a display module (1160).
[0304] Figure 13 is a diagram showing the form in which relationship information between concepts and operations according to various embodiments is stored in a database.
[0305] The capsule database (e.g., the capsule database (1230) of FIG. 12) of the intelligent server (e.g., the intelligent server (1200) of FIG. 12) may store capsules in the form of a CAN (concept action network) (1300). The capsule database may store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network).
[0306] The capsule database may store a plurality of capsules (capsule(A)(1301), capsule(B)(1304)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., capsule(A)(1301)) may correspond to one domain (e.g., location (geo), application). In addition, one capsule may correspond to at least one service provider (e.g., CP 1(1302), CP 2 (1303), CP 3(1306), or CP 4(1305)) for performing a function for a domain related to the capsule. According to one embodiment, one capsule may include at least one operation (1310) and at least one concept (1320) for performing a specified function.
[0307] The natural language platform (e.g., the natural language platform (1220) of FIG. 12) can generate a plan for performing a task corresponding to a received voice input using capsules stored in a capsule database. For example, the planner module of the natural language platform (e.g., the planner module (1225) of FIG. 12) can generate a plan using capsules stored in a capsule database. For example, a plan (1307) can be generated using actions (1301-1, 1301-3) and concepts (1301-2, 1301-4) of capsule A (1301) and actions (1304-1) and concepts (1304-2) of capsule B (1304).
[0308] FIG. 14 is a diagram illustrating a screen for processing voice input received through an intelligent app by an electronic device according to various embodiments.
[0309] The electronic device (1101) can execute an intelligent app to process user input via an intelligent server (e.g., the intelligent server (1200) of FIG. 12).
[0310] According to one embodiment, on the screen (1410), when the electronic device (1101) recognizes a designated voice input (e.g., wake up!) or receives an input via a hardware key (e.g., a dedicated hardware key), the electronic device (1101) may execute an intelligent app for processing the voice input. For example, the electronic device (1101) may execute an intelligent app while the schedule app is running. According to one embodiment, the electronic device (1101) may display an object (e.g., an icon) (1411) corresponding to the intelligent app on a display module (e.g., the display module (1160) of FIG. 11). According to one embodiment, the electronic device (1101) may receive a voice input by a user's speech. For example, the electronic device (1101) may receive a voice input such as "Tell me my schedule for this week!" According to one embodiment, the electronic device (1101) can display a user interface (UI) (1413) (e.g., an input window) of an intelligent app displaying text data of received voice input on a display module (e.g., the display module (1160) of FIG. 11).
[0311] According to one embodiment, on the screen (1420), the electronic device (1101) may display a result corresponding to the received voice input on a display module (e.g., the display module (1160) of FIG. 11). For example, the electronic device (1101) may receive a plan corresponding to the received user input and display 'this week's schedule' on the display module (1160) according to the plan.
[0312] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.
[0313] As described above, the electronic device (101) may include an input interface (320). The electronic device (101) may include a memory (340) that stores instructions and includes one or more storage media. The electronic device (101) may include at least one processor (310) that includes a processing circuit. The instructions, when the at least one processor (310) is individually or collectively executed, may cause the electronic device (101) to receive, through the input interface (320), a first voice signal including a first command based on activation of a virtual assistant. The instructions, when the at least one processor (310) is individually or collectively executed, may cause the electronic device (101) to generate, based on the first voice signal, first identification information indicating a first speaker of the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a function corresponding to the first command of the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to receive, through the input interface (320), a second voice signal including a second command following the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate, based on the second voice signal, second identification information indicating a second speaker of the second voice signal.The above instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information that is greater than or equal to a reference value.
[0314] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to activate the virtual assistant. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a continuous command function of the activated virtual assistant. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to receive the first voice signal and the second voice signal while the continuous command function of the virtual assistant is executed. The activation of the virtual assistant may be performed based on at least one of reception of a voice signal including a wake word through the input interface (320) or reception of an input for activating the virtual assistant.
[0315] According to one embodiment, the first voice signal may be a voice signal first received through the input interface (320) after the virtual assistant is activated.
[0316] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the first speech signal to a first trained model (370) for speaker feature extraction. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate the first identification information through the first trained model (370). The first identification information may include a vector value indicating the first speaker.
[0317] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to refrain from executing the function corresponding to the second command of the second voice signal based on the similarity value being less than the reference value.
[0318] According to one embodiment, the second voice signal may be received through the input interface (320) before a preset time period expires after the execution of the function corresponding to the first command of the first voice signal is completed, or may be received through the input interface (320) while the function corresponding to the first command of the first voice signal is being executed. The function corresponding to the first command of the first voice signal may include text to speech (TTS) synthesized sound reproduction.
[0319] In one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second speech signal to a first trained model (370) for speaker feature extraction. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate the second identification information via the first trained model (370). The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second identification information to a second trained model (375) for speaker verification using the first identification information. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, through the second trained model (375), the similarity value between the first identification information and the second identification information. The first identification information may include a vector value indicating the first speaker. The second identification information may include a vector value indicating the second speaker.
[0320] In one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to store third identification information indicating a registered speaker with respect to the virtual assistant prior to the activation of the virtual assistant. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify a similarity value between the generated first identification information and the third identification information through the second trained model (375) using the third identification information.
[0321] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the received second voice signal to a third trained model (380) for a voice filter using the first identification information or the third identification information, based on the similarity value between the first identification information and the third identification information being greater than or equal to the reference value. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate a filtered second voice signal through the third trained model (380), based on the similarity value between the first identification information and the third identification information being greater than or equal to the reference value. The above instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to provide the filtered second voice signal to the first trained model (370) based on the similarity value between the first identification information and the third identification information being greater than or equal to the reference value.
[0322] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to refrain from providing the received second voice signal to a third trained model (380) for a voice filter using the first identification information or the third identification information, based on the similarity value between the first identification information and the third identification information being less than the threshold value. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second voice signal to the first trained model (370) based on the similarity value between the first identification information and the third identification information being less than the threshold value.
[0323] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify a similarity value between the second identification information and the third identification information through the second trained model (375) using the third identification information. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute the function corresponding to the second command of the second voice signal based on the similarity value between the first identification information and the second identification information being less than the reference value and the similarity value between the first identification information and the second identification information being less than the reference value, and based on the similarity value between the second identification information and the third identification information being greater than or equal to the reference value. The instructions may cause the electronic device (101) to refrain from executing the function corresponding to the second command of the second voice signal based on the similarity value between the first identification information and the second identification information being less than the reference value and the similarity value between the first identification information and the second identification information being less than the reference value, and based on the similarity value between the second identification information and the third identification information being less than the reference value, when the at least one processor (310) is executed individually or collectively.
[0324] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to store the generated first identification information in the memory (340). The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to delete the stored first identification information from the memory (340) upon identifying a specific event. The specific event may include at least one of deactivation of the activated virtual assistant, deactivation of a continuous command function of the virtual assistant, or identification that the first speaker indicated by the first identification information matches a speaker registered with respect to the virtual assistant.
[0325] In one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second voice signal to a fourth trained model (385) for voice activity detection. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to detect, via the fourth trained model (385), a voice segment of the second command from the second voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate the second identification information based on the second command within the voice segment.
[0326] As described above, a method performed by an electronic device (101) having an input interface (320) may include an operation of receiving, through the input interface (320), a first voice signal including a first command based on activation of a virtual assistant. The method may include an operation of generating, based on the first voice signal, first identification information indicating a first speaker of the first voice signal. The method may include an operation of executing a function corresponding to the first command of the first voice signal. The method may include an operation of receiving, through the input interface (320), a second voice signal including a second command following the first voice signal. The method may include an operation of generating, based on the second voice signal, second identification information indicating a second speaker of the second voice signal. The method may include an operation of executing a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information being greater than or equal to a reference value.
[0327] The non-transitory computer-readable storage medium as described above may store one or more programs including instructions that, when individually or collectively executed by at least one processor (310) of an electronic device (101) having an input interface (320), cause the electronic device (101) to receive, through the input interface (320), a first voice signal including a first command based on activation of a virtual assistant. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to generate, based on the first voice signal, first identification information indicating a first speaker of the first voice signal. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to execute a function corresponding to the first command of the first voice signal. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to receive, through the input interface (320), a second voice signal including a second command following the first voice signal.The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to generate second identification information indicating a second speaker of the second voice signal based on the second voice signal. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to execute a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information being greater than or equal to a reference value.
[0328] As described above, the electronic device (101) may include an input interface (320). The electronic device (101) may include a memory (340) that stores instructions and includes one or more storage media. The electronic device (101) may include at least one processor (310) that includes a processing circuit. The instructions, when the at least one processor (310) is individually or collectively executed, may cause the electronic device (101) to receive, through the input interface (320), a first voice signal including a first command based on activation of a virtual assistant. The instructions, when the at least one processor (310) is individually or collectively executed, may cause the electronic device (101) to generate identification information indicating a first speaker of the first voice signal based on the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a function corresponding to the first command of the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to receive, through the input interface (320), a second voice signal including a second command following the first voice signal. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify, based on the identification information and the second voice signal, whether the first speaker corresponds to the second speaker of the second voice signal.The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a function corresponding to the second command of the second speech signal upon identifying that the first speaker corresponds to the second speaker. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to refrain from executing the function corresponding to the second command of the second speech signal upon identifying that the first speaker does not correspond to the second speaker.
[0329] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to activate the virtual assistant. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to execute a continuous command function of the activated virtual assistant. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to receive the first voice signal and the second voice signal while the continuous command function of the virtual assistant is executed. The activation of the virtual assistant may be performed based on at least one of reception of a voice signal including a wake word through the input interface (320) or reception of an input for activating the virtual assistant.
[0330] According to one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the first speech signal to a first trained model (370) for speaker feature extraction. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate the identification information through the first trained model (370). The identification information may include a vector value indicating the first speaker.
[0331] In one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second speech signal to the first trained model (370). The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to generate, via the first trained model (370), other identification information indicative of the second speaker. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the other identification information to a second trained model (375) for speaker verification using the identification information. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify, through the second trained model (375), a similarity value between the identification information and the other identification information. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify, based on the similarity value being greater than or equal to a threshold value, that the first speaker corresponds to the second speaker. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify, based on the similarity value being less than the threshold value, that the first speaker does not correspond to the second speaker. The other identification information may include a vector value indicating the second speaker.
[0332] In one embodiment, the instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to provide the second voice signal to a third trained model (385) for voice activity detection using the identification information. The instructions, when the at least one processor (310) is executed individually or collectively, may cause the electronic device (101) to identify, through the third trained model (385), that the first speaker corresponds to the second speaker upon detection of a voice segment of the second command in the second voice signal. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to identify, from the third trained model (385), that the first speaker does not correspond to the second speaker, based on the speech segment of the second command in the second speech signal not being detected.
[0333] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0334] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0335] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0336] Various embodiments of the present document may be implemented as software (e.g., a program (1140)) including one or more instructions stored in a storage medium (e.g., an internal memory (1136) or an external memory (1138)) readable by a machine (e.g., an electronic device (1101)). For example, a processor (e.g., a processor (1120)) of the machine (e.g., an electronic device (1101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0337] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0338] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, An input interface configured to receive sound data; A memory storing instructions and including one or more storage media; and At least one processor comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Upon activation of the virtual assistant, a first voice signal including a first command is received through the input interface; Based on the first voice signal, generate first identification information corresponding to a first speaker of the first voice signal; Executes a function corresponding to the first command of the first voice signal; Receive a second voice signal including a second command following the first voice signal through the input interface; Based on the second voice signal, generating second identification information corresponding to the second speaker of the second voice signal; and Causing to execute a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information that is greater than or equal to a reference value. Electronic devices.
2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Activate the above virtual assistant; Executes the continuous command function of the above activated virtual assistant; and In a state where the continuous command function of the virtual assistant is executed, causing the first voice signal and the second voice signal to be received, The virtual assistant is activated based on at least one of receiving a voice signal including a call command through the input interface or receiving an input for activating the virtual assistant. Electronic devices.
3. In claim 1, The first voice signal is a voice signal including a command first received through the input interface after the virtual assistant is activated. Electronic devices.
4. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: providing the first speech signal to a first trained model configured to perform speaker feature extraction; and By means of the first trained model, the first identification information is generated, The first identification information includes a vector value corresponding to the first speaker. Electronic devices.
5. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the similarity value being less than the reference value, causing the function corresponding to the second command of the second voice signal to be refrained from being executed, Electronic devices.
6. In claim 1, The second voice signal is: After the function corresponding to the first command of the first voice signal is executed, before the preset time period expires, it is received through the input interface, or In a state where the function corresponding to the first command of the first voice signal is being executed, it is received through the input interface, and The function corresponding to the first command of the first voice signal is configured to reproduce a TTS (text to speech) synthesized sound. Electronic devices.
7. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Providing the second speech signal to a first trained model configured to perform speaker feature extraction; Through the first trained model, the second identification information is generated; Providing the second identification information to a second trained model configured to perform speaker verification using the first identification information; and By means of the second trained model, the similarity value between the first identification information and the second identification information is identified, The first identification information includes a vector value corresponding to the first speaker, and The second identification information includes a vector value corresponding to the second speaker. Electronic devices.
8. In claim 7, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Before the activation of the virtual assistant, store third identification information corresponding to the registered speaker with respect to the virtual assistant; and By means of the second trained model, a similarity value between the first identification information and the third identification information is identified, Electronic devices.
9. In claim 8, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the similarity value between the first identification information and the third identification information that is greater than or equal to the reference value: Providing the second voice signal to a third trained model configured to perform voice filtering using the first identification information or the third identification information; Generating a filtered second speech signal through the third trained model; and causing the filtered second voice signal to be provided to the first trained model, Electronic devices.
10. In claim 8, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the similarity value between the first identification information and the third identification information that is less than the reference value: Refrain from providing the second voice signal to a third trained model configured to perform voice filtering using the first identification information or the third identification information; and causing the second voice signal to be provided to the first trained model, Electronic devices.
11. In claim 8, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identifying a similarity value between the second identification information and the third identification information through the second trained model using the third identification information; Executing the function corresponding to the second command of the second voice signal based on the similarity value between the first identification information and the second identification information that is less than the reference value and the similarity value between the second identification information and the third identification information that is greater than the reference value; and Based on the similarity value between the first identification information and the second identification information that is less than the reference value and the similarity value between the second identification information and the third identification information that is less than the reference value, causing the function corresponding to the second command of the second voice signal to be refrained from executing, Electronic devices.
12. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: storing the first identification information in the memory; and Based on a specific event, causing the stored first identification information to be deleted from the memory, The above specific events are: At least one of deactivating the activated virtual assistant, deactivating the continuous command function of the virtual assistant, or identifying that the first speaker corresponding to the first identification information matches a speaker registered with respect to the virtual assistant. Electronic devices.
13. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Providing the second voice signal to a fourth trained model configured to perform voice activity detection; Through the fourth trained model, detecting the speech section of the second command from the second speech signal; and Causing the second identification information to be generated based on the second command within the detected voice section. Electronic devices.
14. A method performed by an electronic device having an input interface configured to receive sound data, An action of receiving, through the input interface, a first voice signal including a first command based on activation of a virtual assistant; An operation of generating first identification information corresponding to a first speaker of the first voice signal based on the first voice signal; An operation of executing a function corresponding to the first command of the first voice signal; An operation of receiving, through the input interface, a second voice signal including a second command following the first voice signal; An operation of generating second identification information corresponding to a second speaker of the second voice signal based on the second voice signal; and An operation including executing a function corresponding to the second command of the second voice signal based on a similarity value between the first identification information and the second identification information that is greater than or equal to a reference value. method.
15. In a non-transitory computer-readable storage medium, when individually or collectively executed by at least one processor of an electronic device having an input interface configured to receive sound data, said electronic device: Upon activation of the virtual assistant, a first voice signal including a first command is received through the input interface; Based on the first voice signal, generate first identification information corresponding to a first speaker of the first voice signal; Executes a function corresponding to the first command of the first voice signal; Receive a second voice signal including a second command following the first voice signal through the input interface; Based on the second voice signal, generating second identification information corresponding to the second speaker of the second voice signal; and storing one or more programs including instructions that cause a function corresponding to the second command of the second voice signal to be executed based on a similarity value between the first identification information and the second identification information that is greater than or equal to a reference value; Non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Speaker identification method and speaker identification device
JP2016206660A
Functional beverage using Mentha piperascens and manufacturing method thereof
KR1020210106662A
Refrigerator
KR102163741B1
Speaker recognition apparatus and operation method thereof
KR102621897B1
Bird strike test apparatus for personal airial vehicle
KR102913212B1