Electronic device and method for determining speaker by using external electronic device
By integrating a microphone and processor into electronic devices to generate wake words and speaker identification results, and working in conjunction with an external server, the problem of speaker identification in multi-device environments is solved, achieving accurate user identification and automated device functions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-03-10
AI Technical Summary
In existing technologies, electronic devices struggle to effectively identify the speaker's identity, especially when multiple external devices are working collaboratively, as there is a lack of effective wake words and speaker identification mechanisms.
By integrating a microphone and processor into an electronic device, a wake word and speaker identification result are generated and sent to an external server for further verification and identification. The user's identity is confirmed using the registration information of the external device.
It enables accurate speaker identification in multi-device environments, improves the security and automation of electronic devices, and enhances the user experience.
Smart Images

Figure CN121646809A_ABST
Abstract
Description
Technical Field
[0001] The following description relates to an electronic device and a method for identifying a speaker using external electronic devices. Background Technology
[0002] Electronic devices can acquire sound signals from the outside via a microphone. For example, the sound signal can include the speech or voice produced by a speaker. For example, electronic devices can identify the speaker based on the sound signal.
[0003] The information described above may be provided as relevant technology for the purpose of aiding understanding of this disclosure. No statement or determination is made as to whether anything described above can be used as prior art in relation to this disclosure. Summary of the Invention
[0004] Technical solution
[0005] An electronic device may include communication circuitry. The electronic device may include a microphone. The electronic device may include at least one processor. The at least one processor may be configured to acquire a voice signal via the microphone. The at least one processor may be configured to generate a first recognition result for the wake word from the voice signal including the wake word. The at least one processor may be configured to send the first recognition result to a server connected to one or more external electronic devices capable of recognizing the wake word. The at least one processor may be configured to receive a second recognition result from the server for a user who issued the wake word. The at least one processor may be configured to run a specified function for that user within the electronic device based on the second recognition result. The second recognition result may include identification information indicating the user obtained from one or more external electronic devices registered by the user.
[0006] A method performed by an electronic device may include acquiring a voice signal. The method may include generating a first recognition result for the wake word from the voice signal, including the wake word. The method may include sending the first recognition result to a server connected to one or more external electronic devices capable of recognizing the wake word. The method may include receiving a second recognition result from the server for a user who issued the wake word. The method may include, based on the second recognition result, running a specified function for that user in the electronic device. The second recognition result may include identification information indicating the user obtained from one or more external electronic devices registered by the user.
[0007] A non-transitory computer-readable storage medium, when run by at least one processor of an electronic device including communication circuitry and a microphone, can store one or more programs including instructions that cause the electronic device to acquire a voice signal via the microphone. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to generate a first recognition result instruction for the wake-up word from the voice signal including the wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to send the first recognition result to a server connected to one or more external electronic devices capable of recognizing the wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to receive a second recognition result from the server for a user who issued the wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to perform a specified function for the user based on the second recognition result. The second recognition result may include identification information indicating the user obtained from one or more external electronic devices registered by the user.
[0008] An electronic device may include communication circuitry. The electronic device may include a microphone. The electronic device may include at least one processor. The at least one processor may be configured to acquire a voice signal via the microphone. The at least one processor may be configured to generate a first recognition result for the wake word from the voice signal including the wake word. The at least one processor may be configured to broadcast the first recognition result to one or more external electronic devices capable of recognizing the wake word. The at least one processor may be configured to receive a second recognition result for the wake word and a third recognition result for the user who issued the wake word, broadcast from one or more user-registered external electronic devices. The at least one processor may be configured to run a specified function for the user in the electronic device based on the third recognition result. The third recognition result may include identification information indicating the user.
[0009] A method performed by an electronic device may include acquiring a voice signal. The method may include generating a first recognition result for the wake word from the voice signal, including the wake word. The method may include broadcasting the first recognition result to one or more external electronic devices capable of recognizing the wake word. The method may include receiving a second recognition result for the wake word and a third recognition result for the user who issued the wake word, broadcast from one or more user-registered external electronic devices. The method may include performing a specified function for the user in the electronic device based on the third recognition result. The third recognition result may include identification information indicating the user.
[0010] A non-transitory computer-readable storage medium, when run by at least one processor of an electronic device including communication circuitry and a microphone, can store one or more programs including instructions that cause the electronic device to acquire a voice signal via the microphone. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to generate a first recognition result for a wake-up word from a voice signal including a wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to broadcast the first recognition result to one or more external electronic devices capable of recognizing the wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to receive a second recognition result for a wake-up word and a third recognition result for a user who issued the wake-up word, broadcast from one or more user-registered external electronic devices. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to perform a specified function for the user based on the third recognition result. The third recognition result may include identification information indicating the user.
[0011] An electronic device may include communication circuitry. The electronic device may include a microphone. The electronic device may include at least one processor. The at least one processor may be configured to acquire a speech signal via the microphone. The at least one processor may be configured to identify the user who issued the speech signal from the speech signal, including a wake-up word. The at least one processor may be configured to send a signal including a first identification result for the wake-up word and a second identification result for the user to a server connected to one or more external electronic devices, which are capable of identifying the wake-up word based on recognizing the user as a registered user in the electronic device. The first identification result may include at least one of the following: the accuracy value of the wake-up word in the speech signal recognized by the electronic device, the sound power level of the speech signal including the wake-up word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of acquiring the speech signal. The second identification result may include at least one of the following: identification information of the user who issued the wake-up word, including group information of the user, or the user's identification accuracy value.
[0012] A method performed by an electronic device may include acquiring a speech signal. The method may include identifying a user who issued the speech signal from a speech signal including a wake-up word. The method may include sending a signal including a first identification result for the wake-up word and a second identification result for the user to a server connected to one or more external electronic devices, which are capable of identifying the wake-up word based on the user being a registered user in the electronic device. The first identification result may include at least one of the following: a recognition accuracy value of the wake-up word in the speech signal recognized by the electronic device, a sound power level of the speech signal including the wake-up word, a signal-to-noise ratio (SNR) of the speech signal, or the timing of acquiring the speech signal. The second identification result may include at least one of the following: identification information of the user who issued the wake-up word, including group information of the user, or the user's recognition accuracy value.
[0013] A non-transitory computer-readable storage medium, when run by at least one processor of an electronic device including communication circuitry and a microphone, can store one or more programs including instructions to cause the electronic device to acquire a speech signal via the microphone. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to identify the user issuing the speech signal from a speech signal including a wake-up word. The non-transitory computer-readable storage medium, when run by at least one processor, can store one or more programs including instructions that cause the electronic device to send a signal including a first identification result for the wake-up word and a second identification result for the user to a server connected to one or more external electronic devices, the one or more external electronic devices being able to identify the wake-up word based on the identification that the user is a registered user in the electronic device. The first identification result may include at least one of the following: the identification accuracy value of the wake-up word in the speech signal identified by the electronic device, the sound power level of the speech signal including the wake-up word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of acquiring the speech signal. The second identification result may include at least one of the following: identification information of the user issuing the wake-up word, group information including the user, or the user's identification accuracy value. Attached Figure Description
[0014] Figure 1 This is a block diagram of an electronic device in a network environment according to various embodiments.
[0015] Figure 2a An example of a method for identifying wake words and speakers based on speech signals is shown.
[0016] Figure 2b An example of a method for registering speakers based on speech signals is shown.
[0017] Figure 2c An example of a method is shown where multiple electronic devices acquire speech signals and share the recognition results of the acquired speech signals.
[0018] Figure 3 An example is shown of how electronic devices use external electronic devices to identify the speaker.
[0019] Figure 4 An example of a multi-device wake-up (MDW) environment is shown, which includes an electronic device and one or more external electronic devices.
[0020] Figure 5 An example of the operation flow of a method for an electronic device, including a speaker recognition model, to recognize wake words in a speech signal is shown.
[0021] Figure 6An example of the operation flow of a method for an electronic device to identify wake words of speech signals without a speaker identification model is shown.
[0022] Figure 7 An example of the operational flow of a method for generating a speaker recognition model using an electronic device is shown.
[0023] Figure 8 An example of the operation flow of an electronic device using an external electronic device to identify the speaker is shown.
[0024] Figure 9 It is a block diagram indicating an integrated intelligent system according to various embodiments.
[0025] Figure 10 It is a diagram indicating the form in which information about the relationship between concepts and operations according to various embodiments is stored in a database.
[0026] Figure 11 It is a diagram illustrating the screen of an electronic device, according to various embodiments, processing voice input received through a smart application. Detailed Implementation
[0027] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of another embodiment. Singular expressions may include plural expressions unless the context clearly indicates otherwise. The terms used herein (including technical or scientific terms) may have the same meaning as commonly understood by one of ordinary skill in the art as described in this disclosure. Among the terms used in this disclosure, unless expressly defined herein, terms defined in a general dictionary may be interpreted as having the same or similar meaning as in the context of related art and are not to be interpreted as having an ideal or overly formal meaning. In some cases, even terms defined in this disclosure may not be construed as excluding embodiments of this disclosure.
[0028] In the various embodiments of this disclosure described below, hardware methods will be described as examples. However, since the various embodiments of this disclosure include techniques using both hardware and software, software-based methods are not excluded.
[0029] Furthermore, in this disclosure, the terms "greater than" or "less than" may be used to determine whether a particular condition is met, but this is merely a description of examples and does not exclude descriptions of "greater than or equal to" or "less than or equal to". A condition described as "greater than or equal to" may be replaced by "greater than", a condition described as "less than or equal to" may be replaced by "less than", and a condition described as "greater than or equal to and less than" may be replaced by "greater than and less than or equal to". Additionally, in the following, "A" to "B" refers to at least one of the elements from A (inclusive) to B (inclusive). In the following, "C" and / or "D" means including at least one of "C" and "D", i.e., {"C", "D" and "C" and "D"}.
[0030] Figure 1 This is a block diagram illustrating an electronic device 101 in a network environment 100 according to various embodiments.
[0031] refer to Figure 1 In network environment 100, electronic device 101 can communicate with electronic device 102 via a first network 198 (e.g., a short-range wireless communication network), or with at least one of electronic device 104 or server 108 via a second network 199 (e.g., a long-range wireless communication network). According to an embodiment, electronic device 101 can communicate with electronic device 104 via server 108. According to an embodiment, electronic device 101 may include a processor 120, memory 130, input module 150, sound output module 155, display module 160, audio module 170, sensor module 176, interface 177, connection terminal 178, haptic module 179, camera module 180, power management module 188, battery 189, communication module 190, subscriber identification module (SIM) 196, or antenna module 197. In some embodiments, at least one component (e.g., connection terminal 178) may be omitted from electronic device 101, or one or more other components may be added to electronic device 101. In some embodiments, some of the components (e.g., sensor module 176, camera module 180, or antenna module 197) may be implemented as a single component (e.g., display module 160).
[0032] Processor 120 can execute, for example, software (e.g., program 140) to control at least one other component (e.g., hardware or software component) of electronic device 101 coupled to processor 120, and can perform various data processing or calculations. According to embodiments, as at least part of data processing or calculation, processor 120 can store commands or data received from another component (e.g., sensor module 176 or communication module 190) in volatile memory 132, process the commands or data stored in volatile memory 132, and store the resulting data in non-volatile memory 134. According to embodiments, processor 120 may include a main processor 121 (e.g., a central processing unit (CPU) or application processor (AP)) or an auxiliary processor 123 (e.g., a graphics processing unit (GPU), neural processing unit (NPU), image signal processor (ISP), sensor central processor, or communication processor (CP)) that is operationally independent of or combined with the main processor 121. For example, when electronic device 101 includes a main processor 121 and an auxiliary processor 123, the auxiliary processor 123 may be adapted to consume less power than the main processor 121, or to be dedicated to a specific function. The auxiliary processor 123 may be implemented separately from the main processor 121, or may be implemented as part of the main processor 121.
[0033] When the main processor 121 is inactive (e.g., in sleep) state, the auxiliary processor 123 may control at least some of the functions or states associated with at least one component of the electronic device 101 (other than the main processor 121) (e.g., display module 160, sensor module 176, or communication module 190), or when the main processor 121 is active (e.g., executing an application), the auxiliary processor 123 may, together with the main processor 121, control at least some of the functions or states associated with at least one component of the electronic device 101 (e.g., display module 160, sensor module 176, or communication module 190). According to embodiments, the auxiliary processor 123 (e.g., an image signal processor or a communication processor) may be implemented as part of another component (e.g., camera module 180 or communication module 190) functionally associated with the auxiliary processor 123. According to embodiments, the auxiliary processor 123 (e.g., a neural processing unit) may include hardware structures specified for processing artificial intelligence models. Artificial intelligence models can be generated through machine learning. This learning can be performed, for example, by an electronic device 101 performing artificial intelligence or via a separate server (e.g., server 108). The learning algorithm can include, but is not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The artificial intelligence model can include multiple layers of artificial neural networks. The artificial neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of these, but is not limited thereto. Additionally or alternatively, the artificial intelligence model can include software structures in addition to hardware structures.
[0034] Memory 130 may store various data used by at least one component of electronic device 101 (e.g., processor 120 or sensor module 176). The various data may include, for example, software (e.g., program 140) and input or output data for commands associated therewith. Memory 130 may include volatile memory 132 or non-volatile memory 134.
[0035] Program 140 may be stored as software in memory 130 and may include, for example, an operating system (OS) 142, middleware 144, or application 146.
[0036] Input module 150 can receive commands or data from outside electronic device 101 (e.g., a user) to be used by another component of electronic device 101 (e.g., processor 120). Input module 150 may include, for example, a microphone, mouse, keyboard, keys (e.g., buttons), or digital pen (e.g., stylus).
[0037] The audio output module 155 can output audio signals to the outside of the electronic device 101. The audio output module 155 may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or playing records. The receiver can be used to receive incoming calls. According to an embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0038] Display module 160 can visually provide information to the outside of electronic device 101 (e.g., to a user). Display module 160 may include, for example, a display, a holographic device, or a projector, and control circuitry for controlling a respective one of the display, holographic device, and projector. According to an embodiment, display module 160 may include a touch sensor adapted to detect touch or a pressure sensor adapted to measure the intensity of the force caused by touch.
[0039] The audio module 170 can convert sound into electrical signals and vice versa. According to an embodiment, the audio module 170 can obtain sound via the input module 150, or output sound via the sound output module 155 or headphones of an external electronic device (e.g., electronic device 102) that is directly (e.g., wired) or wirelessly connected to the electronic device 101.
[0040] Sensor module 176 can detect the operating state of electronic device 101 (e.g., power or temperature) or the environmental state outside electronic device 101 (e.g., user state), and then generate an electrical signal or data value corresponding to the detected state. According to embodiments, sensor module 176 may include, for example, a gesture sensor, gyroscope sensor, atmospheric pressure sensor, magnetic sensor, accelerometer, grip sensor, proximity sensor, color sensor, infrared (IR) sensor, biometric sensor, temperature sensor, humidity sensor, or illuminance sensor.
[0041] Interface 177 may support one or more specific protocols used to enable electronic device 101 to connect directly (e.g., wired) or wirelessly to external electronic devices (e.g., electronic device 102). According to embodiments, interface 177 may include, for example, a High Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, a Secure Digital Card (SD) interface, or an audio interface.
[0042] Connection terminal 178 may include a connector, via which electronic device 101 can be physically connected to an external electronic device (e.g., electronic device 102). According to embodiments, connection terminal 178 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0043] The haptic module 179 can convert electrical signals into mechanical stimuli (e.g., vibration or motion) or electrical stimuli that can be recognized by a user through his touch or kinesthesia. According to embodiments, the haptic module 179 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.
[0044] Camera module 180 can capture still or moving images. According to an embodiment, camera module 180 may include one or more lenses, an image sensor, an image signal processor, or a flash.
[0045] The power management module 188 can manage the power supply to the electronic device 101. According to an embodiment, the power management module 188 can be implemented as at least part of, for example, a power management integrated circuit (PMIC).
[0046] Battery 189 can power at least one component of electronic device 101. According to embodiments, battery 189 may include, for example, a non-rechargeable primary battery, a rechargeable rechargeable battery, or a fuel cell.
[0047] Communication module 190 can support the establishment of a direct (e.g., wired) or wireless communication channel between electronic device 101 and external electronic devices (e.g., electronic device 102, electronic device 104, or server 108), and perform communication via the established communication channel. Communication module 190 may include one or more communication processors that can operate independently of processor 120 (e.g., application processor (AP)) and support direct (e.g., wired) or wireless communication. According to embodiments, communication module 190 may include wireless communication module 192 (e.g., cellular communication module, short-range wireless communication module, or Global Navigation Satellite System (GNSS) communication module) or wired communication module 194 (e.g., local area network (LAN) communication module or power line communication (PLC) module). One of these communication modules can communicate with an external electronic device via a first network 198 (e.g., a short-range communication network such as Bluetooth™, Wi-Fi Direct, or Infrared Data Association (IrDA)) or a second network 199 (e.g., a long-range communication network such as a traditional cellular network, 5G network, next-generation communication network, the Internet, or a computer network (e.g., a LAN or a wide area network (WAN)). These various types of communication modules can be implemented as a single component (e.g., a single chip) or as multiple components that are separate from each other (e.g., multiple chips). The wireless communication module 192 can use subscriber information (e.g., International Mobile Subscriber Identity (IMSI)) stored in the subscriber identification module 196 to identify and verify the electronic device 101 in the communication network (e.g., the first network 198 or the second network 199).
[0048] Wireless communication module 192 can support 5G networks and next-generation communication technologies, such as New Radio (NR) access technologies, following 4G networks. NR access technologies can support enhanced mobile broadband (eMBB), massive machine-type communication (mMTC), or ultra-reliable and low-latency communication (URLLC). Wireless communication module 192 can support high-frequency bands (e.g., millimeter-wave bands) to achieve, for example, high data transmission rates. Wireless communication module 192 can support various technologies used to ensure performance in high-frequency bands, such as beamforming, massive MIMO, full-dimensional MIMO (FD-MIMO), array antennas, analog beamforming, or massive antennas. Wireless communication module 192 can support various requirements specified in electronic device 101, external electronic device (e.g., electronic device 104), or network system (e.g., second network 199). According to an embodiment, the wireless communication module 192 may support peak data rates (e.g., 20 Gbps or higher) for implementing eMBB, loss coverage (e.g., 164 dB or lower) for implementing mMTC, or U-plane latency (e.g., 0.5 ms or less for each of the downlink (DL) and uplink (UL), or 1 ms or less round trip) for implementing URLLC.
[0049] Antenna module 197 can transmit or receive signals or power to or from the outside of electronic device 101 (e.g., external electronic device). According to an embodiment, antenna module 197 may include an antenna comprising a radiating element formed of a conductive material or conductive pattern formed in or on a substrate (e.g., a printed circuit board (PCB)). According to an embodiment, antenna module 197 may include multiple antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication scheme used in a communication network (such as a first network 198 or a second network 199) can be selected from the multiple antennas, for example by communication module 190 (e.g., wireless communication module 192). Signals or power can then be transmitted or received between communication module 190 and external electronic device via the selected at least one antenna. According to an embodiment, another component besides the radiating element (e.g., a radio frequency integrated circuit (RFIC)) may be additionally incorporated into antenna module 197.
[0050] According to various embodiments, antenna module 197 can form a millimeter-wave antenna module. According to embodiments, the millimeter-wave antenna module may include: a printed circuit board; an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high-frequency band (e.g., millimeter-wave band); and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top or side surface) of the printed circuit board and capable of transmitting or receiving signals in the specified high-frequency band.
[0051] At least some of the aforementioned components may be coupled to each other and transmit signals (e.g., commands or data) between them via peripheral communication schemes (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industrial processor interface (MIPI)).
[0052] According to an embodiment, commands or data can be sent or received between electronic device 101 and external electronic device 104 via server 108 coupled to a second network 199. Each of electronic devices 102 or 104 can be a device of the same or different type as electronic device 101. According to an embodiment, all or some operations to be performed at electronic device 101 can be performed at one or more of external electronic devices 102, 104, or 108. For example, if electronic device 101 is required to automatically perform a function or service, or to perform a function or service in response to a request from a user or another device, electronic device 101 may request one or more external electronic devices to perform at least a portion of the function or service, instead of performing the function or service itself, or may request one or more external electronic devices to perform at least a portion of the function or service in addition to performing the function or service. Upon receiving the request, one or more external electronic devices may perform at least a portion of the requested function or service, or perform additional functions or services related to the request, and transmit the result of the performance to electronic device 101. Electronic device 101 may provide the result, with or without further processing, as at least part of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technologies can be used, for example. Electronic device 101 can use, for example, distributed computing or mobile edge computing to provide ultra-low latency services. In another embodiment, external electronic device 104 may include Internet of Things (IoT) devices. Server 108 may be an intelligent server using machine learning and / or neural networks. According to embodiments, external electronic device 104 or server 108 may be included in a second network 199. Electronic device 101 can be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology or IoT-related technologies.
[0053] Figure 2a An example of a method for identifying wake words and speakers based on speech signals is shown.
[0054] Figure 2a Example 200 shows a method by which electronic device 101 performs wake word recognition and speaker recognition based on voice signal 210. Figure 2a The electronic device 101 can instruct Figure 1 Example of electronic device 101.
[0055] Referring to Example 200, electronic device 101 can receive voice signal 210. For example, electronic device 101 can receive voice signal 210 via a microphone (e.g., Figure 1 The input module 150 receives the voice signal 210 from outside the electronic device 101. For example, the voice signal 210 may include speech. For example, the voice signal 210 may also include speech, noise, or background sound. For example, speech may be received by the electronic device 101 or an external electronic device (e.g., Figure 1 The voice signal 210 is emitted by the user of the electronic device 102 or 104. For example, the voice signal 210 may be referred to as a sound signal, signal, utterance signal, wake word signal, external sound, or user signal.
[0056] Referring to Example 200, electronic device 101 can perform feature extraction 215 on speech signal 210. For example, electronic device 101 can extract multiple feature values from speech signal 210. Although not shown in Example 200, electronic device 101 can enhance the features of the speech signal 210's utterances before performing feature extraction 215 on speech signal 210. For example, enhancing the features of utterances can include at least one of the following operations: removing noise from speech signal 210, enhancing the utterance portions of speech signal 210, and normalizing the volume of the utterance portions. Electronic device 101 can obtain multiple feature values of speech signal 210 after enhancing the features of utterances. For example, electronic device 101 can obtain multiple feature values based on a Mel-frequency cepstral coefficient (MFCC) transformation algorithm. Electronic device 101 can obtain the spectrum by applying a Fast Fourier Transform (FFT) to each frame of speech signal 210. Electronic device 101 can obtain the frequency domain spectrum by applying an FFT to speech signal 210. Electronic device 101 can obtain the Mel spectrum by applying a Mel filter bank to the spectrum. Electronic device 101 can obtain the Mel spectrum based on a Mel scale indicating the relationship between the frequency domain and the low-frequency band as identified by a human. Electronic device 101 can obtain the MFCC by applying cepstral analysis to the Mel spectrum. The MFCC can be referred to as eigenvalues. For example, electronic device 101 can obtain eigenvalues as part of the total eigenvalues, which are peak values obtained based on cepstral analysis.
[0057] Referring to Example 200, electronic device 101 can identify wake words from speech signal 210 using wake word recognition engine 220. For example, electronic device 101 can identify wake words based on multiple feature values using wake word recognition engine 220. For example, wake word recognition engine 220 can be implemented using hardware, software, or a combination of hardware and software for wake word recognition. For example, electronic device 101 can use wake word recognition engine 220 and wake word recognition model 223 to identify wake words included in speech signal 210. For example, wake word recognition model 223 can include a statistical model trained to recognize (or distinguish) wake words. For example, the statistical model can include an artificial neural network model, a hidden Markov model (HMM), a Gaussian mixture model (GMM), a support vector machine (SVM), and a vector quantizer (VQ). For example, electronic device 101 can implement wake word recognition engine 220 and speaker recognition engine 225 as a recognition algorithm (e.g., speaker-related speech recognition). Electronic device 101 can perform speaker-related speech recognition without distinguishing between wake word recognition and speaker recognition by using a speaker-related speech recognition model. For example, based on speaker-related speech recognition, electronic device 101 can obtain wake word recognition information (e.g., the first recognition result below) and speaker recognition information (e.g., the second recognition result below). For example, a wake word may include a designated word or sentence used to instruct a particular electronic device. The wake word can be used to run the speech recognition function of the instructing particular electronic device. For example, electronic device 101 can determine that a user is calling (or invoking, instructing) electronic device 101 based on recognizing that a wake word is included in speech signal 210. For example, the wake word may be preset by electronic device 101 or set by the user. For example, electronic device 101 can generate (or obtain) a recognition result for the wake word in speech signal 210 based on wake word recognition engine 220 and wake word recognition model 223. Hereinafter, the recognition result for the wake word may be referred to as the first recognition result, wake word recognition result, wake word recognition information, or wake word information. The first discrimination result can be defined as shown in the table below.
[0058] [Table 1]
[0059]
[0060] Referring to the table above, the first discrimination result may include a discrimination score, speech power (or loudness or sound pressure level), signal-to-noise ratio (SNR), and time. However, embodiments of this disclosure are not limited thereto. For example, the first discrimination result may also include only speech power. Alternatively, for example, the first discrimination result may include information on the priority between electronic devices in addition to the information illustrated in the table. For example, priority may indicate the priority in response to a wake word. For example, priority may be set by the user of the electronic device.
[0061] For example, a discrimination score can indicate discrimination information about a wake word. For instance, electronic device 101 can obtain a discrimination score that indicates the degree of discrimination of the wake word in speech signal 210 as a probability value. For instance, electronic device 101 can obtain a discrimination score based on wake word discrimination engine 220. The discrimination score can be referred to as the reliability (or discrimination reliability) of the wake word or a discrimination accuracy value.
[0062] For example, speech power can indicate the sound power level of speech signal 210. For example, electronic device 101 can identify the level of speech signal 210. Speech power or sound power level can be used to identify (or distinguish) the electronic device located closest to the user emitting speech signal 210. For example, speech power can be referred to as speech power, sound power, or sound power level.
[0063] For example, SNR can indicate the ratio between signal (e.g., speech) and noise in speech signal 210. For example, electronic device 101 can identify the noise level of the environment (or space) from which speech signal 210 is emitted based on SNR. For example, time can indicate the timing of acquiring speech signal 210.
[0064] For example, electronic device 101 can generate device information associated with the first identification result. The device information may include information about electronic device 101 that has obtained the first identification result. For example, the device information may include a device name, device type, device identifier (device ID), wake word, and the device's Internet Protocol (IP) address. For example, the device name may include the model name of electronic device 101. For example, the device type may include the type of electronic device 101 (e.g., TV). For example, the device identifier may include a unique identifier for electronic device 101. For example, the wake word may include the name of the wake word set for electronic device 101. The contents of the table above are merely examples for ease of description, and the embodiments of this disclosure are not limited thereto. For example, the device information may be sent or received (i.e., shared) together with the first identification result.
[0065] Referring to Example 200, electronic device 101 can identify the speaker (or user) emitting the speech signal 210 from the speech signal 210 using speaker identification engine 225. For example, speaker identification engine 225 can be implemented by hardware, software, or a combination of hardware and software for speaker identification. For example, electronic device 101 can use speaker identification engine 225 and speaker identification model 228 to identify the speaker. For example, speaker identification model 228 can include a statistical model trained to identify (or distinguish) a speaker. For example, the statistical model can include an artificial intelligence model (AI model) or an artificial neural network model. For example, the speaker can indicate a user registered with electronic device 101. For example, electronic device 101 can register a speaker based on multiple speech signals emitted by the speaker. Registration can instruct electronic device 101 to include storing information about the speaker in electronic device 101. Specific details related to this are described below. Figure 2b As described herein. For example, electronic device 101 can generate (or obtain) a speaker identification result. In the following text, the speaker identification result may be referred to as a second identification result. For example, the second identification result may be referred to as a speaker identification result, speaker identification information, or speaker information.
[0066] Referring to Example 200, electronic device 101 can generate a discrimination result 230. Discrimination result 230 may include a first discrimination result and a second discrimination result. However, embodiments of this disclosure are not limited thereto. For example, if electronic device 101 does not include a speaker discrimination engine 225, discrimination result 230 may only include the first discrimination result. For example, electronic device 101 may implement the wake word discrimination engine 220 and the speaker discrimination engine 225 as a discrimination algorithm (e.g., speaker-related speech discrimination). When electronic device 101 uses speaker-related speech discrimination, the first discrimination result may include speaker discrimination information (i.e., the second discrimination result). Figure 2a The present invention illustrates an example 200 of an electronic device 101 generating (or obtaining) a discrimination result 230, but embodiments thereof are not limited thereto. For example, the electronic device 101 may share the generated discrimination result 230 with one or more external electronic devices. For example, sharing may include sending (or broadcasting) information about the discrimination result 230 to one or more external electronic devices, and receiving discrimination results from one or more external electronic devices. Specific details related to this are described below. Figure 2c As described in the text.
[0067] Figure 2b An example of a method for registering speakers based on speech signals is shown.
[0068] Figure 2bExample 240 shows a method by which electronic device 101 registers a speaker based on voice signal 210. Figure 2b The electronic device 101 can instruct Figure 1 Example of electronic device 101.
[0069] Referring to Example 240, the electronic device 101 can acquire a speech signal 210 and perform feature extraction 215 on the acquired speech signal 210. For more details, please refer to [reference needed]. Figure 2a The content is essentially the same. For example, electronic device 101 can obtain multiple speech signals 210 from the user that include the same words or sentences. For example, the same words or sentences may indicate a wake word. Referring to example 240, electronic device 101 can perform model training 245 based on multiple feature values obtained from the speech signals 210. For example, model training 245 may include training a statistical model (e.g., an artificial intelligence model) for registering speakers for electronic device 101. For example, electronic device 101 can register speakers corresponding to account information by performing model training 245 based on account information using multiple feature values. For example, account information may include information for identifying speakers.
[0070] Referring to Example 240, electronic device 101 can register a speaker by performing model training 245 using a specified number of speech signals, including speech signal 210. The registered speaker can instruct electronic device 101, having already acquired the speech signal 210 emitted by the speaker, to identify the speaker based on the speech signal 210. Example 240 describes an example of performing model training 245 based on a specified number of speech signals; however, embodiments of this disclosure are not limited thereto. For example, electronic device 101 can also perform model training 245 based on speech signals of a specified duration.
[0071] Figure 2c An example of a method is shown where multiple electronic devices acquire speech signals and share the recognition results of the acquired speech signals.
[0072] Figure 2c Example 250 illustrates a method by which each of a plurality of electronic devices 260 acquires a speech signal 210 and shares the discrimination results for the acquired speech signal 210. Figure 2c The electronic device 101 can instruct Figure 1 Example of electronic device 101.
[0073] Referring to Example 250, each of the plurality of electronic devices 260 can receive the voice signal 210. For example, the plurality of electronic devices 260 may include a first electronic device 261, a second electronic device 262, and an Nth electronic device 263. For example, the first electronic device 261 may indicate... Figure 1Examples of electronic devices 101. Multiple electronic devices 260 can refer to electronic devices located in the same space. For example, the same space can be defined based on connection to the same access point (AP) or based on the distance of short-range communication (e.g., Bluetooth or Wi-Fi). For example, multiple electronic devices 260 can include electronic devices connected to the same AP or electronic devices connected based on short-range communication. The same wake word can be set (or used) in multiple electronic devices 260. For example, the wake word can refer to a wake word included in voice signal 210.
[0074] Referring to Example 250, multiple electronic devices 260 can share the recognition result for the speech signal 210. For example, the recognition result may include a first recognition result. Alternatively, for example, the recognition result may include a first recognition result and a second recognition result. Referring to Example 250, a first electronic device 261 can acquire the speech signal 210 emitted by a speaker (or user) and generate a recognition result 271 based on the acquired speech signal 210. The first electronic device 261 can send (or broadcast) the recognition result 271 to a second electronic device 262 and an Nth electronic device 263. The second electronic device 262 and the Nth electronic device 263 can receive (or acquire) the recognition result 271. Furthermore, referring to Example 250, a second electronic device 262 can acquire the speech signal 210 emitted by a speaker (or user) and generate a recognition result 272 based on the acquired speech signal 210. The second electronic device 262 can send (or broadcast) the recognition result 272 to the first electronic device 261 and the Nth electronic device 263. The first electronic device 261 and the Nth electronic device 263 can receive (or obtain) the discrimination result 272. Furthermore, referring to Example 250, the Nth electronic device 263 can obtain the speech signal 210 emitted by the speaker (or user) and generate the discrimination result 273 based on the obtained speech signal 210. The Nth electronic device 263 can send (or broadcast) the discrimination result 273 to the first electronic device 261 and the second electronic device 262. The first electronic device 261 and the second electronic device 262 can receive (or obtain) the discrimination result 273.
[0075] For example, each of the plurality of electronic devices 260 can identify the electronic device indicated by the speech signal 210 based on the obtained discrimination result. For example, each of the plurality of electronic devices 260 can identify the electronic device with the highest speech power level of the speech signal 210 as the electronic device indicated by the speech signal 210 based on the obtained discrimination result. The speech power level can be compared by values corrected (or processed) for the speech signal 210 obtained from each of the plurality of electronic devices 260. The electronic device indicated by the speech signal 210 can indicate the electronic device in which the speaker who issued the speech signal 210 intends to perform a function.
[0076] refer to Figures 2a to 2c The electronic device 101 can acquire voice signals, identify wake words, and, based on processing the acquired voice signals, determine whether the speaker emitting the voice signal 210 is a registered speaker. However, in this example, the electronic device 101 does not include a speaker identification module (e.g., Figure 2a If the speaker identification engine 225 and speaker identification model 228 have not yet generated an acoustic model for registering a speaker in electronic device 101, electronic device 101 may have difficulty identifying whether the speaker emitting voice signal 210 is a registered speaker. In the following, in electronic devices and methods according to embodiments of this disclosure, when multiple electronic devices using the same wake word exist in the same space (e.g., the range in which the wake word emitted by the user has an effect, and the range in which the device (e.g., electronic device 101) is activated by the wake word emitted by the user), some electronic devices may provide speaker identification-based services based on the identification result of the speaker emitting the wake word (e.g., a second identification result). For example, a speaker identification-based service may include the operation of a specified function (such as an application running for a specific speaker or a specific UI display). For example, an application running for a specific speaker may include a software application in which speaker authentication is required. However, embodiments of this disclosure are not limited thereto. For example, a speaker identification-based service may include running at least one instruction stored in electronic device 101, performing any operation, or running a function set by the user. When multiple electronic devices using the same wake word exist in the same space, for example, the electronic device and method according to embodiments of this disclosure can provide a speaker-based service by sharing speaker information (e.g., a second recognition result) with another electronic device among the multiple electronic devices, which is identified by an electronic device among the multiple electronic devices that has completed the registration of the speaker who issued the wake word. The other electronic device may be an electronic device where the speaker registration is incomplete or does not include a speaker recognition module (e.g., a speaker recognition engine and an acoustic model).
[0077] Figure 3 An example is shown of how electronic devices use external electronic devices to identify the speaker.
[0078] Figure 3 Example 300 illustrates a method by which electronic device 101 uses external electronic device 103 to identify speaker 310. Figure 3 The electronic device 101 can instruct Figure 1 Example of electronic device 101.
[0079] Referring to Example 300, external electronic device 103 can receive a voice signal 315 emitted by speaker 310. For example, external electronic device 103 can receive a voice signal 315 including a "wake word". In Example 300, it is described as a "wake word", but embodiments of this disclosure are not limited thereto. For example, a "wake word" can include a specified word or sentence for electronic device 101 and external electronic device 103. Furthermore, for example, voice signal 315 can include not only a wake word but also a command word. For example, a wake word can be used to operate the voice recognition function of electronic device 101 or external electronic device 103. For example, a command can be used to command a specified function according to the command when the voice recognition function of external electronic device 103 or electronic device 101 is activated.
[0080] Referring to Example 300, external electronic device 103 can distinguish a wake word and speaker 310 based on speech signal 315. For example, external electronic device 103 can generate a first discrimination result for the wake word based on speech signal 315. Furthermore, for example, external electronic device 103 can generate a second discrimination result for speaker 310 based on speech signal 315. In this case, external electronic device 103 can instruct an electronic device in which speaker 310 is registered (i.e., in which a speaker discrimination model is generated and information about speaker 310 is stored). External electronic device 103 can send a discrimination result 320, including the first and second discrimination results, to electronic device 101. Figure 3 Example 300 illustrates an example where an external electronic device 103 sends a discrimination result 320 to an electronic device 101, but embodiments of this disclosure are not limited thereto. For example, electronic device 101 can acquire a speech signal 315 emitted by a speaker 310 and generate a discrimination result based on the acquired speech signal 315. Electronic device 101 can then send the discrimination result to external electronic device 103. Therefore, external electronic device 103 can receive the discrimination result. In other words, in Figure 3 In Example 300, external electronic device 103 and electronic device 101 can share the recognition results. In this case, the recognition results sent by electronic device 101 may only include the first recognition result for the wake word. This could be because electronic device 101 does not include a speaker recognition module, or it is an electronic device that has not registered a speaker.
[0081] According to an embodiment, the external electronic device 103 can identify whether the electronic device being called by the speaker 310 is the external electronic device 103 based on the discrimination result 320. The target can be referred to as the indication target. For example, the external electronic device 103 can identify whether the electronic device that the speaker 310 intends to identify by the wake word is the external electronic device 103 based on the first discrimination result of the discrimination result 320. In this case, the external electronic device 103 can identify whether the target is the external electronic device 103 based on the discrimination result received from the electronic device 101 and the first discrimination result of the discrimination result 320. Figure 3 In Example 300, it is assumed that the target called by speaker 310 is electronic device 101. Therefore, external electronic device 103 can identify that the target is not external electronic device 103 based on the first discrimination result of discrimination result 320 and the discrimination result received from electronic device 101.
[0082] According to an embodiment, electronic device 101 can identify whether a target is electronic device 101 based on a first discrimination result of discrimination result 320 and a discrimination result generated by electronic device 101. For example, electronic device 101 can identify the electronic device with the highest speech power level based on a first discrimination result received from external electronic device 103 and a discrimination result identified by electronic device 101. Figure 3 In Example 300, the electronic device with the highest speech power level can be electronic device 101.
[0083] According to an embodiment, electronic device 101 can identify whether the discrimination result 320 includes a second discrimination result. For example, if the discrimination result 320 includes a second discrimination result, electronic device 101 can provide a speaker discrimination service. According to an embodiment, electronic device 101 can operate a specified function. For example, the specified function can be included in the speaker discrimination service. For example, electronic device 101 can display a user interface (UI) 330 as a specified function. For example, UI 330 can include a visual object indicating the sentence "Hello, user". The user can point to speaker 310.
[0084] Referring to Example 300, even if electronic device 101 does not include a speaker identification module (or speaker identification model) for identifying speaker 310, or is an electronic device that has not registered information about speaker 310, electronic device 101 can still provide speaker identification services based on speaker information received from external electronic device 103. Referring to the above description, speaker identification services can be provided not only in devices used only by a specific user, but also in electronic device 101, which is available to multiple users. For example, a device available to multiple users can be referred to as a public device or a shared device. For example, a device used only by a specific user can be referred to as a personal device. However, embodiments of this disclosure are not limited thereto. For example, embodiments of this disclosure can be applied even if the device is used only by a specific user, and if another user is registered and speaker identification services are provided for that user.
[0085] Figure 4 An example of a multi-device wake-up (MDW) environment is shown, which includes an electronic device and one or more external electronic devices.
[0086] Figure 4 An example of an MDW environment 400 including electronic device 101, external electronic device 103, and external electronic device 105 is shown. For example, the MDW environment 400 may indicate a situation where multiple electronic devices have the same wake word set (or used). For example, multiple electronic devices can distinguish the wake word. For example, multiple electronic devices may include electronic device 101, external electronic device 103, and external electronic device 105. Figure 4 The MDW environment 400 described herein is merely an example for ease of description, and the embodiments disclosed herein are not limited thereto. For example, the MDW environment 400 may include a greater number of electronic devices, or it may include only two electronic devices.
[0087] refer to Figure 4According to embodiments, the MDW environment 400 may include an electronic device 101, an external electronic device 103, an external electronic device 105, and a server 107. For example, electronic device 101 may refer to an example of a device that does not include a speaker identification module. For example, external electronic device 103 may refer to an example of a device that includes a speaker identification module and has completed speaker registration. For example, external electronic device 105 may refer to an example of a device that includes a speaker identification module but has not completed speaker registration. The device that has completed speaker registration may refer to a device that generates a speaker identification model for the speaker. For example, server 107 may refer to an example of an external electronic device connected to and managing electronic device 101, external electronic device 103, and external electronic device 105. For example, server 107 may collect identification results received from each of electronic device 101, external electronic device 103, and external electronic device 105, and provide information about the identification results to a specific electronic device that is the target of the indication based on the collected identification results. In this example, an example of server 107 providing information to a specific electronic device that is the target of the indication is described, but embodiments of this disclosure are not limited thereto. For example, server 107 can also provide information about the recognition results to all electronic devices connected to server 107. According to an embodiment, the information about the recognition results provided by server 107 may include at least one of a first recognition result for the wake word and a second recognition result for the speaker received from each of the electronic devices connected to server 107. For example, server 107 may obtain a first recognition result received from electronic device 101, a first and second recognition result received from external electronic device 103, and a first recognition result received from external electronic device 105, and may provide information including the obtained recognition results.
[0088] According to an embodiment, electronic device 101 may include a microphone 411, a preprocessing unit 413, a wake word recognition engine 415, and an MDW manager 417. Although not described in... Figure 4 As shown, however, electronic device 101 may include at least one processor and at least one communication circuit. For example, at least one processor may control at least one communication circuit, microphone 411, preprocessing unit 413, wake word recognition engine 415, and MDW manager 417. For example, at least one communication circuit may be used by electronic device 101 to perform communication with external electronic devices 103 and 105 and server 107. However, embodiments of this disclosure are not limited thereto. For example, electronic device 101 may also include other components besides microphone 411, preprocessing unit 413, wake word recognition engine 415, MDW manager 417, at least one processor, and at least one communication circuit.
[0089] At least one processor of the electronic device 101 according to an embodiment may include hardware components for processing data based on one or more instructions. The hardware components for processing data may include, for example, an arithmetic and logic unit (ALU), a floating-point unit (FPU), and a field-programmable gate array (FPGA). As an example, the hardware components for processing data may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing unit (DSP), and / or a neural processing unit (NPU). At least one processor may include... Figure 1 At least a portion of the processor 120.
[0090] For example, at least one processor of electronic device 101 may include various processing circuits and / or multiple processors. For instance, the term "processor" as used in this document (including the claims) may include various processing circuits comprising at least one processor, and one or more of the at least one processor may be configured to individually and / or jointly perform the various functions described below in a distributed manner. As used below, in cases where "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms are not limited to the examples and include cases where one processor performs a portion of the referenced function and other processors perform another portion of the referenced function, and / or cases where one processor can perform all of the referenced functions. Furthermore, at least one processor may include, for example, a combination of processors performing, in a distributed manner, the various functions listed / disclosed. At least one processor may execute program instructions to implement or perform various functions.
[0091] For example, electronic device 101 can receive voice signals from the outside via microphone 411. The voice signal may include a wake-up word, for example. Alternatively, the voice signal may include both a wake-up word and a command word. The wake-up word may include a designated word or sentence registered in electronic device 101.
[0092] For example, electronic device 101 can perform speech signal processing through preprocessing unit 413. For instance, electronic device 101 can extract multiple feature values of the speech signal through preprocessing unit 413 and enhance the features of the speech portion of the speech signal. Specific details related to this can be applied in a substantially similar manner. Figure 2a The content.
[0093] For example, electronic device 101 can identify (or distinguish) wake words included in the speech signal based on wake word recognition engine 415. For example, electronic device 101 can generate a first recognition result for the wake word identified based on wake word recognition engine 415.
[0094] For example, electronic device 101 can share the first identification result through MDW manager 417. For example, electronic device 101 can identify external electronic devices 103 and 105 in MDW environment 400 through MDW manager 417. For example, electronic device 101 can send (or provide) the first identification result to external electronic devices 103 and 105. Furthermore, for example, electronic device 101 can receive (or obtain) the identification result from each of external electronic devices 103 and 105 through MDW manager 417. Alternatively, for example, electronic device 101 can send (or provide) the first identification result to server 107 in MDW environment 400 through MDW manager 417. Furthermore, for example, electronic device 101 can receive (or obtain) information about the identification results of external electronic devices 103 and 105 from server 107 in MDW environment 400 through MDW manager 417.
[0095] For example, electronic device 101 can disable the voice service until a device is identified as the target of instruction in the MDW environment 400 by the MDW manager 417. Alternatively, for example, if another device in the MDW environment 400 (e.g., external electronic device 103 or external electronic device 105) is not identified as the target of instruction by the MDW manager 417, electronic device 101 can disable the voice service. In other words, electronic device 101 can control the process to wait without activating the voice service until the target of instruction is identified. For example, if electronic device 101 in the MDW environment 400 is identified as the target of instruction by the MDW manager 417, electronic device 101 can process command words included in subsequently input voice signals by activating the voice service.
[0096] Although not in Figure 4 As shown, but the electronic device 101 according to an embodiment may also include a memory. The memory may include hardware components for storing data and / or instructions input to and / or output from at least one processor. The memory may include, for example, volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM)). Volatile memory may include at least one of, for example, dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include at least one of, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multimedia card (eMMC). The memory may include... Figure 1 At least a portion of the memory 130.
[0097] Although not in Figure 4 As shown, however, in the memory of the electronic device 101 according to an embodiment, one or more instructions (or command words) instructing the processor 501 of the electronic device 101 to perform calculations and / or operations on data may be stored. The collection of one or more instructions may be referred to as a program, firmware, operating system, process, routine, subroutine, and / or application. In the following, an application installed in an electronic device (e.g., electronic device 101) may mean that one or more instructions provided in the form of an application are stored in memory, and that one or more applications are stored by the processor of the electronic device in an executable format (e.g., a file with an extension specified by the operating system of the electronic device 101).
[0098] According to an embodiment, the external electronic device 103 may include a microphone 431, a preprocessing unit 433, a wake-word recognition engine 435, a speaker recognition engine 437, a speaker registration module 439, and an MDW manager 441. Although not explicitly stated... Figure 4 As shown, however, the external electronic device 103 may include a memory, at least one processor, and at least one communication circuit. For example, the at least one processor may control at least one communication circuit, a microphone 431, a preprocessing unit 433, a wake-word recognition engine 435, a speaker recognition engine 437, a speaker registration module 439, and an MDW manager 441. For example, at least one communication circuit may be used by the external electronic device 103 to perform communication with electronic devices 101, 105, and 107. However, embodiments of this disclosure are not limited thereto. For example, the external electronic device 103 may also include other components besides the microphone 431, preprocessing unit 433, wake-word recognition engine 435, speaker recognition engine 437, speaker registration module 439, MDW manager 441, at least one processor, and at least one communication circuit.
[0099] For example, external electronic device 103 can acquire voice signals from the outside via microphone 431. For example, the voice signal may include a wake-up word. Alternatively, for example, the voice signal may include a wake-up word and a command word. The wake-up word may include a designated word or sentence registered in external electronic device 103.
[0100] For example, the external electronic device 103 can perform speech signal processing through the preprocessing unit 433. For example, the external electronic device 103 can extract multiple feature values of the speech signal through the preprocessing unit 433 and enhance the features of the speech portion of the speech signal. Specific details related to this can be applied in a substantially similar manner. Figure 2a The content.
[0101] For example, external electronic device 103 can identify wake words included in the speech signal based on wake word recognition engine 435. For example, external electronic device 103 can generate a first recognition result for the wake word identified based on wake word recognition engine 435. The first recognition result can refer to the example in Table 1.
[0102] For example, external electronic device 103 can identify the speaker emitting the speech signal based on speaker identification engine 437. Although Figure 4 An external electronic device 103 including a speaker identification engine 437 is shown, but embodiments of this disclosure are not limited thereto. For example, the speaker identification engine 437 may include a speaker identification model trained for a specific speaker, such as speaker identification model 461 of the external electronic device 105. For example, a speaker may include a speaker registered based on a speaker registration module 439 included in the external electronic device 103. For example, a registered speaker may include at least one user. For example, a registered speaker may be registered in a state associated with account information. For example, account information may be provided to the speaker registration module 439 from an account manager (not shown). For example, account information may include a user's account identifier (or ID). For example, the external electronic device 103 may identify whether the speaker emitting the voice signal is a registered speaker. For example, the external electronic device 103 may generate (or obtain) a second identification result for a speaker through the speaker identification engine 437. The second identification result may be defined as shown in the table below.
[0103] [Table 2]
[0104]
[0105] Referring to the table above, the second discrimination result may include speaker group, speaker ID, and confidence score. However, embodiments of this disclosure are not limited thereto. For example, the second discrimination result may also include only speaker ID. Alternatively, for example, the first discrimination result may include priority information between electronic devices in addition to the information illustrated in the table.
[0106] For example, a speaker ID can indicate an identifier used to identify the speaker emitting the voice signal. For example, the speaker ID can be set to an encrypted value or a specified value provided from server 107. For example, the speaker ID can be referred to as a speaker account or speaker identification information. For example, a speaker group can indicate a group that includes the speaker. For example, a speaker group can include the group to which the speaker ID belongs (e.g., family and friends).
[0107] For example, a confidence score can indicate discriminative information about the speaker. For instance, external electronic device 103 can obtain a confidence score that indicates the degree of speaker discriminability in the speech signal as a probability value. For example, external electronic device 103 can obtain a confidence score based on speaker discrimination engine 437. The confidence score can be referred to as the speaker's reliability.
[0108] Referring to the table above, the external electronic device 103 can generate a first discrimination result and device information together with a second discrimination result. For example, the first discrimination result may include a discrimination score, speech power, signal-to-noise ratio (SNR), and time. For example, the device information may include information about the external electronic device 103 that has obtained the first discrimination result. For example, the device information may include the device name, device type, device identifier (device ID), wake word, and the device's Internet Protocol (IP) address. Specific details related to this can be found in essentially the same manner. Figure 2a The content.
[0109] For example, external electronic device 103 can share the first and second identification results through MDW manager 441. For example, external electronic device 103 can identify electronic device 101 and external electronic device 105 in MDW environment 400 through MDW manager 441. For example, external electronic device 103 can send (or provide) the first and second identification results to electronic device 101 and external electronic device 105. Furthermore, for example, external electronic device 103 can receive (or obtain) the identification results from each of electronic device 101 and external electronic device 105 through MDW manager 441. Alternatively, for example, external electronic device 103 can send (or provide) the first and second identification results to server 107 in MDW environment 400 through MDW manager 441. Furthermore, for example, external electronic device 103 can receive (or obtain) information about the identification results of electronic device 101 and external electronic device 105 from server 107 in MDW environment 400 through MDW manager 441.
[0110] According to an embodiment, the external electronic device 105 may include a microphone 451, a preprocessing unit 453, a wake-word recognition engine 455, a speaker recognition engine 457, a speaker registration module 459, a speaker recognition model 461, and an MDW manager 463. Although not explicitly stated... Figure 4As shown, however, external electronic device 105 may include a memory, at least one processor, and at least one communication circuit. For example, at least one processor may control at least one communication circuit, microphone 451, preprocessing unit 453, wake-word recognition engine 455, speaker recognition engine 457, speaker registration module 459, speaker recognition model 461, and MDW manager 463. For example, at least one communication circuit may be used by external electronic device 105 to perform communication with electronic device 101, external electronic device 103, and server 107. However, embodiments of this disclosure are not limited thereto. For example, external electronic device 105 may also include other components besides microphone 451, preprocessing unit 453, wake-word recognition engine 455, speaker recognition engine 457, speaker registration module 459, speaker recognition model 461, MDW manager 463, at least one processor, and at least one communication circuit.
[0111] For example, external electronic device 105 can acquire voice signals from the outside via microphone 451. For example, the voice signal may include a wake-up word. Alternatively, for example, the voice signal may include a wake-up word and a command word. The wake-up word may include a designated word or sentence registered in external electronic device 105.
[0112] For example, the external electronic device 105 can perform speech signal processing through the preprocessing unit 453. For instance, the external electronic device 105 can extract multiple feature values from the speech signal through the preprocessing unit 453 and enhance the features of the speech portion of the speech signal. Specific details related to this can be applied in a substantially similar manner. Figure 2a The content.
[0113] For example, external electronic device 105 can identify wake words included in the speech signal based on wake word recognition engine 455. For example, external electronic device 105 can generate a first recognition result for the wake word identified based on wake word recognition engine 455.
[0114] For example, external electronic device 105 can identify whether a speaker identification model 461 has been generated. For example, external electronic device 105 can identify whether a speaker identification model 461 has been generated through speaker registration module 459. For example, a speaker identification model 461 can be generated in association with account information. For example, a speaker identification model 461 can be generated for a specific user (or speaker).
[0115] For example, if the external electronic device 105 recognizes that a speaker identification model 461 has been generated for the speaker emitting the voice signal, the external electronic device 105 can use the speaker identification model 461 to identify the speaker emitting the voice signal based on the speaker identification engine 457. For example, the speaker may include a speaker registered based on the speaker registration module 459 included in the external electronic device 105. For example, a registered speaker may include at least one user. For example, a registered speaker may be registered in a state associated with account information. For example, account information may be provided to the speaker registration module 459 from an account manager (not shown). For example, account information may include the user's account identifier (or ID). For example, the external electronic device 105 can identify whether the speaker emitting the voice signal is a registered speaker. For example, the external electronic device 105 can generate (or obtain) a second identification result for the speaker through the speaker identification engine 457. The second identification result can be applied to essentially the same content as in Table 2 above.
[0116] Unlike the above description, for example, external electronic device 105 can train speaker identification model 461 even if it is found that no speaker identification model 461 has been generated for the speaker emitting the speech signal. For example, external electronic device 105 can train speaker identification model 461 based on the identification results obtained through MDW manager 463. For example, external electronic device 105 can generate speaker identification model 461 for the speaker emitting the speech signal by performing training of speaker identification model 461. The identification results obtained through MDW manager 463 may include speaker information (e.g., a second identification result) identified by external electronic device 103. External electronic device 105 can identify the speaker emitting the speech signal based on the speaker information. External electronic device 105 can perform speaker identification model 461 training for the identified speaker based on the speech signal obtained by external electronic device 105. In other words, external electronic device 105 can identify the speaker emitting the speech signal using the second discrimination result discerned by external electronic device 103, and can train a speaker discrimination model 461 for the speaker based on the speech signal directly obtained by external electronic device 105. For example, training can be performed based on a specified number of speech signals or speech signals with a specified duration. In other words, if the speaker discrimination model 461 is trained with a smaller number of speech signals than the specified number of speech signals, or if the speaker discrimination model 461 is trained with speech signals shorter than the specified duration, external electronic device 105 can recognize that the speaker discrimination model 461 has not yet been generated.
[0117] For example, external electronic device 105 can share a first discrimination result (or a first discrimination result and a second discrimination result) through MDW manager 463. For example, external electronic device 105 can identify electronic device 101 and external electronic device 103 in MDW environment 400 through MDW manager 463. For example, external electronic device 105 can send (or provide) the first discrimination result to electronic device 101 and external electronic device 103. Alternatively, in the case of generating speaker discrimination model 461, external electronic device 105 can send (or provide) the first discrimination result and the second discrimination result to electronic device 101 and external electronic device 103 in MDW environment 400 through MDW manager 463. Furthermore, for example, external electronic device 105 can receive (or obtain) the discrimination result from each of electronic device 101 and external electronic device 103 through MDW manager 463. Alternatively, for example, external electronic device 105 can send (or provide) the first discrimination result to server 107 in MDW environment 400 through MDW manager 463. Alternatively, after the speaker identification model 461 has been generated, the external electronic device 105 can send (or provide) the first and second identification results to the server 107 in the MDW environment 400 via the MDW manager 463. Furthermore, for example, the external electronic device 105 can receive (or obtain) information about the identification results of electronic devices 101 and 103 from the server 107 in the MDW environment 400 via the MDW manager 463.
[0118] According to an embodiment, server 107 can generate information about the discrimination results based on discrimination results obtained from each of electronic device 101, external electronic device 103, and external electronic device 105. For example, server 107 can obtain a first discrimination result generated by electronic device 101. Furthermore, for example, server 107 can obtain a first discrimination result and a second discrimination result generated by external electronic device 103. Furthermore, for example, server 107 can obtain a first discrimination result generated by external electronic device 105. According to an embodiment, server 107 can identify a target indicated by a voice signal (i.e., an indication target) based on the obtained discrimination results. For example, server 107 can identify a device as an indication target by comparing the sound power levels of the discrimination results. For example, server 107 can provide information to the identified device (e.g., electronic device 101). For example, the information may include at least a portion of the discrimination results. However, embodiments of this disclosure are not limited thereto. According to an embodiment, server 107 may also send information including the identification result to all of the electronic devices 101, external electronic devices 103 and external electronic devices 105 connected to server 107.
[0119] According to an embodiment, server 107 can identify electronic devices included in MDW environment 400 based on the identification result, and can also share information about the identified devices with electronic devices in MDW environment 400.
[0120] According to an embodiment, when server 107 includes an MDW manager, server 107 can identify the device as the target of indication based on the wake word discrimination result received from each of the MDW managers 417, 441, and 463. Unlike the above description, when server 107 does not include an MDW manager, each of the MDW managers 417, 441, and 463 can identify the device as the target of indication based on both the directly obtained wake word discrimination result and the received wake word discrimination result. For example, by comparing the signal-to-noise ratio (SNR) included in the wake word discrimination result, the device with the highest SNR can be identified as the target of indication.
[0121] Figure 5 An example of the operation flow of a method for an electronic device, including a speaker recognition model, to recognize wake words in a speech signal is shown.
[0122] Figure 5 At least a part of the method can be derived from Figure 4 The method is executed by an external electronic device 103. For example, at least a portion of the method may be controlled by at least one processor of the external electronic device 103. In the following embodiments, each of the operations may be executed sequentially, but is not required to be executed sequentially. For example, the order of each of the operations may be changed, and at least two operations may be executed in parallel.
[0123] In operation 500, the external electronic device 103 can acquire voice signals. For example, Figure 5 External electronic device 103 can instruct a device therein that generates a speaker discrimination model for the speaker emitting the speech signal, such as Figure 4 For example, external electronic device 103 can obtain voice signals via microphone 431.
[0124] In operation 505, external electronic device 103 can identify whether the voice signal includes a wake-up word. For example, external electronic device 103 can extract feature values from the voice signal and identify whether the voice signal includes a wake-up word based on the extracted feature values. For example, the wake-up word can indicate a wake-up word commonly set for external electronic device 103 and devices in the MDW environment. For example, the MDW environment may include external electronic device 103 and electronic device 101. However, embodiments of this disclosure are not limited thereto. For example, the MDW environment may also include external electronic device 103, electronic device 101, and another external electronic device 105.
[0125] In operation 505, if the voice signal includes a wake-up word, the external electronic device 103 can execute operation 510. Unlike the above description, in operation 505, if the voice signal does not include a wake-up word, the external electronic device 103 can again execute operation 500.
[0126] In operation 510, the external electronic device 103 can identify whether the speaker emitting the voice signal is a registered speaker. For example, the external electronic device 103 can identify whether the speaker emitting the voice signal is a registered speaker for the external electronic device 103. For example, the external electronic device 103 can identify whether the speaker emitting the voice signal is a registered speaker by comparing the feature values of the voice signal with information about registered speakers.
[0127] In operation 510, if the speaker emitting the voice signal is a registered speaker, the external electronic device 103 can execute operation 515. Unlike the above description, in operation 510, if the speaker emitting the voice signal is not a registered speaker, the external electronic device 103 can execute operation 500 again.
[0128] In operation 515, external electronic device 103 can identify whether it is an MDW environment. For example, external electronic device 103 can identify whether it is an MDW environment by identifying one or more connectable devices based on at least one communication circuit included in external electronic device 103. For example, one or more connectable devices may indicate a device in which a wake word is set. Alternatively, for example, external electronic device 103 can identify whether it is an MDW environment based on a server (e.g., for the MDW environment) used for the MDW environment. Figure 4 The server (107) obtains information to identify whether it is an MDW environment. For example, the information may include information about one or more devices managed by the server. For example, a non-MDW environment may indicate that only external electronic device 103 exists as a device capable of receiving voice signals, or that only external electronic device 103 exists as a device to which a wake word is set. The situation where the external electronic device 103 is a device capable of receiving voice signals may include a state where no other devices are nearby or a state where a signal may not be available (e.g., a power outage). Referring to the above description, in the case where multiple devices include external electronic device 103 with a wake word set, the MDW environment may include multiple devices with a wake word set and a server for the multiple devices. For ease of description, in the following examples, an MDW environment including a server is described as exemplary, but embodiments of this disclosure should not be construed as limited thereto.
[0129] exist Figure 5In the example, external electronic device 103 is described as performing operation 515 after operation 510, but embodiments of this disclosure are not limited thereto. According to embodiments, external electronic device 103 may repeatedly perform operation 515 at each specified time point. For example, external electronic device 103 may repeatedly identify whether it is an MDW environment at each of the aforementioned specified time points. For example, the specified time points may be set periodically or non-periodically.
[0130] In operation 515, if the external electronic device 103 identifies an MDW environment, the external electronic device 103 can execute operation 520. Contrary to the above description, in operation 515, if the external electronic device 103 identifies a non-MDW environment, the external electronic device 103 can execute operation 530.
[0131] In operation 520, the external electronic device 103 can share the recognition results. For example, the external electronic device 103 can generate a first recognition result and a second recognition result for the speech signal. For example, the first recognition result can indicate the recognition result for a wake word included in the speech signal. For example, the second recognition result can indicate the recognition result for the speaker who issued the speech signal.
[0132] According to an embodiment, the external electronic device 103 may provide (or send) identification results to a server connected to the external electronic device 103. Furthermore, the external electronic device 103 may obtain (or receive) information about the identification results of one or more devices from the server. Alternatively, the external electronic device 103 may provide (or send) identification results to one or more identified devices. Furthermore, the external electronic device 103 may obtain (or receive) the identification results of one or more devices.
[0133] In operation 525, the external electronic device 103 can identify whether it is the target of the wake word instruction. For example, the external electronic device 103 can identify whether the target of the wake word instruction is the external electronic device 103. In other words, the external electronic device 103 can identify whether the device responding to the wake word (or running a function) is the external electronic device 103.
[0134] According to an embodiment, the external electronic device 103 can identify whether the target of the wake word is the external electronic device 103 based on the recognition result generated by the external electronic device 103 and the recognition results of one or more devices shared through operation 520. For example, the external electronic device 103 can identify the target based on the sound power level. For example, the device with the highest sound power level among the sound power levels included in the recognition results generated by the external electronic device 103 and the recognition results of one or more devices shared through operation 520 can be identified as the target.
[0135] In operation 525, external electronic device 103 may perform operation 530 based on the identification that the indicated target is external electronic device 103. Contrary to the above description, external electronic device 103 may perform operation 500 based on the identification that the indicated target is not external electronic device 103.
[0136] In operation 530, the external electronic device 103 may operate in a mode for receiving voice signals. For example, the external electronic device 103 may operate in a mode for waiting to receive voice signals. For example, the mode may include a microphone included in the external electronic device 103 (e.g., Figure 4 The microphone 431 is in an activated state. The voice signal received while the mode is running may include command words. For example, based on the processing of the voice signal including the wake-up word identified in operation 505 by the external electronic device 103, it is expected that the voice signal including the command word will be received by the external electronic device 103.
[0137] Although not in Figure 5 As shown, however, the external electronic device 103 can also provide a speaker identification service based on recognizing a voice signal emitted by a registered speaker and including a wake word. For example, the external electronic device 103 can operate a specified function. For example, the specified function can be included in the speaker identification service. For example, as a specified function, the external electronic device 103 may include displaying a specified user interface (UI) (e.g., Figure 3 The UI 330 may be used to run an application that is permitted for a specific speaker (e.g., in cases where the voice signal includes a wake word and a command word for running the application). For example, applications permitted for a specific speaker may include software applications that require speaker authentication. However, embodiments of this disclosure are not limited thereto.
[0138] Figure 6 An example of the operation flow of a method for an electronic device to identify wake words of speech signals without a speaker identification model is shown.
[0139] Figure 6 At least a part of the method can be derived from Figure 4 The method is executed by electronic device 101. For example, at least a portion of the method may be controlled by at least one processor of electronic device 101. However, embodiments of this disclosure are not limited thereto. At least a portion of the method may be controlled by... Figure 4The method is executed by an external electronic device 105. For example, at least a portion of the method may be controlled by at least one processor of the external electronic device 105. In the following embodiments, each of the operations may be executed sequentially, but not necessarily sequentially. For example, the order of each of the operations may be changed, and at least two operations may be executed in parallel. The method executed by electronic device 101 is described below, but substantially the same method may also be applied to external electronic device 105.
[0140] In operation 600, electronic device 101 can acquire voice signals. For example, Figure 6 Electronic device 101 may indicate electronic devices that do not include a speaker identification model for identifying the speaker emitting the speech signal, such as... Figure 4 As described herein. Alternatively, for example, Figure 6 External electronic device 105 can instruct electronic device 105, which includes a speaker identification model for identifying the speaker emitting the speech signal, such as... Figure 4 The speech signal is not registered as described in the text. For example, electronic device 101 may acquire the speech signal via microphone 411.
[0141] In operation 605, electronic device 101 can identify whether the voice signal includes a wake-up word. For example, electronic device 101 can extract feature values from the voice signal and identify whether the voice signal includes a wake-up word based on the extracted feature values. For example, the wake-up word can indicate a wake-up word commonly set for electronic device 101 and devices in the MDW environment. For example, the MDW environment may include electronic device 101 and one or more external electronic devices (e.g., Figure 4 External electronic devices 103 or Figure 4 External electronic device 105). However, embodiments of this disclosure are not limited thereto.
[0142] In operation 605, if the voice signal includes a wake-up word, electronic device 101 can execute operation 610. Unlike the above description, in operation 605, if the voice signal does not include a wake-up word, electronic device 101 can again execute operation 600.
[0143] In operation 610, electronic device 101 can identify whether it is an MDW environment. For example, electronic device 101 can identify whether it is an MDW environment by recognizing one or more connectable devices based on at least one communication circuit included in electronic device 101. For example, one or more connectable devices may indicate a device in which a wake word is set. Alternatively, for example, electronic device 101 can identify whether it is an MDW environment based on a server (e.g., for the MDW environment) from a server used for the MDW environment. Figure 4The information obtained by server 107 identifies whether it is an MDW environment. For example, the information may include information about one or more devices managed by the server. For example, a non-MDW environment may indicate that only electronic device 101 exists as a device capable of receiving voice signals, or that only electronic device 101 exists as a device to which a wake word is set. In the case that electronic device 101 is a device capable of receiving voice signals, it may include a state that no other devices are nearby or a state that may not be able to receive signals (e.g., a power outage).
[0144] exist Figure 6 In the example, electronic device 101 is described as performing operation 610 after operation 605, but embodiments of this disclosure are not limited thereto. According to embodiments, electronic device 101 may repeatedly perform operation 610 at each specified time point. For example, electronic device 101 may repeatedly identify whether it is an MDW environment at each of the aforementioned specified time points. For example, the specified time points may be set periodically or non-periodically.
[0145] In operation 610, if electronic device 101 identifies an MDW environment, electronic device 101 can execute operation 615. Unlike the above description, in operation 610, if electronic device 101 identifies a non-MDW environment, electronic device 101 can execute operation 635.
[0146] In operation 615, electronic device 101 can share the discrimination results. For example, electronic device 101 can generate a first discrimination result for the speech signal. For example, the first discrimination result can indicate the discrimination result for a wake word included in the speech signal.
[0147] According to an embodiment, electronic device 101 may provide (or send) identification results to a server connected to electronic device 101. Furthermore, electronic device 101 may obtain (or receive) information from the server regarding identification results for one or more devices. The identification results for one or more devices may include a first identification result for a wake-up word and / or a second identification result for a speaker. For example, the second identification result may indicate the identification result for a speaker emitting a voice signal. Alternatively, electronic device 101 may provide (or send) identification results to one or more identified devices. Furthermore, electronic device 101 may obtain (or receive) identification results from one or more devices.
[0148] In operation 620, electronic device 101 can identify whether it is the target of the wake word instruction. For example, electronic device 101 can identify whether the target of the wake word instruction is electronic device 101. In other words, electronic device 101 can identify whether the device responding to the wake word (or running a function) is electronic device 101.
[0149] According to an embodiment, electronic device 101 can identify whether the target of the wake word is electronic device 101 based on the identification result generated by electronic device 101 and the identification results of one or more devices shared through operation 615. For example, electronic device 101 can identify the target based on the sound power level. For example, the device with the highest sound power level among the sound power levels included in the identification results generated by electronic device 101 and the identification results of one or more devices shared through operation 615 can be identified as the target.
[0150] In operation 620, electronic device 101 may perform operation 625 based on the identification indicating that the target is electronic device 101. Contrary to the above description, electronic device 101 may perform operation 600 based on the identification indicating that the target is not electronic device 101.
[0151] In operation 625, electronic device 101 can identify whether a speaker identification result has been received. For example, electronic device 101 can identify whether the identification results received during sharing from one or more devices include a second speaker identification result. For example, if a second identification result is included, electronic device 101 can identify that a speaker identification result has been received. Unlike the above description, if a second identification result is not included, electronic device 101 can identify that a speaker identification result has not yet been received.
[0152] In operation 625, if the identification result of one or more devices includes the second identification result, electronic device 101 may perform operation 630. Unlike the above description, in operation 625, if the identification result of one or more devices does not include the second identification result, electronic device 101 may perform operation 635.
[0153] In operation 630, electronic device 101 can run a specified function. For example, when a second discrimination result is included, electronic device 101 can identify the speaker emitting the speech signal based on the second discrimination result. For example, electronic device 101 can run a specified function in electronic device 101 for the speaker emitting the speech signal. For example, the specified function can be included in a speaker discrimination service. For example, as a specified function, electronic device 101 can include displaying a specified user interface (UI) (e.g., Figure 3 The UI 330 may be used to run an application permitted for a specific speaker (e.g., where the voice signal includes a wake word and a command word for running the application). For example, an application permitted for a specific speaker may include a software application that requires speaker authentication. However, embodiments of this disclosure are not limited thereto.
[0154] In operation 635, electronic device 101 may operate a mode for receiving voice signals. For example, electronic device 101 may operate a mode for waiting to receive voice signals. For example, the mode may include a microphone included in electronic device 101 (e.g., Figure 4 The microphone 411 is in an activated state. Voice signals received while the mode is running may include command words. For example, based on the processing performed by electronic device 101 on a voice signal including the wake-up word identified in operation 605, it is expected that a voice signal including command words will be received by electronic device 101.
[0155] In operation 640, electronic device 101 can operate functions based on command words. For example, electronic device 101 can receive voice signals in an operating mode. Electronic device 101 can identify command words included in the voice signals. Electronic device 101 can identify functions based on command words. For example, it can be based on what will be described later. Figure 9 At least a portion of the natural language processing (e.g., the Automatic Speech Recognition (ASR) module 921, the Natural Language Understanding module 923, the Planner module 925, the Natural Language Generator module 927, and the Text-to-Speech module 929) of the natural language platform 920 performs the recognition function. For example, the function of the command word can be associated with a speaker recognition service. For example, in the case of the speech signal "Tell me today's schedule," the function may include recognizing and displaying the user's schedule. Therefore, the electronic device 101 can display the user's schedule via a display.
[0156] According to an embodiment, when a function is executed according to a command word, electronic device 101 can receive personal information about the user (or speaker) from an external electronic device based on identification information about the user (e.g., a second identification result). For example, personal information may include the user's schedule information, phone number, and multimedia content. For example, the external electronic device may also be a device not included in the MDW environment 400.
[0157] Figure 7 An example of the operational flow of a method for generating a speaker recognition model using an electronic device is shown.
[0158] Figure 7 At least a part of the method can be derived from Figure 4 The method is executed by an external electronic device 105. For example, at least a portion of the method may be controlled by at least one processor of the external electronic device 105. However, embodiments of this disclosure are not limited thereto. For example, in the following embodiments, each of the operations may be executed sequentially, but not necessarily sequentially. For example, the order of each of the operations may be changed, and at least two operations may be executed in parallel.
[0159] In operation 700, the external electronic device 105 can acquire voice signals. For example, Figure 7 External electronic device 105 can instruct electronic device 105, which includes a speaker identification model for identifying the speaker emitting the speech signal, such as... Figure 4 However, the speaker emitting the voice signal is not registered. For example, an external electronic device 105 can obtain the voice signal via microphone 451.
[0160] Despite Figure 7 Not shown in the diagram, but according to an embodiment, the external electronic device 105 can identify whether the voice signal obtained in operation 700 includes a wake-up word, identify whether it is an MDW environment, and share the identification results, such as... Figure 6 As shown. Furthermore, according to an embodiment, the external electronic device 105 can identify whether the discrimination results received from one or more devices in the MDW environment include a discrimination result for the speaker emitting the voice signal (e.g., a second discrimination result). Specific details related thereto can be referenced substantially similarly. Figure 6 The operation.
[0161] In operation 705, the external electronic device 105 can identify whether a speaker identification model has been generated. For example, the external electronic device 105 can identify whether a speaker identification model has been generated for the speaker emitting the speech signal.
[0162] In operation 705, if a speaker identification model is generated, the external electronic device 105 can perform operation 710. Unlike the above description, in operation 705, if a speaker identification model is not generated, the external electronic device 105 can perform operation 715.
[0163] In operation 710, the external electronic device 105 can perform speaker identification and run specified functions. For example, in the case of generating a speaker identification model, the external electronic device 105 can identify whether the speaker emitting the speech signal is a speaker registered with the external electronic device 105. Specific details related to this can be referenced substantially the same way. Figure 5 Operation 510. External electronic device 105 can run a specified function based on the identification that the speaker emitting the voice signal is a registered speaker. For example, the specified function may be included in a speaker identification service. For example, as a specified function, external electronic device 105 may include displaying a specified user interface (UI) (e.g., Figure 3 The UI 330 may be used to run an application permitted for a specific speaker (e.g., where the voice signal includes a wake word and a command word for running the application). For example, an application permitted for a specific speaker may include a software application that requires speaker authentication. However, embodiments of this disclosure are not limited thereto.
[0164] In operation 715, the external electronic device 105 can perform training of a speaker identification model. For example, without generating a speaker identification model, the external electronic device 105 can perform training to generate a speaker identification model. For example, the external electronic device 105 can identify the speaker emitting the speech signal based on the identification results received from one or more devices. For example, the external electronic device 105 can perform training on the identified speaker based on the speech signal obtained by the external electronic device 105.
[0165] According to embodiments, a speaker identification model can be trained based on a speaker identification algorithm. For example, the speaker identification algorithm may include algorithms based on Gaussian mixture models (GMM) / hidden Markov models (HMM), algorithms based on deep neural networks (DNN), or algorithms based on identifier vectors (i-vectors).
[0166] According to an embodiment, in the absence of a generated speaker identification model, the external electronic device 105 may display a UI asking whether to perform training for generating the speaker identification model. The external electronic device 105 may perform training based on input to the UI. According to an embodiment, the external electronic device 105 may store speech signals emitted by a speaker identified based on discrimination results received from one or more devices. The speech signals may include speech signals directly obtained by the external electronic device 105. For example, the external electronic device 105 may perform training if the speech signals emitted by the speaker are a specified number or more, or a specified duration or longer. Upon completion of training based on a specified number or more, or a specified duration or longer speech signals, the external electronic device 105 may display a UI notifying of training completion.
[0167] According to embodiments, the speech signal used in training can be a signal that meets specified criteria. For example, the speech signal may include a signal with an SNR greater than or equal to a threshold SNR. However, embodiments of this disclosure are not limited thereto. SNR can be an example indicating the signal quality of the speech signal. Alternatively, for example, the speech signal may include a signal whose signal distance is within a reference distance stored some time ago, or a signal whose signal characteristics are less than or equal to a reference value. For example, the signal distance may include spectral distance and feature domain distance. Signal characteristics may include similarity and coherence.
[0168] According to embodiments, taking into account distortions that may be generated by the preprocessing process, based on the signal flowing in via the microphone (e.g., microphone 451) of the external electronic device 105 through which sound generated by the external electronic device 105 itself is played, the external electronic device 105 may perform processing (e.g., acoustic echo cancellation), perform noise suppression, or define operations such as beamforming, wherein beamforming amplifies the signal in a specific direction. The voice signal may include a voice signal recognized by each of the operations or a combination of operations. According to embodiments, each of the voice signals may include a wake word, a wake word and a command word, or a command word.
[0169] According to an embodiment, the external electronic device 105 can be based on a speaker recognition engine (e.g., Figure 4 The speaker identification engine 457) and the speaker registration module (e.g., Figure 4 The speaker registration module 459) performs training. Furthermore, the external electronic device 105 may include a wake-word recognition engine (e.g., different from the speaker recognition engine) that is different from the speaker recognition engine. Figure 4 The wake word recognition engine 455). According to an embodiment, the speaker recognition engine and the wake word recognition engine can be implemented as recognition engines. For example, the recognition engine can be called a speaker-related utterance recognition engine. In this case, the speech signal used in training can be a speech signal including the wake word. According to an embodiment, the external electronic device 105 can use a speaker recognition model trained using the speaker-related utterance recognition engine as an acoustic model for recognizing the wake word. In this case, the acoustic model for recognizing the wake word can be used based on the user's input to a specified UI displayed through the external electronic device 105.
[0170] According to an embodiment, the user of the external electronic device 105 may not be able to perceive the training. For example, the training may be performed automatically in the background and repeated a specified number of times or for a specified duration.
[0171] Figure 8 An example of the operation flow of an electronic device using an external electronic device to identify the speaker is shown.
[0172] Figure 8 At least a part of the method can be derived from Figure 4 The method is executed by electronic device 101. For example, at least a portion of the method may be controlled by at least one processor of electronic device 101. However, embodiments of this disclosure are not limited thereto. At least a portion of the method may be controlled by... Figure 4The method is executed by an external electronic device 105. For example, at least a portion of the method may be controlled by at least one processor of the external electronic device 105. In the following embodiments, each of the operations may be executed sequentially, but not necessarily sequentially. For example, the order of each of the operations may be changed, and at least two operations may be executed in parallel. The method executed by electronic device 101 is described below, but substantially the same method may also be applied to external electronic device 105.
[0173] In operation 800, electronic device 101 can acquire voice signals. For example, Figure 8 Electronic device 101 may indicate electronic devices that do not include a speaker identification model for identifying the speaker emitting the speech signal, such as... Figure 4 As described herein. Alternatively, for example, Figure 8 External electronic device 105 can instruct electronic device 105, which includes a speaker identification model for identifying the speaker emitting the speech signal, such as... Figure 4 The speech signal is not registered as described in the text. For example, electronic device 101 may acquire the speech signal via microphone 411.
[0174] In operation 810, electronic device 101 can generate a first recognition result for a wake-up word. According to an embodiment, electronic device 101 can identify whether a voice signal includes a wake-up word. For example, electronic device 101 can extract feature values from the voice signal and identify whether the voice signal includes a wake-up word based on the extracted feature values. For example, the wake-up word can indicate a wake-up word commonly set for electronic device 101 and devices in the MDW environment. For example, the MDW environment may include electronic device 101 and one or more external electronic devices (e.g., Figure 4 External electronic devices 103 or Figure 4 External electronic device 105). However, embodiments of this disclosure are not limited thereto.
[0175] According to an embodiment, electronic device 101 can generate a first discrimination result for a wake word included in a speech signal. For example, the first discrimination result may include a discrimination score, speech power, signal-to-noise ratio (SNR), and time. However, embodiments of this disclosure are not limited thereto. For example, the first discrimination result may also include only speech power. Alternatively, for example, the first discrimination result may include information about the priority between electronic devices in addition to the information illustrated in the table. For example, the priority may indicate the priority in response to a wake word. For example, the priority may be set by the user of the electronic device.
[0176] For example, a discrimination score can indicate discrimination information about a wake word. For example, electronic device 101 can obtain a discrimination score that indicates the degree of discrimination of a wake word relative to a speech signal as a probability value. For example, electronic device 101 can obtain a discrimination score. The discrimination score can be referred to as the reliability (or discrimination reliability) of the wake word.
[0177] For example, speech power can indicate the level of sound power in a speech signal. For example, electronic device 101 can identify the level of the speech signal. Speech power or sound power level can be used to identify (or distinguish) the electronic device located closest to the user emitting the speech signal. For example, speech power can be referred to as speech power, sound power, or sound power level.
[0178] For example, SNR can indicate the ratio between signal (e.g., speech) and noise in a speech signal. For example, electronic device 101 can identify the noise level of the environment (or space) from which the speech signal is emitted based on SNR. For example, time can indicate the timing of acquiring the speech signal.
[0179] For example, electronic device 101 can generate device information associated with the first identification result. The device information may include information about electronic device 101 that has obtained the first identification result. For example, the device information may include a device name, device type, device identifier (device ID), wake word, and the device's Internet Protocol (IP) address. For example, the device name may include the model name of electronic device 101. For example, the device type may include the type of electronic device 101 (e.g., TV). For example, the device identifier may include a unique identifier for electronic device 101. For example, the wake word may include the name of the wake word set for electronic device 101. The contents of the table above are merely examples for ease of description, and the embodiments of this disclosure are not limited thereto. For example, the device information may be sent or received (i.e., shared) together with the first identification result.
[0180] In operation 820, electronic device 101 can send the first identification result to a server connected to one or more external electronic devices that have been set with a wake word.
[0181] For example, electronic device 101 can identify whether it is an MDW environment. For example, electronic device 101 can identify one or more connectable external electronic devices based on at least one communication circuit of electronic device 101. For example, electronic device 101 can identify whether it is an MDW environment based on identifying one or more external electronic devices. For example, one or more connectable external electronic devices can indicate the device to which the wake word is set. Alternatively, for example, electronic device 101 can identify whether it is an MDW environment based on a server (e.g., for the MDW environment) from a server used for the MDW environment. Figure 4The information obtained by server 107 identifies whether it is an MDW environment. For example, the information may include information about one or more external electronic devices managed by the server. According to an embodiment, electronic device 101 may provide (or send) a first identification result to a server connected to electronic device 101.
[0182] For example, a non-MDW environment could indicate a situation where only electronic device 101 exists as a device capable of receiving voice signals, or a situation where only electronic device 101 exists as a device to which a wake word is set. When electronic device 101 is the device capable of receiving voice signals, it could include a state where no other devices are nearby or a state where signals may not be available (e.g., a power outage).
[0183] In operation 830, electronic device 101 may receive a second identification result for the user who issued the wake word from the server. The user may be referred to as the speaker who issued the wake word. For example, electronic device 101 may obtain (or receive) information from the server regarding the second identification result of one or more external electronic devices. However, embodiments of this disclosure are not limited thereto. For example, electronic device 101 may include a first identification result for the wake word and / or a second identification result for the speaker from the server.
[0184] In operation 840, electronic device 101 may run a specified function for a user based on the second identification result. For example, electronic device 101 may identify the user based on the second identification result. For example, electronic device 101 may run a specified function for the user. For example, the specified function may be included in a speaker identification service. For example, as a specified function, external electronic device 105 may include displaying a specified user interface (UI) (e.g., Figure 3 The UI 330 may be used to run an application permitted for a specific speaker (e.g., where the voice signal includes a wake word and a command word for executing the application). For example, an application permitted for a specific speaker may include a software application that requires speaker authentication. However, embodiments of this disclosure are not limited thereto.
[0185] Although not in Figure 8As shown, however, according to an embodiment, electronic device 101 can identify whether it is the target of the wake word indication. For example, electronic device 101 can identify whether the target of the wake word indication is electronic device 101. In other words, electronic device 101 can identify whether the device responding to the wake word (or running a function) is electronic device 101. For example, electronic device 101 can identify whether the target of the wake word indication is electronic device 101 based on a discrimination result generated by electronic device 101 and a first discrimination result received from one or more external electronic devices. For example, electronic device 101 can identify the indication target based on the sound power level.
[0186] According to an embodiment, electronic device 101 can operate a mode for receiving voice signals. For example, electronic device 101 can operate a mode for waiting to receive voice signals after performing a specified function. For example, the mode may include a microphone included in electronic device 101 (e.g., Figure 4 The microphone 411 is in an activated state. When the mode is running, the received voice signal may include command words.
[0187] Figure 9 This is a block diagram illustrating an integrated intelligent system according to various embodiments.
[0188] refer to Figure 9 The integrated intelligent system according to the embodiments may include an electronic device 101 and an intelligent server 900 (e.g., Figure 1 Server 108) and server 990 (for example, Figure 1 Server 108).
[0189] The electronic device 101 according to the embodiment may be an internet-connected terminal device (or electronic device), and may be, for example, a mobile phone, smartphone, personal digital assistant (PDA), laptop computer, TV, white goods, wearable device, HMD, or smart speaker.
[0190] According to the illustrated embodiment, electronic device 101 may include interface 177, input module 150, audio output module 155, display module 160, memory 130, or processor 120. The components listed above may be operatively or electrically connected to each other.
[0191] According to an embodiment, interface 177 can be configured to connect to an external device to send and receive data. According to an embodiment, input module 150 can receive sound (e.g., user voice) and convert it into an electrical signal. According to an embodiment, audio output module 155 can output the electrical signal as sound (e.g., speech).
[0192] The display module 160 according to an embodiment can be configured to display images or videos. The display module 160 according to an embodiment can also display the graphical user interface (GUI) of a running app (or application). The display module 160 according to an embodiment can receive touch input via a touch sensor. For example, the display module 160 can receive text input via a touch sensor on a keyboard area displayed on the screen.
[0193] The memory 130 according to the embodiment may store a client module 151, a software development kit (SDK) 153, and multiple applications 146. The client module 151 and the SDK 153 may be configured to perform general functions within a framework (or solution program). Furthermore, the client module 151 or the SDK 153 may be configured to handle user input (e.g., voice input, text input, touch input).
[0194] According to an embodiment, a plurality of applications 146 stored in memory 130 may be programs for performing specified functions. According to an embodiment, the plurality of applications 146 may include a first application 146-1 and a second application 146-2. According to an embodiment, each of the plurality of applications 146 may include a plurality of operations for performing the specified functions. For example, the applications may include an alarm clock application, a messaging application, and / or a calendar application. According to an embodiment, the plurality of applications 146 may be run by processor 120 to sequentially perform at least a portion of the plurality of operations.
[0195] According to the embodiment, the processor 120 can control the overall operation of the electronic device 101. For example, the processor 120 can be electrically connected to the interface 177, the input module 150, the audio output module 155, and the display module 160 to perform specified operations.
[0196] The processor 120 according to an embodiment can also run programs stored in memory 130 to perform specified functions. For example, the processor 120 can perform the following operations for processing user input by running at least one of client module 151 or SDK 153. For example, the processor 120 can control the operation of multiple applications 146 through SDK 153. The following operations, described as operations of client module 151 or SDK 153, can be operations run by the processor 120.
[0197] According to an embodiment, client module 151 can receive user input. For example, client module 151 can receive voice signals corresponding to user voice detected by input module 150. Alternatively, client module 151 can receive touch input detected by display module 160. Alternatively, client module 151 can receive text input detected by keyboard or on-screen keyboard. Furthermore, various types of user input detected by input modules included in or connected to electronic device 101 can be received. Client module 151 can send the received user input to intelligent server 900. Client module 151 can also send status information of electronic device 101 along with the received user input to intelligent server 900. For example, the status information may be application running status information.
[0198] According to the embodiment, the client module 151 can receive a result corresponding to the received user input. For example, when the intelligent server 900 can calculate a result corresponding to the received user input, the client module 151 can receive a result corresponding to the received voice input. The client module 151 can display the received result on the display module 160. Furthermore, the client module 151 can output the received result as audio through the sound output module 155.
[0199] According to an embodiment, the client module 151 can receive a plan corresponding to the received user input. The client module 151 can display the results of multiple operations performed according to the plan on the display module 160. For example, the client module 151 can sequentially display the results of multiple operations on the display and output audio through the sound output module 155. In another example, the electronic device 101 can display only a portion of the results of performing multiple operations (e.g., the result of the last operation) on the display module 160 and output it as audio through the sound output module 155.
[0200] According to an embodiment, client module 151 can receive a request from intelligent server 900 for obtaining information needed to calculate the result corresponding to user input. According to an embodiment, client module 151 can send the necessary information to intelligent server 900 in response to the request.
[0201] According to an embodiment, client module 151 can send result information obtained by running multiple operations according to a plan to intelligent server 900. Intelligent server 900 can use the result information to identify that the received user input has been processed correctly.
[0202] According to an embodiment, client module 151 may include a voice recognition module. According to an embodiment, client module 151 may use the voice recognition module to recognize voice input that performs limited functions. For example, client module 151 may execute a smart application for processing voice input to perform organic operations upon a specified input (e.g., wake-up!).
[0203] According to an embodiment, the intelligent server 900 can receive information related to user voice input from the electronic device 101 via a communication network. According to an embodiment, the intelligent server 900 can convert the data related to the received voice input into text data. According to an embodiment, the intelligent server 900 can generate a plan for performing tasks corresponding to the user's voice input based on the text data.
[0204] According to an embodiment, the plan can be generated by an artificial intelligence (AI) system. The AI system can be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)). Alternatively, it can be a combination of the above or another AI system. According to an embodiment, a plan can be selected from a set of predefined plans, or a plan can be generated in real time in response to a user request. For example, the AI system can select at least one plan from a plurality of predefined plans.
[0205] According to an embodiment, the intelligent server 900 can send the results of the generated plan to the electronic device 101, or send the generated plan to the electronic device 101. According to an embodiment, the electronic device 101 can display the results of the plan on the display module 160. According to an embodiment, the electronic device 101 can display the results of running the planned operation on the display module 160.
[0206] The intelligent server 900 according to the embodiment may include a front-end 910, a natural language platform 920, a capsule database 930, a runtime engine 940, a user interface 950, a management platform 960, a big data platform 970, or an analysis platform 980.
[0207] According to an embodiment, the front end 910 can receive user input received from the electronic device 101. The front end 910 can send a response corresponding to the user input.
[0208] According to an embodiment, the natural language platform 920 may include an automatic speech recognition module (ASR module) 921, a natural language understanding module (NLU module) 923, a planner module 925, a natural language generator module (NLG module) 927, or a text-to-speech module (TTS module) 929.
[0209] According to an embodiment, the ASR module 921 can convert voice input received from the electronic device 101 into text data. According to an embodiment, the NLU module 923 can identify the user's intent using the text data of the voice input. For example, the NLU module 923 can identify the user's intent by performing syntactic or semantic analysis on the user input in the form of text data. According to an embodiment, the NLU module 923 can identify the meaning of words extracted from the user input using linguistic features (e.g., grammatical elements) of morphemes or phrases, and can determine the user's intent by matching the meaning of the identified words with the intent. The NLU module 923 can obtain intent information corresponding to the user's speech. The intent information can be information indicating the user's intent determined by interpreting the text data. The intent information can include information indicating that the user intends to use the operation or function of the device.
[0210] According to an embodiment, the planner module 925 can generate a plan using the intent and parameters determined in the NLU module 923. According to an embodiment, the planner module 925 can determine multiple domains required to perform a task based on the determined intent. The planner module 925 can determine multiple actions included in each of the multiple domains determined based on the intent. According to an embodiment, the planner module 925 can determine the parameters necessary to perform the multiple determined actions, or the result values output by performing the multiple actions. Parameters and result values can be defined as concepts of a specified format (or class). Therefore, a plan can include multiple actions and multiple concepts determined by the user's intent. The planner module 925 can determine the relationships between the multiple actions and multiple concepts in a step-by-step (or hierarchical) manner. For example, the planner module 925 can determine the execution order of the multiple actions determined based on the user's intent based on multiple concepts. In other words, the planner module 925 can determine the execution order of the multiple actions based on the parameters required to perform the multiple actions and the results output by performing the multiple actions. Therefore, the planner module 925 can generate a plan that includes association information (e.g., ontology) between the multiple actions and multiple concepts. The planner module 925 can use information stored in the capsule database 930 to generate plans, which stores a set of relationships between concepts and actions.
[0211] The NLG module 927 according to an embodiment can convert specified information into text form. The information converted into text form can be in the form of natural language speech. The TTS module 929 according to an embodiment can convert information in text form into information in speech form.
[0212] According to an embodiment, some or all of the functions of the natural language platform 920 may also be implemented in the electronic device 101.
[0213] The capsule database 930 can store information about the relationships between actions and multiple concepts corresponding to multiple domains. According to an embodiment, a capsule may include multiple action objects (or action information) and concept objects (or concept information) included in a plan. According to an embodiment, the capsule database 930 may store multiple capsules in the form of a Concept Action Network (CAN). According to an embodiment, multiple capsules may be stored in a function registry included in the capsule database 930.
[0214] The capsule database 930 may include a strategy registry storing strategy information required to determine a plan corresponding to a voice input. The strategy information may include reference information for determining a plan when multiple plans exist corresponding to the user input. According to an embodiment, the capsule database 930 may include a follow-up registry storing information on follow-up actions for suggesting subsequent actions to the user in specified situations. For example, a follow-up action may include a follow-up voice. According to an embodiment, the capsule database 930 may include a layout registry storing layout information of information output through the electronic device 101. According to an embodiment, the capsule database 930 may include a vocabulary registry storing vocabulary information included in the capsule information. According to an embodiment, the capsule database 930 may include a dialogue registry storing information about dialogues (or interactions) with the user. The capsule database 930 may update objects stored via developer tools. For example, the developer tools may include a function editor for updating action objects or concept objects. The developer tools may include a vocabulary editor for updating vocabulary. The developer tools may include a strategy editor for generating and registering strategies for determining plans. The developer tools may include a dialogue editor for generating dialogues with the user. The developer tools may include a follow-up editor for activating follow-up goals and editing follow-up voice prompts. Subsequent goals can be determined based on currently set objectives, user preferences, or environmental conditions. In this embodiment, the capsule database 930 can also be implemented in the electronic device 101.
[0215] The execution engine 940 according to the embodiment can calculate results using the generated plan. The client user interface 950 can send the calculated results to the electronic device 101. Therefore, the electronic device 101 can receive the results and provide the received results to the user. The management platform 960 according to the embodiment can manage the information used in the intelligent server 900. The big data platform 970 according to the embodiment can collect user data. The analysis platform 980 according to the embodiment can manage the quality of service (QoS) of the intelligent server 900. For example, the analysis platform 980 can manage the components and processing speed (or efficiency) of the intelligent server 900.
[0216] According to an embodiment, the service server 990 may include CP service A 991, CP service B 992, and CP service C 993. According to an embodiment, the service server 990 may provide specified services (e.g., food ordering or hotel reservation) to the electronic device 101. According to an embodiment, the service server 990 may be a server operated by a third party. According to an embodiment, the service server 990 may provide the intelligent server 900 with information for generating plans corresponding to received user input. The provided information may be stored in the capsule database 930. Furthermore, the service server 990 may provide the intelligent server 900 with result information based on the plans.
[0217] exist Figure 9 In the integrated intelligent system, electronic device 101 can provide various intelligent services to the user in response to user input. For example, user input may include input via physical buttons, touch input, or voice input.
[0218] In this embodiment, electronic device 101 may provide voice recognition services through a smart application (or voice recognition application) stored therein. In this case, for example, electronic device 101 may identify user voice or speech input received via a microphone and provide the user with services corresponding to the identified voice input.
[0219] In this embodiment, electronic device 101 may perform a specified operation based on the received voice input, either alone or in conjunction with a smart server and / or a service server. For example, electronic device 101 may run an application corresponding to the received voice input and perform the specified operation through the running application.
[0220] In this embodiment, when electronic device 101 provides services together with intelligent server 900 and / or service server 990, electronic device 101 can detect user voice using input module 150 and generate a signal (or voice data) corresponding to the detected user voice. Electronic device 101 can send the voice data to intelligent server 900 using interface 177.
[0221] In response to voice input received from electronic device 101, the intelligent server 900 according to an embodiment can generate a plan for performing a task corresponding to the voice input, or the result of performing an operation according to the plan. For example, the plan may include multiple actions for performing the task corresponding to the user's voice input, and multiple concepts associated with the multiple actions. Concepts may define parameters input by performing the multiple actions, or result values output by performing the multiple actions. The plan may include association information between the multiple actions and the multiple concepts.
[0222] According to the embodiment, the electronic device 101 can receive a response using interface 177. The electronic device 101 can output voice signals generated internally to the outside using audio output module 155, or it can output images generated internally to the outside using display module 160.
[0223] Figure 10 This is a diagram illustrating the relationship information between concepts and actions according to various embodiments, stored in a database.
[0224] Intelligent servers (e.g., Figure 9 The capsule database of the intelligent server 900 (e.g., Figure 9 The capsule database (930) can store capsules in the form of a conceptual action network (CAN). The capsule database can store actions and the parameters required for actions to process tasks corresponding to the user's voice input in the form of a conceptual action network (CAN).
[0225] The capsule database can store multiple capsules (e.g., capsule A 1001, capsule B 1004) corresponding to each of multiple domains (e.g., applications). According to an embodiment, a capsule (e.g., capsule A 1001) may correspond to a domain (e.g., location (geography), application). Furthermore, a capsule may correspond to at least one service provider (e.g., CP 1 1002, CP 2 1003, CP 3 1006, or CP 4 1005) for performing functions associated with the domain. According to an embodiment, a capsule may include at least one action 1010 and at least one concept 1020 for performing a specified function.
[0226] Natural language platforms (e.g., Figure 9 The natural language platform (920) can generate plans for performing tasks corresponding to speech input received by capsules and stored in a capsule database. For example, the planner module of the natural language platform (e.g., Figure 9 The planner module 925 can generate a plan by using capsules stored in the capsule database. For example, plan 1007 can be generated using actions 1001-1 and 1001-3 of capsule A 1001 with concepts 1001-2 and 1001-4, and actions 1004-1 of capsule B 1004 with concept 1004-2.
[0227] Figure 11 This is a diagram illustrating a screen of an electronic device, according to various embodiments, processing voice input received via a smart application.
[0228] Electronic device 101 can run intelligent applications to access intelligent servers (e.g., Figure 9 The intelligent server 900 processes user input.
[0229] According to an embodiment, on screen 1110, when a specified voice input (e.g., wake-up!) is detected or input is received via a hardware key (e.g., a dedicated hardware key), electronic device 101 can run a smart application for processing voice input. For example, electronic device 101 can run the smart application while running a calendar application. According to an embodiment, electronic device 101 can also run the smart application on a display module (e.g., Figure 1 The electronic device 101 displays objects (e.g., icons) 1111 corresponding to the smart application on the display module 160. According to an embodiment, the electronic device 101 can receive voice input via a user's voice. For example, the electronic device 101 can receive voice input such as "Tell me my schedule for this week!". According to an embodiment, the electronic device 101 can display objects (e.g., icons) 1111 on the display module 160. Figure 1 The display module 160 displays the user interface (UI) 1113 (e.g., an input window) of the smart application, in which the text data of the received voice input is displayed.
[0230] According to an embodiment, in screen 1120, electronic device 101 can be displayed on display module (e.g., Figure 1 The display module 160 displays the result corresponding to the received voice input. For example, the electronic device 101 can receive a plan corresponding to the received user input and display "this week's schedule" on the display module 160 according to the plan.
[0231] The electronic device 101 or 105 described above may include communication circuitry. The electronic device 101 or 105 may include a memory, which includes one or more storage media storing instructions. The electronic device 101 or 105 may include a microphone 411 or 451. The electronic device 101 or 105 may include at least one processor containing processing circuitry. When operated individually or jointly by at least one processor, instructions may cause the electronic device 101 or 105 to acquire a voice signal via microphone 411 or 451. When operated individually or jointly by at least one processor, instructions may cause the electronic device 101 or 105 to generate a first recognition result for the wake-up word from the voice signal including the wake-up word. When operated individually or jointly by at least one processor, instructions may cause the electronic device 101 or 105 to send the first recognition result to a server 107 connected to one or more external electronic devices capable of recognizing the wake-up word. When operated individually or jointly by at least one processor, instructions may cause the electronic device 101 or 105 to receive a second recognition result for the user who issued the wake-up word from the server 107. When executed individually or jointly by at least one processor, the instructions can cause electronic device 101 or 105 to perform a user-specified function based on a second identification result. The second identification result may include user identification information obtained from one or more external electronic devices, specifically from an external electronic device 103 registered by the user.
[0232] According to an embodiment, when operated individually or jointly by at least one processor, the instructions can cause electronic device 101 or 105 to extract multiple feature values from a speech signal. When operated individually or jointly by at least one processor, the instructions can cause electronic device 101 or 105 to identify a wake word based on the multiple feature values. The wake word can be used to operate each speech recognition function of electronic device 101 or 105 and one or more external electronic devices.
[0233] According to an embodiment, the first discrimination result may include at least one of the following: the discrimination accuracy value of the wake word in the speech signal recognized by the electronic device 101 or 105, the sound power level of the speech signal including the wake word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of obtaining the speech signal.
[0234] According to an embodiment, the second discrimination result may include a third discrimination result for wake words sent from one or more external electronic devices and a fourth discrimination result including identification information. The third discrimination result may be generated based on another voice signal including a wake word obtained from one or more external electronic devices. The fourth discrimination result may be generated based on another voice signal obtained from an external electronic device registered by the user.
[0235] According to an embodiment, the fourth discrimination result may include at least one of the following: the identification information of the user who issued the wake word, the user's group information, or the user's discrimination accuracy value.
[0236] According to an embodiment, when operated individually or jointly by at least one processor, instructions may cause electronic device 101 or 105 to identify whether the target of the wake word indication is an electronic device based on a first discrimination result and a second discrimination result. When operated individually or jointly by at least one processor, instructions may cause electronic device 101 or 105 to operate a mode for receiving voice signals including command words based on the identification that the target is electronic device 101 or 105.
[0237] According to an embodiment, the specified function may include displaying a user interface for guiding electronic device 101 or 105 to identify the user or running at least one of a software application that requires user authentication.
[0238] According to an embodiment, when operated by at least one processor individually or jointly, the instructions can cause electronic device 101 or 105 to identify whether a speaker identification model for identifying a user identified based on a second identification result has been generated. When operated by at least one processor individually or jointly, the instructions can cause electronic device 101 or 105 to perform the following operation: based on the identification that no speaker identification model has been generated, to train the speaker identification model for the user using the first identification result.
[0239] According to an embodiment, a speaker identification model can be trained based on a specified number of speech signals for a user. The specified number can be identified based on the signal quality of the speech signals.
[0240] The method performed by electronic device 101 or 105 as described above may include acquiring a voice signal. The method may include generating a first recognition result for the wake word from the voice signal, which includes the wake word. The method may include sending the first recognition result to a server 107 connected to one or more external electronic devices capable of recognizing the wake word. The method may include receiving a second recognition result from the server 107 for the user who issued the wake word. The method may include operating a user-specific function in electronic device 101 or 105 based on the second recognition result. The second recognition result may include user identification information obtained from an external electronic device 103 registered by the user among one or more external electronic devices.
[0241] When operated individually or jointly by at least one processor of electronic device 101 or 105, including communication circuitry and microphones 411 or 451, the non-transitory computer-readable storage medium described above may store one or more programs including instructions that cause electronic device 101 or 105 to acquire a voice signal via microphones 411 or 451. When operated individually or jointly by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to generate a first recognition result for a wake-up word from a voice signal including a wake-up word. When operated individually or jointly by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to send the first recognition result to a server 107 connected to one or more external electronic devices capable of recognizing wake-up words. When operated individually or jointly by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to receive a second recognition result for a user who issued the wake-up word from server 107. When executed by at least one processor individually or jointly, a non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to perform a user-specified function within electronic device 101 or 105 based on a second identification result. The second identification result may include identification information indicating the user obtained from an external electronic device 103 registered with the user in one or more external electronic devices.
[0242] The electronic device 101 or 105 described above may include communication circuitry. The electronic device 101 or 105 may include microphone 411 or 451. The electronic device 101 or 105 may include at least one processor. The at least one processor may be configured to acquire a voice signal via microphone 411 or 451. The at least one processor may be configured to generate a first recognition result for the wake word from the voice signal including the wake word. The at least one processor may be configured to broadcast the first recognition result to one or more external electronic devices capable of recognizing the wake word. The at least one processor may be configured to receive a second recognition result for a wake word broadcast from an external electronic device 103 registered by a user among one or more external electronic devices, and a third recognition result for the user who issued the wake word. The at least one processor may be configured to run a user-specific function in the electronic device 101 or 105 based on the third recognition result. The third recognition result may include identification information indicating the user.
[0243] According to an embodiment, at least one processor can be configured to extract multiple feature values from a speech signal. At least one processor can be configured to recognize a wake-up utterance based on the multiple feature values. The wake-up utterance can be used to operate each speech recognition function of electronic device 101 or 105 and one or more external electronic devices.
[0244] According to an embodiment, the first discrimination result may include at least one of the following: the discrimination accuracy value of the wake word in the speech signal recognized by the electronic device 101 or 105, the sound power level of the speech signal including the wake word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of obtaining the speech signal.
[0245] According to an embodiment, each of the second and third discrimination results can be generated based on another voice signal, including a wake word, obtained from an external electronic device.
[0246] According to an embodiment, the third discrimination result may include at least one of the following: the identification information of the user who issued the wake word, the user's group information, or the user's discrimination accuracy value.
[0247] According to an embodiment, at least one processor may be configured to identify whether the target of the wake-up word is an electronic device based on a first discrimination result, a second discrimination result, and a third discrimination result. At least one processor may be configured to operate a mode for receiving voice signals including command words based on the identification that the target is an electronic device 101 or 105.
[0248] According to an embodiment, the specified function may include displaying a user interface for guiding electronic device 101 or 105 to identify the user, or running at least one of a software application that requires user authentication.
[0249] According to an embodiment, at least one processor can be configured to identify whether a speaker identification model has been generated for identifying a user identified based on a third identification result. At least one processor can be configured to, based on the identification that no speaker identification model has been generated, train a speaker identification model for a user using a first identification result.
[0250] According to an embodiment, a speaker identification model can be trained based on a specified number of speech signals for a user. The specified number can be identified based on the signal quality of the speech signals.
[0251] The method performed by electronic device 101 or 105 as described above may include acquiring a voice signal. The method may include generating a first recognition result for the wake word from the voice signal including the wake word. The method may include broadcasting the first recognition result to one or more external electronic devices capable of recognizing the wake word. The method may include receiving a second recognition result for a wake word broadcast from an external electronic device 103 registered by a user among one or more external electronic devices, and a third recognition result for the user who issued the wake word. The method may include operating a user-specific function in electronic device 101 or 105 based on the third recognition result. The third recognition result may include identification information indicating the user.
[0252] When operated by at least one processor of electronic device 101 or 105, including communication circuitry and microphones 411 or 451, the non-transitory computer-readable storage medium described above may store one or more programs including instructions that cause electronic device 101 or 105 to acquire a voice signal via microphones 411 or 451. When operated by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to generate a first recognition result for a wake-up word from a voice signal including a wake-up word. When operated by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to broadcast the first recognition result to one or more external electronic devices capable of recognizing the wake-up word. When operated by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 101 or 105 to receive a second recognition result for a wake-up word broadcast from an external electronic device 103 registered by a user among one or more external electronic devices, and a third recognition result for a user who issued the wake-up word. A non-transitory computer-readable storage medium, when run by at least one processor, may store one or more programs including instructions that cause electronic device 101 or 105 to perform a user-specified function within electronic device 101 or 105 based on a third identification result. The third identification result may include identification information instructing the user.
[0253] The electronic device 103 described above may include communication circuitry. The electronic device 103 may include a microphone 431. The electronic device 103 may include at least one processor. The at least one processor may be configured to acquire a voice signal via the microphone 431. The at least one processor may be configured to identify the user who issued the voice signal from the voice signal including a wake-up word. The at least one processor may be configured to send a signal including a first identification result for the wake-up word and a second identification result for the user to a server 107 connected to one or more external electronic devices 101 and 105, which are capable of identifying the wake-up word based on the identification that the user is a registered user in the electronic device 103. The first identification result may include at least one of the following: the accuracy value of the wake-up word identification in the voice signal identified by the electronic device 103, the sound power level of the voice signal including the wake-up word, the signal-to-noise ratio (SNR) of the voice signal, or the timing of acquiring the voice signal. The second identification result may include at least one of the following: identification information of the user who issued the wake-up word, including group information of the user, or the user's identification accuracy value.
[0254] According to an embodiment, at least one processor may be configured to identify whether the target of the wake word is an electronic device 103 based on a first discrimination result and a second discrimination result. At least one processor may be configured to, based on the identification that the target is an electronic device 103, operate a mode for receiving voice signals including command words.
[0255] The method performed by electronic device 103 as described above may include acquiring a speech signal. The method may include identifying the user who issued the speech signal from a speech signal including a wake-up word. The method may include sending a signal including a first identification result for the wake-up word and a second identification result for the user to a server 107 connected to one or more external electronic devices 101 and 105, which are capable of identifying the wake-up word based on the identification that the user is a registered user in electronic device 103. The first identification result may include at least one of the recognition accuracy value of the wake-up word in the speech signal recognized by electronic device 103, the sound power level of the speech signal including the wake-up word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of acquiring the speech signal. The second identification result may include identification information of the user who issued the wake-up word, including at least one of user group information or the user's recognition accuracy value.
[0256] When operated by at least one processor of electronic device 103, which includes communication circuitry and microphone 431, the non-transitory computer-readable storage medium described above may store one or more programs including instructions to cause electronic device 103 to acquire a speech signal via microphone 431. When operated by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 103 to identify the user emitting the speech signal from a speech signal including a wake-up word. When operated by at least one processor, the non-transitory computer-readable storage medium may store one or more programs including instructions that cause electronic device 103 to send a signal including a first identification result of the wake-up word and a second identification result of the user to a server 107 connected to one or more external electronic devices 101 and 105, which are capable of identifying the wake-up word based on the identification that the user is a registered user in electronic device 103. The first identification result may include at least one of the following: the accuracy value of the wake-up word identification in the speech signal identified by electronic device 103, the sound power level of the speech signal including the wake-up word, the signal-to-noise ratio (SNR) of the speech signal, or the timing of acquiring the speech signal. The second discrimination result may include at least one of the following: the user's identification information who issued the wake word, the user's group information, or the user's discrimination accuracy value.
[0257] The electronic device according to various embodiments can be one of a variety of types of electronic devices. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer equipment, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. According to embodiments of this disclosure, the electronic device is not limited to those described above.
[0258] It should be understood that the various embodiments of this disclosure and the terminology used therein are not intended to limit the technical features set forth herein to the specific embodiments, but rather to include various changes, equivalents, or substitutions to the respective embodiments. Regarding the description of the drawings, similar reference numerals may be used to refer to similar or related elements. It should be understood that, unless the relevant context clearly indicates otherwise, the singular form of the noun corresponding to an item may include one or more things. As used herein, each of the phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may include any or all possible combinations of the items listed together in the corresponding phrase. As used herein, terms such as “first” and “second” or “first” and “second” may be used simply to distinguish the respective component from another component and do not limit the components in other respects (e.g., importance or order). It will be understood that, whether the terms “operably” or “communically” are used or not, if an element (e.g., a first element) is referred to as being “coupled” or “connected” to another element (e.g., a second element), it means that the element can be directly (e.g., wiredly) coupled to the other element, wirelessly coupled to the other element, or coupled to the other element via a third element.
[0259] As used in conjunction with various embodiments of this disclosure, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with other terms (e.g., "logic," "logic block," "part," or "circuit"). A module may be a single integral component suitable for performing one or more functions, or its smallest unit or part. For example, according to an embodiment, a module may be implemented as an application-specific integrated circuit (ASIC).
[0260] The various embodiments set forth herein can be implemented as software (e.g., program 140) including one or more instructions stored in a storage medium (e.g., internal memory 136 or external memory 138) that can be read by a machine (e.g., electronic device 101). For example, under the control of a processor, a processor (e.g., processor 120) of the machine (e.g., electronic device 101) can invoke and execute at least one instruction from one or more instructions stored in the storage medium, with or without the use of one or more other components. This allows the machine to operate to perform at least one function according to the invoked at least one instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Wherein, the term "non-transitory" means only that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but the term does not distinguish between cases where data is stored semi-permanently in the storage medium and cases where data is temporarily stored in the storage medium.
[0261] According to embodiments, methods according to various embodiments of this disclosure may be included and provided in a computer program product. The computer program product can be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)) or via an app store (e.g., the Play Store). TM Online distribution (e.g., download or upload) or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as the memory of a manufacturer's server, an app store's server, or a relay server.
[0262] According to various embodiments, each of the above components (e.g., a module or program) may include a single entity or multiple entities, and some of the multiple entities may be arranged separately in different components. According to various embodiments, one or more of the above components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, according to various embodiments, the integrated component may still perform one or more functions of each of the multiple components in the same or similar manner as the corresponding components in the multiple components before integration. According to various embodiments, operations performed by a module, program, or other component may be performed sequentially, in parallel, repeatedly, or heuristically, or one or more operations may be performed in a different order or omitted, or one or more other operations may be added.
Claims
1. An electronic device (101, 105) comprising: a memory including one or more storage mediums that store instructions; a communication circuitry; a microphone (411, 451); and at least one processor including a processing circuitry; wherein the instructions, when executed by the at least one processor, cause the electronic device (101, 105) to: obtain a voice signal via the microphone (411, 451); generate a first recognition result for a wake-up word from the voice signal including the wake-up word; transmit the first recognition result to a server (107) connected to one or more external electronic devices capable of recognizing the wake-up word; receive a second recognition result for a user who uttered the wake-up word from the server (107); and based on the second recognition result, execute a designated function for the user, wherein the second recognition result includes identification information indicating the user obtained from an external electronic device (103) registered by the user among the one or more external electronic devices. 2.The electronic device (101, 105) of claim 1, wherein the instructions, when executed by the at least one processor, cause the electronic device (101, 105) to: wherein extract a plurality of feature values from the voice signal; and recognize the wake-up word based on the plurality of feature values, wherein the wake-up word is used to execute each voice recognition function of the electronic device (101, 105) and the one or more external electronic devices. 3.The electronic device (101, 105) of claim 1, wherein the first recognition result includes at least one of a recognition accuracy value of the wake-up word in the voice signal recognized by the electronic device (101, 105), a sound power level of the voice signal including the wake-up word, a signal-to-noise ratio (SNR) of the voice signal, or a timing at which the voice signal is obtained. 4.The electronic device (101, 105) of claim 1, wherein wherein the second recognition result includes a third recognition result for the wake-up word and a fourth recognition result including identification information, respectively transmitted from the one or more external electronic devices, wherein the third recognition result is generated based on another voice signal including the wake-up word obtained by the one or more external electronic devices, respectively, and wherein, wherein the fourth recognition result is generated based on the another voice signal obtained by the external electronic device registered by the user. 5.The electronic device (101, 105) of claim 4, wherein the fourth recognition result includes at least one of the identification information of the user who uttered the wake-up word, group information including the user, or a recognition accuracy value of the user. 6.The electronic device (101, 105) of claim 1, wherein, wherein the instructions, when executed by the at least one processor, cause the electronic device (101, 105) to: wherein identifying whether an indication target is the electronic device based on the first recognition result and the second recognition result; and based on identifying that the indication target is the electronic device (101, 105), executing a mode for receiving a voice signal including a command word.
7. The electronic device (101, 105) of claim 1, wherein, the designated function includes at least one of displaying a user interface for guiding the electronic device (101, 105) to recognize the user, or executing a software application requiring authentication of the user.
8. The electronic device (101, 105) of claim 1, wherein the instructions, when executed by the at least one processor, cause the electronic device (101, 105) to: identify whether a speaker recognition model for recognizing the user identified based on the second recognition result is generated; and based on identifying that the speaker recognition model is not generated, implement training of the speaker recognition model for the user by using the first recognition result.
9. The electronic device (101, 105) of claim 8, wherein the training of the speaker recognition model is implemented based on a designated number of voice signals for the user, and wherein the designated number is identified based on a signal quality of the voice signals.
10. A method performed by an electronic device (101, 105), comprising: obtaining a voice signal via a microphone (411, 451); generating a first recognition result for a wake-up word from the voice signal including the wake-up word; transmitting the first recognition result to a server (107) connected to one or more external electronic devices capable of recognizing the wake-up word; receiving a second recognition result for a user who uttered the wake-up word from the server (107); and based on the second recognition result, executing a designated function for the user, wherein the second recognition result includes identification information indicating the user obtained from an external electronic device (103) registered with the user among the one or more external electronic devices.
11. The method of claim 10, comprising: extracting a plurality of feature values from the voice signal; and identifying the wake-up word based on the plurality of feature values, wherein the wake-up word is used to execute each voice recognition function of the electronic device (101, 105) and the one or more external electronic devices.
12. The method of claim 10, wherein, the first recognition result includes at least one of a recognition accuracy value of the wake-up word in the voice signal identified by the electronic device (101, 105), a sound power level of the voice signal including the wake-up word, a signal-to-noise ratio (SNR) of the voice signal, or a timing at which the voice signal is obtained.
13. The method of claim 10, wherein the second recognition result includes a third recognition result for the wake-up word and a fourth recognition result including the identification information respectively transmitted from the one or more external electronic devices, wherein the third recognition result is generated based on another speech signal including a wake-up word obtained by the one or more external electronic devices, respectively, and wherein the fourth recognition result is generated based on the another speech signal obtained by the external electronic device registered by the user. 14.The method of claim 13, wherein the fourth recognition result includes at least one of the identification information of the user who uttered the wake-up word, group information including the user, or a recognition accuracy value of the user. 15.A non-transitory computer-readable storage medium storing one or more programs including instructions, when executed by at least one processor of an electronic device (101, 105) including a communication circuit and a microphone (411, 451) individually or collectively, cause the electronic device (101, 105) to: obtain a speech signal via the microphone (411, 451); generate a first recognition result for a wake-up word from the speech signal including the wake-up word; transmit the first recognition result to a server (107) connected to one or more external electronic devices capable of recognizing the wake-up word; receive a second recognition result for a user who uttered the wake-up word from the server (107); and based on the second recognition result, execute a designated function for the user, wherein the second recognition result includes identification information of the user indicated from an external electronic device (103) registered by the user among the one or more external electronic devices.