Electronic device and system comprising same
The integration of voice recognition technology in electronic devices allows for convenient and secure user authentication by registering and identifying users based on their voice, addressing the inefficiencies and security issues of manual account entry.
Patent Information
- Application Number
- PCT/KR2024/011222
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-07-31
- Publication Date
- 2025-08-14
AI Technical Summary
Existing electronic devices require manual entry of account information for user authentication, leading to inconvenience and security risks, especially when multiple users share a device.
An electronic device and system that utilize voice recognition technology to register, identify, and log in users based on their voice, enabling seamless account switching and optimized user interface adjustments.
Facilitates convenient and secure user authentication by eliminating the need for manual account entry, allowing multiple users to access their accounts effortlessly and enhancing security through voice-based identification.
Smart Images

Figure KR2024011222_14082025_PF_FP_ABST
Abstract
Description
Electronic devices and systems including the same
[0001] The present disclosure relates to an electronic device and a system including the same, and more particularly, to an electronic device utilizing voice recognition technology and a system including the same.
[0002] With recent technological advancements, research on voice recognition technology, which processes speech, is actively underway. In particular, research on voice recognition technology, which originated with smartphones, is now being conducted widely in various fields related to user convenience, such as home appliances used in homes and offices, as well as in vehicles.
[0003] Voice recognition technology is commonly used when a user controls an electronic device using their voice. For example, when a user utters a command to control an electronic device, the electronic device can directly recognize and process the user's voice and operate according to the corresponding command. Alternatively, the device can transmit the voice to a voice processing server and then operate according to the corresponding command received from the server.
[0004] Meanwhile, the services and functions provided through electronic devices are becoming increasingly diverse. Furthermore, users register accounts for various services and then log in with their registered accounts to access them. In this case, service providers utilize the user information managed for each account to provide optimal features or information tailored to the user.
[0005] Traditionally, when attempting to log in to a service, users would manually enter their account information, such as their account ID (identification number) and / or password. However, this presents a significant inconvenience, requiring users to manually enter their account information for each service. Furthermore, maintaining a logged-in status to eliminate the inconvenience of entering account information can lead to security issues, such as third parties accessing the user's account information. Furthermore, when multiple users share a single electronic device, this presents a problem: each user must enter their own account information to log in each time they use the service.
[0006] The present disclosure aims to solve the above-mentioned and other problems.
[0007] Another object is to provide an electronic device and a system including the same that can register identification information about a user's voice to a user's account.
[0008] Another object is to provide an electronic device and a system including the same that can identify a user based on the user's voice.
[0009] Another object is to provide an electronic device and a system including the same that can log in to an account of a user identified based on the user's voice.
[0010] Another object is to provide an electronic device and a system including the same that can perform actions optimized for an account of a user identified based on the user's voice.
[0011] Another object is to provide an electronic device and a system including the same that can provide a user interface optimized for the input method used by the user.
[0012] Another object is to provide an electronic device and a system including the same, which can switch to an account of a user identified based on the user's voice while an application is running.
[0013] In order to achieve the above object, an electronic device according to one embodiment of the present disclosure includes a network interface unit for communicating with a server; a user input interface unit for transmitting a signal corresponding to an input; and a control unit, wherein, when a voice signal corresponding to a voice input is received while a predetermined application linked to a user account is running, the control unit determines whether a user who has uttered the voice input corresponds to a first user account logged into the server, and when the user who has uttered the voice input does not correspond to the first user account, obtains an application account corresponding to the user who has uttered the voice input based on a second user account corresponding to the user who has uttered the voice input, and accesses a predetermined service provided by the predetermined application based on the obtained application account through the network interface unit.
[0014] In order to achieve the above object, a system according to one embodiment of the present disclosure includes an electronic device and a server, wherein the electronic device transmits data including a voice signal corresponding to a voice input uttered by a user to the server while a predetermined application linked to a user account is running, and determines whether the user who uttered the voice input corresponds to a first user account logged into the server based on a result of processing the voice signal received from the server, and if the user who uttered the voice input does not correspond to the first user account, acquires an application account corresponding to the user who uttered the voice input based on a second user account corresponding to the user who uttered the voice input, and accesses a predetermined service provided by the predetermined application based on the acquired application account through the network interface unit, and the server generates identification information for the voice signal included in the data received from the electronic device, determines predetermined identification information corresponding to the identification information for the voice signal from identification information mapped to user identification information corresponding to the user account stored in a database of the server, and determines predetermined user identification information mapped to the predetermined identification information. The result of processing the included voice signal can be transmitted to the electronic device.
[0015] The effects of the electronic device and the system including the same according to the present disclosure are described as follows.
[0016] According to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.
[0017] According to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.
[0018] According to at least one embodiment of the present disclosure, a user can log in to an account identified based on the user's voice.
[0019] According to at least one embodiment of the present disclosure, an action optimized for an account of a user identified based on the user's voice can be performed.
[0020] According to at least one embodiment of the present disclosure, a user interface optimized for an input method used by a user can be provided.
[0021] According to at least one embodiment of the present disclosure, when the application is running, it is possible to switch to an account of a user identified based on the user's voice.
[0022] Further scope of the applicability of the present disclosure will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present disclosure will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present disclosure, are given by way of example only.
[0023] FIG. 1 is a diagram illustrating a system according to one embodiment of the present disclosure.
[0024] Figure 2 is an internal block diagram of the electronic device of Figure 1.
[0025] Figure 3 is a drawing referenced in the description of the server of Figure 1.
[0026] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.
[0027] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0028] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an electronic device according to one embodiment of the present disclosure.
[0029] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0030] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.
[0031] FIGS. 9 to 13 are drawings for reference in explaining a process for registering identification information for a user's voice to a user's account according to one embodiment of the present disclosure.
[0032] FIG. 14 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0033] FIGS. 15 to 28 are drawings for reference in explanation of responses corresponding to a user's voice according to various embodiments of the present disclosure.
[0034] Hereinafter, the present disclosure will be described in detail with reference to the drawings. In the drawings, portions irrelevant to the description are omitted to clearly and concisely describe the present disclosure, and the same reference numerals are used for identical or extremely similar portions throughout the specification.
[0035] The suffixes "module" and "part" used in the following description are given solely for the convenience of writing this specification and do not impart any particularly significant meaning or role to the components themselves. Therefore, the terms "module" and "part" may be used interchangeably.
[0036] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0037] Additionally, while terms such as "first" and "second" may be used in this specification to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another.
[0038] FIG. 1 is a diagram illustrating a system according to various embodiments of the present invention.
[0039] Referring to FIG. 1, the system (10) may include an electronic device (100) and / or a server (400).
[0040] The electronic device (100) can transmit and receive data to and from at least one server (400). For example, the electronic device (100) can transmit and receive data to and from at least one server (400) via a network (300) such as the Internet.
[0041] According to one embodiment, at least one server (400) may include a server that performs voice recognition, a server that processes data using a super-giant artificial intelligence model, a server that provides content, etc.
[0042] The electronic device (100) may include a video display device (100a), an air conditioner (100b), a refrigerator (100c), an air purifier (100d), a washing machine (100e), a vehicle (100f), etc. In the present disclosure, the electronic device (100) is described as an example of a video display device (100a), but the present invention is not limited thereto.
[0043] The image display device (100a) may be a device that processes and outputs an image. The image display device (100a) is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.
[0044] The video display device (100a) can receive a broadcast signal, process it, and output the processed broadcast image. When the video display device (100a) receives a broadcast signal, the video display device (100a) may correspond to a broadcast receiving device.
[0045] The video display device (100a) can receive broadcast signals wirelessly via an antenna, or can receive broadcast signals wired via a cable. For example, the video display device (100a) can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.
[0046] Figure 2 is an internal block diagram of the electronic device of Figure 1.
[0047] Referring to FIG. 2, the electronic device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).
[0048] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulator unit (120).
[0049] Meanwhile, unlike the drawing, the electronic device (100) may include only the broadcast receiving unit (105) and the external device interface unit (130) among the broadcast receiving unit (105), the external device interface unit (130), and the network interface unit (135). That is, the electronic device (100) may not include the network interface unit (135).
[0050] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.
[0051] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).
[0052] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.
[0053] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.
[0054] The demodulation unit (120) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (110).
[0055] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.
[0056] The stream signal output from the demodulation unit (120) can be input to the control unit (170). The control unit (170) can output an image through the display (180) and output an audio through the audio output unit (185) after performing demultiplexing, image / audio signal processing, etc.
[0057] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).
[0058] The external device interface unit (130) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.
[0059] In addition, the external device interface unit (130) can establish a communication network with various remote control devices (200) to receive a control signal related to the operation of the electronic device (100) from the remote control device (200) or transmit data related to the operation of the electronic device (100) to the remote control device (200).
[0060] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit can include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).
[0061] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, in mirroring mode, the external device interface unit (130) may receive device information, running application information, application images, etc. from the mobile terminal.
[0062] The external device interface unit (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.
[0063] The network interface unit (135) can provide an interface for connecting the electronic device (100) to a wired / wireless network including the Internet.
[0064] The network interface unit (135) may include a communication module (not shown) for connection to a wired / wireless network. For example, the network interface unit (135) may include a communication module for WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), etc.
[0065] The network interface unit (135) can transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.
[0066] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and information related thereto provided by a content provider or network provider via a network.
[0067] The network interface unit (135) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.
[0068] The network interface unit (135) can select and receive a desired application from among applications open to the public through a network.
[0069] The storage unit (140) may store programs for signal processing and control within the control unit (170), or may store processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored application programs upon request from the control unit (170).
[0070] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (170).
[0071] The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).
[0072] The storage unit (140) can store information about a specific broadcast channel through a channel memory function such as a channel map.
[0073] Although the storage unit (140) of FIG. 2 is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).
[0074] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.
[0075] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.
[0076] For example, a user input signal such as power on / off, channel selection, screen setting, etc. may be transmitted / received from a remote control device (200), a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. may be transmitted to the control unit (170), a user input signal input from a sensor unit (not shown) that senses a user's gesture may be transmitted to the control unit (170), or a signal from the control unit (170) may be transmitted to the sensor unit.
[0077] The input unit (160) may be provided on one side of the main body of the electronic device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.
[0078] The input unit (160) can receive various user commands related to the operation of the electronic device (100) and transmit a control signal corresponding to the input command to the control unit (170).
[0079] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.
[0080] The control unit (170) may include at least one processor, and may control the overall operation of the electronic device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.
[0081] The control unit (170) can demultiplex a stream input through the tuner unit (110), the demodulator unit (120), the external device interface unit (130), or the network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.
[0082] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).
[0083] The display (180) may include a display panel (not shown) having a plurality of pixels.
[0084] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display (180) may convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (170) to generate driving signals for the plurality of pixels.
[0085] The display (180) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display, etc., and may also be a 3D display. The 3D display (180) can be divided into a glasses-free type and a glasses type.
[0086] Meanwhile, the display (180) is configured as a touch screen and can be used as an input device in addition to an output device.
[0087] The audio output unit (185) receives a signal processed by the control unit (170) and outputs it as voice.
[0088] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (170) can also be input to an external output device through the external device interface unit (130).
[0089] The voice signal processed in the control unit (170) can be output as sound to the audio output unit (185). In addition, the voice signal processed in the control unit (170) can be input to an external output device through the external device interface unit (130).
[0090] Although not shown in FIG. 2, the control unit (170) may include a demultiplexing unit, an image processing unit, etc.
[0091] In addition, the control unit (170) can control the overall operation within the electronic device (100). For example, the control unit (170) can control the tuner unit (110) to select (tune) a broadcast corresponding to a channel selected by the user or a previously stored channel.
[0092] In addition, the control unit (170) can control the electronic device (100) by a user command or internal program input through the user input interface unit (150).
[0093] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a moving image, and may be a 2D image or a 3D image.
[0094] Meanwhile, the control unit (170) can cause a predetermined 2D object to be displayed within an image displayed on the display (180). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.
[0095] Meanwhile, the electronic device (100) may further include a camera (not shown). The camera can capture images of the user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the electronic device (100) above the display (180) or positioned separately. Image information captured by the camera can be input to the control unit (170).
[0096] The control unit (170) can recognize the user's location based on the image captured by the camera. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the electronic device (100). In addition, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.
[0097] The control unit (170) can detect the user's gesture based on an image captured from the camera unit, a signal detected from the sensor unit, or a combination thereof.
[0098] The power supply unit (190) can supply power to the entire electronic device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a system on chip (SOC), a display (180) for displaying images, and an audio output unit (185) for outputting audio.
[0099] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.
[0100] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (150) and display or output the same as voice on the remote control device (200).
[0101] Meanwhile, the electronic device (100) described above may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.
[0102] Meanwhile, the block diagram of the electronic device (100) illustrated in FIG. 2 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the electronic device (100) actually implemented.
[0103] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.
[0104] Figure 3 is a drawing referenced in the description of the server of Figure 1.
[0105] Referring to FIG. 3, the server (400) may include a relay server (410), an STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), a user identification server (440), and / or an account server (450). In the present disclosure, the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) are described as being distinct from each other, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) may be configured as one server.
[0106] The relay server (410) can communicate with the electronic device (100). The relay server (410) can transfer data between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100). The relay server (410) can store at least a portion of the data transferred between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100).
[0107] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the electronic device (100) via the relay server (410). The STT server (420) may also be referred to as an ASR (Automatic Speech Recognition) server.
[0108] The STT server (420) can improve the accuracy of speech-to-text conversion using a language model. The language model can refer to a model that can calculate the probability of a sentence or the probability of a subsequent word given previous words. For example, the language model can include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. In other words, the STT server (420) can use the language model to determine whether text data converted from speech data has been appropriately converted, thereby improving the accuracy of conversion into text data.
[0109] The NLP server (430) can receive text data. Based on the received text data, the NLP server (430) can perform intent analysis on the text data. The NLP server (430) can transmit intent analysis information indicating the results of the intent analysis to the electronic device (100) via the relay server (410).
[0110] According to one embodiment, the NLP server (430) may sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate intent analysis information. The morphological analysis step is a step of classifying text data corresponding to speech uttered by a user into morphemes, which are the smallest units having meaning, and determining which part of speech each classified morpheme has. The syntax analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and to determine what kind of relationship exists between each of the classified phrases. Through the syntax analysis step, the subject, object, and modifiers of the speech uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the speech uttered by the user using the results of the syntax analysis step. Specifically, the speech act analysis step is a step of determining the intent of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.
[0111] The user identification server (440) can receive voice data. Based on the voice data, the user identification server (440) can extract voice features. Here, the voice features may include the voice waveform, the voice frequency band, the voice power spectrum, etc. Extracting voice features will be described below with reference to FIGS. 4 and 5.
[0112] The user identification server (440) can obtain a feature vector of a voice from the features of the voice. The user identification server (440) can obtain a feature vector of a voice from the features of the voice based on a linear predictive coefficient, a cepstrum, a Mel Frequency Cepstral Coefficient (MFCC), and filter bank energy.
[0113] The user identification server (440) can determine the similarity between a plurality of feature vectors. The user identification server (440) can determine the similarity between a plurality of feature vectors using cosine similarity, Euclidean similarity, etc. In the present disclosure, the similarity between a first voice input and a second voice input is calculated based on cosine similarity, but is not limited thereto. For example, a first vector corresponding to a first text and a second vector corresponding to a second text may be generated. In this case, the cosine similarity between the first vector and the second vector may be calculated based on the following mathematical expression 1.
[0114]
[0115] Here, is the dot product of two vectors, and can mean the magnitude of two vectors. That is, cosine similarity can be calculated as the value obtained by dividing the inner product of two vectors by the product of the magnitudes of each vector. Cosine similarity can range from -1 to 1, and the closer it is to 1, the more similar the two vectors can be judged to be.
[0116] The user identification server (440) can determine whether the users who uttered the voice are the same user based on the similarity between multiple feature vectors. For example, if the similarity between the first feature vector corresponding to the first voice input and the second feature vector corresponding to the second voice input is above a predetermined standard, the user identification server (440) can determine that the user who uttered the first voice input and the user who uttered the second voice input are the same user.
[0117] According to one embodiment, the user identification server (440) may obtain a vector by processing a feature vector of a voice using an algorithm such as a GMM (Gaussian mixture model) supervector, i-vector, d-vector, x-vector, etc. The user identification server (440) may determine whether the user who uttered the voice is the same based on the similarity between a first vector processed corresponding to a first feature vector and a second vector processed corresponding to a second feature vector.
[0118] The user identification server (440) can store voice data. The user identification server (440) can store data regarding the voice print (hereinafter, voice print information). Here, the voice print information can include a voice feature vector and / or a vector obtained by processing the voice feature vector.
[0119] The user identification server (440) can store a database for voices. The database for voices can include unique identification information (hereinafter, device identification information) corresponding to an electronic device (100), unique identification information (hereinafter, user identification information) corresponding to a user account, voice data mapped to user identification information, voice print information mapped to user identification information, etc.
[0120] Device identification information, user identification information, voice data, voice print information, etc. included in the voice database may be stored in a user identification server (440) in association with each other. For example, at least one piece of device identification information, multiple voice data, and / or multiple voice print information may be mapped to the user identification information. In other words, it may be interpreted that the device identification information, voice data, voice print information, etc. are mapped to a user account and stored in the user identification server (440). In the present disclosure, an example in which multiple voice data and multiple voice print information are all mapped to the user identification information included in the voice database will be described.
[0121] The user identification server (440) can update the voiceprint information contained in the voice database based on the voice data contained in the voice database. For example, the user identification server (440) can generate voiceprint information corresponding to the voice data contained in the voice database using an algorithm different from the algorithm previously used. In this case, the user identification server (440) can change the voiceprint information contained in the voice database to the newly generated voiceprint information.
[0122] The account server (450) can manage data regarding user accounts. The account server (450) can manage user account IDs, passwords, user identification information, device identification information mapped to the user account, and whether or not to agree to terms and conditions related to various functions.
[0123] The account server (450) can store a database of user accounts. The database of user accounts may include the user account ID, password, user identification information, device identification information mapped to the user account, the registration date and time of the user account, whether or not the user has agreed to terms and conditions related to various functions, and the date and time of agreement to the terms and conditions.
[0124] The account server (450) can communicate with the electronic device (100). For example, the account server (450) can create and register a user account based on data from the electronic device (100). For example, the account server (450) can approve login to a user account based on an ID and password received from the electronic device (100).
[0125] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.
[0126] Referring to FIG. 4, the server (400) may include a preprocessing unit (460), a controller (470), a communication unit (480), and / or a database (490).
[0127] The preprocessing unit (460) can preprocess voice received through the communication unit (480) or voice stored in the database (490).
[0128] The preprocessing unit (460) may be implemented as a separate chip from the controller (470) or as a chip included in the controller (470).
[0129] The preprocessing unit (460) can receive a voice signal (spoken by a user) and filter out noise signals from the voice signal before converting the received voice signal into text data.
[0130] When a preprocessing unit (460) is provided in the electronic device (100), it can recognize a trigger word for activating voice recognition of the electronic device (100). The preprocessing unit (460) converts the trigger word received through the user input interface unit (150) into text data, and if the converted text data corresponds to a previously stored trigger word, it can be determined that the trigger word has been recognized.
[0131] The preprocessing unit (460) can convert the noise-removed voice signal into a power spectrum.
[0132] Power spectrum can be a parameter that indicates which frequency components are included in the waveform of a temporally varying voice signal and at what magnitude.
[0133] The power spectrum shows the distribution of the squared amplitude values of the waveform of a voice signal according to frequency. This is explained with reference to Figure 5.
[0134] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0135] Referring to FIG. 5, a voice signal (510) is illustrated. The voice signal (460) may be received from an external device or may be a signal pre-stored in memory (170).
[0136] The x-axis of the voice signal (510) represents time, and the y-axis can represent the amplitude size.
[0137] The power spectrum processing unit (463) can convert a voice signal (510) whose x-axis is the time axis into a power spectrum (520) whose x-axis is the frequency axis. The power spectrum processing unit (463) can convert the voice signal (510) into a power spectrum (520) using a fast Fourier transform (FFT). The x-axis of the power spectrum (520) represents frequency, and the y-axis represents the square value of the amplitude.
[0138] Referring again to FIG. 4, the functions of the preprocessing unit (460) and the controller (470) described in FIG. 4 can also be performed in the NLP server (430).
[0139] The preprocessing unit (460) may include a wave processing unit (461), a frequency processing unit (462), a power spectrum processing unit (463), a speech to text (STT) conversion unit (464), etc.
[0140] The wave processing unit (461) can extract the waveform of the voice.
[0141] The frequency processing unit (462) can extract the frequency band of the voice.
[0142] The power spectrum processing unit (463) can extract the power spectrum of the voice.
[0143] Power spectrum can be a parameter that indicates which frequency components are included in a given temporally varying waveform and at what magnitude.
[0144] The speech-to-text (STT) conversion unit (464) can convert speech into text. The speech-to-text conversion unit (464) can convert speech in a specific language into text in that language.
[0145] The controller (470) can control the overall operation of the server (400). The controller (470) can include a voice analysis unit (471), a text analysis unit (472), a feature clustering unit (473), a text mapping unit (474), and / or a voice synthesis unit (475).
[0146] The voice analysis unit (471) can extract voice characteristic information by using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed by the preprocessing unit (460). The voice characteristic information may include one or more of the speaker's gender information, the speaker's voice (or tone), pitch, the speaker's speech style, the speaker's speaking speed, and the speaker's emotion. In addition, the voice characteristic information may further include the speaker's voice tone.
[0147] The text analysis unit (472) can extract key phrases from the text converted by the voice-to-text conversion unit (464). If the text analysis unit (472) detects a difference in tone between phrases in the converted text, it can extract the phrases with different tones as key phrases. If the text analysis unit (472) changes the frequency band between phrases by more than a preset band, it can determine that the tone has changed. The text analysis unit (472) can also extract key words within the phrases of the converted text. Key words may be nouns present within the phrases, but this is merely an example.
[0148] The feature clustering unit (473) can classify the speaker's speech type using the characteristic information of the voice extracted from the voice analysis unit (471). The feature clustering unit (473) can classify the speaker's speech type by assigning weights to each of the type items constituting the characteristic information of the voice. The feature clustering unit (473) can classify the speaker's speech type using the attention technique of the deep learning model.
[0149] The text mapping unit (474) can translate a text converted into a first language into a text in a second language. The text mapping unit (474) can map a text translated into a second language to a text in the first language. The text mapping unit (474) can map the main expressions constituting the text in the first language to the corresponding expressions in the second language. The text mapping unit (474) can map the utterance types corresponding to the main expressions constituting the text in the first language to the expressions in the second language. This is to apply the classified utterance types to the expressions in the second language.
[0150] The voice synthesis unit (475) can generate a synthesized voice by applying the speech type and the speaker's tone classified by the feature clustering unit (473) to the main expression phrases of the text translated into a second language by the text mapping unit (474).
[0151] The controller (470) can determine the user's speech characteristics using one or more of the transmitted text data or power spectrum (520).
[0152] User speech characteristics may include the user's gender, the user's pitch, the user's timbre, the user's speech topic, the user's speech rate, the user's voice volume, etc.
[0153] The controller (470) can obtain the frequency of the voice signal (510) and the amplitude corresponding to the frequency by using the power spectrum (520).
[0154] The controller (470) can determine the gender of the user who has spoken using the frequency band of the power spectrum (470). For example, the controller (470) can determine the gender of the user as male if the frequency band of the power spectrum (520) is within a preset first frequency band range.
[0155] The controller (470) can determine the user's gender as female if the frequency band of the power spectrum (520) is within a preset second frequency band range. Here, the second frequency band range may be larger than the first frequency band range.
[0156] The controller (470) can determine the pitch of a sound by using the frequency band of the power spectrum (520). For example, the controller (470) can determine the pitch of a sound based on the amplitude within a specific frequency band range.
[0157] The controller (470) can determine the user's tone by using the frequency band of the power spectrum (520). For example, the controller (470) can determine a frequency band with an amplitude greater than a certain size among the frequency bands of the power spectrum (520) as the user's main vocal range, and determine the determined main vocal range as the user's tone.
[0158] The controller (470) can determine the user's speaking speed through the number of syllables spoken per unit time from the converted text data.
[0159] For the converted text data of the controller (470), the user's speech topic can be determined using the Bag-Of-Word Model technique.
[0160] The Bag-Of-Word Model technique extracts frequently used words based on their frequency within a sentence. Specifically, the Bag-Of-Word Model technique extracts unique words from a sentence, expresses the frequency of each extracted word as a vector, and determines the characteristics of the utterance topic. For example, if words like "running" and "physical strength" frequently appear in the controller (470) text data, the user's utterance topic can be classified as exercise.
[0161] The controller (470) can determine the topic of a user's speech from text data using a known text categorization technique. The controller (470) can extract keywords from the text data to determine the topic of the user's speech.
[0162] The controller (470) can determine the user's vocal volume by considering amplitude information across the entire frequency band. For example, the controller (470) can determine the user's vocal volume based on the average or weighted average of the amplitudes across each frequency band of the power spectrum.
[0163] The communication unit (480) can communicate with an external server via wired or wireless communication. The communication unit (480) can communicate with an electronic device (100) via wired or wireless communication.
[0164] The database (490) can store speech in a first language included in the content. The database (490) can store a synthesized speech in which the speech in the first language is converted into speech in a second language. The database (490) can store a first text corresponding to the speech in the first language and a second text in which the first text is translated into a second language. The database (490) can store various learning models required for speech recognition.
[0165] Meanwhile, the control unit (170) of the electronic device (100) illustrated in FIG. 2 may be equipped with the preprocessing unit (460) and the controller (470) illustrated in FIG. 4. That is, the control unit (170) of the electronic device (100) may perform the functions of the preprocessing unit (460) and the controller (470).
[0166] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.
[0167] That is, the voice recognition and synthesis process of FIG. 6 may be performed by the control unit (170) of the electronic device (100) without going through the server.
[0168] Referring to FIG. 6, the processor (180) of the electronic device (100) may include an STT engine (610), an NLP engine (620), and a voice synthesis engine (630). Each engine may be either hardware or software.
[0169] The STT engine (610) can perform the function of the STT server (420) of FIG. 5. That is, the STT engine (610) can convert voice data into text data.
[0170] The NLP engine (620) can perform the function of the NLP server (430) of FIG. 5. That is, the NLP engine (620) can obtain intent analysis information indicating the speaker's intent from converted text data.
[0171] The speech synthesis engine (630) can perform the functions of a speech synthesis server. The speech synthesis engine (630) can search for syllables or words corresponding to given text data from a database, synthesize a combination of the searched syllables or words, and generate a synthesized speech.
[0172] The speech synthesis engine (630) may include a preprocessing engine (631) and a TTS engine (632).
[0173] The preprocessing engine (631) can preprocess text data before generating synthetic speech. Specifically, the preprocessing engine (631) performs tokenization, which divides the text data into tokens, which are meaningful units. After tokenization, the preprocessing engine (631) can perform a cleansing operation to remove unnecessary characters and symbols to remove noise. Thereafter, the preprocessing engine (631) can integrate word tokens with different expression methods to generate the same word token. Thereafter, the preprocessing engine (631) can remove meaningless word tokens (stopwords).
[0174] The TTS engine (632) can synthesize voice corresponding to preprocessed text data and generate a synthesized voice.
[0175] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0176] Referring to FIG. 7, the electronic device (100) can determine, in operation S701, whether a user account is logged into the server (400). For example, a user can log into the server (400) with a user account by entering the ID and password of the user account.
[0177] According to one embodiment, when a user first logs into a server (400) using an electronic device (100), the electronic device (100) may include user identification information corresponding to the user account in a user list. For example, when three different user accounts log into the server (400) using the electronic device (100), the user list stored in the electronic device (100) may include three different user identification information.
[0178] In operation S702, the electronic device (100) can determine whether voice-related identification information (hereinafter, "Voice ID") is registered for a user account logged into the server (400). Here, the Voice ID may include voiceprint information stored in the user identification server (440). For example, the server (400) may transmit to the electronic device (100) whether a voice ID is registered for a user account logged into the server (400).
[0179] According to one embodiment, the server (400) can determine whether a voice ID is registered based on whether voice print information is mapped to user identification information, which is unique identification information corresponding to a user account logged into the server (400). In this case, if the voice ID corresponds to an unregistered user account, the number of voice print information mapped to the user identification information may be 0.
[0180] According to one embodiment, the server (400) may determine that a voice ID is registered if the number of voice prints mapped to the user identification information is two or more, a predetermined number, and may determine that the voice ID is not registered if the number is less than the predetermined number. For example, in the case of a user account with a registered voice ID, six different voice prints may be mapped to the user identification information. For example, in the case of a user account with an unregistered voice ID, five or fewer voice prints may be mapped to the user identification information.
[0181] According to one embodiment, a flag value indicating whether a voice ID is registered may be mapped to user identification information stored in the server (400). Here, the user identification information to which the flag value is mapped may be stored in the user identification server (440) and / or the account server (450). The server (400) may determine whether a voice ID is registered based on the flag value mapped to the user identification information. For example, if the voice ID is an unregistered user account, the flag value mapped to the user identification information may be 0, and if the voice ID is a registered user account, the flag value mapped to the user identification information may be 1.
[0182] In operation S703, if a voice ID has not been registered for a user account, the electronic device (100) may initiate a process for registering a voice ID. For example, when the electronic device (100) initiates a process for registering a voice ID, it may transmit data including device identification information, user identification information, a value indicating the start of voice ID registration, etc., to the server (400).
[0183] The electronic device (100) can output preset text in operation S704. The electronic device (100) can output any one of a plurality of preset texts. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) can output the preset text through the display (180).
[0184] According to one embodiment, the server (400) may transmit any one of a plurality of preset texts to the electronic device (100) in a preset order. At this time, the electronic device (100) may output the preset text received from the server (400).
[0185] The electronic device (100) can determine, in operation S705, whether a voice is input for a preset text. For example, the electronic device (100) can determine whether a voice is input through a microphone included in the input unit (160) within a preset time period. At this time, a voice signal corresponding to the voice input through the microphone can be transmitted to the control unit (170) through the user input interface unit (150). For example, the electronic device (100) can determine whether data including a voice signal corresponding to a voice spoken by a user is received from the remote control device (200) within a preset time period.
[0186] In operation S706, when a voice is input for a preset text, the electronic device (100) can transmit voice data containing a voice signal corresponding to the voice to the server (400). At this time, the electronic device (100) can transmit device identification information, user identification information, a language code indicating the type of language, etc., together with the voice data to the server (400).
[0187] The server (400) can convert a voice signal included in voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponds to a preset text. For example, the server (400) can determine whether the two texts correspond based on the similarity between the text converted from the voice signal and the preset text.
[0188] The server (400) can generate voiceprint information corresponding to the voice signal when the text converted from the voice signal corresponds to a preset text. The server (400) can map the voiceprint information generated for the preset text to user identification information and store it. The server (400) can map the voice data received for the preset text to user identification information and store it.
[0189] The electronic device (100), in operation S707, can determine whether the processing of the voice for the preset text was successful based on the response received from the server (400). For example, if the text converted from the voice signal and the preset text correspond to each other, the server (400) can notify the electronic device (100) of the success of the processing of the voice. For example, if the voiceprint information corresponding to the voice signal is generated, the server (400) can notify the electronic device (100) of the success of the processing of the voice.
[0190] Meanwhile, in operation S708, if voice input for a preset text is not performed, or if voice processing for a preset text fails, the electronic device (100) may determine whether to retry voice input. For example, the electronic device (100) may retry voice input based on a user input requesting retry voice input. In this case, the electronic device (100) may output the preset text again.
[0191] In operation S709, if the voice processing for a preset text is successful, the electronic device (100) can determine whether the processing for all texts is complete. For example, if the voice processing for a predetermined number of six preset texts is successful, the processing for all texts can be completed. Meanwhile, if the processing for five preset texts is complete, the electronic device (100) can output the last preset text.
[0192] The electronic device (100) may complete the process of registering a voice ID when processing of all texts is completed in operation S710. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) may output a screen indicating completion of voice ID registration through the display (180). For example, the electronic device (100) may transmit data indicating completion of voice ID registration to the account server (450).
[0193] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.
[0194] Referring to FIG. 8, the electronic device (100) can perform a login to the server (400) using a user account in operation S801.
[0195] The electronic device (100) can initiate a process of registering a voice ID in operation S802.
[0196] The electronic device (100) can output a first text among a plurality of preset texts in operation S803.
[0197] The electronic device (100) can receive a first voice for a first text in operation S804.
[0198] The electronic device (100) can transmit first voice data including a voice signal corresponding to the first voice to the server (400) in operation S805.
[0199] The server (400), in operation S806, can process the first voice for the first text based on the first voice data received from the electronic device (100). The server (400) can convert a voice signal corresponding to the first voice included in the first voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the first voice and the first text correspond to each other.
[0200] The server (400) may, in operation S807, notify the electronic device (100) of the completion of processing for the first voice. For example, the server (400) may notify the electronic device (100) of the success of processing for the first voice based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.
[0201] Meanwhile, the server (400) can generate first voiceprint information for the first voice based on the voice signal corresponding to the first voice, based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.
[0202] The server (400) may, in operation S808, store first voice data and first voiceprint information for the first voice. The server (400) may store the first voice data and the first voiceprint information by mapping them to user identification information corresponding to a logged-in user account.
[0203] The electronic device (100) can sequentially output the second text to the fifth text. The electronic device (100) can sequentially receive the second voice to the fifth voice, which correspond to the second text to the fifth text, respectively. The electronic device (100) can sequentially transmit the second voice data to the fifth voice data, which correspond to the second voice to the fifth voice, respectively, to the server (400).
[0204] The server (400) can process the second to fifth voices based on the second to fifth voice data received from the electronic device (100), respectively. In addition, the server (400) can sequentially generate and store second to fifth voice information corresponding to the second to fifth voices, respectively.
[0205] The electronic device (100) can output a sixth text among a plurality of preset texts in operation S809.
[0206] The electronic device (100) can receive a sixth voice for a sixth text in operation S810.
[0207] The electronic device (100) can transmit sixth voice data including a voice signal corresponding to the sixth voice to the server (400) in operation S811.
[0208] The server (400) can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device (100) in operation S812. The server (400) can convert a voice signal corresponding to the sixth voice included in the sixth voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.
[0209] The server (400) can, in operation S813, notify the electronic device (100) of the completion of processing for the sixth voice.
[0210] Meanwhile, the server (400) can generate sixth voice information for the sixth voice based on the voice signal corresponding to the sixth voice, if the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.
[0211] The server (400) may store the sixth voice data and the sixth voiceprint information for the sixth voice in operation S814. The server (400) may store the sixth voice data and the sixth voiceprint information by mapping them to user identification information corresponding to the logged-in user account. At this time, the user identification information corresponding to the logged-in user account may be mapped with six different voice data and multiple voiceprint information.
[0212] The electronic device (100) may complete the process of registering a voice ID in operation S815. For example, the electronic device (100) may complete the process of registering a voice ID based on the completion of processing of a predetermined number of six different preset texts.
[0213] Referring to FIG. 9, if a user account is not logged in to the server (400), the electronic device (100) may output a login screen (900) related to logging in to the server (400) through the display (180). The login screen (900) may include an object (910) indicating a non-logged-in state, a login object (920) for executing login, etc. When a user selects the login object (920) using a pointer (205) corresponding to the remote control device (200), the electronic device (100) may output a screen for inputting an ID and password. At this time, the user may input the ID and password of the user account to log in to the server (400) with the user account.
[0214] Referring to FIG. 10, if a voice ID is not registered in a user account logged in to the server (400), the electronic device (100) can output a first account screen (1000) corresponding to the user account with the unregistered voice ID. The first account screen (1000) can include an object (1010) representing a logged-in user account, an object (1020) corresponding to the registration of a voice ID, etc. When a user selects an object (1020) corresponding to the registration of a voice ID using a pointer (205), the electronic device (100) can initiate a process of registering a voice ID.
[0215] Meanwhile, referring to FIG. 11, if a voice ID is already registered in a user account logged in to the server (400), the electronic device (100) can output a second account screen (1100) corresponding to the user account in which the voice ID is registered. The second account screen (1100) can include an object (1110) representing the logged-in user account, a re-registration object (1120) corresponding to re-registration of the voice ID, a deletion object (1130) corresponding to deletion of the voice ID, an activation object (1140) corresponding to use of a function related to the voice ID, etc. The user can select the activation object (1140) using the pointer (205) to activate or deactivate use of the function related to the voice ID.
[0216] Referring to FIG. 12, when an object (1020) corresponding to the registration of a voice ID is selected on the first account screen (1000), or when a re-registration object (1120) is selected on the second account screen (1100), the electronic device (100) can output a start screen (1200) for initiating the registration of a voice ID. When a user selects a start object (1210) using a pointer (205), the electronic device (100) can output a text screen for outputting preset text.
[0217] Referring to FIG. 13, the electronic device (100) can output a text screen (1300) that outputs any one of a plurality of preset texts. The text screen (1300) can include preset text (1301), a text sequence (1302), a termination object (1310) that terminates a process of registering a voice ID, an input object (1320) that receives voice, etc.
[0218] When the user selects the end object (1310) using the pointer (205), the process of registering a voice ID may be terminated. For example, when the process of registering a voice ID is terminated, all data stored in the server (400) while the process of registering a voice ID is in progress may be deleted.
[0219] When a user selects an input object (1320) using a pointer (205), the electronic device (100) can receive voice for the text.
[0220] According to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a text screen (1300) is displayed, the electronic device (100) can receive a voice for the text based on the user input of pressing the predetermined button received from the remote control device (200).
[0221] Meanwhile, according to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a process for registering a voice ID is in progress, the electronic device (100) may stop the process for registering a voice ID based on a user input for pressing the predetermined button received from the remote control device (200). At this time, the user input for pressing the predetermined button (e.g., a voice input button) included in the remote control device (200) may correspond to a user input for initiating voice recognition for a voice received through the remote control device (200). The electronic device (100) may perform an operation related to voice recognition on voice data including a voice signal received from the remote control device (200).
[0222] FIG. 14 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0223] Referring to FIG. 14, the electronic device (100) can connect to the server (400) through the network interface unit (145) in operation S1401.
[0224] The electronic device (100) can receive a voice in operation S1402. For example, the electronic device (100) can receive a voice through a microphone included in the input unit (160). At this time, a voice signal corresponding to the voice input through the microphone can be transmitted to the control unit (170) through the user input interface unit (150). For example, the electronic device (100) can receive data including a voice signal corresponding to a voice spoken by a user from a remote control device (200).
[0225] In operation S1403, when a voice is input, the electronic device (100) can transmit voice data corresponding to the input voice to the server (400). At this time, the electronic device (100) can transmit device identification information, a user list, a language code indicating the type of language, etc., together with the voice data to the server (400).
[0226] The electronic device (100) may, in operation S1404, receive the result of processing the voice from the server (400). For example, the result of processing the voice may include text corresponding to the voice, intent analysis information resulting from performing natural language processing on the voice, user identification information corresponding to the voice, etc.
[0227] The server (400) can generate voiceprint information for a voice input into the electronic device (100) based on voice data received from the electronic device (100). The server (400) can search a database for voiceprint information (hereinafter, candidate voiceprint information) corresponding to user identification information included in a user list received from the electronic device (100). The server (400) can determine whether the candidate voiceprint information and the generated voiceprint information correspond to each other. The server (400) can determine user identification information, to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped, among the candidate voiceprint information, as user identification information corresponding to the voice input into the electronic device (100). Meanwhile, if there is no candidate voiceprint information corresponding to the generated voiceprint information, the server (400) can determine that there is no user identification information corresponding to the voice input into the electronic device (100).
[0228] The server (400) can transmit the result of processing voice data received from the electronic device (100) to the electronic device (100). For example, the server (400) can transmit to the electronic device (100) text converted from voice data received from the electronic device (100), intent analysis information indicating the result of performing intent analysis on the converted text, user identification information corresponding to the voice, and whether or not user identification information corresponding to the voice exists.
[0229] The electronic device (100) can determine, in operation S1405, whether the input voice is a voice requiring user identification based on the result of voice processing.
[0230] According to one embodiment, the electronic device (100) may determine whether an input voice is a voice requiring user identification based on the type of command corresponding to the voice. For example, the electronic device (100) may determine that the input voice is a voice requiring user identification based on the fact that the input voice corresponds to a command related to a user account (e.g., logging in to a user account, switching user accounts, etc.). For example, the electronic device (100) may determine that the input voice is a voice requiring user identification based on the fact that the input voice corresponds to a command that utilizes information about the user such as usage history, viewing history, and preferred genre (e.g., content search, content recommendation, speech recommendation, external device connection, etc.). For example, the electronic device (100) may determine that the input voice is not a voice requiring user identification based on the fact that the input voice corresponds to a command that confirms general information unrelated to the user (e.g., time, weather, etc.). For example, the electronic device (100) may determine that the input voice is not a voice requiring user identification based on whether the input voice corresponds to a command for adjusting the settings of the electronic device (100) (e.g., volume, screen brightness, etc.).
[0231] According to one embodiment, the electronic device (100) can determine whether the input voice is a voice requiring user identification based on the type of application corresponding to the voice. For example, the electronic device (100) can determine that the input voice is a voice requiring user identification based on whether the input voice corresponds to an application linked to a user account (e.g., an over-the-top media service (OTT) service, a social network service (SNS), etc.). For example, the electronic device (100) can determine that the input voice is not a voice requiring user identification based on whether the input voice corresponds to an application not linked to a user account (e.g., terrestrial broadcasting, weather, etc.).
[0232] In operation S1406, if the input voice is a voice requiring user identification, the electronic device (100) can determine whether the user who spoke the voice is a user preset to use a voice ID based on user identification information corresponding to the voice. For example, the electronic device (100) can store whether the use of a voice ID-related function is activated for each piece of user identification information included in a user list. At this time, if the use of a voice ID-related function is activated for the user identification information corresponding to the voice included in the result of voice processing, the electronic device (100) can determine that the user who spoke the voice is a user preset to use a voice ID.
[0233] Meanwhile, if there is no user identification information corresponding to the voice based on the result of processing the voice, the electronic device (100) may determine that the user who spoke the voice is not a user who has been preset to use a voice ID.
[0234] In operation S1407, if the user who has spoken the voice is a user who has been preset to use a voice ID, the electronic device (100) can determine whether the user who has spoken the voice corresponds to the currently logged-in user account. For example, if the user identification information corresponding to the voice corresponds to the user identification information corresponding to the currently logged-in user account, the electronic device (100) can determine that the user who has spoken the voice corresponds to the currently logged-in user account.
[0235] In operation S1408, the electronic device (100) can determine whether an application linked to a user account is running. For example, the electronic device (100) can determine whether an application running in the foreground is an application linked to a user account (e.g., an OTT service, SNS, etc.). In this case, the application running in the foreground can correspond to a screen output through the display (180).
[0236] In operation S1409, if an application linked to a user account is not running in the foreground, the electronic device (100) may determine whether a user account switch is required. The electronic device (100) may determine whether a user account switch is required based on intent analysis information included in the voice processing result. Meanwhile, if no user account is currently logged into the server (400), the electronic device (100) may determine that a user account switch is required.
[0237] According to one embodiment, the electronic device (100) can determine whether a switch to a user account is required based on the type of command corresponding to the voice.
[0238] For example, the electronic device (100) may determine that a switch to a user account is required based on whether the input voice corresponds to a command related to the user account (e.g., logging in to a user account, switching, etc.).
[0239] For example, the electronic device (100) may determine that a switch to a user account is necessary based on whether the input voice corresponds to a command to execute an application (e.g., OTT service, SNS, etc.) linked to the user account.
[0240] For example, the electronic device (100) may determine that a switch to a user account is necessary based on whether the input voice corresponds to a command (e.g., content search, content recommendation, speech recommendation, external device connection, etc.) that utilizes information about the user, such as usage history, viewing history, and preferred genre.
[0241] In operation S1410, if a switch is required for the currently logged-in user account, the electronic device (100) can switch the user account logged into the server (400). For example, the electronic device (100) can perform a logout for the first user account currently logged into the server (400). Once the logout for the first user account is completed, the electronic device (100) can log into the server (400) with a second user account corresponding to the user identification information corresponding to the voice.
[0242] Meanwhile, in operation S1411, the electronic device (100) can determine whether the input voice is a predetermined command related to the execution of an application linked to the user account when the application linked to the user account is running. Here, the predetermined command may include a command corresponding to switching the user account, a command corresponding to terminating the currently running application, etc.
[0243] For example, the electronic device (100) may determine that a switch to a user account is required based on whether the input voice corresponds to a command related to the user account (e.g., logging in to a user account, switching, etc.).
[0244] For example, the electronic device (100) may determine that termination of a first application linked to a currently running user account is required based on the input voice corresponding to a command to terminate the application.
[0245] For example, the electronic device (100) may determine that termination of the currently running first application is required based on the fact that the input voice corresponds to a command to run a second application that is different from the first application linked to the currently running user account.
[0246] For example, the electronic device (100) may determine that a switch to a user account or termination of a currently running application is not required based on whether the input voice corresponds to a command (e.g., content search, content recommendation, speech recommendation, external device connection, etc.) that utilizes information about the user, such as usage history, viewing history, and preferred genre.
[0247] In operation S1412, when a switch to a user account or termination of a currently running application is required, the electronic device (100) can switch the user account logged into the server (400) while the application linked to the user account is running.
[0248] According to one embodiment, the electronic device (100) may perform a logout for a first user account currently logged into the server (400). Once the logout for the first user account is completed, the electronic device (100) may log into the server (400) with a second user account corresponding to the user identification information corresponding to the voice. At this time, the electronic device (100) may terminate an application linked to the user account running in the foreground and then perform a switch for the user account logged into the server (400).
[0249] When the electronic device (100) completes logging into the server (400) using a second user account, the electronic device (100) can re-execute the application linked to the user account that was running prior to the user account switch. At this time, the electronic device (100) can utilize the application account corresponding to the second user account logged into the server (400). Here, the application account may be an account used to access services provided by the application.
[0250] An application account corresponding to a user account may be stored in the storage unit (140) and / or the server (400). For example, when the electronic device (100) completes logging in to the server (400) using a second user account, the electronic device (100) may obtain an application account corresponding to the second user account from the storage unit (140) and / or the server (400).
[0251] According to one embodiment, the electronic device (100), while currently logged in to the server (400) with a first user account, may acquire an application account corresponding to a second user account from the storage unit (140) and / or the server (400). At this time, the electronic device (100) may access a service provided by the first application with the application account corresponding to the second user account by using a predetermined API (application programming interface) corresponding to the first application linked to the currently running user account. That is, the electronic device (100), while currently logged in to the server (400) with the first user account, may access a service provided by the first application with the application account corresponding to the second user account.
[0252] Meanwhile, the electronic device (100) may log in to the server (400) with the second user account when access to the service provided by the first application using the application account corresponding to the second user account is completed. Alternatively, the electronic device (100) may log in to the server (400) with the second user account while accessing the service provided by the first application using the application account corresponding to the second user account. At this time, the electronic device (100) may log in to the server (400) with the second user account in the background. That is, the electronic device (100) may log in to the server (400) with the second user account while the screen corresponding to the first application is output through the display (180).
[0253] Meanwhile, according to one embodiment, when an application linked to a user account is running, the electronic device (100) may switch the user account logged into the server (400) based on whether the user who uttered the voice corresponds to a user account currently logged into the server (400), regardless of whether the input voice is a preset command.
[0254] The electronic device (100) may obtain information about the user in operation S1413. For example, the electronic device (100) may obtain the usage history stored in the storage unit (140) related to the user of the currently logged-in user account. For example, the electronic device (100) may obtain the name, nickname, icon, photo, age, gender, preferred genre, etc. stored in the server (400) related to the currently logged-in user account.
[0255] The electronic device (100) can perform an operation according to a voice in operation S1414. For example, the electronic device (100) can perform an operation according to a command corresponding to an input voice based on intent analysis information included in the result of processing the voice.
[0256] According to one embodiment, the electronic device (100) may provide a voice-related user interface (hereinafter, "voice UI") when a user logs into a server (400) with a user account based on the voice spoken by the user. For example, the voice UI may include numbers mapped to objects included on the screen.
[0257] Meanwhile, when a user logs into a server (400) with a user account based on an input other than voice, the electronic device (100) may provide a user interface (hereinafter, “general UI”) different from the voice UI. For example, when a user logs into the server (400) with a user account by entering the ID and password of the user account via a remote control device (200), the electronic device (100) may provide a general UI.
[0258] According to one embodiment, the electronic device (100) may provide a voice UI when executing an application or function based on a voice spoken by the user. On the other hand, the electronic device (100) may provide a general UI when executing an application or function based on an input other than voice.
[0259] Referring to drawing reference numeral 1501 of FIG. 15, the electronic device (100) can output a first home screen (1500). The first home screen (1500) may be a home screen output when logged in to the server (400) with a first user account.
[0260] When the electronic device (100) receives a voice spoken by a user, the electronic device (100) may output text (1510) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400). The electronic device (100) may determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command related to the user account.
[0261] Referring to drawing reference numeral 1502 of FIG. 15, a voice ID may be registered in the user account of the user who uttered the voice, and use of a voice ID-related function may be activated. Furthermore, the user who uttered the voice may not correspond to the first user account currently logged in. In this case, the electronic device (100) may output a first response (1520) that includes a user indicator and notifies that login has been performed with the user account of the user who uttered the voice.
[0262] The electronic device (100) can switch the user account logged into the server (400) from the currently logged-in first user account to the second user account of the user who spoke the voice. For example, the electronic device (100) can log out the currently logged-in first user account and then log into the server (400) using the second user account of the user who spoke the voice. According to one embodiment, the electronic device (100) can output a user switching screen related to the user switching while switching user accounts. For example, the user switching screen can include an object representing the first user account, an object representing the second user account, an object representing the switching of user accounts, and the like.
[0263] Referring to FIG. 16, the electronic device (100) may output a second home screen (1600) corresponding to the second user account when logged in with the second user account. The second home screen (1600) may be different from the first home screen (1500). That is, the style of the home screen output from the electronic device (100), such as the type, number, and arrangement of objects included in the home screen, may be different for each user account.
[0264] Referring to FIG. 17, the electronic device (100) can output a first home screen (1500). When the electronic device (100) receives a voice spoken by a user, the electronic device (100) can output text (1710) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400). The electronic device (100) can determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command to search for music content. In this case, the command to search for content may be a command that utilizes information about the user.
[0265] Referring to the drawing symbol 1801 of FIG. 18, the electronic device (100) may output a second response (1810) that notifies the search for music content without a user indicator. For example, if a voice ID is not registered in the user account of the user who spoke the voice, if the user who spoke the voice is not a user preset to use a voice ID, or if the user who spoke the voice corresponds to a first user account that is currently logged in, the electronic device (100) may output a second response (1810) that does not include a user indicator.
[0266] The electronic device (100) may obtain information about a user of a first user account and / or setting values corresponding to the first user account from the storage unit (140) and / or the server (400). Based on the information about the user and / or setting values obtained in correspondence with the first user account, the electronic device (100) may determine an application related to music content corresponding to the user of the first user account. For example, if a history of using a first content application related to music content is stored in the storage unit (140) of the electronic device (100) with respect to the user of the first user account, the electronic device (100) may determine to execute the first content application. For example, if a first content application is preset with an application setting value related to music content in the server (400) with respect to the user of the first user account, the electronic device (100) may determine to execute the first content application.
[0267] Referring to the reference numeral 1802 of FIG. 18, when the use of the voice ID related function is activated for the second user account of the user who has spoken the voice, if the second user account is different from the first user account that is currently logged in, the electronic device (100) may switch the user account logged in to the server (400) from the first user account that is currently logged in to the second user account of the user who has spoken the voice. In addition, the electronic device (100) may output a first response (1820) that notifies the search for music content, including a user indicator.
[0268] The electronic device (100) may obtain information about a user of a second user account and / or setting values corresponding to the second user account from the storage unit (140) and / or the server (400). Based on the information about the user and / or setting values obtained in response to the second user account, the electronic device (100) may determine an application related to music content corresponding to the user of the second user account. For example, if a history of using a second content application related to music content is stored in the storage unit (140) of the electronic device (100) with respect to the user of the second user account, the electronic device (100) may determine to execute the second content application. For example, if a second content application is preset with application setting values related to music content in the server (400) with respect to the user of the second user account, the electronic device (100) may determine to execute the second content application.
[0269] Referring to FIG. 19, the electronic device (100) can output a first home screen (1500). When the electronic device (100) receives a voice spoken by a user, the electronic device (100) can output text (1910) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400). The electronic device (100) can determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command to execute a specific OTT service application. Here, the OTT service application may be an application linked to a user account.
[0270] Referring to the reference numeral 2001 of FIG. 20, the electronic device (100) can execute a specific OTT service application. At this time, the electronic device (100) can access the specific OTT service using a first OTT service account corresponding to a first user account logged into the server (400). For example, if a voice ID is not registered in the user account of the user who has spoken the voice, if the user who has spoken the voice is not a user preset to use a voice ID, or if the user who has spoken the voice corresponds to a first user account that is currently logged in, the electronic device (100) can access the specific OTT service using the first OTT service account corresponding to the first user account.
[0271] The electronic device (100) can output an OTT service screen (2010) corresponding to the first OTT service account. The OTT service screen (2010) corresponding to the first OTT service account can include an object (2011) representing the first OTT service account, an object (2015) for content corresponding to the first OTT service account, and the like.
[0272] Meanwhile, the electronic device (100) may provide a voice UI (2013) through the OTT service screen (2010) based on the execution of a specific OTT service application according to the voice spoken by the user. For example, the voice UI (2013) may be a user interface that displays a number mapped to each piece of content corresponding to a first OTT service account. In this case, when a voice is input to select one of the numbers mapped to each piece of content, the electronic device (100) may output the content corresponding to the selected number through the display (180).
[0273] Referring to reference numeral 2002 of FIG. 20, the electronic device (100) can execute a specific OTT service application. At this time, the electronic device (100) can access the specific OTT service using a second OTT service account corresponding to a second user account different from a first user account. For example, if the use of a voice ID-related function is activated for the second user account of the user who has spoken the voice, and the second user account is different from the currently logged-in first user account, the electronic device (100) can switch the user account logged into the server (400) from the currently logged-in first user account to the second user account of the user who has spoken the voice. In addition, the electronic device (100) can access the specific OTT service using a second OTT service account corresponding to the second user account logged into the server (400).
[0274] The electronic device (100) can output an OTT service screen (2020) corresponding to a second OTT service account. The OTT service screen (2020) corresponding to the second OTT service account can include an object (2021) representing the second OTT service account, an object (2025) for content corresponding to the second OTT service account, etc.
[0275] Meanwhile, the electronic device (100) can log in to a server (400) with a user account according to the voice spoken by the user, and provide a voice UI (2023) through an OTT service screen (2020) based on the execution of a specific OTT service application.
[0276] Referring to FIG. 21, the electronic device (100) can output a first home screen (1500). When receiving a voice spoken by a user, the electronic device (100) can output text (2110) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400).
[0277] The electronic device (100) may determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command to connect a speaker among external devices. In this case, the command to connect the external device may be a command utilizing information about the user.
[0278] Referring to drawing symbol 2201 of FIG. 22, the electronic device (100) can output a pop-up screen (2210) for a speaker previously registered in the electronic device (100).
[0279] For example, if a voice ID is not registered in the user account of the user who made the voice call, or if the use of a voice ID related function is disabled for the user account of the user who made the voice call, the electronic device (100) may output a pop-up screen (2210) for a speaker already registered in the electronic device (100).
[0280] For example, if the use of the voice ID related function is activated for the user account of the user who has spoken the voice, and at least one of the speakers already registered in the electronic device (100) is set for public use, the electronic device (100) may output a pop-up screen (2210) for the speaker already registered in the electronic device (100). At this time, if the second user account of the user who has spoken the voice is different from the first user account that is currently logged in, the electronic device (100) may switch the user account logged into the server (400) from the first user account that is currently logged in to the second user account of the user who has spoken the voice.
[0281] Referring to the reference numeral 2202 of FIG. 22, the electronic device (100) may output a pop-up screen (2220) for a speaker corresponding to a user who has spoken a voice. For example, if the use of a voice ID related function is activated for the user account of the user who has spoken a voice, and all speakers already registered in the electronic device (100) are set for personal use, the electronic device (100) may output a pop-up screen (2220) for a speaker corresponding to the user who has spoken a voice. At this time, if the second user account of the user who has spoken a voice is different from the first user account that is currently logged in, the electronic device (100) may switch the user account logged into the server (400) from the first user account that is currently logged in to the second user account of the user who has spoken a voice.
[0282] Referring to FIG. 23, the electronic device (100) can output a first home screen (1500). When receiving a voice spoken by a user, the electronic device (100) can output text (2310) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400).
[0283] The electronic device (100) may determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command to connect the user's terminal among external devices. In this case, the command to connect the external device may be a command utilizing information about the user.
[0284] Referring to drawing symbol 2401 of FIG. 24, the electronic device (100) can output a pop-up screen (2410, 2420) for a terminal previously registered in the electronic device (100).
[0285] For example, if a voice ID is not registered in the user account of the user who made the voice call, or if the use of a voice ID related function is disabled for the user account of the user who made the voice call, the electronic device (100) may output a pop-up screen (2410, 2420) for a terminal already registered in the electronic device (100).
[0286] Referring to drawing reference numeral 2402 of FIG. 24, the electronic device (100) can output a screen (2400) corresponding to the terminal of the user who has spoken the voice. The screen (2400) corresponding to the terminal of the user who has spoken the voice may include an object (2430) corresponding to the screen of the terminal of the user who has spoken the voice.
[0287] For example, if the voice ID related function is activated for the user account of the user who has spoken the voice, and the user who has spoken the voice corresponds to the first user account that is currently logged in, the electronic device (100) can output a screen (2400) corresponding to the terminal of the user who has spoken the voice without switching the user account for the server (400).
[0288] For example, if the voice ID related function is activated for the second user account of the user who made the voice call, and the second user account is different from the first user account that is currently logged in, the electronic device (100) may switch the user account logged into the server (400) from the first user account that is currently logged in to the second user account of the user who made the voice call. In addition, after logging in to the second user account, the electronic device (100) may output a screen (2400) corresponding to the terminal of the user who made the voice call.
[0289] Referring to drawing reference numeral 2501 of FIG. 25, the electronic device (100) can output an OTT service screen (2500) corresponding to the OTT service when an OTT service application is executed. Here, the OTT service application may be an application linked to a user account.
[0290] When the electronic device (100) receives a voice spoken by a user while the OTT service screen (2500) is running, the electronic device (100) can output text (2510) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400).
[0291] The electronic device (100) can determine that the voice spoken by the user corresponds to a command for switching the user account based on the result of processing the voice received from the server (400).
[0292] Referring to drawing reference numeral 2502 of FIG. 25, the electronic device (100) may have a voice ID registered in the user account of the user who spoke the voice, and the use of voice ID-related functions may be activated. Furthermore, the user who spoke the voice may not correspond to the first user account currently logged in. In this case, the electronic device (100) may output a first response (2520) that includes a user indicator and notifies that the user account of the user who spoke the voice has been switched.
[0293] Referring to FIG. 26, the electronic device (100), while currently logged into the server (400) with a first user account, can access a service provided by an OTT service application using a second OTT service account corresponding to a second user account. For example, while the OTT service application is running, the electronic device (100) can access the OTT service using a second OTT service account using a predetermined API corresponding to the OTT service application.
[0294] When the electronic device (100) is connected to an OTT service using a second OTT service account, it can output an OTT service screen (2600) corresponding to the second OTT service account. The OTT service screen (2600) corresponding to the second OTT service account can include an object (2611) representing the second OTT service account, an object (2615) for content corresponding to the second OTT service account, and the like.
[0295] Meanwhile, referring to drawing reference numeral 2701 of FIG. 27, the electronic device (100) can switch the user account logged into the server (400) from the currently logged-in first user account to the second user account of the user who uttered the voice. For example, the electronic device (100) can log out the currently logged-in first user account and then log into the server (400) using the second user account of the user who uttered the voice.
[0296] The electronic device (100) may output a user switching screen (2700) related to user switching while switching user accounts. The user switching screen (2700) may include an object (2710) representing a first user account, an object (2720) representing a second user account, an object (2730) representing switching user accounts, and the like.
[0297] Referring to reference numeral 2702 of FIG. 27, when the electronic device (100) completes logging into the server (400) with a second user account, the electronic device (100) can re-execute the OTT service application. At this time, the electronic device (100) can access the OTT service with a second OTT service account corresponding to the second user account logged into the server (400). When accessing the OTT service with the second OTT service account, the electronic device (100) can output an OTT service screen (2600) corresponding to the second OTT service account.
[0298] Meanwhile, referring to drawing reference numeral 2801 of FIG. 28, the electronic device (100) can output an OTT service screen (2500) corresponding to the OTT service when the OTT service application is executed.
[0299] When the electronic device (100) receives a voice spoken by a user while the OTT service screen (2500) is displayed, the electronic device (100) can output text (2810) corresponding to the voice spoken by the user based on the result of processing the voice received from the server (400).
[0300] The electronic device (100) may determine, based on the result of processing the voice received from the server (400), that the voice spoken by the user corresponds to a command for recommending video content. In this case, if the voice spoken by the user corresponds to a command for recommending video content while the OTT service application is running, the electronic device (100) may determine that no user account switching or termination of the currently running application is required.
[0301] Referring to drawing reference numeral 2802 of FIG. 28, the electronic device (100) may output a second response (2820) that informs of a recommendation of video content without a user indicator. At this time, since no switching of user accounts or termination of currently running applications is required, no switching of user accounts for the server (400) or switching of accounts for OTT services may occur.
[0302] As described above, according to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.
[0303] Additionally, according to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.
[0304] Additionally, according to at least one embodiment of the present disclosure, the user can log in to an account identified based on the user's voice.
[0305] Additionally, according to at least one embodiment of the present disclosure, an operation optimized for an account of a user identified based on the user's voice can be performed.
[0306] Additionally, according to at least one embodiment of the present disclosure, a user interface optimized for an input method used by a user can be provided.
[0307] Additionally, according to at least one embodiment of the present disclosure, when the application is running, it is possible to switch to an account of a user identified based on the user's voice.
[0308] Referring to FIGS. 1 to 28, an electronic device (100) according to one aspect of the present disclosure includes a network interface unit (135) for communicating with a server (400); a user input interface unit (150) for transmitting a signal corresponding to an input; and a control unit (170). When a voice signal corresponding to a voice input is received while a predetermined application linked to a user account is running, the control unit (170) determines whether a user who has uttered the voice input corresponds to a first user account logged into the server (400), and when the user who has uttered the voice input does not correspond to the first user account, acquires an application account corresponding to the user who has uttered the voice input based on a second user account corresponding to the user who has uttered the voice input, and accesses a predetermined service provided by the predetermined application based on the acquired application account through the network interface unit (135).
[0309] In addition, according to one aspect of the present disclosure, if the user who uttered the voice input does not correspond to the first user account, the control unit (170) determines whether the command corresponding to the voice input is a predetermined command set in advance in relation to the execution of the predetermined application, and if the command corresponding to the voice input is the predetermined command, the control unit can obtain an application account corresponding to the user who uttered the voice input based on the second user account.
[0310] In addition, according to one aspect of the present disclosure, the control unit (170) may include at least one of a command related to the user account and a command related to termination of the specified application.
[0311] In addition, according to one aspect of the present disclosure, if the command corresponding to the voice input is not the predetermined command, the control unit (170) can maintain the login of the first user account for the server (400) and, while the predetermined application is running, perform an operation corresponding to the voice input related to the first user account.
[0312] In addition, according to one aspect of the present disclosure, the control unit (170) can, while the first user account is logged in to the server (400), access the predetermined service using the acquired application account using an application programming interface (API) corresponding to the predetermined application, and log in to the server (400) using the second user account in response to access to the predetermined service using the acquired application account.
[0313] In addition, according to one aspect of the present disclosure, the server (400) stores an application account corresponding to a predetermined user account, and the control unit (170) can obtain an application account corresponding to the second user account from the server (400) while the first user account is logged in to the server (400).
[0314] In addition, according to one aspect of the present disclosure, the control unit (170) can log in to the server (400) with the second user account, obtain an application account corresponding to the second user account while the second user account is logged in to the server (400), and re-execute the predetermined application to access the predetermined service with the obtained application account.
[0315] In addition, according to one aspect of the present disclosure, a display (180) is further included, and the control unit (170) can output a screen corresponding to the switching of the user account logged into the server (400) through the display (180) after the execution of the predetermined application is terminated and the second user account is logged in to the server (400) from the termination of the execution of the predetermined application until the re-execution of the predetermined application.
[0316] In addition, according to one aspect of the present disclosure, the control unit (170) further includes a memory (140) that stores a user list including at least one user identification information corresponding to a user account that has a history of logging into the server (400), and the control unit (170) transmits the user list to the server (400) together with the data including the voice signal, and compares the first user identification information corresponding to the first user account with the second user identification information corresponding to the voice input received from the server (400), thereby determining whether the user who uttered the voice input corresponds to the first user account.
[0317] A system (10) according to one aspect of the present disclosure includes an electronic device (100) and a server (400), wherein the electronic device (100) transmits data including a voice signal corresponding to a voice input uttered by a user to the server (400) while a predetermined application linked to a user account is running, and determines whether the user who uttered the voice input corresponds to a first user account logged into the server (400) based on a result of processing the voice signal received from the server (400), and if the user who uttered the voice input does not correspond to the first user account, acquires an application account corresponding to the user who uttered the voice input based on a second user account corresponding to the user who uttered the voice input, and accesses a predetermined service provided by the predetermined application through the network interface unit (135) based on the acquired application account, and the server (400) generates identification information for the voice signal included in the data received from the electronic device (100), and From the identification information mapped to the user identification information corresponding to the user account stored in the database (490) of the server (400), predetermined identification information corresponding to the identification information for the voice signal can be determined, and the result of processing the voice signal including the predetermined user identification information mapped to the predetermined identification information can be transmitted to the electronic device (100).
[0318] In addition, according to one aspect of the present disclosure, if the user who uttered the voice input does not correspond to the first user account, the electronic device (100) determines whether the command corresponding to the voice input is a predetermined command set in advance in relation to the execution of the predetermined application, and if the command corresponding to the voice input is the predetermined command, the electronic device can obtain an application account corresponding to the user who uttered the voice input based on the second user account.
[0319] In addition, according to one aspect of the present disclosure, the electronic device (100) can, while the first user account is logged in to the server (400), access the predetermined service with the acquired application account using an application programming interface (API) corresponding to the predetermined application, and log in to the server (400) with the second user account in response to access to the predetermined service using the acquired application account.
[0320] In addition, according to one aspect of the present disclosure, the electronic device (100) may request transmission of an application account corresponding to the second user account while the first user account is logged in to the server (400), and the server (400) may transmit the application account corresponding to the second user account to the electronic device (100).
[0321] In addition, according to one aspect of the present disclosure, the electronic device (100) can log in to the server (400) with the second user account, obtain an application account corresponding to the second user account while the second user account is logged in to the server (400), and re-execute the predetermined application to access the predetermined service with the obtained application account.
[0322] In addition, according to one aspect of the present disclosure, the electronic device (100) may log in to the server (400) with the second user account after terminating execution of the predetermined application, and output a screen corresponding to the switching of the user account logged in to the server (400) through the display (180) from the termination of execution of the predetermined application to the re-execution of the predetermined application.
[0323] The attached drawings are only intended to facilitate understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure.
[0324] Meanwhile, the operating method of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices that store data that can be read by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across network-connected computer systems, so that the processor-readable code can be stored and executed in a distributed manner.
[0325] In addition, although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. Network interface section that communicates with the server; A user input interface unit that transmits a signal corresponding to the input; and Including a control unit, The above control unit, When a voice signal corresponding to a voice input is received while a predetermined application linked to a user account is running, it is determined whether the user who uttered the voice input corresponds to a first user account logged in to the server, If the user who uttered the voice input does not correspond to the first user account, an application account corresponding to the user who uttered the voice input is obtained based on a second user account corresponding to the user who uttered the voice input, An electronic device characterized in that it accesses a predetermined service provided by a predetermined application through the network interface unit based on the acquired application account.
2. In paragraph 1, The above control unit, If the user who uttered the above voice input does not correspond to the first user account, it is determined whether the command corresponding to the voice input is a preset command related to the execution of the above specified application, An electronic device characterized in that, if the command corresponding to the voice input is the predetermined command, an application account corresponding to the user who uttered the voice input is acquired based on the second user account.
3. In paragraph 1, The above prescribed order is, An electronic device characterized in that it includes at least one of a command related to the user account and a command related to terminating the specified application.
4. In paragraph 2, The above control unit, If the command corresponding to the voice input is not the predetermined command, the login of the first user account to the server is maintained, An electronic device characterized in that, while the above-mentioned application is running, it performs an action corresponding to the voice input related to the first user account.
5. In paragraph 1, The above control unit, While the first user account is logged in to the server, the API (application programming interface) corresponding to the predetermined application is used to access the predetermined service with the acquired application account, An electronic device characterized in that, in response to access to the specified service using the acquired application account, the device logs in to the server using the second user account.
6. In paragraph 5, The above server stores an application account corresponding to a given user account, The above control unit, An electronic device characterized in that, while the first user account is logged in to the server, an application account corresponding to the second user account is acquired from the server.
7. In paragraph 1, The above control unit, Log in to the server with the second user account, While the second user account is logged in to the server, an application account corresponding to the second user account is acquired, An electronic device characterized in that it re-executes the above-mentioned application and accesses the above-mentioned service with the above-mentioned acquired application account.
8. In paragraph 7, Including more displays, The above control unit, After terminating the execution of the above-mentioned application, log in to the server with the second user account, An electronic device characterized in that it outputs a screen corresponding to the switching of user accounts logged into the server through the display from the end of execution of the above-mentioned application to the re-execution of the above-mentioned application.
9. In paragraph 1, Further comprising a memory storing a user list including at least one user identification information corresponding to a user account that has logged in to the server, The above control unit, Transmitting the above user list to the server together with data including the above voice signal, An electronic device characterized in that it compares first user identification information corresponding to the first user account with second user identification information corresponding to the voice input received from the server to determine whether the user who uttered the voice input corresponds to the first user account.
10. In a system including electronic devices and servers, The above electronic device, When a specific application linked to a user account is running, data containing a voice signal corresponding to the voice input spoken by the user is transmitted to the server, Based on the result of processing the voice signal received from the server, it is determined whether the user who uttered the voice input corresponds to a first user account logged in to the server, If the user who uttered the voice input does not correspond to the first user account, an application account corresponding to the user who uttered the voice input is obtained based on a second user account corresponding to the user who uttered the voice input, Based on the application account acquired above, access to a predetermined service provided by the predetermined application through the network interface unit, The above server, Generate identification information for the voice signal included in the data received from the electronic device, Determine predetermined identification information corresponding to the identification information for the voice signal from the identification information mapped to the user identification information corresponding to the user account stored in the database of the above server, A system characterized in that it transmits the result of processing the voice signal including the predetermined user identification information mapped to the predetermined identification information to the electronic device.
11. In paragraph 10, The above electronic device, If the user who uttered the above voice input does not correspond to the first user account, it is determined whether the command corresponding to the voice input is a preset command related to the execution of the above specified application, A system characterized in that, if the command corresponding to the voice input is the predetermined command, an application account corresponding to the user who uttered the voice input is acquired based on the second user account.
12. In paragraph 10, The above electronic device, While the first user account is logged in to the server, the API (application programming interface) corresponding to the predetermined application is used to access the predetermined service with the acquired application account, A system characterized in that, in response to access to the specified service using the acquired application account, the system logs in to the server using the second user account.
13. In paragraph 11, The above electronic device, When the first user account is logged in to the server, a request is made to transfer an application account corresponding to the second user account, The above server, A system characterized in that it transmits an application account corresponding to the second user account to the electronic device.
14. In paragraph 10, The above electronic device, Log in to the server with the second user account, While the second user account is logged in to the server, an application account corresponding to the second user account is acquired, A system characterized in that the system re-executes the above-mentioned application and accesses the above-mentioned service using the acquired application account.
15. In paragraph 14, The above electronic device, After terminating the execution of the above-mentioned application, log in to the server with the second user account, A system characterized in that it outputs a screen corresponding to the switching of user accounts logged into the server through a display from the end of execution of the above-mentioned predetermined application to the re-execution of the above-mentioned predetermined application.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing system
JP7342862B2
Conveyor system of automation line
KR1020220046200A
Electronic apparatus and method for therof
KR102429582B1
Method for authenticating user based on voice command and electronic dvice thereof
KR102483834B1
electronic apparatus, method for contolling mobile apparatus by electronic apparatus and computer-readable recording medium
KR102582332B1