Electronic device, server, and system comprising same

The system addresses the inconvenience and security issues of manual account entry by using voice recognition for user authentication, providing secure and seamless login through voice-based identification.

WO2025170132A1PCT designated stage Publication Date: 2025-08-14LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/011228
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2024-07-31
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing voice recognition technologies require manual entry of account information for user authentication, leading to inconvenience and security risks, especially in multi-user environments.

Method used

An electronic device and server system that registers and identifies users based on their voice, allowing seamless login and maintaining continuity of identification information through voice recognition processes.

Benefits of technology

Enables secure and convenient user authentication by voice, eliminating the need for manual account entry and maintaining user identity across sessions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024011228_14082025_PF_FP_ABST
    Figure KR2024011228_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an electronic device, a server, and a system comprising same. An electronic device according to an embodiment of the present disclosure may include: a display; an external device interface unit for communicating with a remote control device; a network interface unit for communicating with a server; a user input interface unit for transmitting a signal corresponding to a user input; and a control unit, wherein the control unit outputs, while performing a process of registering identification information related to a voice with respect to a user account used to log in to the server, pre-configured text related to the registration of the identification information through the display, when a voice signal corresponding to the pre-configured text is received through the user input interface unit, transmits data including the voice signal to the server, completes the process of registering the identification information on the basis of processing of the voice signal corresponding to the pre-configured text by the server, and suspends the process of registering the identification information when a predetermined input related to voice recognition is received from the remote control device while performing the process of registering the identification information.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices, servers and systems including the same

[0001] The present disclosure relates to an electronic device, a server, and a system including the same, and more particularly, to an electronic device, a server, and a system including the same utilizing voice recognition technology.

[0002] With recent technological advancements, research on voice recognition technology, which processes speech, is actively underway. In particular, research on voice recognition technology, which originated with smartphones, is now being conducted widely in various fields related to user convenience, such as home appliances used in homes and offices, as well as in vehicles.

[0003] Voice recognition technology is commonly used when a user controls an electronic device using their voice. For example, when a user utters a command to control an electronic device, the electronic device can directly recognize and process the user's voice and operate according to the corresponding command. Alternatively, the device can transmit the voice to a voice processing server and then operate according to the corresponding command received from the server.

[0004] Meanwhile, the services and functions provided through electronic devices are becoming increasingly diverse. Furthermore, users register accounts for various services and then log in with their registered accounts to access them. In this case, service providers utilize the user information managed for each account to provide optimal features or information tailored to the user.

[0005] Traditionally, when attempting to log in to a service, users would manually enter their account information, such as their account ID (identification number) and / or password. However, this presents a significant inconvenience, requiring users to manually enter their account information for each service. Furthermore, maintaining a logged-in status to eliminate the inconvenience of entering account information can lead to security issues, such as third parties accessing the user's account information. Furthermore, when multiple users share a single electronic device, this presents a problem: each user must enter their own account information to log in each time they use the service.

[0006] The present disclosure aims to solve the above-mentioned and other problems.

[0007] Another object is to provide an electronic device, a server and a system including the same, which can register identification information about a user's voice to a user's account.

[0008] Another object is to provide an electronic device, a server and a system including the same that can identify a user based on the user's voice.

[0009] Another object is to provide an electronic device, a server and a system including the same, which can log in to an account of a user identified based on the user's voice.

[0010] Another object is to provide an electronic device, a server and a system including the same, which can maintain continuity of identification information for a user who has previously attempted to register, during the process of registering identification information for a user's voice to the user's account.

[0011] In order to achieve the above object, an electronic device according to one embodiment of the present disclosure includes: a display; an external device interface unit communicating with a remote control device; a network interface unit communicating with a server; a user input interface unit transmitting a signal corresponding to a user input; and a control unit, wherein the control unit, while performing a process of registering identification information related to voice for a user account logged in to the server, outputs a preset text related to registration of the identification information through the display, and when a voice signal corresponding to the preset text is received through the user input interface unit, transmits data including the voice signal to the server, and completes the process of registering the identification information based on the server's processing of the voice signal corresponding to the preset text, and while performing the process of registering the identification information, when a predetermined input related to voice recognition is received from the remote control device, the process of registering the identification information can be stopped.

[0012] In order to achieve the above object, according to one embodiment of the present disclosure, a server includes a communication unit communicating with an electronic device; a database; and a controller, wherein the controller, while performing a process of registering voice-related identification information for a user account logged into the server, converts a voice signal included in data received from the electronic device into text, generates identification information for the voice signal based on a correspondence between the converted text and a text preset in relation to registration of the identification information, maps the generated identification information to user identification information corresponding to the user account, and stores the mapped identification information in the database, and while performing the process of registering the identification information, when a notification of interruption of the process of registering the identification information is received from the electronic device, the identification information mapped to the user identification information can be maintained or deleted.

[0013] In order to achieve the above object, a system according to one embodiment of the present disclosure includes an electronic device and a server, wherein the electronic device, while performing a process of registering identification information related to voice for a user account logged into the server, outputs a preset text in relation to the registration of the identification information, and when a voice signal corresponding to the preset text is received, transmits data including the voice signal to the server, and completes the process of registering the identification information based on the processing of the voice signal corresponding to the preset text by the server, and while performing the process of registering the identification information, when a predetermined input related to voice recognition is received from a remote control device, stops the process of registering the identification information, and the server converts the voice signal included in data received from the electronic device into text, and based on a correspondence between the converted text and the preset text, generates identification information for the voice signal, and stores the generated identification information in a database by mapping it to user identification information corresponding to the user account, and while performing the process of registering the identification information, when a notification of the interruption of the process of registering the identification information is received from the electronic device, stores the mapped identification information in a database. The above identification information may be maintained or deleted.

[0014] The effects of the electronic device, server and system including the same according to the present disclosure are described as follows.

[0015] According to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.

[0016] According to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.

[0017] According to at least one embodiment of the present disclosure, a user can log in to an account identified based on the user's voice.

[0018] According to at least one embodiment of the present disclosure, in the process of registering identification information for a user's voice to a user's account, continuity of identification information for a user who has previously attempted to register can be maintained.

[0019] Further scope of the applicability of the present disclosure will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present disclosure will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present disclosure, are given by way of example only.

[0020] FIG. 1 is a diagram illustrating a system according to one embodiment of the present disclosure.

[0021] Figure 2 is an internal block diagram of the electronic device of Figure 1.

[0022] Figure 3 is a drawing referenced in the description of the server of Figure 1.

[0023] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.

[0024] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.

[0025] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an electronic device according to one embodiment of the present disclosure.

[0026] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.

[0027] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.

[0028] FIGS. 9 to 16 are drawings for reference in explaining a process for registering identification information for a user's voice to a user's account according to one embodiment of the present disclosure.

[0029] FIGS. 17A and 17B are flowcharts of a method of operating an electronic device according to another embodiment of the present disclosure.

[0030] FIGS. 18 to 20 are flowcharts of a method of operating a system according to various embodiments of the present disclosure.

[0031] Hereinafter, the present disclosure will be described in detail with reference to the drawings. In the drawings, portions irrelevant to the description are omitted to clearly and concisely describe the present disclosure, and the same reference numerals are used for identical or extremely similar portions throughout the specification.

[0032] The suffixes "module" and "part" used in the following description are given solely for the convenience of writing this specification and do not impart any particularly significant meaning or role to the components themselves. Therefore, the terms "module" and "part" may be used interchangeably.

[0033] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0034] Additionally, while terms such as "first" and "second" may be used in this specification to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another.

[0035] FIG. 1 is a diagram illustrating a system according to various embodiments of the present invention.

[0036] Referring to FIG. 1, the system (10) may include an electronic device (100) and / or a server (400).

[0037] The electronic device (100) can transmit and receive data to and from at least one server (400). For example, the electronic device (100) can transmit and receive data to and from at least one server (400) via a network (300) such as the Internet.

[0038] According to one embodiment, at least one server (400) may include a server that performs voice recognition, a server that processes data using a super-giant artificial intelligence model, a server that provides content, etc.

[0039] The electronic device (100) may include a video display device (100a), an air conditioner (100b), a refrigerator (100c), an air purifier (100d), a washing machine (100e), a vehicle (100f), etc. In the present disclosure, the electronic device (100) is described as an example of a video display device (100a), but the present invention is not limited thereto.

[0040] The image display device (100a) may be a device that processes and outputs an image. The image display device (100a) is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.

[0041] The video display device (100a) can receive a broadcast signal, process it, and output the processed broadcast image. When the video display device (100a) receives a broadcast signal, the video display device (100a) may correspond to a broadcast receiving device.

[0042] The video display device (100a) can receive broadcast signals wirelessly via an antenna, or can receive broadcast signals wired via a cable. For example, the video display device (100a) can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.

[0043] Figure 2 is an internal block diagram of the electronic device of Figure 1.

[0044] Referring to FIG. 2, the electronic device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).

[0045] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulator unit (120).

[0046] Meanwhile, unlike the drawing, the electronic device (100) may include only the broadcast receiving unit (105) and the external device interface unit (130) among the broadcast receiving unit (105), the external device interface unit (130), and the network interface unit (135). That is, the electronic device (100) may not include the network interface unit (135).

[0047] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.

[0048] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).

[0049] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.

[0050] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.

[0051] The demodulation unit (120) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (110).

[0052] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.

[0053] The stream signal output from the demodulation unit (120) can be input to the control unit (170). The control unit (170) can output an image through the display (180) and output an audio through the audio output unit (185) after performing demultiplexing, image / audio signal processing, etc.

[0054] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).

[0055] The external device interface unit (130) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.

[0056] In addition, the external device interface unit (130) can establish a communication network with various remote control devices (200) to receive a control signal related to the operation of the electronic device (100) from the remote control device (200) or transmit data related to the operation of the electronic device (100) to the remote control device (200).

[0057] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit can include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).

[0058] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, in mirroring mode, the external device interface unit (130) may receive device information, running application information, application images, etc. from the mobile terminal.

[0059] The external device interface unit (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.

[0060] The network interface unit (135) can provide an interface for connecting the electronic device (100) to a wired / wireless network including the Internet.

[0061] The network interface unit (135) may include a communication module (not shown) for connection to a wired / wireless network. For example, the network interface unit (135) may include a communication module for WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), etc.

[0062] The network interface unit (135) can transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.

[0063] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and information related thereto provided by a content provider or network provider via a network.

[0064] The network interface unit (135) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.

[0065] The network interface unit (135) can select and receive a desired application from among applications open to the public through a network.

[0066] The storage unit (140) may store programs for signal processing and control within the control unit (170), or may store processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored application programs upon request from the control unit (170).

[0067] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (170).

[0068] The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).

[0069] The storage unit (140) can store information about a specific broadcast channel through a channel memory function such as a channel map.

[0070] Although the storage unit (140) of FIG. 2 is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).

[0071] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.

[0072] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.

[0073] For example, a user input signal such as power on / off, channel selection, screen setting, etc. may be transmitted / received from a remote control device (200), a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. may be transmitted to the control unit (170), a user input signal input from a sensor unit (not shown) that senses a user's gesture may be transmitted to the control unit (170), or a signal from the control unit (170) may be transmitted to the sensor unit.

[0074] The input unit (160) may be provided on one side of the main body of the electronic device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.

[0075] The input unit (160) can receive various user commands related to the operation of the electronic device (100) and transmit a control signal corresponding to the input command to the control unit (170).

[0076] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.

[0077] The control unit (170) may include at least one processor, and may control the overall operation of the electronic device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.

[0078] The control unit (170) can demultiplex a stream input through the tuner unit (110), the demodulator unit (120), the external device interface unit (130), or the network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.

[0079] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).

[0080] The display (180) may include a display panel (not shown) having a plurality of pixels.

[0081] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display (180) may convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (170) to generate driving signals for the plurality of pixels.

[0082] The display (180) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display, etc., and may also be a 3D display. The 3D display (180) can be divided into a glasses-free type and a glasses type.

[0083] Meanwhile, the display (180) is configured as a touch screen and can be used as an input device in addition to an output device.

[0084] The audio output unit (185) receives a signal processed by the control unit (170) and outputs it as voice.

[0085] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (170) can also be input to an external output device through the external device interface unit (130).

[0086] The voice signal processed in the control unit (170) can be output as sound to the audio output unit (185). In addition, the voice signal processed in the control unit (170) can be input to an external output device through the external device interface unit (130).

[0087] Although not shown in FIG. 2, the control unit (170) may include a demultiplexing unit, an image processing unit, etc.

[0088] In addition, the control unit (170) can control the overall operation within the electronic device (100). For example, the control unit (170) can control the tuner unit (110) to select (tune) a broadcast corresponding to a channel selected by the user or a previously stored channel.

[0089] In addition, the control unit (170) can control the electronic device (100) by a user command or internal program input through the user input interface unit (150).

[0090] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a moving image, and may be a 2D image or a 3D image.

[0091] Meanwhile, the control unit (170) can cause a predetermined 2D object to be displayed within an image displayed on the display (180). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.

[0092] Meanwhile, the electronic device (100) may further include a camera (not shown). The camera can capture images of the user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the electronic device (100) above the display (180) or positioned separately. Image information captured by the camera can be input to the control unit (170).

[0093] The control unit (170) can recognize the user's location based on the image captured by the camera. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the electronic device (100). In addition, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.

[0094] The control unit (170) can detect the user's gesture based on an image captured from the camera unit, a signal detected from the sensor unit, or a combination thereof.

[0095] The power supply unit (190) can supply power to the entire electronic device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a system on chip (SOC), a display (180) for displaying images, and an audio output unit (185) for outputting audio.

[0096] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.

[0097] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (150) and display or output the same as voice on the remote control device (200).

[0098] Meanwhile, the electronic device (100) described above may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.

[0099] Meanwhile, the block diagram of the electronic device (100) illustrated in FIG. 2 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the electronic device (100) actually implemented.

[0100] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0101] Figure 3 is a drawing referenced in the description of the server of Figure 1.

[0102] Referring to FIG. 3, the server (400) may include a relay server (410), an STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), a user identification server (440), and / or an account server (450). In the present disclosure, the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) are described as being distinct from each other, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) may be configured as one server.

[0103] The relay server (410) can communicate with the electronic device (100). The relay server (410) can transfer data between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100). The relay server (410) can store at least a portion of the data transferred between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100).

[0104] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the electronic device (100) via the relay server (410). The STT server (420) may also be referred to as an ASR (Automatic Speech Recognition) server.

[0105] The STT server (420) can improve the accuracy of speech-to-text conversion using a language model. The language model can refer to a model that can calculate the probability of a sentence or the probability of a subsequent word given previous words. For example, the language model can include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. In other words, the STT server (420) can use the language model to determine whether text data converted from speech data has been appropriately converted, thereby improving the accuracy of conversion into text data.

[0106] The NLP server (430) can receive text data. Based on the received text data, the NLP server (430) can perform intent analysis on the text data. The NLP server (430) can transmit intent analysis information indicating the results of the intent analysis to the electronic device (100) via the relay server (410).

[0107] According to one embodiment, the NLP server (430) may sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate intent analysis information. The morphological analysis step is a step of classifying text data corresponding to speech uttered by a user into morphemes, which are the smallest units having meaning, and determining which part of speech each classified morpheme has. The syntax analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and to determine what kind of relationship exists between each of the classified phrases. Through the syntax analysis step, the subject, object, and modifiers of the speech uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the speech uttered by the user using the results of the syntax analysis step. Specifically, the speech act analysis step is a step of determining the intent of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.

[0108] The user identification server (440) can receive voice data. Based on the voice data, the user identification server (440) can extract voice features. Here, the voice features may include the voice waveform, the voice frequency band, the voice power spectrum, etc. Extracting voice features will be described below with reference to FIGS. 4 and 5.

[0109] The user identification server (440) can obtain a feature vector of a voice from the features of the voice. The user identification server (440) can obtain a feature vector of a voice from the features of the voice based on a linear predictive coefficient, a cepstrum, a Mel Frequency Cepstral Coefficient (MFCC), and filter bank energy.

[0110] The user identification server (440) can determine the similarity between a plurality of feature vectors. The user identification server (440) can determine the similarity between a plurality of feature vectors using cosine similarity, Euclidean similarity, etc. In the present disclosure, the similarity between a first voice input and a second voice input is calculated based on cosine similarity, but is not limited thereto. For example, a first vector corresponding to a first text and a second vector corresponding to a second text may be generated. In this case, the cosine similarity between the first vector and the second vector may be calculated based on the following mathematical expression 1.

[0111]

[0112] Here, is the dot product of two vectors, and can mean the magnitude of two vectors. That is, cosine similarity can be calculated as the value obtained by dividing the inner product of two vectors by the product of the magnitudes of each vector. Cosine similarity can range from -1 to 1, and the closer it is to 1, the more similar the two vectors can be judged to be.

[0113] The user identification server (440) can determine whether the users who uttered the voice are the same user based on the similarity between multiple feature vectors. For example, if the similarity between the first feature vector corresponding to the first voice input and the second feature vector corresponding to the second voice input is above a predetermined standard, the user identification server (440) can determine that the user who uttered the first voice input and the user who uttered the second voice input are the same user.

[0114] According to one embodiment, the user identification server (440) may obtain a vector by processing a feature vector of a voice using an algorithm such as a GMM (Gaussian mixture model) supervector, i-vector, d-vector, x-vector, etc. The user identification server (440) may determine whether the user who uttered the voice is the same based on the similarity between a first vector processed corresponding to a first feature vector and a second vector processed corresponding to a second feature vector.

[0115] The user identification server (440) can store voice data. The user identification server (440) can store data regarding the voice print (hereinafter, voice print information). Here, the voice print information can include a voice feature vector and / or a vector obtained by processing the voice feature vector.

[0116] The user identification server (440) can store a database for voices. The database for voices can include unique identification information (hereinafter, device identification information) corresponding to an electronic device (100), unique identification information (hereinafter, user identification information) corresponding to a user account, voice data mapped to user identification information, voice print information mapped to user identification information, etc.

[0117] Device identification information, user identification information, voice data, voice print information, etc. included in the voice database may be stored in a user identification server (440) in association with each other. For example, at least one piece of device identification information, multiple voice data, and / or multiple voice print information may be mapped to the user identification information. In other words, it may be interpreted that the device identification information, voice data, voice print information, etc. are mapped to a user account and stored in the user identification server (440). In the present disclosure, an example in which multiple voice data and multiple voice print information are all mapped to the user identification information included in the voice database will be described.

[0118] The user identification server (440) can update the voiceprint information contained in the voice database based on the voice data contained in the voice database. For example, the user identification server (440) can generate voiceprint information corresponding to the voice data contained in the voice database using an algorithm different from the algorithm previously used. In this case, the user identification server (440) can change the voiceprint information contained in the voice database to the newly generated voiceprint information.

[0119] The account server (450) can manage data regarding user accounts. The account server (450) can manage user account IDs, passwords, user identification information, device identification information mapped to the user account, and whether or not to agree to terms and conditions related to various functions.

[0120] The account server (450) can store a database of user accounts. The database of user accounts may include the user account ID, password, user identification information, device identification information mapped to the user account, the registration date and time of the user account, whether or not the user has agreed to terms and conditions related to various functions, and the date and time of agreement to the terms and conditions.

[0121] The account server (450) can communicate with the electronic device (100). For example, the account server (450) can create and register a user account based on data from the electronic device (100). For example, the account server (450) can approve login to a user account based on an ID and password received from the electronic device (100).

[0122] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.

[0123] Referring to FIG. 4, the server (400) may include a preprocessing unit (460), a controller (470), a communication unit (480), and / or a database (490).

[0124] The preprocessing unit (460) can preprocess voice received through the communication unit (480) or voice stored in the database (490).

[0125] The preprocessing unit (460) may be implemented as a separate chip from the controller (470) or as a chip included in the controller (470).

[0126] The preprocessing unit (460) can receive a voice signal (spoken by a user) and filter out noise signals from the voice signal before converting the received voice signal into text data.

[0127] When a preprocessing unit (460) is provided in the electronic device (100), it can recognize a trigger word for activating voice recognition of the electronic device (100). The preprocessing unit (460) converts the trigger word received through the user input interface unit (150) into text data, and if the converted text data corresponds to a previously stored trigger word, it can be determined that the trigger word has been recognized.

[0128] The preprocessing unit (460) can convert the noise-removed voice signal into a power spectrum.

[0129] Power spectrum can be a parameter that indicates which frequency components are included in the waveform of a temporally varying voice signal and at what magnitude.

[0130] The power spectrum shows the distribution of the squared amplitude values ​​of the waveform of a voice signal according to frequency. This is explained with reference to Figure 5.

[0131] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.

[0132] Referring to FIG. 5, a voice signal (510) is illustrated. The voice signal (460) may be received from an external device or may be a signal pre-stored in memory (170).

[0133] The x-axis of the voice signal (510) represents time, and the y-axis can represent the amplitude size.

[0134] The power spectrum processing unit (463) can convert a voice signal (510) whose x-axis is the time axis into a power spectrum (520) whose x-axis is the frequency axis. The power spectrum processing unit (463) can convert the voice signal (510) into a power spectrum (520) using a fast Fourier transform (FFT). The x-axis of the power spectrum (520) represents frequency, and the y-axis represents the square value of the amplitude.

[0135] Referring again to FIG. 6, the functions of the preprocessing unit (460) and the controller (470) described in FIG. 6 can also be performed in the NLP server (430).

[0136] The preprocessing unit (460) may include a wave processing unit (461), a frequency processing unit (462), a power spectrum processing unit (463), a speech to text (STT) conversion unit (464), etc.

[0137] The wave processing unit (461) can extract the waveform of the voice.

[0138] The frequency processing unit (462) can extract the frequency band of the voice.

[0139] The power spectrum processing unit (463) can extract the power spectrum of the voice.

[0140] Power spectrum can be a parameter that indicates which frequency components are included in a given temporally varying waveform and at what magnitude.

[0141] The speech-to-text (STT) conversion unit (464) can convert speech into text. The speech-to-text conversion unit (464) can convert speech in a specific language into text in that language.

[0142] The controller (470) can control the overall operation of the server (400). The controller (470) can include a voice analysis unit (471), a text analysis unit (472), a feature clustering unit (473), a text mapping unit (474), and / or a voice synthesis unit (475).

[0143] The voice analysis unit (471) can extract voice characteristic information by using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed by the preprocessing unit (460). The voice characteristic information may include one or more of the speaker's gender information, the speaker's voice (or tone), pitch, the speaker's speech style, the speaker's speaking speed, and the speaker's emotion. In addition, the voice characteristic information may further include the speaker's voice tone.

[0144] The text analysis unit (472) can extract key phrases from the text converted by the voice-to-text conversion unit (464). If the text analysis unit (472) detects a difference in tone between phrases in the converted text, it can extract the phrases with different tones as key phrases. If the text analysis unit (472) changes the frequency band between phrases by more than a preset band, it can determine that the tone has changed. The text analysis unit (472) can also extract key words within the phrases of the converted text. Key words may be nouns present within the phrases, but this is merely an example.

[0145] The feature clustering unit (473) can classify the speaker's speech type using the characteristic information of the voice extracted from the voice analysis unit (471). The feature clustering unit (473) can classify the speaker's speech type by assigning weights to each of the type items constituting the characteristic information of the voice. The feature clustering unit (473) can classify the speaker's speech type using the attention technique of the deep learning model.

[0146] The text mapping unit (474) can translate a text converted into a first language into a text in a second language. The text mapping unit (474) can map a text translated into a second language to a text in the first language. The text mapping unit (474) can map the main expressions constituting the text in the first language to the corresponding expressions in the second language. The text mapping unit (474) can map the utterance types corresponding to the main expressions constituting the text in the first language to the expressions in the second language. This is to apply the classified utterance types to the expressions in the second language.

[0147] The voice synthesis unit (475) can generate a synthesized voice by applying the speech type and the speaker's tone classified by the feature clustering unit (473) to the main expression phrases of the text translated into a second language by the text mapping unit (474).

[0148] The controller (470) can determine the user's speech characteristics using one or more of the transmitted text data or power spectrum (520).

[0149] User speech characteristics may include the user's gender, the user's pitch, the user's timbre, the user's speech topic, the user's speech rate, the user's voice volume, etc.

[0150] The controller (470) can obtain the frequency of the voice signal (510) and the amplitude corresponding to the frequency by using the power spectrum (520).

[0151] The controller (470) can determine the gender of the user who has spoken using the frequency band of the power spectrum (470). For example, the controller (470) can determine the gender of the user as male if the frequency band of the power spectrum (520) is within a preset first frequency band range.

[0152] The controller (470) can determine the user's gender as female if the frequency band of the power spectrum (520) is within a preset second frequency band range. Here, the second frequency band range may be larger than the first frequency band range.

[0153] The controller (470) can determine the pitch of a sound by using the frequency band of the power spectrum (520). For example, the controller (470) can determine the pitch of a sound based on the amplitude within a specific frequency band range.

[0154] The controller (470) can determine the user's tone by using the frequency band of the power spectrum (520). For example, the controller (470) can determine a frequency band with an amplitude greater than a certain size among the frequency bands of the power spectrum (520) as the user's main vocal range, and determine the determined main vocal range as the user's tone.

[0155] The controller (470) can determine the user's speaking speed through the number of syllables spoken per unit time from the converted text data.

[0156] For the converted text data of the controller (470), the user's speech topic can be determined using the Bag-Of-Word Model technique.

[0157] The Bag-Of-Word Model technique extracts frequently used words based on their frequency within a sentence. Specifically, the Bag-Of-Word Model technique extracts unique words from a sentence, expresses the frequency of each extracted word as a vector, and determines the characteristics of the utterance topic. For example, if words like "running" and "physical strength" frequently appear in the controller (470) text data, the user's utterance topic can be classified as exercise.

[0158] The controller (470) can determine the topic of a user's speech from text data using a known text categorization technique. The controller (470) can extract keywords from the text data to determine the topic of the user's speech.

[0159] The controller (470) can determine the user's vocal volume by considering amplitude information across the entire frequency band. For example, the controller (470) can determine the user's vocal volume based on the average or weighted average of the amplitudes across each frequency band of the power spectrum.

[0160] The communication unit (480) can communicate with an external server via wired or wireless communication. The communication unit (480) can communicate with an electronic device (100) via wired or wireless communication.

[0161] The database (490) can store speech in a first language included in the content. The database (490) can store a synthesized speech in which the speech in the first language is converted into speech in a second language. The database (490) can store a first text corresponding to the speech in the first language and a second text in which the first text is translated into a second language. The database (490) can store various learning models required for speech recognition.

[0162] Meanwhile, the control unit (170) of the electronic device (100) illustrated in FIG. 2 may be equipped with the preprocessing unit (460) and the controller (470) illustrated in FIG. 4. That is, the control unit (170) of the electronic device (100) may perform the functions of the preprocessing unit (460) and the controller (470).

[0163] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.

[0164] That is, the voice recognition and synthesis process of FIG. 6 may be performed by the control unit (170) of the electronic device (100) without going through the server.

[0165] Referring to FIG. 6, the processor (180) of the electronic device (100) may include an STT engine (610), an NLP engine (620), and a voice synthesis engine (630). Each engine may be either hardware or software.

[0166] The STT engine (610) can perform the function of the STT server (420) of FIG. 5. That is, the STT engine (610) can convert voice data into text data.

[0167] The NLP engine (620) can perform the function of the NLP server (430) of FIG. 5. That is, the NLP engine (620) can obtain intent analysis information indicating the speaker's intent from converted text data.

[0168] The speech synthesis engine (630) can perform the functions of a speech synthesis server. The speech synthesis engine (630) can search for syllables or words corresponding to given text data from a database, synthesize a combination of the searched syllables or words, and generate a synthesized speech.

[0169] The voice synthesis engine (630) may include a preprocessing engine (631) and a TTS engine (632).

[0170] The preprocessing engine (631) can preprocess text data before generating synthetic speech. Specifically, the preprocessing engine (631) performs tokenization, which divides the text data into tokens, which are meaningful units. After tokenization, the preprocessing engine (631) can perform a cleansing operation to remove unnecessary characters and symbols to remove noise. Thereafter, the preprocessing engine (631) can integrate word tokens with different expression methods to generate the same word token. Thereafter, the preprocessing engine (631) can remove meaningless word tokens (stopwords).

[0171] The TTS engine (632) can synthesize voice corresponding to preprocessed text data and generate a synthesized voice.

[0172] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.

[0173] Referring to FIG. 7, the electronic device (100) can determine, in operation S701, whether a user account is logged into the server (400). For example, a user can log into the server (400) with a user account by entering the ID and password of the user account.

[0174] According to one embodiment, when a user first logs into a server (400) using an electronic device (100), the electronic device (100) may include user identification information corresponding to the user account in a user list. For example, when three different user accounts log into the server (400) using the electronic device (100), the user list stored in the electronic device (100) may include three different user identification information.

[0175] In operation S702, the electronic device (100) can determine whether voice-related identification information (hereinafter, "Voice ID") is registered for a user account logged into the server (400). Here, the Voice ID may include voiceprint information stored in the user identification server (440). For example, the server (400) may transmit to the electronic device (100) whether a voice ID is registered for a user account logged into the server (400).

[0176] According to one embodiment, the server (400) can determine whether a voice ID is registered based on whether voice print information is mapped to user identification information, which is unique identification information corresponding to a user account logged into the server (400). In this case, if the voice ID corresponds to an unregistered user account, the number of voice print information mapped to the user identification information may be 0.

[0177] According to one embodiment, the server (400) may determine that a voice ID is registered if the number of voice prints mapped to the user identification information is two or more, a predetermined number, and may determine that the voice ID is not registered if the number is less than the predetermined number. For example, in the case of a user account with a registered voice ID, six different voice prints may be mapped to the user identification information. For example, in the case of a user account with an unregistered voice ID, five or fewer voice prints may be mapped to the user identification information.

[0178] According to one embodiment, a flag value indicating whether a voice ID is registered may be mapped to user identification information stored in the server (400). Here, the user identification information to which the flag value is mapped may be stored in the user identification server (440) and / or the account server (450). The server (400) may determine whether a voice ID is registered based on the flag value mapped to the user identification information. For example, if the voice ID is an unregistered user account, the flag value mapped to the user identification information may be 0, and if the voice ID is a registered user account, the flag value mapped to the user identification information may be 1.

[0179] In operation S703, if a voice ID has not been registered for a user account, the electronic device (100) may initiate a process for registering a voice ID. For example, when the electronic device (100) initiates a process for registering a voice ID, it may transmit data including device identification information, user identification information, a value indicating the start of voice ID registration, etc., to the server (400).

[0180] The electronic device (100) can output preset text in operation S704. The electronic device (100) can output any one of a plurality of preset texts. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) can output the preset text through the display (180).

[0181] According to one embodiment, the server (400) may transmit any one of a plurality of preset texts to the electronic device (100) in a preset order. At this time, the electronic device (100) may output the preset text received from the server (400).

[0182] The electronic device (100) can determine, in operation S705, whether a voice is input for a preset text. For example, the electronic device (100) can determine whether a voice is input through a microphone included in the input unit (160) within a preset time period. At this time, a voice signal corresponding to the voice input through the microphone can be transmitted to the control unit (170) through the user input interface unit (150). For example, the electronic device (100) can determine whether data including a voice signal corresponding to a voice spoken by a user is received from the remote control device (200) within a preset time period.

[0183] In operation S706, when a voice is input for a preset text, the electronic device (100) can transmit voice data containing a voice signal corresponding to the voice to the server (400). At this time, the electronic device (100) can transmit device identification information, user identification information, a language code indicating the type of language, etc., together with the voice data to the server (400).

[0184] The server (400) can convert a voice signal included in voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponds to a preset text. For example, the server (400) can determine whether the two texts correspond based on the similarity between the text converted from the voice signal and the preset text.

[0185] The server (400) can generate voiceprint information corresponding to the voice signal when the text converted from the voice signal corresponds to a preset text. The server (400) can map the voiceprint information generated for the preset text to user identification information and store it. The server (400) can map the voice data received for the preset text to user identification information and store it.

[0186] The electronic device (100), in operation S707, can determine whether the processing of the voice for the preset text was successful based on the response received from the server (400). For example, if the text converted from the voice signal and the preset text correspond to each other, the server (400) can notify the electronic device (100) of the success of the processing of the voice. For example, if the voiceprint information corresponding to the voice signal is generated, the server (400) can notify the electronic device (100) of the success of the processing of the voice.

[0187] Meanwhile, in operation S708, if voice input for a preset text is not performed, or if voice processing for a preset text fails, the electronic device (100) may determine whether to retry voice input. For example, the electronic device (100) may retry voice input based on a user input requesting retry voice input. In this case, the electronic device (100) may output the preset text again.

[0188] In operation S709, if the voice processing for a preset text is successful, the electronic device (100) can determine whether the processing for all texts is complete. For example, if the voice processing for a predetermined number of six preset texts is successful, the processing for all texts can be completed. Meanwhile, if the processing for five preset texts is complete, the electronic device (100) can output the last preset text.

[0189] The electronic device (100) may complete the process of registering a voice ID when processing of all texts is completed in operation S710. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) may output a screen indicating completion of voice ID registration through the display (180). For example, the electronic device (100) may transmit data indicating completion of voice ID registration to the account server (450).

[0190] According to one embodiment, the electronic device (100) can log in to the server (400) using a user account with a registered voice ID based on a voice input into the electronic device (100).

[0191] When a voice is input, the electronic device (100) can transmit voice data corresponding to the input voice to the server (400). At this time, the electronic device (100) can transmit device identification information, a user list, a language code indicating the type of language, etc., along with the voice data to the server (400).

[0192] The server (400) can generate voiceprint information for a voice input into the electronic device (100) based on voice data received from the electronic device (100). The server (400) can search a database for voiceprint information (hereinafter, candidate voiceprint information) corresponding to user identification information included in a user list received from the electronic device (100). The server (400) can determine whether the candidate voiceprint information and the generated voiceprint information correspond to each other. The server (400) can determine user identification information, to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped, among the candidate voiceprint information, as user identification information corresponding to the voice input into the electronic device (100). Meanwhile, if there is no candidate voiceprint information corresponding to the generated voiceprint information, the server (400) can determine that there is no user identification information corresponding to the voice input into the electronic device (100).

[0193] The server (400) can transmit the result of processing voice data received from the electronic device (100) to the electronic device (100). For example, the server (400) can transmit text converted from voice data received from the electronic device (100), the result of performing intent analysis on the converted text, user identification information corresponding to the voice, etc. to the electronic device (100).

[0194] The electronic device (100) can perform an operation corresponding to a voice input into the electronic device (100) based on the result of processing voice data received from the server (400). For example, if the user identification information corresponding to the voice does not correspond to a user account currently logged into the server (400), the electronic device (100) can log into the server (400) with a user account corresponding to the user identification information corresponding to the voice. For example, if the user identification information corresponding to the voice corresponds to a user account currently logged into the server (400) or if there is no user identification information corresponding to the voice, the electronic device (100) can maintain the login status of the user account currently logged into the server (400).

[0195] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.

[0196] Referring to FIG. 8, the electronic device (100) can perform a login to the server (400) using a user account in operation S801.

[0197] The electronic device (100) can initiate a process of registering a voice ID in operation S802.

[0198] The electronic device (100) can output a first text among a plurality of preset texts in operation S803.

[0199] The electronic device (100) can receive a first voice for a first text in operation S804.

[0200] The electronic device (100) can transmit first voice data including a voice signal corresponding to the first voice to the server (400) in operation S805.

[0201] The server (400), in operation S806, can process the first voice for the first text based on the first voice data received from the electronic device (100). The server (400) can convert a voice signal corresponding to the first voice included in the first voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the first voice and the first text correspond to each other.

[0202] The server (400) may, in operation S807, notify the electronic device (100) of the completion of processing for the first voice. For example, the server (400) may notify the electronic device (100) of the success of processing for the first voice based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.

[0203] Meanwhile, the server (400) can generate first voiceprint information for the first voice based on the voice signal corresponding to the first voice, based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.

[0204] The server (400) may, in operation S808, store first voice data and first voiceprint information for the first voice. The server (400) may store the first voice data and the first voiceprint information by mapping them to user identification information corresponding to a logged-in user account.

[0205] The electronic device (100) can sequentially output the second text to the fifth text. The electronic device (100) can sequentially receive the second voice to the fifth voice, which correspond to the second text to the fifth text, respectively. The electronic device (100) can sequentially transmit the second voice data to the fifth voice data, which correspond to the second voice to the fifth voice, respectively, to the server (400).

[0206] The server (400) can process the second to fifth voices based on the second to fifth voice data received from the electronic device (100), respectively. In addition, the server (400) can sequentially generate and store second to fifth voice information corresponding to the second to fifth voices, respectively.

[0207] The electronic device (100) can output a sixth text among a plurality of preset texts in operation S809.

[0208] The electronic device (100) can receive a sixth voice for a sixth text in operation S810.

[0209] The electronic device (100) can transmit sixth voice data including a voice signal corresponding to the sixth voice to the server (400) in operation S811.

[0210] The server (400) can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device (100) in operation S812. The server (400) can convert a voice signal corresponding to the sixth voice included in the sixth voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.

[0211] The server (400) can, in operation S813, notify the electronic device (100) of the completion of processing for the sixth voice.

[0212] Meanwhile, the server (400) can generate sixth voice information for the sixth voice based on the voice signal corresponding to the sixth voice, if the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.

[0213] The server (400) may store the sixth voice data and the sixth voiceprint information for the sixth voice in operation S814. The server (400) may store the sixth voice data and the sixth voiceprint information by mapping them to user identification information corresponding to the logged-in user account. At this time, the user identification information corresponding to the logged-in user account may be mapped with six different voice data and multiple voiceprint information.

[0214] The electronic device (100) may complete the process of registering a voice ID in operation S815. For example, the electronic device (100) may complete the process of registering a voice ID based on the completion of processing of a predetermined number of six different preset texts.

[0215] Referring to FIG. 9, if a user account is not logged in to the server (400), the electronic device (100) may output a login screen (900) related to logging in to the server (400) through the display (180). The login screen (900) may include an object (910) indicating a non-logged-in state, a login object (920) for executing login, etc. When a user selects the login object (920) using a pointer (205) corresponding to the remote control device (200), the electronic device (100) may output a screen for inputting an ID and password. At this time, the user may input the ID and password of the user account to log in to the server (400) with the user account.

[0216] Referring to FIG. 10, if a voice ID is not registered in a user account logged in to the server (400), the electronic device (100) can output a first account screen (1000) corresponding to the user account with the unregistered voice ID. The first account screen (1000) can include an object (1010) representing a logged-in user account, an object (1020) corresponding to the registration of a voice ID, etc. When a user selects an object (1020) corresponding to the registration of a voice ID using a pointer (205), the electronic device (100) can initiate a process of registering a voice ID.

[0217] Meanwhile, referring to FIGS. 11A to 11C, if a voice ID is already registered in a user account logged in to the server (400), the electronic device (100) may output a second account screen (1100) corresponding to the user account in which the voice ID is registered. The second account screen (1100) may include an object (1110) representing the logged-in user account, a re-registration object (1120) corresponding to re-registration of the voice ID, a deletion object (1130) corresponding to deletion of the voice ID, an activation object (1140) corresponding to use of a function related to the voice ID, etc. The user may select the activation object (1140) using the pointer (205) to activate or deactivate use of the function related to the voice ID.

[0218] When a user selects a re-registration object (1120) using a pointer (205), the electronic device (100) can output a first notification screen (1121) notifying that data for a previously registered voice ID is deleted from the server (400) due to re-registration of the voice ID. The user can re-register the voice ID by selecting a confirmation object (1123) using the pointer (205), or maintain the previously registered voice ID by selecting a cancel object (1125).

[0219] Meanwhile, when a user selects a delete object (1130) using a pointer (205), the electronic device (100) can output a second notification screen (1131) that notifies that data for a pre-registered voice ID is deleted from the server (400) due to the deletion of the voice ID. The user can delete the voice ID by selecting a confirmation object (1133) using the pointer (205), or maintain the pre-registered voice ID by selecting a cancel object (1135).

[0220] Referring to FIG. 12, when an object (1020) corresponding to the registration of a voice ID is selected on the first account screen (1000), or when a re-registration object (1120) is selected on the second account screen (1100), the electronic device (100) can output a start screen (1200) for initiating the registration of a voice ID. When a user selects a start object (1210) using a pointer (205), the electronic device (100) can output a text screen for outputting preset text.

[0221] Referring to FIG. 13, the electronic device (100) can output a text screen (1300) that outputs any one of a plurality of preset texts. The text screen (1300) can include preset text (1301), a text sequence (1302), a termination object (1310) that terminates a process of registering a voice ID, an input object (1320) that receives voice, etc.

[0222] When the user selects the end object (1310) using the pointer (205), the process of registering a voice ID may be terminated. For example, when the process of registering a voice ID is terminated, all data stored in the server (400) while the process of registering a voice ID is in progress may be deleted.

[0223] When a user selects an input object (1320) using a pointer (205), the electronic device (100) can receive voice for the text.

[0224] According to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a text screen (1300) is displayed, the electronic device (100) can receive a voice for the text based on the user input of pressing the predetermined button received from the remote control device (200).

[0225] Meanwhile, according to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a process for registering a voice ID is in progress, the electronic device (100) may stop the process for registering a voice ID based on a user input for pressing the predetermined button received from the remote control device (200). At this time, the user input for pressing the predetermined button (e.g., a voice input button) included in the remote control device (200) may correspond to a user input for initiating voice recognition for a voice received through the remote control device (200). The electronic device (100) may perform an operation related to voice recognition on voice data including a voice signal received from the remote control device (200).

[0226] Referring to FIG. 14, if the processing of voice to text is successful, the electronic device (100) may output a success screen (1400) indicating the success of the processing of text. The success screen (1400) may include an object (1410) indicating the success of the processing of voice to text.

[0227] Meanwhile, referring to FIG. 15, if voice for text is not input or if voice processing for text fails, the electronic device (100) may output a failure screen (1500) indicating a failure in voice processing for text. The failure screen (1500) may include a termination object (1510) that terminates the process of registering a voice ID, a re-input object (1520) that attempts to receive voice again, etc. If the user selects the re-input object (1520) using the pointer (205), the electronic device (100) may receive voice for text again.

[0228] Referring to FIG. 16, when processing for all texts is completed, the electronic device (100) can output a completion screen (1600) indicating completion of voice ID registration. The completion screen (1600) can include an object (1610) indicating a user account in which a voice ID is registered, a completion object (1620) for completing the process of registering a voice ID, etc. When a user selects the completion object (1620) using a pointer (205), the electronic device (100) can complete the process of registering a voice ID.

[0229] FIGS. 17A and 17B are flowcharts illustrating an operating method of an electronic device according to another embodiment of the present disclosure. Any details that overlap with those described in FIG. 7 will be omitted for brevity.

[0230] Referring to FIG. 17a, the electronic device (100) can determine, in operation S1701, whether a user account is logged in to the server (400).

[0231] The electronic device (100) can determine, in operation S1702, whether a voice ID is registered for a user account logged into the server (400).

[0232] The electronic device (100), in operation S1703, may initiate a process for registering a voice ID if a voice ID is not registered for the user account.

[0233] The electronic device (100) can determine, in operation S1704, whether the voice ID is a temporarily stored user account. For example, the server (400) can transmit to the electronic device (100) whether a voice ID is temporarily stored for a user account logged into the server (400).

[0234] According to one embodiment, the server (400) can determine whether to temporarily store a voice ID based on whether voice print information is mapped to user identification information, which is unique identification information corresponding to a user account logged into the server (400). In the case of a user account for which a voice ID is temporarily stored, a minimum number of voice print information or more may be mapped to the user identification information. In this case, the minimum number may be less than a predetermined number of 6.

[0235] According to one embodiment, the server (400) may determine whether to temporarily store a voice ID based on a flag value mapped to the user identification information. For example, if the voice ID is a temporarily stored user account, the flag value mapped to the user identification information may be 2.

[0236] The electronic device (100) can output a preset text in operation S1705. For example, if the voice ID is not temporarily stored, the server (400) can transmit the first text among a plurality of preset texts to the electronic device (100).

[0237] The electronic device (100) can determine, in operation S1706, whether a voice for a preset text is input.

[0238] In operation S1707, when a voice is input for a preset text, the electronic device (100) can transmit voice data including a voice signal corresponding to the voice to the server (400).

[0239] The electronic device (100) can determine, in operation S1708, whether the processing of voice for the preset text was successful based on the response received from the server (400).

[0240] Meanwhile, in operation S1709, the electronic device (100) can determine whether to retry inputting voice if voice for the preset text is not input or if voice processing for the preset text fails.

[0241] The electronic device (100), in operation S1710, can determine whether processing of all texts is completed if the processing of voice for the preset text is successful.

[0242] The electronic device (100) can complete the process of registering a voice ID when processing of all texts is completed in operation S1711.

[0243] Meanwhile, referring to FIG. 17b, in operation S1712, if the voice ID is temporarily stored, the electronic device (100) may output any one of a plurality of preset texts (hereinafter, “unsaved text”), excluding the text corresponding to the voiceprint information mapped to the user identification information. For example, if the voiceprint information corresponding to the first to third texts is mapped to the user identification information, which is unique identification information corresponding to the user account logged in to the server (400), the server (400) may transmit the fourth text to the electronic device (100) as an unsaved text. At this time, the electronic device (100) may output the fourth text received from the server (400).

[0244] The electronic device (100) can determine, in operation S1713, whether voice is input for unsaved text.

[0245] In operation S1714, when a voice is input for an unsaved text, the electronic device (100) can transmit voice data including a voice signal corresponding to the voice to the server (400).

[0246] In operation S1715, the electronic device (100) can determine whether the previous user who temporarily stored the voice ID and the current user who spoke the voice for the unsaved text are the same. For example, the electronic device (100) can determine whether the previous user and the current user are the same based on the result of the determination of whether the users are the same, received from the server (400).

[0247] The server (400) can convert a voice signal included in voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponds to an unsaved text. For example, the server (400) can determine whether the two texts correspond based on the similarity between the text converted from the voice signal and the unsaved text.

[0248] The server (400) may generate voiceprint information (hereinafter, “unstored voiceprint information”) corresponding to the voice of the unstored text based on the voice signal included in the voice data received from the electronic device (100) when the text converted from the voice signal corresponds to the unstored text. The server (400) may determine whether at least one of the voiceprint information mapped to the user identification information corresponds to the unstored voiceprint information. For example, when the voiceprint information corresponding to the first text to the third text is mapped to the user identification information, the server (400) may calculate the similarity between the first voiceprint information corresponding to the first text and the unstored voiceprint information. At this time, the server (400) may determine that the first voiceprint information and the unstored voiceprint information correspond to each other when the similarity between the first voiceprint information and the unstored voiceprint information is equal to or higher than a predetermined standard. Additionally, the server (400) can determine that the previous user and the current user are the same if the first voiceprint information and the unsaved voiceprint information correspond to each other.

[0249] In operation S1716, if the previous user and the current user are different, the electronic device (100) may determine not to use data related to the previous user stored in the server (400). For example, if the electronic device (100) is a video display device (100a), the electronic device (100) may output a screen indicating that the previous user and the current user are different through the display (180).

[0250] The server (400) may maintain the voiceprint information mapped to the user identification information if the previous user and the current user are the same. For example, the server (400) may map the voiceprint information corresponding to the first to third texts to the user identification information and store the unsaved voiceprint information as the voiceprint information corresponding to the fourth text. In addition, the server (400) may sequentially transmit the fifth and sixth texts to the electronic device (100).

[0251] The server (400) may delete voiceprint information mapped to user identification information if the previous user and the current user are different. At this time, the server (400) may also delete voice data mapped to the user identification information. For example, the server (400) may delete voiceprint information corresponding to the first to third texts mapped to the user identification information, and store unsaved voiceprint information by mapping it to the user identification information as voiceprint information corresponding to the fourth text. In addition, the server (400) may sequentially transmit the remaining texts, excluding the fourth text, among a plurality of preset texts to the electronic device (100).

[0252] Meanwhile, in operation S1717, the electronic device (100) may determine whether to retry voice input if voice input for the unsaved text is not performed or if voice processing for the unsaved text fails. For example, the electronic device (100) may retry voice input based on a user input requesting retry voice input. In this case, the electronic device (100) may output the unsaved text again.

[0253] FIG. 18 is a flowchart illustrating a method of operating a system when a voice ID is not registered and temporarily stored, according to one embodiment of the present disclosure.

[0254] Referring to FIG. 18, the electronic device (100) can perform a login to the server (400) using a user account in operation S1801.

[0255] The electronic device (100) can initiate a process of registering a voice ID in operation S1802.

[0256] The electronic device (100) can output a first text among a plurality of preset texts in operation S1803.

[0257] The electronic device (100) can receive a first voice for a first text in operation S1804.

[0258] The electronic device (100) can transmit first voice data including a voice signal corresponding to the first voice to the server (400) in operation S1805.

[0259] The server (400) can process the first voice for the first text based on the first voice data received from the electronic device (100) in operation S1806.

[0260] The server (400) can, in operation S1807, notify the electronic device (100) of the completion of processing for the first voice.

[0261] The server (400) can store first voice data and first voiceprint information for the first voice in operation S1808.

[0262] The electronic device (100) can output a second text among a plurality of preset texts in operation S1809.

[0263] The electronic device (100) can receive a second voice for a second text in operation S1810.

[0264] The electronic device (100) can transmit second voice data including a voice signal corresponding to the second voice to the server (400) in operation S1811.

[0265] The server (400) can process the second voice for the second text based on the second voice data received from the electronic device (100) in operation S1812.

[0266] The server (400) can notify the electronic device (100) of the completion of processing for the second voice in operation S1813.

[0267] The server (400) can store second voice data and second voiceprint information for the second voice in operation S1814.

[0268] The electronic device (100) can notify the server (400) of the suspension of voice ID registration in operation S1815.

[0269] For example, when a user presses a power button included in a remote control device (200), the electronic device (100) can turn off the power of the electronic device (100) based on a signal for a user input of pressing the power button received from the remote control device (200). At this time, the electronic device (100) can notify the server (400) of the suspension of voice ID registration based on the power being turned off.

[0270] For example, when a user presses a button related to a predetermined function (e.g., an OTT service button) included in a remote control device (200), the electronic device (100) may execute a predetermined function based on a signal regarding a user input of pressing the button related to the predetermined function (e.g., an OTT service button) received from the remote control device (200). At this time, the electronic device (100) may notify the server (400) of the suspension of voice ID registration based on the execution of the predetermined function. Meanwhile, when the execution of the predetermined function is terminated, the electronic device (100) may output a notification notifying the user that the registration of the voice ID is suspended. The electronic device (100) may determine whether to resume the registration of the suspended voice ID based on the user input. When resuming the registration of the suspended voice ID, the electronic device (100) may initiate a process of registering the voice ID.

[0271] For example, when a user presses a voice recognition-related button (e.g., a voice input button) included in a remote control device (200), the electronic device (100) may execute a voice recognition function based on a signal for a user input of pressing the voice recognition-related button (e.g., a voice input button) received from the remote control device (200). At this time, the electronic device (100) may notify the server (400) of the suspension of voice ID registration based on the execution of the voice recognition function. Meanwhile, when the execution of the voice recognition function is terminated, the electronic device (100) may output a notification notifying the user that the registration of the voice ID is suspended. The electronic device (100) may determine whether to resume the registration of the suspended voice ID based on the user input. When resuming the registration of the suspended voice ID, the electronic device (100) may initiate a process of registering the voice ID.

[0272] In operation S1816, if the electronic device (100) notifies that voice ID registration has been suspended, the server (400) can check the number of voiceprint information stored and mapped to user identification information. At this time, the server (400) can delete voice data and voiceprint information stored and mapped to user identification information based on the fact that the number of voiceprint information stored and mapped to user identification information is two, which is less than the minimum number (e.g., three).

[0273] Meanwhile, the server (400) can maintain voice data and voiceprint information mapped to user identification information and stored based on the number of voiceprint information mapped to user identification information being greater than or equal to the minimum number (e.g., 3).

[0274] FIG. 19 is a flowchart illustrating a method of operating a system when a voice ID is temporarily stored, according to one embodiment of the present disclosure.

[0275] Referring to FIG. 19, the electronic device (100) can perform a login to the server (400) using a user account in operation S1901.

[0276] The electronic device (100) can initiate a process of registering a voice ID in operation S1902.

[0277] In operation S1903, the server (400) may transmit information about the unsaved text to the electronic device (100) based on the voice ID being temporarily stored. For example, the server (400) may transmit the fourth text as an unsaved text to the electronic device (100) based on the mapping of the voice print information corresponding to the first to third texts to the user identification information, which is unique identification information corresponding to the logged-in user account.

[0278] The electronic device (100) can output the fourth text, which is an unsaved text, in operation S1904.

[0279] The electronic device (100) can receive a fourth voice for a fourth text in operation S1905.

[0280] The electronic device (100) can transmit fourth voice data including a voice signal corresponding to the fourth voice to the server (400) in operation S1906.

[0281] The server (400) can process the fourth voice for the fourth text based on the fourth voice data received from the electronic device (100) in operation S1907.

[0282] The server (400) can determine whether at least one of the fourth voiceprint information generated based on the fourth voice data and the voiceprint information mapped to the user identification information corresponds to each other. For example, the server (400) can determine whether the first voiceprint information and the fourth voiceprint information correspond to each other based on the similarity between the first voiceprint information corresponding to the first text and the fourth voiceprint information.

[0283] In operation S1908, the server (400) may notify the electronic device (100) of the result of determining that the previous user who temporarily stored the voice ID and the user who uttered the fourth voice are the same. For example, the server (400) may determine that the previous user and the current user are the same based on the correspondence between the first and fourth voice print information.

[0284] The server (400) can, in operation S1909, notify the electronic device (100) of the completion of processing for the fourth voice.

[0285] The server (400) may store fourth voice data and fourth voiceprint information for the fourth voice in operation S1910. The server (400) may store the fourth voice data and fourth voiceprint information by mapping them to user identification information corresponding to a logged-in user account.

[0286] The electronic device (100) can output a fifth text. The electronic device (100) can receive a fifth voice corresponding to each of the fifth texts. The electronic device (100) can transmit fifth voice data corresponding to the fifth voice to the server (400).

[0287] The server (400) can process the fifth voice based on the fifth voice data received from the electronic device (100). In addition, the server (400) can generate and store fifth voice information corresponding to the fifth voice.

[0288] The electronic device (100) can output a sixth text among a plurality of preset texts in operation S1911.

[0289] The electronic device (100) can receive the sixth voice for the sixth text in operation S1912.

[0290] The electronic device (100) can transmit sixth voice data including a voice signal corresponding to the sixth voice to the server (400) in operation S1913.

[0291] The server (400) can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device (100) in operation S1914.

[0292] The server (400) can notify the electronic device (100) of the completion of processing for the sixth voice in operation S1915.

[0293] The server (400) can store the sixth voice data and the sixth voiceprint information for the sixth voice in operation S1916.

[0294] The electronic device (100) can complete the process of registering a voice ID in operation S1917.

[0295] FIG. 20 is a flowchart of a method of operating a system when a voice ID is temporarily stored, according to another embodiment of the present disclosure.

[0296] Referring to FIG. 20, the electronic device (100) can perform a login to the server (400) with a user account in operation S2001.

[0297] The electronic device (100) can initiate a process of registering a voice ID in operation S1902.

[0298] The server (400) may transmit information about the unsaved text to the electronic device (100) based on the temporarily stored voice ID in operation S2003. For example, the server (400) may transmit the fourth text as an unsaved text to the electronic device (100) based on the mapping of the voice print information corresponding to the first to third texts to the user identification information, which is unique identification information corresponding to the logged-in user account.

[0299] The electronic device (100) can output the fourth text, which is an unsaved text, in operation S2004.

[0300] The electronic device (100) can receive a fourth voice for a fourth text in operation S2005.

[0301] The electronic device (100) can transmit fourth voice data including a voice signal corresponding to the fourth voice to the server (400) in operation S2006.

[0302] The server (400) can process the fourth voice for the fourth text based on the fourth voice data received from the electronic device (100) in operation S2007.

[0303] In operation S2008, the server (400) may notify the electronic device (100) of the determination that the previous user who temporarily stored the voice ID and the user who uttered the fourth voice are different from each other. For example, the server (400) may determine that the previous user and the current user are different based on the fact that the first and fourth voice print information do not correspond to each other. In this case, the server (400) may delete the voice print information mapped to the user identification information, which is the unique identification information corresponding to the logged-in user account.

[0304] The server (400) can, in operation S2009, notify the electronic device (100) of the completion of processing for the fourth voice.

[0305] The server (400) may store fourth voice data and fourth voiceprint information for the fourth voice in operation S2010. The server (400) may store the fourth voice data and fourth voiceprint information by mapping them to user identification information corresponding to a logged-in user account.

[0306] Meanwhile, the server (400) can transmit any one of the preset texts, excluding the fourth text, to the electronic device (100). In the present disclosure, an example is provided in which the fifth and sixth texts are sequentially transmitted to the electronic device (100), and then the first to third texts are transmitted to the electronic device (100).

[0307] The electronic device (100) can output a fifth text. The electronic device (100) can receive a fifth voice corresponding to each of the fifth texts. The electronic device (100) can transmit fifth voice data corresponding to the fifth voice to the server (400).

[0308] The server (400) can process the fifth voice based on the fifth voice data received from the electronic device (100). In addition, the server (400) can generate and store fifth voice information corresponding to the fifth voice.

[0309] The electronic device (100) can output a sixth text among a plurality of preset texts in operation S2011.

[0310] The electronic device (100) can receive the sixth voice for the sixth text in operation S2012.

[0311] The electronic device (100) can transmit sixth voice data including a voice signal corresponding to the sixth voice to the server (400) in operation S2013.

[0312] The server (400) can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device (100) in operation S2014.

[0313] The server (400) can, in operation S2015, notify the electronic device (100) of the completion of processing for the sixth voice.

[0314] The server (400) can store the sixth voice data and the sixth voiceprint information for the sixth voice in operation S2016.

[0315] The electronic device (100) can output a first text among a plurality of preset texts in operation S2017.

[0316] The electronic device (100) can receive a first voice for a first text in operation S2018.

[0317] The electronic device (100) can transmit first voice data including a voice signal corresponding to the first voice to the server (400) in operation S2019.

[0318] The server (400) can process the first voice for the first text based on the first voice data received from the electronic device (100) in operation S2020.

[0319] The server (400) can, in operation S2021, notify the electronic device (100) of the completion of processing for the first voice.

[0320] The server (400) can store first voice data and first voiceprint information for the first voice in operation S2022.

[0321] The electronic device (100) can output a second text. The electronic device (100) can receive a second voice corresponding to each of the second texts. The electronic device (100) can transmit second voice data corresponding to the second voice to the server (400).

[0322] The server (400) can process the second voice based on the second voice data received from the electronic device (100). In addition, the server (400) can generate and store second voice information corresponding to the second voice.

[0323] The electronic device (100) can output a third text among a plurality of preset texts in operation S2023.

[0324] The electronic device (100) can receive a third voice for a third text in operation S2024.

[0325] The electronic device (100) can transmit third voice data including a voice signal corresponding to the third voice to the server (400) in operation S2025.

[0326] The server (400) can process the third voice for the third text based on the third voice data received from the electronic device (100) in operation S2026.

[0327] The server (400) can, in operation S2027, notify the electronic device (100) of the completion of processing for the third voice.

[0328] The server (400) can store third voice data and third voiceprint information for the third voice in operation S2028.

[0329] The electronic device (100) can complete the process of registering a voice ID in operation S2029.

[0330] As described above, according to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.

[0331] Additionally, according to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.

[0332] Additionally, according to at least one embodiment of the present disclosure, the user can log in to an account identified based on the user's voice.

[0333] Additionally, according to at least one embodiment of the present disclosure, in the process of registering identification information for a user's voice to a user's account, continuity of identification information for a user who has previously attempted to register can be maintained.

[0334] Referring to FIGS. 1 to 20, an electronic device (100) according to one aspect of the present disclosure includes: a display (180); an external device interface unit (130) communicating with a remote control device (200); a network interface unit (135) communicating with a server (400); a user input interface unit (150) transmitting a signal corresponding to a user input; And it includes a control unit (170), and the control unit (170) outputs a preset text related to the registration of the identification information through the display (180) while performing a process of registering identification information related to a voice for a user account logged in to the server (400), and when a voice signal corresponding to the preset text is received through the user input interface unit (150), transmits data including the voice signal to the server (400), and completes the process of registering the identification information based on the processing of the voice signal corresponding to the preset text by the server (400), and when a predetermined input related to voice recognition is received from the remote control device (200) while performing the process of registering the identification information, the process of registering the identification information can be stopped.

[0335] Additionally, according to one aspect of the present disclosure, the identification information may include a feature vector for a voice print of the voice.

[0336] In addition, according to one aspect of the present disclosure, the control unit (170) may transmit data including a voice signal corresponding to the predetermined input, received from the remote control device (200) after the predetermined input is received, to the server (400), and perform an operation corresponding to the predetermined input based on processing by the server (400) of the voice signal corresponding to the predetermined input.

[0337] In addition, according to one aspect of the present disclosure, when the control unit (170) initiates a process of registering the identification information for the user account, it determines whether the identification information is temporarily stored in the server (400), and if the identification information is not temporarily stored in the server (400), it can output a predetermined number of multiple texts in stages, and if the identification information is temporarily stored in the server (400), it can output a portion of the multiple texts except for the text corresponding to the identification information temporarily stored in the server (400) in stages.

[0338] In addition, according to one aspect of the present disclosure, the control unit (170) may output a notification of the interruption of the process of registering the identification information through the display (180) based on the termination of the operation corresponding to the predetermined input, and may determine whether to start the process of registering the identification information based on the user input received through the user input interface unit (150).

[0339] According to one aspect of the present disclosure, a server (400) includes a communication unit (480) that communicates with an electronic device (100); a database (490); and a controller (470). The controller (470) converts a voice signal included in data received from the electronic device (100) into text while performing a process of registering voice-related identification information for a user account logged into the server (400), generates identification information for the voice signal based on a correspondence between the converted text and a text preset in relation to the registration of the identification information, maps the generated identification information to user identification information corresponding to the user account, and stores the mapped identification information in the database (490). While performing the process of registering the identification information, when a notification of interruption of the process of registering the identification information is received from the electronic device (100), the identification information mapped to the user identification information can be maintained or deleted.

[0340] Additionally, according to one aspect of the present disclosure, the identification information may include a feature vector for a voice print of the voice.

[0341] In addition, according to one aspect of the present disclosure, the controller (470) may maintain the identification information mapped to the user identification information when the number of the identification information mapped to the user identification information is equal to or greater than a preset minimum number, and may delete the identification information mapped to the user identification information when the number of the identification information mapped to the user identification information is less than the preset minimum number.

[0342] In addition, according to one aspect of the present disclosure, when the controller (470) initiates a process of registering the identification information for the user account, it determines whether the identification information is mapped to the user identification information, and if the identification information is not mapped to the user identification information, it can step-by-step transmit a predetermined number of multiple texts to the electronic device (100), and if the identification information is mapped to the user identification information, it can step-by-step transmit a portion of the multiple texts, excluding the text corresponding to the identification information mapped to the user identification information, to the electronic device (100).

[0343] In addition, according to one aspect of the present disclosure, when the controller (470) initiates a process of registering the identification information for the user account, the controller (470) determines whether the identification information is mapped to the user identification information, and if the identification information is mapped to the user identification information, transmits a first text, excluding the text corresponding to the identification information mapped to the user identification information, among the plurality of texts, to the electronic device (100), and generates identification information for a first voice signal corresponding to the first text included in data received from the electronic device (100), and if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information corresponds to each other, maintains the identification information mapped to the user identification information, and if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information does not correspond to each other, deletes the identification information mapped to the user identification information.

[0344] Additionally, according to one aspect of the present disclosure, the controller (470) can map identification information for the first voice signal to the user identification information and store it in the database (490).

[0345] In addition, according to one aspect of the present disclosure, the controller (470) may transmit to the electronic device (100) a result of determining that the current user and the previous user are the same if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information corresponds to each other, and may transmit to the electronic device (100) a result of determining that the current user and the previous user are different if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information does not correspond to each other.

[0346] In addition, according to one aspect of the present disclosure, the controller (470) can map the voice signal to the user identification information and store it in the database (490) based on the correspondence between the converted text and the preset text.

[0347] In addition, according to one aspect of the present disclosure, the controller (470) may generate identification information corresponding to the voice signal mapped to the user identification information using a first algorithm, and change the identification information generated using a second algorithm mapped to the user identification information into identification information generated using the first algorithm.

[0348] A system (10) according to one aspect of the present disclosure includes an electronic device (100) and a server (400), wherein the electronic device (100) outputs a preset text in relation to the registration of the identification information while performing a process of registering identification information related to voice for a user account logged into the server (400), and when a voice signal corresponding to the preset text is received, transmits data including the voice signal to the server (400), and completes the process of registering the identification information based on processing by the server (400) of the voice signal corresponding to the preset text, and while performing the process of registering the identification information, stops the process of registering the identification information when a predetermined input related to voice recognition is received from a remote control device (200), and the server (400) converts the voice signal included in the data received from the electronic device (100) into text, and generates identification information for the voice signal based on a correspondence between the converted text and the preset text, and stores the generated identification information corresponding to the user account. When a notification regarding the interruption of the process of registering the identification information is received from the electronic device (100) while the process of mapping the identification information to the user and storing it in the database (490) and registering the identification information is performed, the identification information mapped to the user identification information can be maintained or deleted.

[0349] The attached drawings are only intended to facilitate understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure.

[0350] Meanwhile, the operating method of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices that store data that can be read by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across network-connected computer systems, so that the processor-readable code can be stored and executed in a distributed manner.

[0351] In addition, although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, display; An external device interface unit that communicates with a remote control device; Network interface section that communicates with the server; A user input interface unit that transmits a signal corresponding to user input; and Including a control unit, The above control unit, During the process of registering voice-related identification information for a user account logged into the above server, a preset text is output in relation to the registration of the identification information through the above display, When a voice signal corresponding to the preset text is received through the user input interface unit, data including the voice signal is transmitted to the server, Based on the processing of the voice signal corresponding to the above preset text by the server, the process of registering the identification information is completed, An electronic device characterized in that, while performing a process of registering the above identification information, if a predetermined input related to voice recognition is received from the remote control device, the process of registering the above identification information is stopped.

2. In paragraph 1, An electronic device characterized in that the above identification information includes a feature vector for a voice print of the voice.

3. In paragraph 1, The above control unit, After the predetermined input is received, data including a voice signal corresponding to the predetermined input received from the remote control device is transmitted to the server, An electronic device characterized in that it performs an operation corresponding to the predetermined input based on the processing of the server for a voice signal corresponding to the predetermined input.

4. In paragraph 1, The above control unit, When initiating a process of registering the identification information for the user account, determine whether the identification information is temporarily stored on the server, If the above identification information is not temporarily stored on the above server, a predetermined number of multiple texts are output in stages, An electronic device characterized in that, when the identification information is temporarily stored in the server, a portion of the plurality of texts is output in stages, excluding the text corresponding to the identification information temporarily stored in the server.

5. In paragraph 1, The above control unit, Based on the termination of the operation corresponding to the above-mentioned predetermined input, a notification regarding the interruption of the process of registering the above-mentioned identification information is output through the above-mentioned display, An electronic device characterized in that it determines whether to initiate a process for registering the identification information based on the user input received through the user input interface unit.

6. On the server, A communication unit that communicates with electronic devices; database; and Includes a controller, The above controller, While performing the process of registering voice-related identification information for a user account logged into the above server, converting a voice signal included in data received from the electronic device into text, Based on the correspondence between the converted text and the preset text related to the registration of the identification information, identification information for the voice signal is generated, The above-mentioned generated identification information is mapped to user identification information corresponding to the user account and stored in the database, A server characterized in that, while performing a process of registering the above identification information, if a notification of the interruption of the process of registering the above identification information is received from the electronic device, the identification information mapped to the user identification information is maintained or deleted.

7. In paragraph 6, A server characterized in that the above identification information includes a feature vector for the voice print of the voice.

8. In paragraph 6, The above controller, If the number of identification information mapped to the user identification information is greater than or equal to the preset minimum number, the identification information mapped to the user identification information is maintained, A server characterized in that, if the number of the identification information mapped to the user identification information is less than the minimum number, the identification information mapped to the user identification information is deleted.

9. In paragraph 6, The above controller, When initiating a process of registering the identification information for the user account, it is determined whether the identification information is mapped to the user identification information, If the above identification information is not mapped to the above user identification information, a predetermined number of multiple texts are sequentially transmitted to the electronic device, A server characterized in that, when the user identification information is mapped to the user identification information, a portion of the plurality of texts, excluding the text corresponding to the user identification information mapped to the user identification information, is transmitted to the electronic device in stages.

10. In paragraph 6, The above controller, When initiating a process of registering the identification information for the user account, it is determined whether the identification information is mapped to the user identification information, If the user identification information is mapped to the user identification information, the first text, excluding the text corresponding to the user identification information mapped to the user identification information, is transmitted to the electronic device among the plurality of texts. Generate identification information for a first voice signal corresponding to the first text included in data received from the electronic device, If at least one of the identification information for the first voice signal and the identification information mapped to the user identification information corresponds to each other, the identification information mapped to the user identification information is maintained, A server characterized in that, if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information does not correspond to each other, the identification information mapped to the user identification information is deleted.

11. In paragraph 10, The above controller, A server characterized in that it stores identification information for the first voice signal in the database by mapping it to the user identification information.

12. In paragraph 10, The above controller, If at least one of the identification information for the first voice signal and the identification information mapped to the user identification information corresponds to each other, the result of determining that the current user and the previous user are the same is transmitted to the electronic device, A server characterized in that, if at least one of the identification information for the first voice signal and the identification information mapped to the user identification information does not correspond to each other, the result of determining that the current user and the previous user are different is transmitted to the electronic device.

13. In paragraph 6, The above controller, A server characterized in that, based on the correspondence between the converted text and the preset text, the voice signal is mapped to the user identification information and stored in the database.

14. In paragraph 13, The above controller, Using the first algorithm, identification information corresponding to the voice signal mapped to the user identification information is generated, A server characterized in that it changes the identification information generated using the second algorithm, which is mapped to the user identification information, into identification information generated using the first algorithm.

15. In a system including electronic devices and servers, The above electronic device, During the process of registering voice-related identification information for a user account logged into the above server, a preset text is output in relation to the registration of the identification information, When a voice signal corresponding to the above preset text is received, data including the voice signal is transmitted to the server, Based on the processing of the voice signal corresponding to the above preset text by the server, the process of registering the identification information is completed, During the process of registering the above identification information, if a predetermined input related to voice recognition is received from the remote control device, the process of registering the above identification information is stopped, The above server, Converting the voice signal included in the data received from the electronic device into text, Generate identification information for the voice signal based on the correspondence between the converted text and the preset text, The above-mentioned generated identification information is mapped to user identification information corresponding to the user account and stored in a database, A system characterized in that, while performing a process of registering the above identification information, if a notification of the interruption of the process of registering the above identification information is received from the electronic device, the identification information mapped to the user identification information is maintained or deleted.

Citation Information

Patent Citations

  • User authentication device and method

    JP2010102383A

  • Artificial intelligence based voiceprint login method and device

    KR1020160147280A

  • Communication method, apparatus and system based on voiceprint

    KR1020170003366A

  • Smart movement system for the elder

    KR1020210010779A

  • Conveyor system of automation line

    KR1020220046200A