Electronic device, server, and system including the same
Patent Information
- Application Number
- CN202480087367.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-07-31
- Publication Date
- 2026-09-04
AI Technical Summary
然而,用户针对各种服务逐一输入账户信息是不方便的
[0017] The advantages of the electronic device, server, and system including the electronic device and the server according to this disclosure will now be described below.
Smart Images

Figure CN122700518A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to electronic devices, servers, and systems including such electronic devices and servers, and more specifically, to electronic devices, servers, and systems including such electronic devices and servers that utilize speech recognition technology. Background Technology
[0002] With recent technological advancements, research on speech recognition technology for processing speech is actively underway. In particular, research on speech recognition technology, starting with smartphones, is being conducted extensively in various fields related to user convenience, such as vehicles and home appliances used in homes and offices.
[0003] Speech recognition technology is typically used when a user controls an electronic device using their voice. For example, when a user issues a command to control an electronic device, the device can either directly recognize and process the user's voice and operate according to the command associated with that voice, or it can send the voice to a server that processes voice and then operate according to the command associated with that voice received from the server.
[0004] Meanwhile, the services and functions offered through electronic devices are becoming increasingly diversified. Furthermore, users register accounts for various services and then log in using those accounts to access those services. In this scenario, service providers use the user information managed for each account to provide customized features or information tailored to that user.
[0005] Traditionally, when attempting to log in to use a service, users are required to directly enter account information, such as an account identifier (ID) and / or password. However, it is inconvenient for users to enter account information one by one for each service. Furthermore, if users remain logged in to eliminate the inconvenience of entering account information, security issues such as unauthorized access to user account information may arise. Additionally, when multiple users share an electronic device, there is the problem of multiple users needing to enter their account information and log in each time they use the service. Summary of the Invention
[0006] Technical issues
[0007] One purpose of this disclosure is to address the aforementioned and other issues.
[0008] Another object of this disclosure is to provide an electronic device, a server, and a system including the electronic device and the server, capable of registering identification information about a user's voice into a user account.
[0009] Another object of this disclosure is to provide an electronic device, a server, and a system including the electronic device and the server that can identify a user based on the user's voice.
[0010] Another object of this disclosure is to provide an electronic device, a server, and a system including the electronic device and the server, capable of logging in using a user account identified by the user's voice.
[0011] Another object of this disclosure is to provide an electronic device, a server, and a system including the electronic device and the server, capable of maintaining the continuity of identification information for users who have previously attempted to register during the process of registering identification information about a user's voice to a user account.
[0012] Technical solution
[0013] To achieve the above objectives, an electronic device according to one embodiment of the present disclosure includes: a display, an external device interface configured to communicate with a remote control device, a network interface configured to communicate with a server, a user input interface configured to send signals related to user input, and a controller, wherein the controller is configured to: display preset text related to the registration of the identification information on the display while performing a process of registering voice-related identification information for a user account logged into the server; send data containing the voice signal to the server when a voice signal related to the preset text is received through the user input interface; complete the process of registering the identification information based on the server's processing of the voice signal related to the preset text; and interrupt the process of registering the identification information when a predetermined input related to voice recognition is received from the remote control device during the process of registering the identification information.
[0014] To achieve the above objectives, a server according to one embodiment of the present disclosure includes: a communication unit configured to communicate with an electronic device, a database, and a controller, wherein the controller is configured to: while performing a process of registering voice-related identification information for a user account logged into the server, convert voice signals included in data received from the electronic device into text, generate voice signal-related identification information based on the converted text and a preset text related to the registration of identification information, map the generated identification information to user identification information related to the user account, store the mapped information in the database, and maintain or delete the identification information mapped to the user identification information when a notification regarding an interruption in the registration of identification information is received from the electronic device during the registration of identification information process.
[0015] To achieve the above objectives, a system according to one embodiment of the present disclosure includes: an electronic device; and a server, wherein the electronic device is configured to: display preset text related to the registration of the identification information while performing a process of registering voice-related identification information for a user account logged into the server; send data containing the voice signal to the server when a voice signal related to the preset text is received; complete the process of registering the identification information based on the server's processing of the voice signal related to the preset text; and stop the process of registering the identification information when a predetermined input related to voice recognition is received from a remote control device during the process of registering the identification information; and the server is configured to: convert the voice signal included in the data received from the electronic device into text; generate identification information related to the voice signal based on the correlation between the converted text and the preset text; map the generated identification information to user identification information related to the user account; store the mapped information in a database; and maintain or delete the identification information mapped to the user identification information when a notification of interruption in the process of registering the identification information is received from the electronic device during the process of registering the identification information.
[0016] Beneficial effects
[0017] The advantages of the electronic device, server, and system including the electronic device and the server according to this disclosure will now be described below.
[0018] According to at least one embodiment of this disclosure, identification information about a user's voice can be registered in a user account.
[0019] According to at least one embodiment of this disclosure, a user can be identified based on the user's voice.
[0020] According to at least one embodiment of this disclosure, it is possible to log in to a user's account based on the user's voice identifier.
[0021] According to at least one embodiment of this disclosure, the continuity of identification information about users who have previously attempted to register can be maintained during the process of registering identification information about user voice to a user account.
[0022] The additional scope of this disclosure will become clear from the detailed description given below. However, it should be understood that the detailed description and specific implementations (such as the preferred embodiments of this disclosure) are given by way of example only, as various changes and modifications within the spirit and scope of this disclosure will become clear to those skilled in the art from the detailed description. Attached Figure Description
[0023] Figure 1 These are diagrams illustrating a system according to an embodiment of the present disclosure; Figure 2 yes Figure 1 Internal block diagram of an electronic device; Figure 3 In describing Figure 1 The server is the reference diagram; Figure 4 This is a block diagram illustrating a server configuration according to an embodiment of the present disclosure; Figure 5 This is a diagram illustrating an example of converting a speech signal into a power spectrum according to an embodiment of the present disclosure; Figure 6 This is a block diagram illustrating the configuration of a controller for speech recognition and synthesis in an electronic device according to an embodiment of the present disclosure; Figure 7 This is a flowchart of a method for operating an electronic device according to an embodiment of the present disclosure; Figure 8 This is a flowchart of an operating system method according to an embodiment of the present disclosure; Figures 9 to 16 The figures are referenced in describing the process of registering identification information about a user's voice into a user account, according to embodiments of this disclosure; Figure 17a and Figure 17b This is a flowchart of a method for operating an electronic device according to another embodiment of the present disclosure; and Figures 18 to 20 This is a flowchart of a method for operating a system according to various embodiments of the present disclosure. Detailed Implementation
[0024] The present disclosure will now be described in detail with reference to the accompanying drawings. In the drawings, for the sake of clarity and conciseness, parts irrelevant to the description have been omitted, and throughout the specification, identical or very similar parts are indicated by the same reference numerals.
[0025] The suffixes “module” and “section” used for components in the following description are merely for convenience in writing this specification and have no particularly important meaning or function. Therefore, the terms “module” and “section” are used interchangeably.
[0026] In this disclosure, it will also be understood that the terms “comprising” or “including” specify the presence of the stated features, figures, steps, operations, components, parts or combinations thereof, but do not exclude the presence or addition of one or more other features, figures, steps, operations, components, parts or combinations thereof.
[0027] Furthermore, in this specification, the terms "first" and / or "second" are used to describe various components, but such components are not limited to those described by these terms. These terms are used to distinguish one component from another.
[0028] Figure 1 This is a diagram illustrating various embodiments of a system according to the present disclosure.
[0029] refer to Figure 1 System 10 may include electronic device 100 and / or server 400.
[0030] Electronic device 100 can send data to / receive data from at least one server 400. For example, electronic device 100 can send data to / receive data from at least one server 400 via a network 300 such as the Internet.
[0031] According to the implementation, at least one server 400 may include a server that performs speech recognition, a server that processes data using a massive artificial intelligence model, a server that provides content, etc.
[0032] Electronic device 100 may include image display device 100a, air conditioner 100b, refrigerator 100c, air purifier 100d, washing machine 100e, vehicle 100f, etc. Although electronic device 100 is image display device 100a in this disclosure, this disclosure is not limited thereto.
[0033] The image display device 100a can be a device for processing and outputting images. There are no particular limitations on the image display device 100a as long as it can output a screen related to video signals, such as a TV, laptop computer, or monitor.
[0034] The image display device 100a can receive broadcast signals, process the broadcast signals, and output the processed broadcast image. When the image display device 100a receives a broadcast signal, the image display device 100a can correspond to a broadcast receiving device.
[0035] The image display device 100a can receive broadcast signals wirelessly via an antenna or via a cable. For example, the image display device 100a can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, and Internet Protocol Television (IPTV) broadcast signals.
[0036] Figure 2 yes Figure 1 Internal block diagram of an electronic device.
[0037] refer to Figure 2 The electronic device 100 may include a broadcast receiver 105, an external device interface 130, a network interface 135, a storage device 140, a user input interface 150, an input unit 160, a controller 170, a display 180, an audio output unit 185, and / or a power supply 190.
[0038] The broadcast receiver 105 may include a tuner 110 and a demodulator 120.
[0039] Meanwhile, the electronic device 100 may include only the broadcast receiver 105, the external device interface 130, and the network interface 135. That is, the electronic device 100 may not include the network interface 135.
[0040] Tuner 110 can select a broadcast signal associated with a user-selected channel or broadcast signals from all previously stored channels from broadcast signals received via an antenna (not shown) or a cable (not shown). Tuner 110 can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.
[0041] For example, if the selected broadcast signal is a digital broadcast signal, tuner 110 can convert the selected broadcast signal into a digital IF signal (DIF), and if the selected broadcast signal is an analog broadcast signal, it can convert the selected broadcast signal into an analog baseband video or audio signal (CVBS / SIF). That is, tuner 110 can process both digital and analog broadcast signals. The analog baseband video or audio signal (CVBS / SIF) output from tuner 110 can be directly input to controller 170.
[0042] Meanwhile, the tuner 110 can sequentially select all the stored broadcast signals from the received broadcast signals through the channel storage function, and convert the selected broadcast signals into intermediate frequency signals or baseband video or audio signals.
[0043] Tuner 110 may include multiple tuners to receive broadcast signals from multiple channels. Alternatively, a single tuner may be used to simultaneously receive broadcast signals from multiple channels.
[0044] Demodulator 120 can receive the digital IF signal (DIF) converted by tuner 110 and perform demodulation operation.
[0045] Demodulator 120 can output a streaming signal TS after performing demodulation and channel decoding. Here, the streaming signal can be a multiplexed video signal, audio signal, or data signal.
[0046] The streaming signal output from demodulator 120 can be input to controller 170. After performing demultiplexing and video / audio signal processing, controller 170 can output video through display 180 and audio through audio output unit 185.
[0047] The external device interface 130 can send data to / receive data from a connected external device. For this purpose, the external device interface 130 may include an A / V input / output section (not shown).
[0048] The external device interface 130 can be connected to external devices such as digital multifunction disc (DVD) players, Blu-ray players, game consoles, cameras, camcorders, laptop computers, set-top boxes, etc. via wired / wireless means, and can also perform input / output operations related to external devices.
[0049] In addition, the external device interface 130 can establish a communication network with various remote control devices 200 to receive control signals related to the operation of the electronic device 100 from the remote control device 200, or to send data related to the operation of the electronic device 100 to the remote control device 200.
[0050] The A / V input / output unit can receive video and audio signals from external devices. For example, the A / V input / output unit may include Ethernet terminals, USB terminals, Composite Video Blanking Synchronization (CVBS) terminals, component terminals, S-Video terminals (analog), Digital Vision Interface (DVI) terminals, High-Definition Multimedia Interface (HDMI) terminals, Mobile High-Definition Link (MHL) terminals, RGB terminals, D-SUB terminals, IEEE 1394 terminals, SPDIF terminals, LCD HD terminals, etc. Digital signals input through these terminals can be sent to the controller 170. Here, analog signals input through CVBS terminals and S-Video terminals can be converted into digital signals by an analog-to-digital converter (not shown) and sent to the controller 170.
[0051] The external device interface 130 may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. The external device interface 130 can exchange data with a neighboring mobile terminal via the wireless communication unit. For example, the external device interface 130 can receive device information, running application information, application images, etc., from the mobile terminal in mirror mode.
[0052] The external device interface 130 can use Bluetooth, radio frequency identification (RFID), Infrared Data Association (IrDA), Ultra Wideband (UWB), ZigBee, etc. to perform short-range wireless communication.
[0053] Network interface 135 can provide an interface for connecting electronic device 100 to wired / wireless networks, including the Internet.
[0054] Network interface 135 may include a communication module (not shown) for connecting to a wired / wireless network. For example, network interface 135 may include a communication module for wireless LAN (WLAN) (Wi-Fi), wireless broadband (WiBro), global microwave access interoperability (WiMax), and high-speed downlink packet access (HSDPA).
[0055] Network interface 135 can send data to or receive data from other users or other electronic devices through the connected network or another network linked to the connected network.
[0056] Network interface 135 can receive web content or data provided by content providers or network operators. That is, network interface 135 can receive content such as movies, advertisements, games, VOD, and broadcasts, as well as related information, provided by content providers or network providers via the network.
[0057] Network interface 135 can receive firmware update information and update files provided by network operators, and can send data to the Internet, content providers, or network operators.
[0058] Network interface 135 can select and receive desired applications from applications that are publicly accessible via the network.
[0059] Storage device 140 can store programs for processing and controlling each signal in controller 170, and can store processed video, audio, or data signals. For example, storage device 140 can store application programs designed to perform various tasks that can be processed by controller 170, and can selectively provide some of the stored applications according to the request of controller 170.
[0060] There are no special restrictions on the program stored in storage device 140, as long as it can be executed by controller 170.
[0061] Storage device 140 can perform the function of temporarily storing video, voice or data signals received from an external device via external device interface 130.
[0062] Storage device 140 can store information about a predetermined broadcast channel through channel storage functions such as channel mapping.
[0063] although Figure 2 An embodiment in which the storage device 140 and the controller 170 are separately configured is illustrated, but the scope of this disclosure is not limited thereto, and the storage device 140 may be included in the controller 170.
[0064] Storage device 140 may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of this disclosure, "storage device" and "memory" may be used interchangeably.
[0065] User input interface 150 can send signals input by the user to controller 170, or send signals from controller 170 to the user.
[0066] For example, the user input interface 150 can send user input signals such as power on / off, channel selection, and screen settings to / receive such user input signals from the remote control device 200, send user input signals input via local keys (not shown) (such as power key, channel key, volume key, and settings key) to the controller 170, send user input signals input via a sensor (not shown) that senses user gestures to the controller 170, or send signals from the controller 170 to the sensor.
[0067] The input section 160 may be disposed on one side of the main body of the electronic device 100. For example, the input section 160 may include a touchpad, physical buttons, etc.
[0068] The input unit 160 can receive various user commands related to the operation of the electronic device 100 and send control signals related to the input commands to the controller 170.
[0069] The input unit 160 may include at least one microphone (not shown) and may receive user voice through the microphone.
[0070] The controller 170 may include at least one processor, and the processor may be used to control the overall operation of the electronic device 100. Here, the processor may be a general-purpose processor such as a central processing unit (CPU). Alternatively, the processor may be a special-purpose device, such as an ASIC, or other hardware-based processor.
[0071] The controller 170 can demultiplex the streams input through the tuner 110, demodulator 120, external device interface 130, or network interface 135, or process the demultiplexed signals to generate and output signals for video or audio output.
[0072] The display 180 can convert video signals, data signals, OSD signals and control signals processed by the controller 170, or video signals, data signals and control signals received from the external device interface 130, to generate drive signals.
[0073] The display 180 may include a display panel (not shown) having a plurality of pixels.
[0074] The multiple pixels disposed in the display panel may include RGB subpixels. Alternatively, the multiple pixels disposed in the display panel may include RGBW subpixels. The display 180 may convert video signals, data signals, OSD signals, control signals, etc., processed by the controller 170 to generate drive signals for the multiple pixels.
[0075] The display 180 can be a plasma display panel (PDP), a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, or a flexible display, and can also be a 3D display. The 3D display 180 can be classified as either glasses-free or glasses-based.
[0076] Meanwhile, the display 180 can be configured as a touch screen and used as an input device in addition to being an output device.
[0077] The audio output unit 185 receives the audio signal processed by the controller 170 and outputs the audio signal as audio.
[0078] The video signal processed by controller 170 can be input to display 180 and displayed as an image associated with the video signal. Additionally, the video signal processed by controller 170 can be input to an external output device via external device interface 130.
[0079] The audio signal processed by the controller 170 can be output as sound from the audio output unit 185. Additionally, the audio signal processed by the controller 170 can be input to an external output device via the external device interface 130.
[0080] Despite Figure 2 Although not illustrated, controller 170 may include a demultiplexer, an image processor, etc.
[0081] Furthermore, the controller 170 can control the overall operation of the electronic device 100. For example, the controller 170 can control the tuner 110 to select (tune to) a broadcast associated with a user-selected channel or a previously stored channel.
[0082] Additionally, the controller 170 can control the electronic device 100 using user commands or internal programs input through the user input interface 150.
[0083] Meanwhile, the controller 170 can control the display 180 to display images. Here, the images displayed on the display 180 can be still images or videos, and can be 2D or 3D images.
[0084] Furthermore, the controller 170 can cause a predetermined 2D object to be displayed in an image displayed on the monitor 180. For example, the object can be at least one of a connected web page screen (newspaper, magazine, etc.), electronic program guide (EPG), various menus, widgets, icons, still images, videos, or text.
[0085] Additionally, the electronic device 100 may also include an imaging device (not shown). This imaging device can capture images of the user. The imaging device can be implemented as a single camera, but this disclosure is not limited thereto; it can also be implemented as multiple cameras. Furthermore, the imaging device can be embedded in the electronic device 100 and located on top of the display 180, or it can be set up separately. Image information captured by the imaging device can be input to the controller 170.
[0086] The controller 170 can identify the user's position based on the image captured by the imaging device. For example, the controller 170 can determine the distance (z-axis coordinate) between the user and the electronic device 100. Furthermore, the controller 170 can determine the x-axis and y-axis coordinates of the user's position on the display 180.
[0087] The controller 170 can detect the user's posture based on the image captured by the imaging device, each signal detected by the sensor, or a combination thereof.
[0088] The power supply 190 can supply corresponding power to the entire electronic device 100. Specifically, the power supply 190 can supply power to the controller 170, which can be implemented as a system-on-a-chip (SOC), the display 180 for displaying images, and the audio output unit 185 for audio output.
[0089] Specifically, power supply 190 may include a converter (not shown) that converts AC power to DC power and a DC / DC converter (not shown) that converts DC power levels.
[0090] The remote control device 200 can send user input to the user input interface 150. For this purpose, the remote control device 200 can use Bluetooth, radio frequency (RF) communication, infrared communication, ultra-wideband (UWB), ZigBee, etc. Additionally, the remote control device 200 can receive video, audio, or data signals output from the user input interface 150, and display or output such video, audio, or data signals.
[0091] The aforementioned electronic device 100 may be a fixed or mobile digital broadcast receiver capable of receiving digital broadcasts.
[0092] at the same time, Figure 2The block diagram of the electronic device 100 shown is only a block diagram of an embodiment of this disclosure, and the components in the block diagram may be integrated, added, or omitted according to the specifications of the actual implemented electronic device 100.
[0093] That is, two or more components can be combined into one component, or a component can be subdivided into two or more components as needed. Furthermore, the function performed by each block is for the purpose of describing an implementation of this disclosure, and the specific operation or apparatus does not limit the scope of this disclosure.
[0094] Figure 3 In describing Figure 1 The server is shown in the diagram for reference.
[0095] refer to Figure 3 Server 400 may include relay server 410, speech-to-text (STT) server 420, natural language processing (NLP) server 430, user identity server 440, and / or account server 450. Although relay server 410, STT server 420, NLP server 430, user identity server 440, and account server 450 are distinguished from each other in this disclosure, this disclosure is not limited thereto. For example, two or more of relay server 410, STT server 420, NLP server 430, user identity server 440, and account server 450 may be configured as a single server.
[0096] Relay server 410 can communicate with electronic device 100. Relay server 410 can send data between STT server 420, NLP server 430, user identification server 440 and electronic device 100. Relay server 410 can store at least some of the data sent between STT server 420, NLP server 430, user identification server 440 and electronic device 100.
[0097] STT server 420 can receive audio data. STT server 420 can convert the audio data into text data. STT server 420 can send the text data to electronic device 100 via relay server 410. STT server 420 can be referred to as an automatic speech recognition (ASR) server.
[0098] STT server 420 can use a language model to improve the accuracy of speech-to-text conversion. A language model can refer to a model that can calculate the probability of a sentence or the probability of the next word appearing given previous words. For example, the language model can include probabilistic language models such as unigram models, bigram models, and N-gram models. That is, STT server 420 can determine whether the text data has been properly converted from audio data and improve the accuracy of the conversion to text data accordingly.
[0099] NLP server 430 can receive text data. NLP server 430 can perform intent analysis on the received text data. NLP server 430 can send the intent analysis information of the pointer pattern analysis results to electronic device 100 via relay server 410.
[0100] According to the implementation, the NLP server 430 can generate intent analysis information by sequentially performing lexical analysis, syntactic analysis, speech behavior analysis, and dialogue processing steps on text data. The lexical analysis step classifies text data related to the user's spoken words into lexical units (the smallest units with meaning) and determines which part of the speech each classified lexical unit corresponds to. The syntactic analysis step uses the results of the lexical analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and determines the relationships between the classified phrases. Through the syntactic analysis step, the subject, object, and modifiers of the user's spoken words can be determined. The speech behavior analysis step uses the results of the syntactic analysis step to analyze the intent of the user's spoken words. Specifically, the speech behavior analysis step determines the intent of the sentence (such as whether the user is asking a question, making a request, or simply expressing emotion). The dialogue processing step determines whether to respond to the user's speech, provide a response, or request additional information.
[0101] User identification server 440 can receive audio data. User identification server 440 can extract speech features based on the audio data. Here, speech features may include the speech waveform, speech frequency band, speech power spectrum, etc. The extraction of speech features will refer to... Figure 4 and Figure 5 Described later.
[0102] User identification server 440 can obtain a voice feature vector from voice features. User identification server 440 can obtain a voice feature vector from voice features based on linear prediction coefficients, cepstrum, Mel-frequency cepstrum coefficients (MFCC), and filter bank energy.
[0103] User identification server 440 can determine the similarity between multiple feature vectors. User identification server 440 can use cosine similarity, Euclidean similarity, etc., to determine the similarity between multiple feature vectors. Although an example of calculating the similarity between a first voice input and a second voice input based on cosine similarity is described in this disclosure, the method for determining similarity is not limited to this. For example, a first vector associated with a first text and a second vector associated with a second text can be created. The cosine similarity between the first vector and the second vector can be calculated based on Equation 1 below.
[0104] [Formula 1]
[0105] Here, Let represent the dot product of two vectors, and and Cosine similarity represents the magnitudes of two vectors. Specifically, it can be calculated by dividing the dot product of the two vectors by the product of their magnitudes. The range of cosine similarity is from -1 to 1, and the closer the cosine similarity is to 1, the more similar the two vectors are considered.
[0106] User identification server 440 can determine whether the users who emitted the speech are the same based on the similarity between multiple feature vectors. For example, when the similarity between a first feature vector associated with a first voice input and a second feature vector associated with a second voice input is equal to or greater than a predetermined standard, user identification server 440 can determine that the user who emitted the first voice input and the user who emitted the second voice input are the same.
[0107] According to the implementation, the user identification server 440 can use algorithms such as Gaussian mixture model (GMM), hypervector, i-vector, d-vector, x-vector, etc., to process the voice feature vector to obtain a vector. The user identification server 440 can determine whether the users who made the voices are the same based on the similarity between the first vector obtained by processing the first feature vector and the second vector obtained by processing the second feature vector.
[0108] User identification server 440 may store audio data. User identification server 440 may store data about voiceprints (hereinafter referred to as voiceprint information). Here, voiceprint information may include a voice feature vector and / or a vector obtained by processing the voice feature vector.
[0109] User identification server 440 can store a voice database. The voice database may include unique identification information associated with electronic device 100 (hereinafter referred to as device identification information), unique identification information associated with user account (hereinafter referred to as user identification information), voice data mapped to user identification information, and voiceprint information mapped to user identification information.
[0110] Device identification information, user identification information, audio data, and voiceprint information included in the voice database can be stored in the user identification server 440 in association with each other. For example, at least one piece of device identification information, multiple pieces of audio data, and / or multiple pieces of voiceprint information can be mapped to user identification information. That is, it can be interpreted that device identification information, audio data, and voiceprint information are mapped to user accounts and stored in the user identification server 440. In this disclosure, an example will be described in which multiple pieces of audio data and multiple pieces of voiceprint information are all mapped to the user identification information included in the voice database.
[0111] User identification server 440 can update the voiceprint information included in the voice database based on the audio data included in the voice database. For example, user identification server 440 can use an algorithm different from the previously used algorithm to generate voiceprint information related to the audio data included in the voice database. Here, user identification server 440 can change the voiceprint information included in the voice database to the newly generated voiceprint information.
[0112] Account server 450 can manage data about user accounts. Account server 450 can manage user account IDs, passwords, user identification information, device identification information mapped to user accounts, and whether the user agrees to the terms and conditions related to various functions.
[0113] Account server 450 can store a database of user accounts. The database of user accounts may include user account ID, password, user identification information, device identification information mapped to the user account, user account registration date and time, whether the user agrees to the terms and conditions related to various functions, and the date and time of the user's agreement to the terms and conditions.
[0114] Account server 450 can communicate with electronic device 100. For example, account server 450 can create and register user accounts based on data from electronic device 100. For example, account server 450 can approve user account logins based on IDs and passwords received from electronic device 100.
[0115] Figure 4 This is a block diagram used to describe a server configuration according to an embodiment of the present disclosure.
[0116] refer to Figure 4Server 400 may include preprocessor 460, controller 470, communication unit 480 and / or database 490.
[0117] The preprocessor 460 can preprocess voice received through the communication unit 480 or voice stored in the database 490.
[0118] The preprocessor 460 can be implemented as a separate chip from the controller 470, or it can be implemented as a chip included in the controller 470.
[0119] The preprocessor 460 can receive (user-sent) voice signals and filter out noise signals from the voice signals before converting the received voice signals into text data.
[0120] If a preprocessor 460 is provided in the electronic device 100, the preprocessor 460 can identify a start word for activating speech recognition of the electronic device 100. The preprocessor 460 can convert the start word received through the user input interface 150 into text data, and if the converted text data is text data related to a pre-stored start word, then the start word has been identified.
[0121] The preprocessor 460 can convert the denoised voice signal into a power spectrum.
[0122] The power spectrum can be a parameter that indicates the frequency components included in the time-varying waveform of a voice signal and the magnitude of those frequency components.
[0123] The power spectrum shows the distribution of the squared amplitude values of the frequency response of the voice signal waveform. This will be used as a reference. Figure 5 Describe it.
[0124] Figure 5 This is a diagram illustrating an example of converting a voice signal into a power spectrum according to an embodiment of the present disclosure.
[0125] Figure 5 Voice signal 510 is shown. This voice signal 510 may be a signal received from an external device or a signal previously stored in the memory 170.
[0126] The x-axis of the voice signal 510 represents time, and the y-axis represents amplitude.
[0127] The power spectrum processor 463 can convert the voice signal 510, with the x-axis representing time, into a power spectrum 520, with the x-axis representing frequency. The power spectrum processor 463 can use a Fast Fourier Transform (FFT) to convert the voice signal 510 into a power spectrum 520. The x-axis of the power spectrum 520 represents frequency, and the y-axis represents the square of the amplitude.
[0128] Refer again Figure 4 , Figure 4 The functions of the preprocessor 460 and controller 470 described herein can also be performed in the NLP server 430.
[0129] The preprocessor 460 may include a waveform processor 461, a frequency processor 462, a power spectrum processor 463, a speech-to-text (STT) converter 464, etc.
[0130] Waveform processor 461 can extract the waveform of speech.
[0131] The frequency processor 462 can extract the frequency band of speech.
[0132] The power spectrum processor 463 can extract the power spectrum of speech.
[0133] The power spectrum can be the following parameters: when a time-varying waveform is given, it indicates the frequency components included in the waveform and the magnitude of the frequency components.
[0134] The STT converter 464 can convert speech to text. The STT converter 464 can convert speech in a specific language into text in that language.
[0135] The controller 470 can control the overall operation of the server 400. The controller 470 may include a speech analyzer 471, a text analyzer 472, a feature clustering unit 473, a text mapper 474, and / or a speech synthesizer 475.
[0136] The speech analyzer 471 can extract speech characteristic information using one or more of the waveform, frequency band, and power spectrum of the speech preprocessed in the preprocessor 460. The speech characteristic information may include one or more of the following: the speaker's gender, the speaker's voice (or intonation), the pitch of the voice, the speaker's speaking style, the speaker's speaking rate, and the speaker's emotional state. Additionally, the speech characteristic information may also include the speaker's timbre.
[0137] Text analyzer 472 can extract main expressions from the text converted by STT converter 464. When intonation changes between phrases are detected in the converted text, text analyzer 472 can extract phrases with different intonations as main expression phrases. Text analyzer 472 can determine that intonation has changed when the frequency band change between phrases exceeds a preset frequency band. Text analyzer 472 can extract keywords from phrases in the converted text. Keywords can be nouns present in the phrases, but this is only an example.
[0138] The feature clustering unit 473 can classify the speaker's speech type using the speech characteristic information extracted by the speech analyzer 471. The feature clustering unit 473 can classify the speaker's speech type by assigning weights to each type item constituting the speech characteristic information. The feature clustering unit 473 can use attention techniques from a deep learning model to classify the speaker's speech type.
[0139] Text mapper 474 can translate text converted to the first language into text in the second language. Text mapper 474 can map the translated text to the first language text. Text mapper 474 can map the main expressions constituting the first language text to corresponding phrases in the second language. Text mapper 474 can map the phonetic types related to the main expressions constituting the first language text to phrases in the second language. This is for the purpose of applying the categorized phonetic types to phrases in the second language.
[0140] The speech synthesizer 475 can apply the speech type and speaker intonation classified by the feature clustering unit 473 to the main expression of the text translated into a second language in the text mapper 474 to generate synthesized speech.
[0141] The controller 470 can use one or more of the transmitted text data or power spectrum 520 to determine the user's voice characteristics.
[0142] A user's voice characteristics can include their gender, tone of voice, intonation, speech topic, speech rate, and volume.
[0143] The controller 470 can obtain the frequency of the voice signal 510 and the amplitude corresponding to that frequency.
[0144] The controller 470 can use the frequency band of the power spectrum 470 to determine the gender of the user emitting the voice. For example, if the frequency band of the power spectrum 520 is within a preset first frequency band range, the controller 470 can determine that the user is male.
[0145] If the power spectrum 520 falls within a preset second frequency band, the controller 470 can determine that the user is female. Here, the second frequency band can be higher than the first frequency band.
[0146] The controller 470 can determine the pitch of the speech using the frequency band of the power spectrum 520. For example, the controller 470 can determine the pitch of the speech based on the amplitude within a specific frequency band.
[0147] The controller 470 can use the frequency bands of the power spectrum 520 to determine the user's tone of voice. For example, the controller 470 can determine the frequency bands in the power spectrum 520 with amplitudes equal to or greater than a certain level as the user's main vocal range, and determine the main vocal range as the user's tone of voice.
[0148] The controller 470 can determine the user's speaking rate based on the number of syllables emitted per unit time from the converted text data.
[0149] The controller 470 can use the bag-of-words model technique to determine the user's speech topic based on the converted text data.
[0150] The bag-of-words model is a technique that extracts frequently used words based on their frequency in a sentence. Specifically, it extracts unique words from a sentence and represents the frequency of each extracted word as a vector to determine the features of the speech topic. For example, if words such as "running" and "physical strength" appear frequently in text data, the controller 470 can classify the user's speech topic as "exercise."
[0151] The controller 470 can use known text classification techniques to determine the user's voice topic from text data. The controller 470 can extract keywords from the text data and determine the user's voice topic.
[0152] The controller 470 can determine the user's voice volume by taking into account amplitude information across the entire frequency band. For example, the controller 470 can determine the user's voice volume based on the average or weighted average of the amplitudes in each frequency band of the power spectrum 470.
[0153] The communication unit 480 can communicate with an external server via wired or wireless means. The communication unit 480 can communicate with the electronic device 100 via wired or wireless means.
[0154] Database 490 can store the first language speech included in the content. Database 490 can store synthesized speech that has been converted from the first language speech to the second language speech. Database 490 can store the first text associated with the first language speech and the second text translated from the first text into the second language. Database 490 can store various learning models required for speech recognition.
[0155] at the same time, Figure 2 The controller 170 of the illustrated electronic device 100 may include Figure 4 The preprocessor 460 and controller 470 are illustrated in the example. That is, the controller 170 of the electronic device 100 can perform the functions of the preprocessor 460 and controller 470.
[0156] Figure 6This is a block diagram illustrating the configuration of a controller for speech recognition and synthesis in an image display apparatus according to an embodiment of the present disclosure.
[0157] Right now, Figure 6 The speech recognition and synthesis process illustrated herein can be performed by the controller 170 of the electronic device 100 without the use of a server.
[0158] refer to Figure 6 The processor 170 of the electronic device 100 may include an STT engine 610, an NLP engine 620, and a speech synthesis engine 630. Each engine may be hardware or software.
[0159] STT engine 610 can execute Figure 5 The STT server 420 has the following functionality: the STT engine 610 can convert audio data into text data.
[0160] NLP Engine 620 can execute Figure 5 The NLP server 430 shown has the following functionality: Specifically, the NLP engine 620 can obtain intent analysis information indicating the speaker's intent from the converted text data.
[0161] The speech synthesis engine 630 can perform the functions of a speech synthesis server. The speech synthesis engine 630 can search a database for syllables or words related to given text data and synthesize combinations of the searched syllables or words to generate synthesized speech.
[0162] The speech synthesis engine 630 may include a preprocessing engine 631 and a TTS engine 632.
[0163] The preprocessing engine 631 can preprocess text data before generating synthesized speech. Specifically, the preprocessing engine 631 performs tokenization to divide the text data into tokens, which are meaningful units. After tokenization, the preprocessing engine 631 can perform cleaning operations to remove unnecessary characters and symbols, thereby eliminating noise. Subsequently, the preprocessing engine 631 can generate the same tokens by integrating tokens with different expressions. Afterward, the preprocessing engine 631 can remove meaningless tokens (stop words).
[0164] The TTS engine 632 can synthesize speech related to preprocessed text data and generate synthesized speech.
[0165] Figure 7 This is a flowchart of a method for operating an electronic device according to an embodiment of the present disclosure.
[0166] refer to Figure 7In operation S701, electronic device 100 can determine whether a user account is logged into server 400. For example, a user can log into server 400 by entering a user account ID and password.
[0167] According to the implementation, when a user logs into the server 400 for the first time using the electronic device 100 with a user account, the electronic device 100 can include user identification information associated with the user account in the user list. For example, if three different user accounts log into the server 400 using the electronic device 100, the user list stored in the electronic device 100 can include three different user identification information entries.
[0168] In operation S702, electronic device 100 can determine whether a user account logged into server 400 has registered voice-related identification information (hereinafter referred to as voice ID). Here, voice ID may include voiceprint information stored in user identification server 440. For example, server 400 may send information about whether a user account logged into server 400 has registered a voice ID to electronic device 100.
[0169] According to the implementation method, server 400 can determine whether a voice ID has been registered based on whether the voiceprint information is mapped to user identification information (unique identification information associated with a user account logged into server 400). Here, when no voice ID is registered for a user account, the number of voiceprint information entries mapped to the user identification information can be 0.
[0170] According to the implementation method, if the number of voiceprint information entries mapped to the user identification information is two or more predetermined numbers, the server 400 can determine that the voice ID has been registered; and if the number of voiceprint information entries is less than the predetermined number, it can determine that the voice ID has not been registered. For example, in the case of a user account that has registered a voice ID, six different voiceprint information entries can be mapped to the user identification information. For example, in the case of a user account that has not registered a voice ID, five or fewer voiceprint information entries can be mapped to the user identification information.
[0171] According to the implementation, a flag value indicating whether a voice ID has been registered can be mapped to user identification information stored in server 400. Here, the user identification information with the mapped flag value can be stored in user identification server 440 and / or account server 450. Server 400 can determine whether a voice ID has been registered based on the flag value mapped to the user identification information. For example, in the case of a user account without a registered voice ID, the flag value mapped to the user identification information can be 0, and in the case of a user account with a registered voice ID, the flag value mapped to the user identification information can be 1.
[0172] When a user account has not registered a voice ID, the electronic device 100 can begin the voice ID registration process in operation S703. For example, when the voice ID registration process begins, the electronic device 100 can send data including device identification information, user identification information, and a value indicating the start of voice ID registration to the server 400.
[0173] In operation S704, electronic device 100 can output preset text. Electronic device 100 can output any one of multiple preset texts. For example, when electronic device 100 is an image display device 100a, electronic device 100 can output preset text through display 180.
[0174] According to the implementation method, the server 400 can send any one of a plurality of preset texts to the electronic device 100 in a preset order. Here, the electronic device 100 can output the preset text received from the server 400.
[0175] Electronic device 100 can determine whether voice input regarding a preset text has been received during operation S705. For example, electronic device 100 can determine whether voice input has been received via a microphone included in input unit 160 within a preset time period. Here, the voice-related speech signal input via microphone can be sent to controller 170 via user input interface 150. For example, electronic device 100 can determine whether data containing a voice signal related to the user's speech has been received from remote control device 200 within a preset time period.
[0176] When voice input is given regarding a preset text, the electronic device 100 can send audio data containing voice signals related to the voice to the server 400 during operation S706. Here, the electronic device 100 can send device identification information, user identification information, and a language code indicating the language type along with the audio data to the server 400.
[0177] Server 400 can convert speech signals included in audio data received from electronic device 100 into text. Server 400 can determine whether the text converted from the speech signals corresponds to preset text. For example, server 400 can determine whether they correspond to each other based on the similarity between the text converted from the speech signals and the preset text.
[0178] When the text converted from the voice signal corresponds to a preset text, the server 400 can generate voiceprint information related to the voice signal. The server 400 can map the voiceprint information generated about the preset text to user identification information and store it. The server 400 can also map the audio data received about the preset text to user identification information and store it.
[0179] In operation S707, electronic device 100 can determine whether speech processing of a preset text was successful based on the response received from server 400. For example, if the text converted from the speech signal corresponds to the preset text, server 400 can notify electronic device 100 that speech processing was successful. For example, when voiceprint information related to the speech signal is generated, server 400 can notify electronic device 100 that speech processing was successful.
[0180] Meanwhile, in operation S708, when no voice input is received regarding the preset text or when voice processing regarding the preset text fails, the electronic device 100 can determine whether the user should retry inputting voice. For example, the electronic device 100 can retry inputting voice based on user input from the user retrying to input voice. Here, the electronic device 100 can output the preset text again.
[0181] In operation S709, if the speech processing of the preset text is successful, the electronic device 100 can determine whether the processing of all texts has been completed. For example, if all speech processing of six texts is successful, the processing of all texts can be completed. At the same time, when the processing of five preset texts is completed, the electronic device 100 can output the last preset text.
[0182] In operation S710, when processing of all text is complete, electronic device 100 can end the voice ID registration process. For example, when electronic device 100 is an image display device 100a, electronic device 100 can output a screen indicating that voice ID registration is complete via display 180. For example, electronic device 100 can send data indicating that voice ID registration is complete to account server 450.
[0183] According to one implementation, electronic device 100 can log in to server 400 with a user account that has registered a voice ID based on the voice input to electronic device 100.
[0184] When voice input is received, the electronic device 100 can send audio data related to the input voice to the server 400. Here, the electronic device 100 can send device identification information, a user list, and a language code indicating the language type along with the audio data to the server 400.
[0185] Server 400 can generate voiceprint information about the speech input to electronic device 100 based on audio data received from electronic device 100. Server 400 can search a voice database for voiceprint information (hereinafter referred to as candidate voiceprint information) related to user identification information included in the user list received from electronic device 100. Server 400 can determine whether the candidate voiceprint information is related to the generated voiceprint information. Server 400 can determine the user identification information mapped to the candidate voiceprint information related to the generated voiceprint information as user identification information related to the speech input to electronic device 100. If no candidate voiceprint information is related to the generated voiceprint information, server 400 can determine that no user identification information is related to the speech input to electronic device 100.
[0186] Server 400 can send the processing results of audio data received from electronic device 100 to electronic device 100. For example, server 400 can send text converted from audio data received from electronic device 100, the results of intent analysis performed on the converted text, voice-related user identification information, etc., to electronic device 100.
[0187] Electronic device 100 can perform voice-related operations input to it based on the processing results of audio data received from server 400. For example, if the voice-related user identification information is not related to the user account currently logged into server 400, electronic device 100 can log into server 400 using the user account associated with the voice-related user identification information. For example, if the voice-related user identification information is associated with the user account currently logged into server 400, or if no voice-related user identification information exists, electronic device 100 can maintain the login status of the user account currently logged into server 400.
[0188] Figure 8 This is a flowchart of an operating system method according to an embodiment of the present disclosure.
[0189] refer to Figure 8 In operation S801, electronic device 100 can log in to server 400 using a user account.
[0190] In operation S802, electronic device 100 can begin the process of registering a voice ID.
[0191] In operation S803, the electronic device 100 can output the first text among multiple preset texts.
[0192] In operation S804, electronic device 100 can receive first speech about the first text.
[0193] In operation S805, electronic device 100 can send first audio data containing voice signals related to the first speech to server 400.
[0194] In operation S806, server 400 can process first speech about first text based on first audio data received from electronic device 100. Server 400 can convert the speech signals related to the first speech included in the first audio data received from electronic device 100 into text. Server 400 can determine whether the text converted from the speech signals related to the first speech corresponds to the first text.
[0195] In operation S807, server 400 can notify electronic device 100 that processing of the first speech is complete. For example, server 400 can notify electronic device 100 of successful processing of the first speech based on the fact that the text converted from the speech signal associated with the first speech corresponds to the first text.
[0196] Furthermore, server 400 can generate first voiceprint information about the first speech based on the fact that the text converted from the voice signal associated with the first speech corresponds to the first text.
[0197] In operation S808, server 400 can store first audio data and first voiceprint information related to the first voice. Server 400 can map the first audio data and first voiceprint information to user identification information associated with the logged-in user account and store it.
[0198] Electronic device 100 can output the second to fifth text messages in stages. Electronic device 100 can sequentially receive the second to fifth voice messages associated with the second to fifth text messages. Electronic device 100 can sequentially send the second to fifth audio data associated with the second to fifth voice messages to server 400.
[0199] Server 400 can process second to fifth voice data based on second to fifth audio data received from electronic device 100. Additionally, server 400 can sequentially generate and store second to fifth voiceprint information associated with the second to fifth voice data.
[0200] In operation S809, electronic device 100 can output a sixth text from multiple preset texts.
[0201] In operation S810, electronic device 100 can receive a sixth voice regarding a sixth text.
[0202] In operation S811, electronic device 100 can send sixth audio data containing voice signals related to the sixth speech to server 400.
[0203] In operation S812, server 400 can process sixth speech related to sixth text based on sixth audio data received from electronic device 100. Server 400 can convert the speech signals related to the sixth speech included in the sixth audio data received from electronic device 100 into text. Server 400 can determine whether the text converted from the speech signals related to the sixth speech corresponds to the sixth text.
[0204] In operation S813, server 400 can notify electronic device 100 that processing for the sixth voice is complete.
[0205] At the same time, when the text converted from the voice signal related to the sixth speech corresponds to the sixth text, the server 400 can generate sixth voiceprint information about the sixth speech based on the voice signal related to the sixth speech.
[0206] In operation S814, server 400 can store sixth audio data and sixth voiceprint information related to the sixth voice. Server 400 can map the sixth audio data and sixth voiceprint information to user identification information associated with the logged-in user account and store them. Here, six different audio data and multiple voiceprint information can be mapped to user identification information associated with the logged-in user account.
[0207] In operation S815, electronic device 100 can terminate the voice ID registration process. For example, electronic device 100 can terminate the voice ID registration process based on the completion of processing six different preset texts.
[0208] refer to Figure 9 If a user account is not logged into server 400, electronic device 100 can output a login screen 900 related to logging into server 400 via display 180. Login screen 900 may include an object 910 indicating a non-login state and a login object 920 for performing a login. When the user selects login object 920 using pointer 205 associated with remote control device 200, electronic device 100 can output a screen for entering an ID and password. Here, the user can log into server 400 by entering their user account ID and password.
[0209] refer to Figure 10When a user account logged into server 400 has not registered a voice ID, electronic device 100 can output a first account screen 1000 associated with the user account that has not registered a voice ID. The first account screen 1000 may include an object 1010 indicating the logged-in user account and an object 1020 regarding voice ID registration. When the user selects the object 1020 regarding voice ID registration using pointer 205, electronic device 100 can begin the voice ID registration process.
[0210] refer to Figures 11a to 11c When a voice ID has been registered in a user account logged into server 400, electronic device 100 can display a second account screen 1100 associated with the user account that has registered the voice ID. The second account screen 1100 may include an object 1110 indicating the logged-in user account, a re-registration object 1120 for re-registering the voice ID, a deletion object 1130 for deleting the voice ID, and an activation object 1140 for using functions associated with the voice ID. The user can use pointer 205 to select activation object 1140 to activate or deactivate the functions associated with the voice ID.
[0211] When a user selects re-register object 1120 using pointer 205, the electronic device 100 can display a first notification screen 1121, which indicates that due to the re-registration of the voice ID, the data associated with the previously registered voice ID has been deleted from the server 400. The user can use pointer 205 to select confirm object 1123 to re-register the voice ID, or select cancel object 1125 to maintain the previously registered voice ID.
[0212] When a user selects to delete object 1130 using pointer 205, the electronic device 100 can display a second notification screen 1131, which indicates that due to the deletion of the voice ID, the data associated with the previously registered voice ID has been deleted from the server 400. The user can use pointer 205 to select confirm object 1133 to delete the voice ID, or select cancel object 1135 to retain the previously registered voice ID.
[0213] refer to Figure 12 When the user selects object 1020 for voice ID registration on the first account screen 1000, or selects confirmation object 1123 on the first notification screen 1121, the electronic device 100 can output a start screen 1200 for starting voice ID registration. When the user selects start object 1210 using pointer 205, the electronic device 100 can output a text screen for displaying preset text.
[0214] refer to Figure 13The electronic device 100 can output a text screen 1300 for displaying any one of multiple preset texts. The text screen 1300 may include preset text 1301, text sequence number 1302, end object 1310 for ending the process of registering a voice ID, and input object 1320 for receiving voice.
[0215] The voice ID registration process can end when the user selects the end object 1310 using pointer 205. For example, when the voice ID registration process ends, all data stored on server 400 during the voice ID registration process can be deleted.
[0216] When the user selects input object 1320 using pointer 205, the electronic device 100 can receive voice about the text.
[0217] According to the implementation, when a user presses a predetermined button (e.g., a voice input button) included in the remote control device 200 while the text screen 1300 is displayed, the electronic device 100 can receive voice about the text based on the user input received from the remote control device 200 regarding the pressing of the predetermined button.
[0218] Simultaneously, according to the embodiment, when a user presses a predetermined button (e.g., a voice input button) included in the remote control device 200 during the voice ID registration process, the electronic device 100 can stop the voice ID registration process based on user input received from the remote control device 200 regarding the pressing of the predetermined button. Here, user input pressing the predetermined button (e.g., a voice input button) included in the remote control device 200 can correspond to user input initiating speech recognition of speech received through the remote control device 200. The electronic device 100 can perform speech recognition-related operations on audio data containing voice signals received from the remote control device 200.
[0219] refer to Figure 14 When the speech processing of the text is successful, the electronic device 100 can display a success screen 1400 indicating the success of the text processing. The success screen 1400 may include an object 1410 indicating the success of the speech processing of the text.
[0220] refer to Figure 15 When no voice input is received regarding text or when voice processing regarding text fails, the electronic device 100 may display a failure screen 1500 indicating the failure of voice processing regarding text. The failure screen 1500 may include a termination object 1510 indicating the termination of the voice ID registration process, a re-input object 1520 indicating a retry to receive voice, etc. When the user selects the re-input object 1520 using pointer 205, the electronic device 100 can receive voice input regarding text again.
[0221] refer to Figure 16 When all text processing is complete, the electronic device 100 can display a completion screen 1600 indicating that the voice ID registration is complete. The completion screen 1600 may include an object 1610 indicating the user account for the registered voice ID, a completion object 1620 indicating the completion of the voice ID registration process, etc. When the user selects the completion object 1620 using pointer 205, the electronic device 100 can complete the voice ID registration process.
[0222] Figure 17a and Figure 17b This is a flowchart of a method for operating an electronic device according to another embodiment of the present disclosure. (The details will be omitted.) Figure 7 A detailed description of the redundant parts described in the document.
[0223] refer to Figure 17a In operation S1701, electronic device 100 can determine whether a user account has logged into server 400.
[0224] In operation S1702, electronic device 100 can determine whether a voice ID has been registered for a user account that has logged into server 400.
[0225] If a voice ID has not been registered for the user account, the electronic device 100 can initiate the voice ID registration process in operation S1703.
[0226] In operation S1704, electronic device 100 can determine whether a user account is one that has temporarily stored a voice ID. For example, server 400 can send information about whether a user account logged into server 400 has temporarily stored a voice ID to electronic device 100.
[0227] According to one implementation, server 400 can determine whether a voice ID has been temporarily stored based on whether voiceprint information is mapped to user identification information (unique identification information associated with a user account logged into server 400). In the case of a user account where a voice ID has been temporarily stored, more than a minimum number of voiceprint information entries can be mapped to the user identification information. Here, the minimum number can be less than a predetermined number of 6.
[0228] According to one implementation, server 400 can determine whether a voice ID has been temporarily stored based on a flag value mapped to user identification information. For example, in the case of a user account where a voice ID has been temporarily stored, the flag value mapped to user identification information can be 2.
[0229] In operation S1705, electronic device 100 can output preset text. For example, when the voice ID is not temporarily stored, server 400 can send the first text from multiple preset texts to electronic device 100.
[0230] In operation S1706, the electronic device 100 can determine whether voice input regarding a preset text has been received.
[0231] When voice input is given about a preset text, the electronic device 100 can send audio data containing voice signals related to the voice to the server 400 in operation S1707.
[0232] In operation S1708, electronic device 100 can determine whether speech processing of preset text was successful based on the response received from server 400.
[0233] In operation S1709, when no voice input is received for the preset text or when voice processing for the preset text fails, the electronic device 100 can determine whether to retry the voice input.
[0234] In operation S1710, when the speech processing of the preset text is successful, the electronic device 100 can determine whether the processing of all texts has been completed.
[0235] When all text processing in operation S1711 is complete, the electronic device 100 can complete the process of registering the voice ID.
[0236] refer to Figure 17b When a voice ID is temporarily stored in operation S1712, the electronic device 100 can output any one of multiple preset texts, excluding the text related to the voiceprint information mapped to the user identification information (hereinafter referred to as unsaved text). For example, when the voiceprint information related to the first to third texts is mapped to the user identification information (unique identification information related to the user account logged into the server 400), the server 400 can send the fourth text as unsaved text to the electronic device 100. Here, the electronic device 100 can output the fourth text received from the server 400.
[0237] In operation S1713, the electronic device 100 can determine whether voice input regarding unsaved text has been received.
[0238] When voice input is given regarding unsaved text, electronic device 100 can send audio data containing voice-related voice signals to server 400 in operation S1714.
[0239] In operation S1715, electronic device 100 can determine whether the previous user who temporarily stored the voice ID is the same as the current user who spoke about the unsaved text. For example, electronic device 100 can determine whether the previous user is the same as the current user based on the result of determining whether the users are the same received from server 400.
[0240] Server 400 can convert speech signals included in audio data received from electronic device 100 into text. Server 400 can determine whether the text converted from the speech signal is related to unsaved text. For example, server 400 can determine whether the converted text and unsaved text are related to each other based on the similarity between the text converted from the speech signal and the unsaved text.
[0241] When the text converted from the speech signal is related to the unsaved text, the server 400 can generate voiceprint information (hereinafter referred to as unsaved voiceprint information) related to the speech of the unsaved text based on the speech signal included in the audio data received from the electronic device 100. The server 400 can determine whether at least one of the voiceprint information mapped to the user identification information is related to the unsaved voiceprint information. For example, when voiceprint information related to the first text to the third text is mapped to the user identification information, the server 400 can calculate the similarity between the first voiceprint information related to the first text and the unsaved voiceprint information. Here, when the similarity between the first voiceprint information and the unsaved voiceprint information is greater than a predetermined criterion, the server 400 can determine that the first voiceprint information and the unsaved voiceprint information are related to each other. Additionally, when the first voiceprint information and the unsaved voiceprint information are related to each other, the server 400 can determine that the previous user and the current user are the same.
[0242] If the previous user is different from the current user, the electronic device 100 can determine in operation S1716 not to use the data related to the previous user stored in the server 400. For example, if the electronic device 100 is an image display device 100a, the electronic device 100 can display a screen on the display 180 indicating that the previous user is different from the current user.
[0243] When the previous user is the same as the current user, the server 400 can maintain the voiceprint information mapped to the user identification information. For example, while mapping voiceprint information associated with the first to third texts to the user identification information, the server 400 can map unsaved voiceprint information to the user identification information and store it as voiceprint information associated with the fourth text. Additionally, the server 400 can sequentially send the fifth and sixth texts to the electronic device 100.
[0244] When the previous user is different from the current user, server 400 can delete the voiceprint information mapped to the user identification information. Here, server 400 can also delete the audio data mapped to the user identification information. For example, server 400 can delete the voiceprint information mapped to the user identification information related to the first to third texts, and map the unsaved voiceprint information to the user identification information, storing it as voiceprint information related to the fourth text. Additionally, server 400 can sequentially send multiple preset texts, excluding the fourth text, to electronic device 100.
[0245] In operation S1717, when no voice input is received regarding the unsaved text or when voice processing regarding the unsaved text fails, the electronic device 100 can determine whether to retry the voice input. For example, the electronic device 100 can retry the voice input based on user input for retrying the voice input. Here, the electronic device 100 can output the unsaved text again.
[0246] Figure 18 This is a flowchart of an operating system method according to an embodiment of the present disclosure when a voice ID is not registered and is temporarily stored.
[0247] refer to Figure 18 In operation S1801, electronic device 100 can log in to server 400 using a user account.
[0248] In operation S1802, electronic device 100 can initiate the process of registering a voice ID.
[0249] In operation S1803, the electronic device 100 can output the first text from multiple preset texts.
[0250] In operation S1804, electronic device 100 can receive first speech about the first text.
[0251] In operation S1805, electronic device 100 can send first audio data containing a voice signal related to the first voice to server 400.
[0252] In operation S1806, server 400 can process first speech about first text based on first audio data received from electronic device 100.
[0253] In operation S1807, server 400 can notify electronic device 100 that the processing of the first voice is complete.
[0254] In operation S1808, server 400 can store first audio data and first voiceprint information about the first speech.
[0255] In operation S1809, electronic device 100 can output a second text from multiple preset texts.
[0256] In operation S1810, the electronic device 100 can receive a second voice regarding the second text.
[0257] In operation S1811, electronic device 100 can send second audio data containing voice signals related to the second speech to server 400.
[0258] In operation S1812, server 400 can process second speech about second text based on second audio data received from electronic device 100.
[0259] In operation S1813, server 400 can notify electronic device 100 that the processing of the second voice is complete.
[0260] In operation S1814, server 400 can store second audio data and second voiceprint information about the second voice.
[0261] In operation S1815, electronic device 100 can notify server 400 of the interruption of voice ID registration.
[0262] For example, when a user presses the power button included in the remote control device 200, the electronic device 100 can be shut down based on a signal received from the remote control device 200 related to the user's input of pressing the power button. Here, the electronic device 100 can notify the server 400 of the interruption of voice ID registration based on the power being turned off.
[0263] For example, when a user presses a button associated with a predetermined function (e.g., an OTT service button) included in the remote control device 200, the electronic device 100 can execute the predetermined function based on a signal received from the remote control device 200 related to the user input of pressing the button associated with the predetermined function (e.g., the OTT service button). In this case, the electronic device 100 can notify the server 400 of an interruption in voice ID registration based on the execution of the predetermined function. When the execution of the predetermined function terminates, the electronic device 100 can output a notification to the user regarding the interruption of voice ID registration. The electronic device 100 can determine whether to resume interrupted voice ID registration based on user input. When resumed, the electronic device 100 can begin the voice ID registration process.
[0264] For example, when a user presses a voice recognition-related button (e.g., a voice input button) included in the remote control device 200, the electronic device 100 can perform a voice recognition function based on a signal received from the remote control device 200 related to the user's input after pressing the voice recognition-related button. In this case, the electronic device 100 can notify the server 400 of an interruption in voice ID registration based on the execution of the voice recognition function. When the execution of the voice recognition function terminates, the electronic device 100 can output a notification to the user regarding the interruption of voice ID registration. The electronic device 100 can determine whether to resume interrupted voice ID registration based on user input. When resuming interrupted voice ID registration, the electronic device 100 can initiate the voice ID registration process.
[0265] In operation S1816, when a notification of interruption of voice ID registration is received from electronic device 100, server 400 can check the number of voiceprint information entries stored by mapping to user identification information. Here, if the number of voiceprint information entries stored by mapping to user identification information is 2 (less than the minimum number (e.g., 3)), server 400 can delete the audio data and voiceprint information stored by mapping to user identification information.
[0266] Meanwhile, if the number of voiceprint information entries stored by mapping to user identification information is equal to or greater than the minimum number (e.g., 3), then server 400 can maintain the audio data and voiceprint information stored by mapping to user identification information.
[0267] Figure 19 This is a flowchart of an operating system method according to an embodiment of the present disclosure when a voice ID has been temporarily stored.
[0268] refer to Figure 19 In operation S1901, electronic device 100 can log in to server 400 using a user account.
[0269] In operation S1902, electronic device 100 can initiate the process of registering a voice ID.
[0270] In operation S1903, server 400 may send information about unsaved text to electronic device 100 based on the voice ID having been temporarily stored. For example, server 400 may send a fourth text as unsaved text to electronic device 100 based on the voiceprint information associated with the first to third texts being mapped to user identification information (unique identification information associated with the logged-in user account).
[0271] In operation S1904, electronic device 100 can output a fourth text as unsaved text.
[0272] In operation S1905, the electronic device 100 can receive a fourth voice regarding a fourth text.
[0273] In operation S1906, electronic device 100 can send fourth audio data containing voice signals related to the fourth speech to server 400.
[0274] In operation S1907, server 400 can process fourth speech about fourth text based on fourth audio data received from electronic device 100.
[0275] Server 400 can determine whether the fourth voiceprint information generated based on the fourth audio data is related to at least one voiceprint information mapped to user identification information. For example, server 400 can determine whether the first voiceprint information related to the first text is related to the fourth voiceprint information based on the similarity between the first voiceprint information and the fourth voiceprint information.
[0276] In operation S1908, server 400 may notify electronic device 100 of the result that the previous user who temporarily stored the voice ID is the same as the user who issued the fourth voice. For example, server 400 may determine that the previous user and the current user are the same based on the correlation between the first voiceprint information and the fourth voiceprint information.
[0277] In operation S1909, server 400 can notify electronic device 100 that the processing of the fourth voice is complete.
[0278] In operation S1910, server 400 can store fourth audio data and fourth voiceprint information about the fourth voice. Server 400 can store the fourth audio data and fourth voiceprint information by mapping them to user identification information associated with the logged-in user account.
[0279] Electronic device 100 can output fifth text. Electronic device 100 can receive fifth speech associated with the fifth text. Electronic device 100 can send fifth audio data associated with the fifth speech to server 400.
[0280] Server 400 can process fifth speech based on fifth audio data received from electronic device 100. Additionally, server 400 can generate and store fifth audio information related to the fifth speech.
[0281] In operation S1911, electronic device 100 can output a sixth text from multiple preset texts.
[0282] In operation S1912, electronic device 100 can receive a sixth voice regarding a sixth text.
[0283] In operation S1913, electronic device 100 can send sixth audio data containing voice signals related to the sixth speech to server 400.
[0284] In operation S1914, server 400 can process sixth speech about sixth text based on sixth audio data received from electronic device 100.
[0285] In operation S1915, server 400 can notify electronic device 100 that the processing of the sixth voice is complete.
[0286] In operation S1916, server 400 can store sixth audio data and sixth voiceprint information about the sixth voice.
[0287] In operation S1917, electronic device 100 can complete the process of registering voice ID.
[0288] Figure 20 This is a flowchart of an operating system method according to another embodiment of the present disclosure when a voice ID is temporarily stored.
[0289] refer to Figure 20 In operation S2001, electronic device 100 can log in to server 400 using a user account.
[0290] In operation S2002, electronic device 100 can initiate the process of registering a voice ID.
[0291] In operation S2003, server 400 can send information about unsaved text to electronic device 100 based on the fact that the voice ID has been temporarily stored. For example, server 400 can send a fourth text as unsaved text to electronic device 100 based on the fact that voiceprint information associated with the first to third texts has been mapped to user identification information (unique identification information associated with the logged-in user account).
[0292] In operation S2004, electronic device 100 can output a fourth text as unsaved text.
[0293] In operation S2005, electronic device 100 can receive a fourth voice regarding a fourth text.
[0294] In operation S2006, electronic device 100 can send fourth audio data containing voice signals related to the fourth speech to server 400.
[0295] In operation S2007, server 400 can process fourth speech about fourth text based on fourth audio data received from electronic device 100.
[0296] In operation S2008, server 400 may notify electronic device 100 of the result that the previous user who temporarily stored the voice ID is different from the user who issued the fourth voice message. For example, server 400 may determine that the previous user and the current user are different based on the fact that the first voiceprint information and the fourth voiceprint information are unrelated. In this case, server 400 may delete the voiceprint information mapped to identification information (unique identification information associated with the logged-in user account).
[0297] In operation S2009, server 400 can notify electronic device 100 that the processing of the fourth voice is complete.
[0298] In operation S2010, server 400 can store fourth audio data and fourth voiceprint information related to the fourth voice. Server 400 can store the fourth audio data and fourth voiceprint information by mapping them to user identification information associated with the logged-in user account.
[0299] Furthermore, server 400 can send any text from the preset text other than the fourth text to electronic device 100. In this disclosure, an example will be described in which the fifth and sixth texts are sent to electronic device 100 in sequence, and then the first to third texts are sent to electronic device 100.
[0300] Electronic device 100 can output fifth text. Electronic device 100 can receive fifth speech associated with the fifth text. Electronic device 100 can send fifth audio data associated with the fifth speech to server 400.
[0301] Server 400 can process fifth speech based on fifth audio data received from electronic device 100. Additionally, server 400 can generate and store fifth audio information related to the fifth speech.
[0302] In operation S2011, electronic device 100 can output a sixth text from multiple preset texts.
[0303] In operation S2012, electronic device 100 can receive a sixth voice regarding a sixth text.
[0304] In operation S2013, electronic device 100 can send sixth audio data containing voice signals related to the sixth speech to server 400.
[0305] In operation S2014, server 400 can process sixth speech about sixth text based on sixth audio data received from electronic device 100.
[0306] In operation S2015, server 400 can notify electronic device 100 that the processing of the sixth voice is complete.
[0307] In operation S2016, server 400 can store sixth audio data and sixth voiceprint information about the sixth voice.
[0308] In operation S2017, electronic device 100 can output the first text from multiple preset texts.
[0309] In operation S2018, electronic device 100 can receive first speech about the first text.
[0310] In operation S2019, electronic device 100 can send first audio data containing voice signals related to the first speech to server 400.
[0311] In operation S2020, server 400 can process first speech about first text based on first audio data received from electronic device 100.
[0312] In operation S2021, server 400 can notify electronic device 100 that the processing of the first voice is complete.
[0313] In operation S2022, server 400 can store first audio data and first voiceprint information about the first speech.
[0314] Electronic device 100 can output second text. Electronic device 100 can receive second speech related to the second text. Electronic device 100 can send second audio data associated with the second speech to server 400.
[0315] Server 400 can process the second speech based on the second audio data received from electronic device 100. Additionally, server 400 can generate and store second audio information related to the second speech.
[0316] In operation S2023, electronic device 100 can output a third text from multiple preset texts.
[0317] In operation S2024, electronic device 100 can receive third voice regarding third text.
[0318] In operation S2025, electronic device 100 can send third audio data containing voice signals related to the third voice to server 400.
[0319] In operation S2026, server 400 can process third speech about third text based on third audio data received from electronic device 100.
[0320] In operation S2027, server 400 can notify electronic device 100 that the processing of the third voice is complete.
[0321] In operation S2028, server 400 can store third audio data and third voiceprint information about the third voice.
[0322] In operation S2029, electronic device 100 can complete the process of registering voice ID.
[0323] As described above, according to at least one embodiment of this disclosure, identification information about a user's voice can be registered in a user account.
[0324] Additionally, according to at least one embodiment of this disclosure, a user can be identified based on their voice.
[0325] Additionally, according to at least one embodiment of this disclosure, it is possible to log in to a user's account based on the user's voice identifier.
[0326] Additionally, according to at least one embodiment of this disclosure, continuity of identification information for users who have previously attempted to register can be maintained during the process of registering identification information about a user's voice to a user account.
[0327] Reference Figures 1 to 20 According to one aspect of this disclosure, an electronic device 100 includes: a display 180, an external device interface 130 configured to communicate with a remote control device 200, a network interface 135 configured to communicate with a server 400, a user input interface 150 configured to send signals related to user input, and a controller 170, wherein the controller 170 is configured to: display preset text related to the registration of the identification information on the display 180 while performing a process of registering voice-related identification information for a user account logged into the server 400; send data containing the voice signal to the server 400 when a voice signal related to the preset text is received through the user input interface 150; complete the registration of the identification information based on the processing of the voice signal related to the preset text by the server 400; and interrupt the registration of the identification information process when a predetermined input related to voice recognition is received from the remote control device 200 during the registration of the identification information.
[0328] Furthermore, according to one aspect of this disclosure, the identification information may include a feature vector of the voiceprint of the speech.
[0329] Furthermore, according to one aspect of this disclosure, the controller 170 may, upon receiving the predetermined input, send data received from the remote control device 200 and containing voice signals related to the predetermined input to the server 400, and perform an operation related to the predetermined input based on the server 400's processing of the voice signals related to the predetermined input.
[0330] Furthermore, according to one aspect of this disclosure, the controller 170 may, when initiating the process of registering the identification information for the user account, determine whether the identification information has been temporarily stored in the server 400; if the identification information has not been temporarily stored in the server 400, display a predetermined number of texts in stages; and if the identification information has been temporarily stored in the server 400, display some texts in stages, excluding texts related to the identification information temporarily stored in the server 400.
[0331] Furthermore, according to one aspect of this disclosure, the controller 170 may display a notification of interruption of the process of registering the identification information via the display 180 based on the termination of the operation associated with the predetermined input, and determine whether to start the process of registering the identification information based on the user input received through the user input interface 150.
[0332] According to one aspect of this disclosure, a server 400 includes: a communication unit 480 configured to communicate with an electronic device 100, a database 490, and a controller 470, wherein the controller 470 is configured to: while performing a process of registering voice-related identification information for a user account logged into the server 400, convert voice signals included in data received from the electronic device 100 into text, generate identification information related to the voice signal based on the converted text and a preset text related to the registration of the identification information, map the generated identification information to user identification information related to the user account, store the mapped information in the database 490, and maintain or delete the identification information mapped to the user identification information when a notification of interruption in the registration of the identification information is received from the electronic device 100 during the registration process of the identification information.
[0333] Furthermore, according to one aspect of this disclosure, the identification information may include a feature vector of the voiceprint of the speech.
[0334] Furthermore, according to one aspect of this disclosure, the controller 470 can maintain the identifier information mapped to the user identifier information based on the number of identifier information entries mapped to the user identifier information being equal to or greater than a preset minimum number, and delete the identifier information mapped to the user identifier information based on the number of identifier information entries mapped to the user identifier information being less than the preset minimum number.
[0335] Furthermore, according to one aspect of this disclosure, the controller 470 can, based on the start of the registration process for the identification information for the user account, determine whether the identification information is mapped to the user identification information; based on the fact that the identification information is not mapped to the user identification information, send a predetermined number of texts to the electronic device 100 in stages; and based on the fact that the identification information is mapped to the user identification information, send some texts, excluding texts related to the identification information mapped to the user identification information, to the electronic device 100 in stages.
[0336] Furthermore, according to one aspect of this disclosure, the controller 470 can, based on the commencement of the registration process for the identification information for the user account, determine whether the identification information is mapped to the user identification information; based on the identification information being mapped to the user identification information, send a first text, excluding text related to the identification information mapped to the user identification information, from a plurality of texts to the electronic device 100; generate identification information about a first voice signal related to the first text and included in data received from the electronic device 100; maintain the identification information mapped to the user identification information based on the fact that the identification information about the first voice signal is related to at least one of the identification information mapped to the user identification information; and delete the identification information mapped to the user identification information based on the fact that the identification information about the first voice signal is not related to at least one of the identification information mapped to the user identification information.
[0337] Furthermore, according to one aspect of this disclosure, the controller 470 can map the identification information about the first voice signal to the user identification information and store the mapped information in the database 490.
[0338] Furthermore, according to one aspect of this disclosure, the controller 470 may send a result to the electronic device 100 indicating that the current user is the same as the previous user based on the correlation between the identification information regarding the first voice signal and at least one of the identification information mapped to the user identification information, and send a result to the electronic device 100 indicating that the current user is different from the previous user based on the fact that the identification information regarding the first voice signal is not correlated with at least one of the identification information mapped to the user identification information.
[0339] Furthermore, according to one aspect of this disclosure, the controller 470 can map the voice signal to the user identification information based on the correlation between the converted text and the preset text, and store the mapped signal in the database 490.
[0340] Furthermore, according to one aspect of this disclosure, the controller 470 may use a first algorithm to generate identification information related to the voice signal mapped to the user identification information, and change the identification information mapped to the user identification information and generated using a second algorithm to the identification information generated using the first algorithm.
[0341] According to one aspect of this disclosure, a system 10 includes an electronic device 100 and a server 400, wherein the electronic device 100 is configured to: display preset text related to the registration of the identification information while performing a process of registering voice-related identification information for a user account logged into the server 400; when receiving a voice signal related to the preset text, send data containing the voice signal to the server 400; complete the registration of the identification information based on the server 400's processing of the voice signal related to the preset text; and during the registration of the identification information, when receiving a voice signal related to the preset text from a remote control device 200... When a predetermined input is received, the process of registering the identification information is stopped, and the server 400 is configured to: convert the voice signal included in the data received from the electronic device 100 into text, generate identification information related to the voice signal based on the correlation between the converted text and the preset text, map the generated identification information to user identification information related to the user account, store the mapped information in the database 490, and maintain or delete the identification information mapped to the user identification information when a notification of interruption in the registration of the identification information is received from the electronic device 100 during the registration process.
[0342] The accompanying drawings are provided only to facilitate understanding of the embodiments disclosed in this specification, and the technical concepts disclosed in this specification are not limited to the drawings. It should be understood that this disclosure includes all modifications, equivalents, and substitutions falling within the scope of the ideas and techniques of this disclosure.
[0343] Furthermore, the method of operation according to this disclosure can be implemented as processor-readable code recorded on a processor-readable recording medium. This processor-readable recording medium can include any type of recording device storing processor-readable data. Examples of such a medium include read-only memory (ROM), random access memory (RAM), optical disc read-only memory (CD-ROM), magnetic tape, floppy disk, optical data storage devices, etc., and may also include media implemented in carrier wave form, such as transmission via the Internet. Moreover, the processor-readable recording medium can be distributed across connected computer systems via a network, enabling the processor-readable code to be stored and executed in a distributed manner.
[0344] While preferred embodiments of the present disclosure have been shown and described above, the present disclosure is not limited to the specific embodiments described above. Various modifications may be made by those skilled in the art without departing from the spirit of the present disclosure as claimed in the appended claims, and such modifications should not be construed as departing from the technical concept or scope of the present disclosure.
Claims
1. An electronic device, the electronic device comprising: monitor; An external device interface configured to communicate with a remote control device; A network interface configured to communicate with a server; A user input interface, configured to send signals related to user input; as well as Controller The controller is configured as follows: While performing the registration process for voice-related identification information of a user account logged into the server, preset text related to the registration of the identification information is displayed on the screen. When a voice signal related to the preset text is received through the user input interface, data containing the voice signal is sent to the server. The process of registering the identification information is completed based on the processing of the voice signal related to the preset text by the server. and During the process of registering the identification information, the process of registering the identification information is interrupted when a predetermined input related to voice recognition is received from the remote control device.
2. The electronic device according to claim 1, wherein, The identification information includes a feature vector of the voiceprint related to the speech.
3. The electronic device according to claim 1, wherein, The controller is configured to: After receiving the predetermined input, data received from the remote control device and containing voice signals related to the predetermined input is sent to the server, and Based on the processing of the voice signal associated with the predetermined input by the server, the operation associated with the predetermined input is performed.
4. The electronic device according to claim 1, wherein, The controller is configured to: When the process of registering the identification information for the user account begins, it is determined whether the identification information has been temporarily stored in the server. Based on the fact that the identification information is not temporarily stored in the server, a predetermined number of text entries are displayed in stages, and Based on the fact that the identification information has been temporarily stored in the server, some texts from multiple texts are displayed in stages, excluding the texts related to the identification information that has been temporarily stored in the server.
5. The electronic device according to claim 1, wherein, The controller is configured to: Based on the termination of the operation associated with the predetermined input, a notification regarding the interruption of the process of registering the identification information is displayed on the screen, and The process of determining whether to start registering the identification information based on the user input received through the user input interface.
6. A server, the server comprising: The communication unit is configured to communicate with an electronic device; database; as well as Controller The controller is configured as follows: While performing the process of registering voice-related identification information for user accounts logged into the server, voice signals included in the data received from the electronic device are converted into text; Based on the converted text and preset text related to the registration of the identification information, identification information related to the voice signal is generated; The generated identification information is mapped to user identification information associated with the user account, and the mapped information is stored in the database; and During the process of registering the identification information, if a notification is received from the electronic device regarding an interruption of the registration process, the identification information mapped to the user identification information is maintained or deleted.
7. The server according to claim 6, wherein, The identification information includes a feature vector of the voiceprint related to the speech.
8. The server according to claim 6, wherein, The controller is configured to: Based on the number of identifiers mapped to the user identifier information being equal to or greater than a preset minimum number, the identifier information mapped to the user identifier information is maintained, and If the number of identifiers mapped to the user identifier information is less than the preset minimum number, the identifier information mapped to the user identifier information is deleted.
9. The server according to claim 6, wherein, The controller is configured to: Based on the process of registering the identification information for the user account, it is determined whether the identification information is mapped to the user identification information. Based on the fact that the identification information is not mapped to the user identification information, a predetermined number of texts are sent to the electronic device in stages, and Based on the mapping of the identification information to the user identification information, some texts from multiple texts, excluding those related to the identification information mapped to the user identification information, are sent to the electronic device in stages.
10. The server according to claim 6, wherein, The controller is configured to: Based on the process of registering the identification information for the user account, it is determined whether the identification information is mapped to the user identification information. Based on the mapping of the identification information to the user identification information, a first text, excluding the text related to the identification information mapped to the user identification information, is sent to the electronic device. Generate identification information about a first voice signal that is related to the first text and included in the data received from the electronic device. Based on the correlation between the identification information regarding the first voice signal and at least one piece of identification information mapped to the user identification information, the identification information mapped to the user identification information is maintained, and Based on the fact that the identification information regarding the first voice signal is unrelated to at least one of the identification information mapped to the user identification information, the identification information mapped to the user identification information is deleted.
11. The server according to claim 10, wherein, The controller is configured to map the identification information about the first voice signal to the user identification information and store the mapped information in the database.
12. The server according to claim 10, wherein, The controller is configured to: Based on the correlation between the identification information regarding the first voice signal and at least one piece of identification information mapped to the user identification information, a result determining that the current user is the same as the previous user is sent to the electronic device, and Based on the fact that the identification information regarding the first voice signal is unrelated to at least one of the identification information mapped to the user identification information, the result determining that the current user is different from the previous user is sent to the electronic device.
13. The server according to claim 6, wherein, The controller is configured to: map the voice signal to the user identification information based on the correlation between the converted text and the preset text, and store the mapped signal in the database.
14. The server according to claim 13, wherein, The controller is configured to: The first algorithm is used to generate identification information associated with the voice signal mapped to the user identification information, and The identifier information that was mapped to the user identifier information and generated using the second algorithm is changed to the identifier information generated using the first algorithm.
15. A system comprising: Electronic devices; as well as server, The electronic device is configured as follows: While performing the registration process for voice-related identification information of user accounts logged into the server, preset text related to the registration of the identification information is displayed. When a voice signal related to the preset text is received, data containing the voice signal is sent to the server; The process of registering the identification information is completed based on the processing of the voice signal related to the preset text by the server; and During the process of registering the identification information, when a predetermined input related to voice recognition is received from the remote control device, the process of registering the identification information is stopped, and The server is configured as follows: The voice signal included in the data received from the electronic device is converted into text; Based on the correlation between the converted text and the preset text, identification information related to the speech signal is generated; The generated identification information is mapped to user identification information associated with the user account, and the mapped information is stored in the database; and During the process of registering the identification information, if a notification is received from the electronic device regarding an interruption of the registration process, the identification information mapped to the user identification information is maintained or deleted.