Server and system including the same

The server and system leverage voice recognition to address the inconvenience and security issues of manual account input by registering and identifying users through voice characteristics, improving accuracy and personalization in service delivery.

JP2026020119APending Publication Date: 2026-02-06LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025123032
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-07-23
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing systems require manual input of account information for service login, leading to inconvenience and security risks, especially in shared device scenarios, and lack efficient voice-based user identification and personalized service provision.

Method used

A server and system that utilize voice recognition technology to register and identify users based on their voice characteristics, generating identification information and improving processing accuracy through user-specific databases and usage history integration.

Benefits of technology

Enables secure and convenient user identification and personalized service delivery by registering voice-based identification information, enhancing processing accuracy and updating databases with user usage history.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026020119000001_ABST
    Figure 2026020119000001_ABST
Patent Text Reader

Abstract

One or more example embodiments also provide a server configured to register identification information associated with a voice of a user in an account of the user, and a system including the server.SOLUTION: In the system, the server 400 includes a communication unit that communicates with the electronic device 100 (the image display 100a, the air conditioner, the refrigerator, the air cleaner, the washing machine, the car, or the like), a database that stores a history of use of a text for each characteristic, and a controller. Generating a plurality of first texts corresponding to a voice signal received from an electronic device, generating first identification information corresponding to the voice signal, acquiring a characteristic of a user corresponding to the voice signal based on the first identification information, and determining a second text corresponding to the voice signal among the plurality of first texts based on a first text history in which each of the plurality of first texts is used in association with the characteristic of the user among a history in which texts are used for each characteristic; A result of performing intention analysis on the second text is sent to the electronic device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a server and a system including the same, and more particularly to a server that utilizes voice recognition technology and a system including the same. [Background technology]

[0002] In recent years, with the advancement of technology, active research has been conducted on speech recognition technology for processing speech. In particular, research on speech recognition technology, which began with smartphones, is now being widely applied to various fields related to user convenience, such as home appliances used in homes and offices, as well as vehicles.

[0003] Speech recognition technology is commonly used when a user controls an electronic device using voice. For example, when a user speaks a command to control an electronic device, the electronic device can directly recognize and process the user's voice and operate according to the command corresponding to the voice, or can transmit the voice to a server that processes the voice and then operate according to the command corresponding to the voice received from the server.

[0004] Meanwhile, services or functions provided through electronic devices are becoming increasingly diverse. In addition, users register accounts for various services and then log in to use the services using the registered accounts. In this case, service providers use user information managed for each account to provide functions and information optimally suited to the user.

[0005] Conventionally, when attempting to log in to use a service, a user manually inputs account information, such as an account identification (ID) and / or password. However, this is inconvenient for the user, as the user must input account information for each service. Furthermore, if the user maintains a logged-in state to avoid the inconvenience of inputting account information, security issues may arise, such as other people gaining access to the user's account information. Furthermore, when multiple users share a single electronic device, there is a problem in that the users must input their own account information to log in each time they use a service. Summary of the Invention [Problem to be solved by the invention]

[0006] The present disclosure is directed to solving the above-mentioned problems and other problems.

[0007] Yet another object is to provide a server capable of registering identification information for a user's voice to a user's account, and a system including (having; configuring; building; setting; including; containing; containing; having) the same.

[0008] A further object is to provide a server and a system including the same that can identify a user based on the user's voice.

[0009] Yet another object is to provide a server and a system including the same that can improve the accuracy of the results of processing a user's voice based on the user's characteristics identified from the voice.

[0010] Yet another object is to provide a server and a system including the same that can improve the accuracy of results of processing a user's voice based on the user's characteristics included in data for the user account.

[0011] A further object is to provide a server and a system including the same that can update a database used to process a user's voice using the user's usage history. [Means for solving the problem]

[0012] To achieve the above object, a server according to one embodiment of the present disclosure includes a communication unit that communicates with an electronic device, a database that stores a history of text usage by characteristic, and a controller, wherein the controller generates a plurality of first texts corresponding to a voice signal received from the electronic device, generates first identification information corresponding to the voice signal, acquires a user's characteristic corresponding to the voice signal based on the first identification information, determines a second text from the plurality of first texts corresponding to the voice signal based on a first text history in which each of the plurality of first texts was used in association with the user's characteristic among the history of text usage by characteristic, and transmits a result of intention analysis of the second text to the electronic device.

[0013] To achieve the above object, a system according to one embodiment of the present disclosure includes an electronic device and a server, wherein when a voice signal is received via a user input interface unit, the electronic device transmits data including the voice signal to the server and outputs a result of an intention analysis performed on the voice signal received from the server, and the server generates a plurality of first texts corresponding to the voice signal received from the electronic device, generates first identification information corresponding to the voice signal, acquires user characteristics corresponding to the voice signal based on the first identification information, determines a second text from the plurality of first texts corresponding to the voice signal based on a first text history in which each of the plurality of first texts was used in association with the user characteristic among histories of text usage by characteristic stored in a database, and transmits a result of the intention analysis performed on the second text to the electronic device as a result of the intention analysis performed on the voice signal. [Effects of the Invention]

[0014] The effects of the server and the system including the server according to the present disclosure are as follows.

[0015] According to at least one embodiment of the present disclosure, identification information for a user's voice can be registered with the user's account.

[0016] In accordance with at least one embodiment of the present disclosure, a user can be identified based on the user's voice.

[0017] According to at least one embodiment of the present disclosure, the accuracy of the results of processing a user's voice can be improved based on user characteristics identified from the voice.

[0018] In accordance with at least one embodiment of the present disclosure, the accuracy of the results of processing a user's voice can be improved based on user characteristics contained in data for the user account.

[0019] In accordance with at least one embodiment of the present disclosure, a user's usage history can be used to update the database used to process the user's voice.

[0020] Further scope of applicability of the present disclosure will become apparent from the following detailed description. However, it should be understood that the detailed description and specific embodiments, such as preferred embodiments of the present disclosure, are given by way of example only, since various changes and modifications within the spirit and scope of the present disclosure will be apparent to those skilled in the art. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 illustrates a system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is an internal block diagram of the electronic device of FIG. 1. [Figure 3]FIG. 2 is a diagram referred to in the explanation of the server in FIG. 1. [Figure 4] FIG. 2 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure. [Figure 5] FIG. 10 is a diagram illustrating an example in which an audio signal is converted into a power spectrum according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a block diagram illustrating the configuration of a control unit for speech recognition and synthesis of an electronic device according to an embodiment of the present disclosure. [Figure 7] 1 is a flow chart illustrating a method for operating an electronic device according to an embodiment of the present disclosure. [Figure 8] 1 is a flow chart illustrating a method of operation of a system according to one embodiment of the present disclosure. [Figure 9-13] 1A-1C are diagrams and the like referenced in describing a process for registering identification information for a user's voice to a user's account according to one embodiment of the present disclosure. [Figure 14] 1 is a flow chart illustrating a method of operating a server according to an embodiment of the present disclosure. [Figure 15-21] 1A to 1C are diagrams etc. referred to in the description of providing the results of processing a user's voice according to an embodiment etc. of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0022] The present disclosure will be described in detail below with reference to the drawings. In the drawings, in order to clearly and concisely explain the present disclosure, parts that are not relevant to the description will be omitted, and the same reference numerals will be used throughout the specification to refer to the same or very similar parts.

[0023] The suffixes "module" and "section" for components used in the following description are given simply for the sake of ease of writing this specification, and do not impart any special significance or role to them. Therefore, the terms "module" and "section" may be used interchangeably.

[0024] In this application, the use of terms such as "comprise" or "have" is intended to specify the presence of a stated feature, number, step, operation, component, part, or combination thereof, and is to be understood as not precluding the presence or possible addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0025] Furthermore, although terms such as first and second may be used to describe various elements in this specification, such elements are not limited by such terms. Such terms are used only to distinguish one element from another.

[0026] FIG. 1 illustrates a system according to various embodiments of the present invention.

[0027] As shown in FIG. 1, the system 10 may include an electronic device 100 and / or a server 400 .

[0028] The electronic device 100 can transmit and receive data to and from at least one server 400. For example, the electronic device 100 can transmit and receive data to and from at least one server 400 via a network 300 such as the Internet.

[0029] According to one embodiment, at least one server 400 may include a server that performs voice recognition, a server that processes data using a super-giant artificial intelligence model, a server that provides content, etc.

[0030] The electronic device 100 may include an image display device 100a, an air conditioner 100b, a refrigerator 100c, an air purifier 100d, a washing machine 100e, a vehicle 100f, etc. In the present disclosure, the electronic device 100 will be described as an image display device 100a by way of example, but the present invention is not limited thereto.

[0031] The image display device 100a may be a device that processes and outputs an image. The image display device 100a is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.

[0032] The image display device 100a can receive a broadcast signal, process the signal, and output the processed broadcast image. When the image display device 100a receives a broadcast signal, the image display device 100a can correspond to a broadcast receiving device.

[0033] The image display device 100a can receive broadcast signals wirelessly via an antenna, and can also receive broadcast signals wired via a cable. For example, the image display device 100a can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.

[0034] FIG. 2 is an internal block diagram of the electronic device of FIG.

[0035] As shown in FIG. 2, the electronic device 100 may include a broadcast receiving unit 105, an external device interface unit 130, a network interface unit 135, a storage unit 140, a user input interface unit 150, an input unit 160, a control unit 170, a display 180, an audio output unit 185, and / or a power supply unit 190.

[0036] The broadcast receiving unit 105 may include a tuner unit 110 and a demodulation unit 120 .

[0037] Meanwhile, unlike the drawing, the electronic device 100 may include only the broadcast receiving unit 105 and the external device interface unit 130 among the broadcast receiving unit 105, the external device interface unit 130, and the network interface unit 135. In other words, the electronic device 100 may not include the network interface unit 135.

[0038] The tuner unit 110 can select a broadcast signal corresponding to a channel selected by a user or all pre-stored channels from among broadcast signals received via an antenna (not shown) or a cable (not shown), and convert the selected broadcast signal into an intermediate frequency signal or a baseband image or audio signal.

[0039] For example, if the selected broadcast signal is a digital broadcast signal, the tuner unit 110 converts it into a digital IF signal DIF, and if the selected broadcast signal is an analog broadcast signal, it converts it into an analog baseband image or audio signal CVBS / SIF. That is, the tuner unit 110 can process digital broadcast signals or analog broadcast signals. The analog baseband image or audio signal CVBS / SIF output from the tuner unit 110 can be directly input to the control unit 170.

[0040] Meanwhile, the tuner unit 110 can sequentially select broadcast signals of all broadcast channels stored through the channel storage function from among the received broadcast signals and convert them into intermediate frequency signals or baseband image or audio signals.

[0041] Meanwhile, the tuner unit 110 may include multiple tuners to receive broadcast signals of multiple channels, or may include a single tuner that simultaneously receives broadcast signals of multiple channels.

[0042] The demodulator 120 receives the digital IF signal DIF converted by the tuner 110 and performs a demodulation operation.

[0043] The demodulation unit 120 can output a stream signal TS after demodulation and channel decoding, where the stream signal may be a signal in which an image signal, an audio signal, or a data signal is multiplexed.

[0044] The stream signal output from the demodulator 120 may be input to the controller 170. The controller 170 may perform demultiplexing, image / audio signal processing, etc., and then output an image through the display 180 and output audio through the audio output unit 185.

[0045] The external device interface unit 130 can transmit and receive data to and from a connected external device, and for this purpose, the external device interface unit 130 can include an A / V input / output unit (not shown).

[0046] The external device interface unit 130 can be connected to external devices such as DVDs (Digital Versatile Disks), Blu-rays (registered trademark: the same applies hereinafter), game consoles, cameras, camcorders, computers (notebooks), set-top boxes, etc. via wired or wireless connections, and can also perform input / output operations with the external devices.

[0047] In addition, the external device interface unit 130 can establish a communication network with various remote control devices 200 to receive control signals related to the operation of the electronic device 100 from the remote control devices 200 or transmit data related to the operation of the electronic device 100 to the remote control devices 200.

[0048] The A / V input / output unit may receive image and audio signals from an external device. For example, the A / V input / output unit may include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-Video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface: registered trademark) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input via these terminals may be transmitted to the control unit 170. In this case, analog signals input via the CVBS terminal and the S-Video terminal may be converted into digital signals via an analog-to-digital converter (not shown) and then transmitted to the control unit 170.

[0049] The external device interface unit 130 may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. The external device interface unit 130 can exchange data with a nearby mobile terminal through the wireless communication unit. For example, the external device interface unit 130 can receive device information, information about an application being executed, an application image, etc. from the mobile terminal in mirroring mode.

[0050] The external device interface unit 130 can perform short-range wireless communication using Bluetooth (registered trademark: same below), RFID (Radio Frequency Identification), infrared communication (IrDA, Infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.

[0051] The network interface unit 135 may provide an interface for connecting the electronic device 100 to a wired or wireless network including the Internet.

[0052] The network interface unit 135 may include a communication module (not shown) for connection to a wired or wireless network. For example, the network interface unit 135 may include a communication module for a wireless LAN (WLAN) (Wi-Fi), wireless broadband (Wibro), World Interoperability for Microwave Access (Wimax), high speed downlink packet access (HSDPA), etc.

[0053] The network interface unit 135 can transmit or receive data to or from other users or other electronic devices via the connected network or other networks linked to the connected network.

[0054] The network interface unit 135 can receive web content or data provided by a content provider or a network operator. That is, the network interface unit 135 can receive content and related information such as movies, advertisements, games, VOD, and broadcasting provided by a content provider or a network operator via a network.

[0055] The network interface unit 135 can receive firmware update information and update files provided by a network operator, and can transmit data to the Internet, a content provider, or a network operator.

[0056] The network interface unit 135 can select and receive a desired application from among applications that are open to the public via a network.

[0057] The storage unit 140 may store programs for signal processing and control within the control unit 170, and may also store processed image, audio, or data signals. For example, the storage unit 140 may store application programs designed to perform various tasks that can be processed by the control unit 170, and may selectively provide some of the stored application programs upon request from the control unit 170.

[0058] The programs stored in the storage unit 140 are not particularly limited as long as they can be executed by the control unit 170.

[0059] The storage unit 140 may also function as a temporary storage for image, audio, or data signals received from an external device via the external device interface unit 130 .

[0060] The storage unit 140 can store information about a given broadcast channel through a channel storage function such as a channel map.

[0061] Although FIG. 2 illustrates an embodiment in which the storage unit 140 is provided separately from the control unit 170, the scope of the present invention is not limited thereto, and the storage unit 140 may be included within the control unit 170.

[0062] The storage unit 140 may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) and non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit 140 and memory may be used interchangeably.

[0063] The user input interface unit 150 can transfer a signal input by the user to the control unit 170 or transfer a signal from the control unit 170 to the user.

[0064] For example, the remote control device 200 can send / receive user input signals such as power on / off, channel selection, screen setting, etc., or transmit user input signals input from local keys (not shown) such as a power key, channel key, volume key, setting value, etc. to the control unit 170, or transmit user input signals input from a sensor unit (not shown) that senses user gestures to the control unit 170, or transmit signals from the control unit 170 to the sensor unit.

[0065] The input unit 160 may be provided on one side of the main body of the electronic device 100. For example, the input unit 160 may include a touchpad, physical buttons, etc.

[0066] The input unit 160 may receive various user commands related to the operation of the electronic device 100 and may transmit control signals corresponding to the input commands to the control unit 170 .

[0067] The input unit 160 may include at least one microphone (not shown) and may receive the user's voice via the microphone.

[0068] The control unit 170 may include at least one processor, and may use the processor to control the overall operation of the electronic device 100. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC, or a processor based on other hardware.

[0069] The control unit 170 can demultiplex a stream input via the tuner unit 110, demodulator unit 120, external device interface unit 130, or network interface unit 135, or process the demultiplexed signal to generate and output a signal for image or audio output.

[0070] The display 180 can generate driving signals by converting image signals, data signals, OSD signals, and control signals processed by the control unit 170 or image signals, data signals, and control signals received from the external device interface unit 130.

[0071] Display 180 may include a display panel (not shown) comprising a plurality of pixels.

[0072] The plurality of pixels provided on the display panel may include RGB sub-pixels, or may include RGBW sub-pixels. The display 180 may convert the image signals, data signals, OSD signals, control signals, etc. processed by the controller 170 to generate driving signals for the plurality of pixels.

[0073] The display 180 may be a plasma display panel (PDP), a liquid crystal display (LCD), an organic light emitting diode (OLED), a flexible display, etc., and may also be a 3D display. The 3D display 180 may be classified into a glasses-free type and a glasses-based type.

[0074] On the other hand, the display 180 may be configured as a touch screen and may be used as an input device in addition to an output device.

[0075] The audio output unit 185 receives the signal that has been subjected to audio processing by the control unit 170 and outputs it as audio.

[0076] The image signal processed by the control unit 170 can be input to the display 180 and displayed as an image corresponding to the image signal. The image signal processed by the control unit 170 can also be input to an external output device via the external device interface unit 130.

[0077] The audio signal processed by the control unit 170 may be output as sound to the audio output unit 185. In addition, the audio signal processed by the control unit 170 may be input to an external output device via the external device interface unit 130.

[0078] Although not shown in FIG. 2, the control unit 170 may include a demultiplexing unit, an image processing unit, and the like.

[0079] Additionally, the control unit 170 may control the overall operation of the electronic device 100. For example, the control unit 170 may control the tuner unit 110 to select (tune) a broadcast corresponding to a channel selected by a user or a pre-stored channel.

[0080] The control unit 170 can also control the electronic device 100 according to a user command input through the user input interface unit 150 or an internal program.

[0081] Meanwhile, the control unit 170 can control the display 180 to display an image. At this time, the image displayed on the display 180 may be a still image or a moving image, and may be a 2D image or a 3D image.

[0082] Meanwhile, the control unit 170 can display a predetermined 2D object within the image displayed on the display 180. For example, the object may be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.

[0083] Meanwhile, the electronic device 100 may further include a photographing unit (not shown). The photographing unit may photograph a user. The photographing unit may be implemented as one camera, but is not limited thereto, and may also be implemented as a plurality of cameras. Meanwhile, the photographing unit may be embedded in the electronic device 100 above the display 180 or may be separately disposed. Image information photographed by the photographing unit may be input to the control unit 170.

[0084] The control unit 170 can recognize the user's position based on the image captured by the image capturing unit. For example, the control unit 170 can determine the distance (z-axis coordinate) between the user and the electronic device 100. In addition, the control unit 170 can determine the x-axis coordinate and the y-axis coordinate in the display 180 corresponding to the user's position.

[0085] The control unit 170 can detect a user's gesture based on an image captured by the image capturing unit or a signal sensed by the sensor unit, or a combination thereof.

[0086] The power supply unit 190 may supply power to the entire electronic device 100. In particular, the power supply unit 190 may supply power to the control unit 170, which may be implemented in the form of a system on chip (SOC), the display 180 for displaying images, and the audio output unit 185 for outputting audio.

[0087] Specifically, the power supply unit 190 may include a converter (not shown) that converts AC power into DC power, and a Dc / Dc converter (not shown) that converts the level of the DC power.

[0088] The remote control device 200 can transmit user input to the user input interface unit 150. To this end, the remote control device 200 can use Bluetooth, RF (Radio Frequency) communication, Infrared (Infrared) communication, UWB (Ultra-wideband), ZigBee, etc. Also, the remote control device 200 can receive an image, audio, or data signal output from the user input interface unit 150 and display it on the remote control device 200 or output it as audio.

[0089] Meanwhile, the electronic device 100 may be a fixed or mobile digital broadcast receiver capable of receiving digital broadcasts.

[0090] Meanwhile, the block diagram of the electronic device 100 shown in FIG. 2 is a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the electronic device 100 actually implemented.

[0091] That is, two or more components may be combined into one component, or one component may be divided into two or more components, as needed. Also, the functions performed by each block are for explaining the embodiments of the present invention, and the specific operations and devices thereof do not limit the scope of the present invention.

[0092] FIG. 3 is a diagram referred to in the description of the server in FIG.

[0093] 3, the server 400 may include a relay server 410, a STT (Speech To Text) server 420, a NLP (Natural Language Processing) server 430, a user identification server 440, and / or an account server 450. In the present disclosure, the relay server 410, the STT server 420, the NLP server 430, the user identification server 440, and the account server 450 are described as being separate from one another, but are not limited thereto. For example, two or more of the relay server 410, the STT server 420, the NLP server 430, the user identification server 440, and the account server 450 may be configured in a single server.

[0094] The relay server 410 can communicate with the electronic device 100. The relay server 410 can communicate data between the STT server 420, the NLP server 430, the user identification server 440, and the electronic device 100. The relay server 410 can store at least a portion of the data communicated between the STT server 420, the NLP server 430, the user identification server 440, and the electronic device 100.

[0095] The STT server 420 can receive voice data. The STT server 420 can convert the voice data into text data. The STT server 420 can transmit the text data to the electronic device 100 via the relay server 410. The STT server 420 can also be called an ASR (Automatic Speech Recognition) server.

[0096] The STT server 420 can improve the accuracy of speech-to-text conversion by using a language model. A language model can mean a model that can calculate the probability of a sentence or the probability of a next word appearing given a previous word, etc. For example, the language model can include a probabilistic language model such as a unigram model, a bigram model, an N-gram model, etc. That is, the STT server 420 can use the language model to determine whether the text data converted from speech data is converted appropriately, thereby improving the accuracy of the conversion to text data.

[0097] The NLP server 430 can receive text data. The NLP server 430 can perform an intention analysis on the text data based on the received text data. The NLP server 430 can transmit intention analysis information representing the execution result of the intention analysis to the electronic device 100 via the relay server 410.

[0098] According to one embodiment, the NLP server 430 may generate intention analysis information by sequentially performing a morpheme analysis step, a syntactic analysis step, a speech act analysis step, and a dialogue processing step on the text data. The morpheme analysis step classifies text data corresponding to speech uttered by a user into morphemes, which are the smallest meaningful units, and determines the part of speech of each classified morpheme. The syntactic analysis step uses the results of the morpheme analysis step to divide the text data into noun phrases, verb phrases, adjective phrases, etc., and determines the relationship between each divided phrase. The syntactic analysis step may determine the subject, object, modifier, etc., of the speech uttered by the user. The speech act analysis step uses the results of the syntactic analysis step to analyze the intention of the speech uttered by the user. Specifically, the speech act analysis step determines the intention of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The dialogue processing stage uses the results of the speech act analysis stage to determine whether to reply to the user's utterance, respond, or ask a question requesting additional information.

[0099] The user identification server 440 may receive the voice data. The user identification server 440 may extract voice features based on the voice data. Here, the voice features may include a voice waveform, a voice frequency band, a voice power spectrum, etc. The extraction of voice features will be described below with reference to FIGS. 4 and 5.

[0100] The user identification server 440 can acquire a voice feature vector from the voice features. The user identification server 440 can acquire a voice feature vector from the voice features based on a linear predictive coefficient, a cepstrum, a Mel Frequency Cepstral Coefficient (MFCC), a filter bank energy, etc.

[0101] The user identification server 440 may determine the similarity between multiple feature vectors. The user identification server 440 may determine the similarity between multiple feature vectors using cosine similarity, Euclidean similarity, or the like. In the present disclosure, the similarity between a first speech input and a second speech input is calculated based on cosine similarity, but is not limited thereto. For example, a first vector corresponding to the first text and a second vector corresponding to the second text may be generated. In this case, the cosine similarity between the first vector and the second vector may be calculated based on the following mathematical formula 1:

[0102]

number

[0103] where A·B is the dot product of two vectors, ||A|| and ||B|| may refer to the magnitudes of two vectors. That is, cosine similarity may be calculated by dividing the inner product of two vectors by the product of the magnitudes of each vector. Cosine similarity may range from -1 to 1, and the closer the cosine similarity is to 1, the more similar the two vectors are considered to be.

[0104] User identification server 440 can determine whether the users who spoke the voices are the same based on the similarity between multiple feature vectors. For example, if the similarity between a first feature vector corresponding to a first voice input and a second feature vector corresponding to a second voice input is equal to or greater than a predetermined standard, user identification server 440 can determine that the user who spoke the first voice input and the user who spoke the second voice input are the same.

[0105] According to one embodiment, the user identification server 440 can obtain a vector obtained by processing a feature vector of the voice using an algorithm such as a GMM (Gaussian mixture model) supervector, i-vector, d-vector, or x-vector. The user identification server 440 can determine whether the user who spoke the voice is the same or different based on the similarity between a first vector processed in correspondence with a first feature vector and a second vector processed in correspondence with a second feature vector.

[0106] The user identification server 440 may store voice data. The user identification server 440 may store data on a voice print (hereinafter, referred to as voice print information). Here, the voice print information may include a voice feature vector and / or a vector obtained by processing the voice feature vector.

[0107] The user identification server 440 may store a database for voice, which may include unique identification information corresponding to the electronic device 100 (hereinafter referred to as device identification information), unique identification information corresponding to a user account (hereinafter referred to as user identification information), voice data mapped to the user identification information, and voiceprint information mapped to the user identification information.

[0108] The device identification information, user identification information, voice data, voiceprint information, etc. included in the voice database may be associated with each other and stored in the user identification server 440. For example, at least one device identification information, a plurality of voice data, and / or a plurality of voiceprint information may be mapped to the user identification information. That is, it may be interpreted that the device identification information, voice data, voiceprint information, etc. are mapped to a user account and stored in the user identification server 440. In the present disclosure, a case where a plurality of voice data and a plurality of voiceprint information are both mapped to the user identification information included in the voice database will be described as an example.

[0109] The user identification server 440 can update the voiceprint information included in the voice database based on the voice data included in the voice database. For example, the user identification server 440 can generate voiceprint information corresponding to the voice data included in the voice database using an algorithm different from the previously used algorithm. In this case, the user identification server 440 can change the voiceprint information included in the voice database to the newly generated voiceprint information.

[0110] The account server 450 may manage data related to a user account, such as a user account ID, a password, user identification information, device identification information mapped to the user account, and whether or not the user agrees to terms and conditions related to various functions.

[0111] The account server 450 may store a database for user accounts, which may include user account IDs, passwords, user identification information, device identification information mapped to the user accounts, the date and time of registration of the user accounts, whether or not users have agreed to terms and conditions related to various functions, and the date and time the terms and conditions were agreed to.

[0112] Account server 450 can communicate with electronic device 100. For example, account server 450 can create and register a user account based on data from electronic device 100. For example, account server 450 can approve a user account login based on an ID and password received from electronic device 100.

[0113] FIG. 4 is a block diagram illustrating the configuration of a server according to an embodiment of the present disclosure.

[0114] As shown in FIG. 4, the server 400 may include a preprocessing unit 460, a controller 470, a communication unit 480, and / or a database 490.

[0115] The preprocessing unit 460 can preprocess the voice received via the communication unit 480 or the voice stored in the database 490 .

[0116] The pre-processing unit 460 can be implemented on a chip separate from the controller 470 or on a chip included in the controller 470 .

[0117] The pre-processing unit 460 can receive a voice signal (spoken by a user) and filter noise signals from the voice signal before converting the received voice signal into text data.

[0118] When the pre-processing unit 460 is provided in the electronic device 100, it may recognize a wake-up word for activating voice recognition of the electronic device 100. The pre-processing unit 460 converts the wake-up word received through the user input interface unit 150 into text data, and may determine that the wake-up word has been recognized if the converted text data corresponds to a pre-stored wake-up word.

[0119] The preprocessing unit 460 can convert the noise-removed audio signal into a power spectrum.

[0120] The power spectrum may be a parameter that indicates what frequency components are contained in the waveform of a time-varying audio signal and at what magnitude.

[0121] The power spectrum shows the distribution of the squared magnitude of the waveform of an audio signal according to frequency, which will be explained with reference to FIG.

[0122] FIG. 5 is a diagram illustrating an example in which an audio signal is converted into a power spectrum according to an embodiment of the present disclosure.

[0123] 5, there is shown an audio signal 510. The audio signal 510 may be received from an external device or may be a signal pre-stored in the memory 170.

[0124] The x-axis of the audio signal 510 may represent time and the y-axis may represent amplitude magnitude.

[0125] The power spectrum processing unit 463 can convert the audio signal 510, whose x-axis is the time axis, into a power spectrum 520, whose x-axis is the frequency axis. The power spectrum processing unit 463 can convert the audio signal 510 into the power spectrum 520 using a fast Fourier transform (FFT). The x-axis of the power spectrum 520 represents frequency, and the y-axis represents the squared value of the amplitude.

[0126] As further shown in FIG. 4, the functions of the preprocessor 460 and controller 470 described in FIG. 4 can also be performed by the NLP server 430.

[0127] The preprocessing unit 460 may include a wave processing unit 461, a frequency processing unit 462, a power spectrum processing unit 463, a speech to text (STT) conversion unit 464, and the like.

[0128] The wave processing unit 461 can extract the waveform of the audio.

[0129] The frequency processing unit 462 can extract the frequency band of the audio.

[0130] The power spectrum processing unit 463 can extract the power spectrum of the audio.

[0131] The power spectrum may be a parameter that indicates, when a waveform that varies over time is given, what frequency components are contained in the waveform and at what magnitudes.

[0132] The Speech to Text (STT) conversion unit 464 can convert speech into text, i.e., speech in a specific language into text in that language.

[0133] The controller 470 may control the overall operation of the server 400. The controller 470 may include a speech analyzer 471, a text analyzer 472, a feature clusterer 473, a text mapper 474, and / or a speech synthesizer 475.

[0134] The voice analysis unit 471 can extract voice characteristic information using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed by the preprocessing unit 460. The voice characteristic information can include one or more of the speaker's gender information, the speaker's voice (or timbre, tone), pitch, the speaker's speaking style, the speaker's speaking rate, and the speaker's emotion. The voice characteristic information can also include the speaker's timbre.

[0135] The text analysis unit 472 can extract key phrases from the text converted by the speech-to-text conversion unit 464. When the text analysis unit 472 detects a change in tone between phrases from the converted text, it can extract the phrases with the change in tone as key phrases. When the frequency band between phrases changes by more than a preset band, the text analysis unit 472 can determine that the tone has changed. The text analysis unit 472 can also extract key words from phrases of the converted text. Key words can be nouns present in the phrases, but this is merely an example.

[0136] The feature clustering unit 473 can classify the speech type of a speaker by using the speech characteristic information extracted by the speech analysis unit 471. The feature clustering unit 473 can classify the speech type of a speaker by assigning a weight to each type item constituting the speech characteristic information. The feature clustering unit 473 can classify the speech type of a speaker by using an attention technique of a deep learning model.

[0137] The text mapping unit 474 can translate text converted into a first language into text in a second language. The text mapping unit 474 can map text translated into a second language with text in the first language. The text mapping unit 474 can map main expressions constituting text in the first language to corresponding expressions in the second language. The text mapping unit 474 can map speech types corresponding to main expressions constituting text in the first language to phrases in the second language. This is to apply speech types classified into phrases in the second language.

[0138] The speech synthesis unit 475 can generate synthesized speech by applying the speech type and speaker tone classified by the feature clustering unit 473 to the main expressions of the text translated into the second language by the text mapping unit 474.

[0139] The controller 470 can use one or more of the transmitted text data or the power spectrum 520 to determine the user's speech characteristics.

[0140] The user's speech characteristics may include the user's gender, the user's pitch, the user's timbre, the user's speech subject, the user's speech rate, the user's voice volume, and the like.

[0141] The controller 470 can use the power spectrum 520 to obtain the frequencies of the audio signal 510 and the amplitudes corresponding to the frequencies.

[0142] The controller 470 can determine the gender of the user who spoke the voice using the frequency band of the power spectrum 470. For example, the controller 470 can determine the gender of the user as male if the frequency band of the power spectrum 520 is within a preset first frequency band range.

[0143] Controller 470 may determine the user's gender as female if the frequency band of power spectrum 520 is within a preset second frequency band range, where the second frequency band range may be larger than the first frequency band range.

[0144] The controller 470 can determine the pitch of a voice using the frequency band of the power spectrum 520. For example, the controller 470 can determine the pitch of a sound based on the magnitude of the amplitude within a particular frequency band range.

[0145] Controller 470 can determine the user's tone using the frequency bands of power spectrum 520. For example, controller 470 can determine a frequency band of power spectrum 520 whose amplitude is equal to or greater than a certain value as the user's main frequency range, and determine the determined main frequency range as the user's tone.

[0146] From the converted text data, the controller 470 can determine the user's speaking rate via the number of syllables spoken per unit time.

[0147] The controller 470 can determine the user's topic of speech using a Bag-Of-Word Model technique on the converted text data.

[0148] The Bag-Of-Word Model technique extracts commonly used words based on the frequency of words in a sentence. Specifically, the Bag-Of-Word Model technique extracts unique words in a sentence, expresses the frequency of each extracted word as a vector, and determines the characteristics of the topic of speech. For example, if words such as "running" and "physical strength" appear frequently in the text data, the controller 470 can classify the topic of the user's speech as "exercise."

[0149] The controller 470 can determine the user's topic of speech from the text data using known text categorization techniques, and can extract keywords from the text data to determine the user's topic of speech.

[0150] The controller 470 may determine the user's voice volume based on amplitude information across the entire frequency band, for example, based on the average or weighted average of the amplitude across each frequency band of the power spectrum.

[0151] The communication unit 480 may communicate with an external server via wire or wirelessly, and may communicate with the electronic device 100 via wire or wirelessly.

[0152] Database 490 may store speech in a first language included in content. Database 490 may store synthesized speech in which speech in the first language is converted into speech in a second language. Database 490 may store first text corresponding to speech in the first language and second text in which the first text is translated into a second language. Database 490 may also store various learning models necessary for speech recognition.

[0153] 2 may include the pre-processing unit 460 and the controller 470 shown in FIG. 4. That is, the control unit 170 of the electronic device 100 may perform the functions of the pre-processing unit 460 and the controller 470.

[0154] FIG. 6 is a block diagram illustrating the configuration of a control unit for speech recognition and synthesis of an electronic device according to an embodiment of the present disclosure.

[0155] That is, the voice recognition and synthesis process of FIG. 6 can be performed by the control unit 170 of the electronic device 100 without going through a server.

[0156] 6, the control unit 170 of the electronic device 100 may include an STT engine 610, an NLP engine 620, and a speech synthesis engine 630. Each engine may be either hardware or software.

[0157] The STT engine 610 can perform the function of the STT server 420 in Figure 5. That is, the STT engine 610 can convert voice data into text data.

[0158] The NLP engine 620 can perform the function of the NLP server 430 in Fig. 5. That is, the NLP engine 620 can obtain intention analysis information that indicates the intention of the speaker from the converted text data.

[0159] The speech synthesis engine 630 can function as a speech synthesis server. The speech synthesis engine 630 can search a database for syllables or words corresponding to given text data, and synthesize combinations of the searched syllables or words to generate synthetic speech.

[0160] The speech synthesis engine 630 may include a pre-processing engine 631 and a TTS engine 632 .

[0161] The preprocessing engine 631 can preprocess text data before generating synthetic speech. Specifically, the preprocessing engine 631 performs tokenization, which divides text data into tokens, which are meaningful units. After tokenization, the preprocessing engine 631 can perform a cleansing process to remove unnecessary characters and symbols to remove noise. The preprocessing engine 631 can then combine word tokens expressed in different ways to generate the same word token. The preprocessing engine 631 can then remove meaningless word tokens (non-words, stopwords).

[0162] The TTS engine 632 can synthesize speech corresponding to the preprocessed text data to generate synthesized speech.

[0163] FIG. 7 is a flow chart illustrating a method of operating an electronic device according to one embodiment of the present disclosure.

[0164] 7, in operation S701, the electronic device 100 can determine whether a user account is logged in to the server 400. For example, a user can input the ID and password of the user account to log in to the server 400 with the user account.

[0165] According to one embodiment, when a user first logs in to the server 400 using the electronic device 100 with a user account, the electronic device 100 may include user identification information corresponding to the user account in a user list. For example, when three different user accounts log in to the server 400 using the electronic device 100, the user list stored in the electronic device 100 may include three different user identification information.

[0166] In operation S702, the electronic device 100 may determine whether voice-related identification information (hereinafter, referred to as a voice ID) is registered for a user account logged in to the server 400. Here, the voice ID may include voiceprint information stored in the user identification server 440. For example, the server 400 may notify the electronic device 100 of whether a voice ID is registered for a user account logged in to the server 400.

[0167] According to one embodiment, the server 400 can determine whether or not to register a voice ID based on whether voiceprint information is mapped to user identification information, which is unique identification information corresponding to a user account logged in to the server 400. In this case, if the voice ID is an unregistered user account, the number of voiceprint information mapped to the user identification information may be zero.

[0168] According to one embodiment, the server 400 may determine that a voice ID is registered if the number of pieces of voiceprint information mapped to the user identification information is a predetermined number equal to or greater than two, and may determine that a voice ID is not registered if the number of pieces of voiceprint information mapped to the user identification information is less than the predetermined number. For example, for a user account with a registered voice ID, six different pieces of voiceprint information may be mapped to the user identification information. For example, for a user account with an unregistered voice ID, five or fewer pieces of voiceprint information may be mapped to the user identification information.

[0169] According to one embodiment, a flag value indicating whether a voice ID is registered may be mapped to the user identification information stored in the server 400. Here, the user identification information to which the flag value is mapped may be stored in the user identification server 440 and / or the account server 450. The server 400 may determine whether a voice ID is registered based on the flag value mapped to the user identification information. For example, if the voice ID is an unregistered user account, the flag value mapped to the user identification information may be 0, and if the voice ID is a registered user account, the flag value mapped to the user identification information may be 1.

[0170] In operation S703, if a voice ID is not registered for the user account, the electronic device 100 can start a process to register the voice ID. For example, when starting the process to register the voice ID, the electronic device 100 can send data including device identification information, user identification information, a value indicating the start of voice ID registration, etc. to the server 400.

[0171] In operation S704, the electronic device 100 may output a preset text. The electronic device 100 may output any one of a plurality of preset texts. For example, if the electronic device 100 is an image display device 100a, the electronic device 100 may output the preset text via the display 180.

[0172] According to an embodiment, the server 400 may transmit any one of a plurality of preset texts to the electronic device 100 in a preset order. In this case, the electronic device 100 may output the preset text received from the server 400.

[0173] In operation S705, the electronic device 100 may determine whether voice for a preset text is input. For example, the electronic device 100 may determine whether voice is input through a microphone provided in the input unit 160 within a preset time period. At this time, a voice signal corresponding to the voice input through the microphone may be transmitted to the control unit 170 via the user input interface unit 150. For example, the electronic device 100 may determine whether data including a voice signal corresponding to the voice spoken by the user is received from the remote control device 200 within a preset time period.

[0174] When a voice for a preset text is input in operation S706, the electronic device 100 may transmit voice data including a voice signal corresponding to the voice to the server 400. At this time, the electronic device 100 may transmit device identification information, user identification information, a language code indicating a language type, and the like together with the voice data to the server 400.

[0175] The server 400 can convert a voice signal included in the voice data received from the electronic device 100 into text. The server 400 can determine whether the text converted from the voice signal corresponds to a preset text. For example, the server 400 can determine whether the two texts correspond to each other based on the similarity between the text converted from the voice signal and the preset text.

[0176] When text converted from a voice signal corresponds to a preset text, the server 400 can generate voiceprint information corresponding to the voice signal. The server 400 can map the voiceprint information generated for the preset text to user identification information and store the map. The server 400 can map received voice data for the preset text to user identification information and store the map.

[0177] In operation S707, the electronic device 100 may determine whether the processing of the voice on the preset text has been successful based on the response received from the server 400. For example, the server 400 may notify the electronic device 100 of the success of the processing of the voice when the text converted from the voice signal corresponds to the preset text. For example, the server 400 may notify the electronic device 100 of the success of the processing of the voice when voiceprint information corresponding to the voice signal is generated.

[0178] Meanwhile, in operation S708, if voice input for the preset text is not performed or if processing of the voice for the preset text fails, the electronic device 100 may determine whether to retry voice input. For example, the electronic device 100 may retry voice input based on a user input for retrying voice input. In this case, the electronic device 100 may output the preset text again.

[0179] In operation S709, if the voice processing for the preset text is successful, the electronic device 100 may determine whether the processing for all texts is complete. For example, if the voice processing for all six preset texts is successful, the processing for all texts may be complete. On the other hand, if the processing for five preset texts is completed, the electronic device 100 may output the last preset text.

[0180] When the electronic device 100 has completed processing for all texts in operation S710, the electronic device 100 can complete the process of registering the voice ID. For example, if the electronic device 100 is an image display device 100a, the electronic device 100 can output a screen indicating that the voice ID registration has been completed via the display 180. For example, the electronic device 100 can transmit data indicating that the voice ID registration has been completed to the account server 450.

[0181] FIG. 8 is a flow chart for a method of operation of the system according to one embodiment of the present disclosure.

[0182] As shown in FIG. 8, in operation S801, the electronic device 100 can log in to the server 400 using a user account.

[0183] The electronic device 100 may begin the process of registering a voice ID at operation S802.

[0184] In operation S803, the electronic device 100 can output a first text among a plurality of preset texts.

[0185] The electronic device 100 may receive a first voice for the first text in operation S804.

[0186] In operation S805, the electronic device 100 can transmit, to the server 400, first audio data including an audio signal corresponding to the first audio.

[0187] In operation S806, the server 400 can process the first voice for the first text based on the first voice data received from the electronic device 100. The server 400 can convert the voice signal corresponding to the first voice included in the first voice data received from the electronic device 100 into text. The server 400 can determine whether the text converted from the voice signal corresponding to the first voice and the first text correspond to each other.

[0188] In operation S807, the server 400 can notify the electronic device 100 of the completion of processing for the first voice. For example, the server 400 can notify the electronic device 100 of the success of processing for the first voice based on the correspondence between the text converted from the audio signal corresponding to the first voice and the first text.

[0189] Meanwhile, the server 400 can generate first voiceprint information for the first voice based on the voice signal corresponding to the first voice, based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.

[0190] In operation S808, the server 400 can store the first voice data and first voiceprint information for the first voice. The server 400 can map the first voice data and the first voiceprint information to user identification information corresponding to the logged-in user account and store the mapped information.

[0191] The electronic device 100 can output the second text through the fifth text in a stepwise manner. The electronic device 100 can sequentially receive the second voice through the fifth voice corresponding to the second text through the fifth text, respectively. The electronic device 100 can sequentially transmit the second voice data through the fifth voice data corresponding to the second voice through the fifth voice, respectively, to the server 400.

[0192] The server 400 can process the second to fifth sounds based on the second to fifth sound data received from the electronic device 100. The server 400 can also sequentially generate and store second to fifth sound information corresponding to the second to fifth sounds, respectively.

[0193] In operation S809, the electronic device 100 can output a sixth text among the plurality of preset texts.

[0194] The electronic device 100 may receive a sixth voice for the sixth text in operation S810.

[0195] In operation S811, the electronic device 100 can transmit to the server 400 sixth audio data including an audio signal corresponding to a sixth audio.

[0196] In operation S812, the server 400 can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device 100. The server 400 can convert the voice signal corresponding to the sixth voice included in the sixth voice data received from the electronic device 100 into text. The server 400 can determine whether the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.

[0197] In operation S813, the server 400 can notify the electronic device 100 that the processing for the sixth voice has been completed.

[0198] On the other hand, if the text converted from the audio signal corresponding to the sixth audio corresponds to the sixth text, the server 400 can generate sixth voiceprint information for the sixth audio based on the audio signal corresponding to the sixth audio.

[0199] In operation S814, the server 400 may store sixth voice data and sixth voiceprint information for the sixth voice. The server 400 may map the sixth voice data and sixth voiceprint information to user identification information corresponding to the logged-in user account and store the sixth voice data and sixth voiceprint information. In this case, six different pieces of voice data and six different pieces of voiceprint information may be mapped to the user identification information corresponding to the logged-in user account.

[0200] The electronic device 100 may complete the process of registering a voice ID in operation S815. For example, the electronic device 100 may complete the process of registering a voice ID based on the completion of processing for a predetermined number of six different preset texts.

[0201] 9 , when a user account is not logged in to the server 400, the electronic device 100 may output a login screen 900 related to logging in to the server 400 via the display 180. The login screen 900 may include an object 910 representing a non-login state, a login object 920 for executing login, etc. When a user selects the login object 920 using the pointer 205 corresponding to the remote control device 200, the electronic device 100 may output a screen for inputting an ID and a password. In this case, the user can input the ID and password of the user account to log in to the server 400 with the user account.

[0202] 10 , when a voice ID is not registered for a user account logged in to the server 400, the electronic device 100 can output a first account screen 1000 corresponding to the user account for which a voice ID is not registered. The first account screen 1000 can include an object 1010 representing the logged-in user account, an object 1020 corresponding to registering a voice ID, etc. When the user uses the pointer 205 to select the object 1020 corresponding to registering a voice ID, the electronic device 100 can start the process of registering a voice ID.

[0203] 11, when a voice ID has been registered in advance to a user account logged into the server 400, the electronic device 100 may output a second account screen 1100 corresponding to the user account to which the voice ID has been registered. The second account screen 1100 may include an object 1110 representing the logged-in user account, a re-registration object 1120 corresponding to re-registration of the voice ID, a deletion object 1130 corresponding to deletion of the voice ID, an activation object 1140 corresponding to use of a function related to the voice ID, etc. The user may select the activation object 1140 using the pointer 205 to activate or deactivate use of a function related to the voice ID.

[0204] 12 , when an object 1020 corresponding to voice ID registration is selected on the first account screen 1000, or when a re-registration object 1120 is selected on the second account screen 1100, the electronic device 100 can output a start screen 1200 that starts voice ID registration. When the user selects the start object 1210 using the pointer 205, the electronic device 100 can output a text screen that outputs preset text.

[0205] 13, the electronic device 100 may output a text screen 1300 that displays one of a plurality of preset texts. The text screen 1300 may include preset text 1301, a text order 1302, an end object 1310 that ends the process of registering a voice ID, an input object 1320 that receives voice, and the like.

[0206] The voice ID registration process may end when the user selects the end object 1310 using the pointer 205. For example, when the voice ID registration process ends, all data stored in the server 400 while the voice ID registration process is in progress may be deleted.

[0207] When the user selects the input object 1320 with the pointer 205, the electronic device 100 can receive speech for the text.

[0208] According to one embodiment, when the text screen 1300 is output and the user presses a predetermined button (e.g., a voice input button) included in the remote control device 200, the electronic device 100 can receive voice for the text based on the user input of pressing the predetermined button received from the remote control device 200.

[0209] Meanwhile, according to an embodiment, when a user presses a predetermined button (e.g., a voice input button) included in the remote control device 200 while a voice ID registration process is in progress, the electronic device 100 may interrupt the voice ID registration process based on a user input of pressing the predetermined button received from the remote control device 200. In this case, the user input of pressing the predetermined button (e.g., a voice input button) included in the remote control device 200 may correspond to a user input of starting voice recognition for voice received through the remote control device 200. The electronic device 100 may perform an operation related to voice recognition for voice data including a voice signal received from the remote control device 200.

[0210] FIG. 14 is a flow chart of a method of operating a server according to one embodiment of the present disclosure.

[0211] As shown in FIG. 14, the server 400 may receive data including an audio signal from the electronic device 100 in operation S1410.

[0212] The electronic device 100 can receive a voice input corresponding to a voice uttered by a user. For example, the image display device 100 can receive a voice signal corresponding to the voice input from the remote control device 200. For example, the image display device 100 can receive a voice signal corresponding to the voice input via a microphone provided in the input unit 160.

[0213] In response to receiving a voice input, the electronic device 100 may transmit voice data including a voice signal to the server 400. For example, the electronic device 100 may transmit voice data including a voice signal in a predetermined unit such as a syllable or a word to the server 400. That is, when a user speaks a sentence, the electronic device 100 may transmit a voice signal in a predetermined unit to the server 400 while receiving a voice input corresponding to the sentence or phrase from the remote control device 200.

[0214] In operation S1420, the server 400 can generate voiceprint information for the voice spoken by the user based on the voice data received from the electronic device 100.

[0215] In operation S1430, server 400 can determine whether a user account corresponding to the voice uttered by the user exists. For example, server 400 can compare the generated voiceprint information with voiceprint information stored in database 490 and determine whether voiceprint information corresponding to the generated voiceprint information is stored in database 490. In this case, if voiceprint information corresponding to the generated voiceprint information is stored in database 490, server 400 can determine that the user account corresponding to the voiceprint information is the user account corresponding to the voice uttered by the user.

[0216] According to an embodiment, the electronic device 100 may transmit device identification information, a user list, a language code indicating a language type, and the like, along with voice data, to the server 400. The server 400 may search the database 490 for voiceprint information (hereinafter, referred to as candidate voiceprint information) corresponding to the user identification information included in the user list received from the electronic device 100. The server 400 may determine whether the candidate voiceprint information and the generated voiceprint information correspond to each other. The server 400 may determine, as the user identification information corresponding to the voice input to the electronic device 100, the user identification information to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped among the candidate voiceprint information. Meanwhile, if there is no candidate voiceprint information corresponding to the generated voiceprint information, the server 400 may determine that there is no user identification information corresponding to the voice input to the electronic device 100.

[0217] In operation S1440, if a user account corresponding to the voice uttered by the user exists, the server 400 can acquire data for the user account (hereinafter, account data) corresponding to the voice uttered by the user. The data for the user account can be stored in database 490. For example, the account data stored in database 490 can include a user account ID, a password, user identification information, device identification information mapped to the user account, the user's gender, the user's age, a user's content search history, a content viewing history, an application execution history, a voice utterance history, a user's preferred content genre, etc.

[0218] In operation S1450, server 400 may determine whether the acquired account data includes user characteristics. For example, user characteristics may include the user's age, gender, region, country, etc. If the acquired account data includes user characteristics, server 400 may acquire the user characteristics from the acquired account data.

[0219] In operation S1460, server 400 can acquire user characteristics from the generated voiceprint information. Server 400 can determine the gender, age, etc. corresponding to the voice spoken by the user based on the voice waveform, voice frequency band, voice power spectrum, etc. corresponding to the voiceprint information for the voice spoken by the user generated in operation S1420. For example, if there is no user account corresponding to the voice spoken by the user, server 400 can acquire the user characteristics from the generated voiceprint information. For example, if the account data corresponding to the voice spoken by the user does not include the user characteristics, server 400 can acquire the user characteristics from the generated voiceprint information.

[0220] The server 400 can acquire user characteristics based on voice features corresponding to the generated voiceprint information. According to one embodiment, the server 400 can acquire user characteristics from the generated voiceprint information using a learning model trained through machine learning stored in the database 490. Machine learning refers to a computer learning process that allows the computer to solve problems by learning through data without a human directly instructing the computer on logic. Deep learning refers to an artificial intelligence technology that allows a computer to learn like a human being, using a method of teaching a computer how to think using an artificial neural network (ANN). An artificial neural network (ANN) can be implemented in the form of software or hardware such as a chip. For example, artificial neural networks (ANNs) can include various types of algorithms, such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep belief networks (DBNs).

[0221] In operation S1470, the server 400 can determine text corresponding to the voice uttered by the user (hereinafter, spoken text) based on the characteristics of the user.

[0222] The database 490 of the server 400 may store a history of text usage by characteristics based on the usage history received from the plurality of electronic devices 100. Here, the usage history may include a history of a user searching for a predetermined text, a history of a user viewing content corresponding to the predetermined text, a history of a user selecting a predetermined text corresponding to a voice spoken by the user, etc. For example, the server 400 may organize the usage history received from the plurality of electronic devices 100 into a database by user characteristics such as gender and age.

[0223] The server 400 can generate a group of candidates (hereinafter, candidate texts) for text corresponding to the voice signal received from the electronic device 100. For example, the server 400 can generate multiple candidate texts based on the waveform of syllables included in the voice signal, the relationship between words, the probability of the next word appearing given the previous word, etc. In this case, the server 400 can generate a predetermined number (e.g., three) of candidate texts in descending order of rank using an n-best method.

[0224] The server 400 may determine the utterance text based on the history of use of each of the plurality of candidate texts in relation to the user's characteristics among the history of use of text by characteristics stored in the database 490. For example, the server 400 may determine the priority of each of the plurality of candidate texts in consideration of the frequency of use of each of the plurality of candidate texts according to the user's age, the frequency of use of each of the plurality of candidate texts according to the user's gender, etc. In this case, the server 400 may determine one of the plurality of candidate texts with the highest priority as the utterance text.

[0225] If the acquired account data includes a history of using at least one of the plurality of candidate texts, the server 400 may determine the utterance text based on the history included in the acquired account data. If the acquired account data does not include a history of using at least one of the plurality of candidate texts, the server 400 may determine the utterance text based on a history of using each of the plurality of candidate texts in association with the user's characteristics among a history of using text by characteristic stored in the database 490. If the database 490 does not include a history of using each of the plurality of candidate texts, the server 400 may determine the highest ranked candidate text generated by the n-best method as the utterance text.

[0226] In operation S1480, the server 400 may perform an intention analysis on the spoken text. For example, the server 400 may acquire keywords included in the spoken text, sentence components of the keywords, the intention of the sentence, instructions corresponding to the spoken text, etc. The server 400 may transmit the result of the intention analysis on the spoken text to the electronic device 100.

[0227] According to an embodiment, the server 400 may generate a result of performing an intention analysis on the spoken text based on the type of the spoken text. Here, the type of the spoken text may include a genre, service, function, application, etc. corresponding to a keyword included in the spoken text. In the present disclosure, the genre is described as an example in relation to the type of the spoken text, but is not limited thereto.

[0228] The database 490 of the server 400 may store a history of how text types are used by characteristics based on usage histories received from a plurality of electronic devices 100. The server 400 may generate a result of performing an intention analysis on the spoken text based on a history of how a type of spoken text is used in association with a user's characteristics among the history of how text types are used by characteristics stored in the database 490. For example, the server 400 may determine a priority for multiple types of spoken text by considering the frequency with which the type of spoken text is used depending on the user's gender. In this case, the server 400 may generate a result of performing an intention analysis on the spoken text based on one of the multiple types with the highest priority.

[0229] The server 400 may generate a result of intention analysis performed on the spoken text based on the acquired account data. For example, the server 400 may generate a result of intention analysis performed on the spoken text based on a user's preferred genre included in the acquired account data. For example, if the acquired account data includes a history of using a type of spoken text, the server 400 may generate a result of intention analysis performed on the spoken text based on the history included in the acquired account data. On the other hand, if the acquired account data does not include a user's preferred genre or a history of using a type of spoken text, the server 400 may generate a result of intention analysis performed on the spoken text based on a history of using a type of spoken text in association with a user's characteristics among the history of using a type of text by characteristics stored in the database 490.

[0230] According to one embodiment, the server 400 may transmit data for the candidate text to the electronic device 100. For example, the server 400 may transmit data for the candidate text to the electronic device 100 along with the results of an intention analysis performed on the spoken text.

[0231] The electronic device 100 may provide the user with candidate texts received from the server 400. For example, the electronic device 100 may output an object corresponding to the candidate texts via the display 180. In this case, when the user selects one of the candidate texts, the electronic device 100 may request the server 400 to perform an intention analysis on the selected candidate text. The server 400 may transmit to the electronic device 100 a result of performing an intention analysis on the candidate text selected by the user and received from the electronic device 100.

[0232] According to one embodiment, server 400 may transmit data regarding the type of spoken text to electronic device 100. For example, server 400 may transmit data regarding the type of spoken text to electronic device 100 along with the results of performing an intention analysis on the spoken text.

[0233] The electronic device 100 may provide the type of spoken text received from the server 400 to the user. For example, the electronic device 100 may output an object corresponding to the type of spoken text via the display 180. In this case, when the user selects one of the types of spoken text, the electronic device 100 may request the server 400 to perform an intention analysis on the spoken text associated with the selected type. The server 400 may transmit to the electronic device 100 a result of performing an intention analysis on the spoken text according to the type selected by the user received from the electronic device 100.

[0234] Meanwhile, at least a part of the operations of the server 400 described with reference to FIG.

[0235] 15 and 16, the user 1 can utter voices 1510 and 1610 to search for a specific person. The electronic device 100 can transmit voice data including voice signals corresponding to the voices 1510 and 1610 uttered by the user 1 to the server 400.

[0236] The server 400 can generate multiple candidate texts corresponding to the speech 1510, 1610 uttered by the user 1. For example, the server 400 can generate "Search for Song Ga-in's songs," "Search for Song Ga-in's songs," and "Search for Song A-in's songs" as candidate texts.

[0237] The server 400 can determine one of a plurality of candidate texts as the utterance text based on a history of use of a plurality of candidate texts included in the account data of the user 1. For example, the server 400 can check whether the account data of the user 1 includes a history of use of "Song Gain", a history of use of "Song Gain", and a history of use of "Song Ain". In this case, the server 400 can determine the most frequently used one of "Song Gain", "Song Gain", and "Song Ain" as the utterance text.

[0238] The server 400 may determine one of the plurality of candidate texts as the spoken text based on the history of use of each of the plurality of candidate texts in association with the user characteristics stored in the database 490. For example, if the user 1 is male and in his 30s, the server 400 may check the history of use of "Song Gain," "Song Gain," and "Song Ain" by male users in their 30s in the database 490. In this case, the server 400 may determine the most frequently used one of "Song Gain," "Song Gain," and "Song Ain" as the spoken text.

[0239] 15 , the electronic device 100 can output text 1520 corresponding to speech 1510 spoken by the user 1. For example, the electronic device 100 can output text 1520 corresponding to speech 1510 spoken by the user 1 based on the result of intention analysis performed on the spoken text received from the server 400.

[0240] The electronic device 100 may provide a search result 1530 for the song "Song Ga-in" based on the result of an intention analysis performed on the spoken text received from the server 400. For example, the server 400 may transmit to the electronic device 100 a result of an intention analysis performed on "Search for the song of Song Ga-in" based on the fact that the frequency of the history of the use of "Song Ga-in" is the highest in the account data of User 1. For example, if User 1 is male and in his 30s, the server 400 may transmit to the electronic device 100 a result of an intention analysis performed on "Search for the song of Song Ga-in" based on the fact that the percentage of male users in their 30s who use "Song Ga-in" is 60%, the percentage of those who use "Song Ga-in" is 38%, and the percentage of those who use "Song A-in" is 2%.

[0241] The electronic device 100 may output an object 1540 corresponding to "Song Ga-in" based on data on the candidate text received from the server 400. In this case, if the user selects the object 1540 corresponding to "Song Ga-in," the electronic device 100 may request the server 400 to analyze the user's intention to "find Song Ga-in's songs."

[0242] 16 , the electronic device 100 can output text 1620 corresponding to speech 1610 spoken by the user 1. For example, the electronic device 100 can output text 1620 corresponding to speech 1610 spoken by the user 1 based on the result of intention analysis performed on the spoken text received from the server 400.

[0243] The electronic device 100 may provide a search result 1630 for the song "Song Ga-in" based on the result of an intention analysis performed on the spoken text received from the server 400. For example, the server 400 may transmit to the electronic device 100 the result of an intention analysis performed on "Search for Song Ga-in's songs" based on the fact that the frequency of the history of the use of "Song Ga-in" in the account data of User 1 is the highest. For example, if User 1 is male and in his 60s, the server 400 may transmit to the electronic device 100 the result of an intention analysis performed on "Search for Song Ga-in's songs" based on the fact that the percentage of male users in their 60s who use "Song Ga-in" is 20%, the percentage of those who use "Song Ga-in" is 78%, and the percentage of those who use "Song A-in" is 2%.

[0244] The electronic device 100 may output an object 1640 corresponding to "Song Ga-in" based on data on the candidate text received from the server 400. In this case, if the user selects the object 1640 corresponding to "Song Ga-in," the electronic device 100 may request the server 400 to analyze the user's intention to "find me a song by Song Ga-in."

[0245] 17 and 18, User 2 can utter "Search for New Kids on the Block" 1710, 1810. Electronic device 100 can transmit to server 400 voice data including a voice signal corresponding to the voice 1710, 1810 uttered by User 2.

[0246] The server 400 can generate multiple candidate texts corresponding to "Search for new kids on the block," which is speech 1710, 1810 spoken by the user 2. For example, the server 400 can generate candidate texts such as "Search for new kids on the block," "Search for new kiz on the block," and "Search for you kids on the block."

[0247] Server 400 can determine one of multiple candidate texts as the utterance text based on a history of use of multiple candidate texts included in User 2's account data. For example, server 400 can check whether User 2's account data includes a history of use of "New Kids on the Block," a history of use of "New Kidz on the Block," and a history of use of "You Kids on the Block." In this case, server 400 can determine the most frequently used one of "New Kids on the Block," "New Kidz on the Block," and "You Kids on the Block" as the utterance text.

[0248] The server 400 may determine one of the plurality of candidate texts as the utterance text based on the usage history of each of the plurality of candidate texts in association with the user characteristics stored in the database 490. For example, if the user 2 is female and in her 40s, the server 400 may check the usage history of "New Kids on the Block," "New Kidz on the Block," and "You Kids on the Block" by female users in their 40s in the database 490. In this case, the server 400 may determine the most frequently used one of "New Kids on the Block," "New Kidz on the Block," and "You Kids on the Block" as the utterance text.

[0249] 17 , the electronic device 100 can output text 1720 corresponding to speech 1710 uttered by the user 2. For example, the electronic device 100 can output text 1720 corresponding to speech 1710 uttered by the user 2 based on the result of intention analysis performed on the utterance text received from the server 400.

[0250] The electronic device 100 can provide search results 1730 for "you kids on the block" based on the results of an intention analysis performed on the spoken text received from the server 400. For example, the server 400 can transmit to the electronic device 100 the results of an intention analysis performed on "search for you kids on the block" based on the highest frequency of historical use of "you kids on the block" in the account data of User 2. For example, if user 2 is female and in her 40s, based on the fact that 10% of female users in their 40s have used “New Kids on the Block,” 20% have used “New Kidz on the Block,” and 70% have used “You Kids on the Block,” server 400 can transmit the results of an intention analysis performed on “Search You Kids on the Block” to electronic device 100.

[0251] The electronic device 100 may output an object 1740 corresponding to "New Kiz on the Block" based on data on the candidate text received from the server 400. In this case, if the user selects the object 1740 corresponding to "New Kiz on the Block," the electronic device 100 may request the server 400 to analyze the intention of "Search for New Kiz on the Block."

[0252] 18 , the electronic device 100 can output text 1820 corresponding to speech 1810 uttered by the user 2. For example, the electronic device 100 can output text 1820 corresponding to speech 1810 uttered by the user 2 based on the result of intention analysis performed on the utterance text received from the server 400.

[0253] The electronic device 100 can provide search results 1830 for the song "Kiz on the Block" based on the results of an intention analysis performed on the spoken text received from the server 400. For example, the server 400 can transmit to the electronic device 100 the results of an intention analysis performed on "Search for Kiz on the Block" based on the highest frequency of historical use of "Kiz on the Block" in the account data of User 2. For example, if user 2 is female and in her 60s, based on the fact that 20% of female users in their 60s have used “New Kids on the Block,” 50% have used “New Kidz on the Block,” and 30% have used “You Kids on the Block,” server 400 can transmit the results of an intention analysis performed on “Search for New Kidz on the Block” to electronic device 100.

[0254] The electronic device 100 may output an object 1840 corresponding to "You Kids on the Block" based on data on the candidate text received from the server 400. In this case, if the user selects the object 1840 corresponding to "You Kids on the Block," the electronic device 100 may request the server 400 to analyze the intention of "Search for You Kids on the Block."

[0255] 19 to 21, the user can speak to search for a specific content, "Begin Again." The electronic device 100 can transmit to the server 400 voice data including a voice signal corresponding to the voice spoken by the user.

[0256] The server 400 can determine "Begin Again" as the spoken text from multiple candidate texts corresponding to the speech uttered by the user. The server 400 can generate a result of performing an intent analysis on the search for "Begin Again" based on the type of "Begin Again."

[0257] The server 400 can determine the type of "Begin Again" based on the user's account data. For example, the server 400 can determine the type of "Begin Again" based on the user's preferred genre, which is included in the user's account data. For example, the server 400 can determine the type of "Begin Again" based on the history of use of the "Begin Again" type, which is included in the user's account data. In this case, the server 400 can determine the type of "Begin Again" to be the one that has been used most frequently, out of the "Begin Again" types, "Movie" and "Entertainment."

[0258] The server 400 may determine the type of utterance text based on the history of use of the type of utterance text in relation to the user's characteristics stored in the database 490. For example, if the user is female and in her 30s, the server 400 may check the history of female users in their 30s using "Begin Again," which belongs to the "movie" genre, and "Begin Again," which belongs to the "entertainment" genre, in the database 490. In this case, the server 400 may determine the most frequently used type of "Begin Again," either "movie" or "entertainment," as the utterance text.

[0259] As shown in FIG. 19, the electronic device 100 can output text 1900 corresponding to the speech uttered by the user.

[0260] The electronic device 100 can provide search results 1910 for "Begin Again" in the "movie" genre based on the result of intention analysis performed on the spoken text received from the server 400. For example, the server 400 can transmit to the electronic device 100 search results for "Begin Again" in the "movie" genre based on the fact that the user's preferred genre is "movie." For example, if the gender of user 1 is female and the age is in her 30s, the server 400 can transmit to the electronic device 100 search results for "Begin Again" in the "movie" genre based on the fact that 45% of female users in their 30s have used "Begin Again" in the "movie" genre and 30% have used "Begin Again" in the "entertainment" genre.

[0261] The electronic device 100 may output an object 1920 corresponding to “Entertainment,” which is a type of “Begin Again,” based on data on the type of “Begin Again” received from the server 400. In this case, if the user selects the object 1920 corresponding to “Entertainment,” the electronic device 100 may request the server 400 to search for “Begin Again” in the “Entertainment” genre.

[0262] 20 , the electronic device 100 can provide search results 2010 for "Begin Again" in the "Entertainment" genre based on the results of intention analysis performed on the spoken text received from the server 400. For example, the server 400 can transmit to the electronic device 100 the search results for "Begin Again" in the "Entertainment" genre, based on the fact that the user's preferred genre is "Entertainment." For example, if the user 1 is male and in his twenties, the server 400 can transmit to the electronic device 100 the search results for "Begin Again" in the "Entertainment" genre, based on the fact that 25% of male users in their twenties have used "Begin Again" in the "Movie" genre and 45% have used "Begin Again" in the "Entertainment" genre.

[0263] The electronic device 100 may output an object 2020 corresponding to "Movie", which is a type of "Begin Again", based on data on the type of "Begin Again" received from the server 400. In this case, if the user selects the object 2020 corresponding to "Movie", the electronic device 100 may request the server 400 to search for "Begin Again" in the genre of "Movie".

[0264] As shown in FIG. 21 , server 400 can transmit to electronic device 100 the search results for "Begin Again" in the genre "movies" and "Begin Again" in the genre "entertainment." For example, based on the fact that the user's preferred genres are "movies" and "entertainment," server 400 can transmit to electronic device 100 the search results for "Begin Again" in the genre "movies" and "Begin Again" in the genre "entertainment." For example, when user 1 is male and in his twenties, server 400 can transmit to electronic device 100 the search results for "Begin Again" in the genre "movies" and "Begin Again" in the genre "entertainment," based on the fact that the difference in the frequency of use of each type is less than a predetermined difference (e.g., 5%), assuming that 45% of male users in their thirties use "Begin Again" in the genre "movies" and 40% use "Begin Again" in the genre "entertainment."

[0265] The electronic device 100 can provide search results 2110 for "Begin Again" in the "movie" genre and "Begin Again" in the "entertainment" genre based on the result of intent analysis performed on the spoken text received from the server 400. For example, if the proportion of "Begin Again" in the "movie" genre used is greater than the proportion of "Begin Again" in the "entertainment" genre used, the "movie" related search results can be provided before the "entertainment" related search results.

[0266] The electronic device 100 may output an object 2121 corresponding to “Movie” and an object 2122 corresponding to “Entertainment,” which are types of “Begin Again,” based on data on the type of “Begin Again” received from the server 400. In this case, when the user selects one of the multiple objects 2121 and 2122, the electronic device 100 may request the server 400 to provide a search result for “Begin Again” in the genre selected by the user.

[0267] As noted above, in accordance with at least one embodiment of the present disclosure, identification information for a user's voice can be registered with the user's account.

[0268] Additionally, in accordance with at least one embodiment of the present disclosure, a user can be identified based on the user's voice.

[0269] Furthermore, according to at least one embodiment of the present disclosure, the accuracy of the results of processing a user's voice can be improved based on the user's characteristics identified from the voice.

[0270] Additionally, in accordance with at least one embodiment of the present disclosure, the accuracy of the results of processing a user's voice can be improved based on user characteristics contained in data for the user account.

[0271] Additionally, in accordance with at least one embodiment of the present disclosure, a user's usage history can be used to update a database used to process the user's voice.

[0272] As shown in FIGS. 1 to 21 , a server 400 according to an aspect of the present disclosure includes a communication unit 480 that communicates with the electronic device 100, a database 490 that stores a history of text usage by characteristic, and a controller 470. The controller 470 may generate a plurality of first texts corresponding to a voice signal received from the electronic device 100, generate first identification information corresponding to the voice signal, acquire a user characteristic corresponding to the voice signal based on the first identification information, determine a second text corresponding to the voice signal from the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user characteristic among the history of text usage by characteristic, and transmit a result of intention analysis of the second text to the electronic device 100.

[0273] Furthermore, according to one aspect of the present disclosure, the first identification information may include a feature vector for a voice print of the audio signal.

[0274] According to one aspect of the present disclosure, the user characteristics may include at least one of age and gender.

[0275] Furthermore, according to one aspect of the present disclosure, the database 490 stores identification information corresponding to a user account and account data for the user account, and the controller 470, when second identification information corresponding to the first identification information is stored in the database 490, acquires the user characteristics from the account data corresponding to the second identification information, and when the second identification information is not stored in the database 490 or when the account data corresponding to the second identification information does not include the user characteristics, acquires the user characteristics from the first identification information.

[0276] According to one aspect of the present disclosure, the controller 470 can receive a user list including at least one user identification from the electronic device 100, search the database 490 for third identification corresponding to the user identification included in the user list, and compare the first identification with the third identification to determine the second identification.

[0277] According to one aspect of the present disclosure, the account data for the user account stored in the database 490 includes a history of the user's use of text, and the controller 470 can determine the second text based on the second text history if the account data corresponding to the second identification information includes a second text history using at least one of the plurality of first texts, and can determine the second text based on the first text history if the account data corresponding to the second identification information does not include the second text history.

[0278] According to one aspect of the present disclosure, the database 490 stores a history of how a text type is used according to a characteristic, and the controller 470 determines the second text type based on a first type history in which the second text type is used in association with the user's characteristic among the history of how a text type is used according to the characteristic, and generates a result of performing an intention analysis on the second text based on the determined type.

[0279] According to one aspect of the present disclosure, the account data for the user account includes a history of the user's use of text types, and the controller 470 can determine the second text type based on the second typing history if the account data corresponding to the second identification information includes a second typing history using the second text type, and can determine the second text type based on the first typing history if the account data corresponding to the second identification information does not include the second typing history.

[0280] A system 10 according to an aspect of the present disclosure includes an electronic device 100 and a server 400. When a voice signal is received via a user input interface unit 150, the electronic device 100 transmits data including the voice signal to the server 400 and outputs a result of an intention analysis performed on the voice signal received from the server 400. The server 400 generates a plurality of first texts corresponding to the voice signal received from the electronic device 100, generates first identification information corresponding to the voice signal, acquires a user characteristic corresponding to the voice signal based on the first identification information, determines a second text corresponding to the voice signal from the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user characteristic among histories of text usage by characteristic stored in a database 490, and transmits a result of the intention analysis performed on the second text to the electronic device 100 as a result of the intention analysis performed on the voice signal.

[0281] According to another aspect of the present disclosure, the server 400 may transmit at least one of the remaining first texts, excluding the second text, to the electronic device 100, and the electronic device 100 may output a text object corresponding to at least one of the remaining first texts via the display 180, and may request the server 400 to perform an intention analysis on the text corresponding to the selected text object based on a user input selecting the text object.

[0282] Furthermore, according to one aspect of the present disclosure, the database 490 stores identification information corresponding to a user account and account data for the user account, and the server 400, when second identification information corresponding to the first identification information is stored in the database 490, acquires the user characteristics from the account data corresponding to the second identification information, and when the second identification information is not stored in the database 490 or when the account data corresponding to the second identification information does not include the user characteristics, acquires the user characteristics from the first identification information.

[0283] According to one aspect of the present disclosure, the electronic device 100 transmits a user list including at least one user identification information to the server 400, and the server 400 searches the database 490 for third identification information corresponding to the user identification information included in the user list, and compares the first identification information with the third identification information to determine the second identification information.

[0284] According to one aspect of the present disclosure, the database 490 stores a history of how a text type is used according to a characteristic, and the server 400 determines the second text type based on a first type history in which the second text type is used in association with the user's characteristic among the history of how a text type is used according to the characteristic, and generates a result of an intention analysis of the second text based on the determined type.

[0285] According to one aspect of the present disclosure, the account data for the user account includes a history of the user's use of text types, and if the account data corresponding to the second identification information includes a second typing history using the second text type, the server 400 can determine the second text type based on the second typing history, and if the account data corresponding to the second identification information does not include the second typing history, the server 400 can determine the second text type based on the first typing history.

[0286] According to another aspect of the present disclosure, the server 400 may transmit at least one type corresponding to the second text to the electronic device 100, and the electronic device 100 may output a type object corresponding to the at least one type corresponding to the second text via the display 180, and may request the server 400 to perform an intention analysis on the type corresponding to the selected type object based on a user input for selecting the type object.

[0287] The attached drawings are intended to facilitate understanding of the embodiments disclosed in this specification, and it should be understood that the attached drawings do not limit the technical ideas disclosed in this specification, and include any modifications, equivalents, or alternatives that fall within the idea and technical scope of the present disclosure.

[0288] Meanwhile, the operating method of the present disclosure can be realized as processor-readable code on a processor-readable recording medium. The processor-readable recording medium includes all types of recording devices in which processor-readable data is stored. Examples of the processor-readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc., and also includes being realized in the form of a carrier wave such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed among computer systems connected via a network, so that the processor-readable code can be stored and executed in a distributed manner.

[0289] Furthermore, while the above illustrates and describes preferred embodiments of the present disclosure, the present disclosure is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the technical field to which the invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical ideas and perspectives of the present disclosure.

[0290] [One aspect of the present invention] One aspect of the present invention is as follows. [Claim 1] a server, a communication unit for communicating with the electronic device; A database that stores the history of text usage by characteristics; a controller; The controller generating a plurality of first texts corresponding to the audio signals received from the electronic device; generating first identification information corresponding to the audio signal; Obtaining a user characteristic corresponding to the voice signal based on the first identification information; determining a second text corresponding to the voice signal from among the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user's characteristic among the history of text usage by characteristic; and transmitting a result of performing an intention analysis on the second text to the electronic device. [Claim 2] 2. The server according to claim 1, wherein the first identification information includes a feature vector for a voice print of the audio signal. [Claim 3] 2. The server according to claim 1, wherein the user characteristics include at least one of age and gender. [Claim 4] the database stores identification information corresponding to a user account and account data for the user account; The controller If second identification information corresponding to the first identification information is stored in the database, obtaining characteristics of the user from account data corresponding to the second identification information; The server of claim 1, characterized in that if the second identification information is not stored in the database or if the account data corresponding to the second identification information does not include the user characteristics, the server obtains the user characteristics from the first identification information. [Claim 5] The controller receiving a user list including at least one user identification from the electronic device; searching the database for third identification information corresponding to the user identification information included in the user list; 5. The server of claim 4, further comprising: comparing the first identification information with the third identification information to determine the second identification information. [Claim 6] the account data for the user account stored in the database includes a history of the user's use of text; The controller If the account data corresponding to the second identification information includes a second text history using at least one of the plurality of first texts, determining the second text based on the second text history; 5. The server of claim 4, wherein if the second text history is not included in the account data corresponding to the second identification information, the server determines the second text based on the first text history. [Claim 7] The database stores a history of the use of text types by characteristics; The controller determining the second text type based on a first type history in which the second text type was used in association with the user's characteristic among the history of text type usage according to the characteristic; The server of claim 4, further comprising: generating a result of performing an intention analysis on the second text based on the determined type. [Claim 8] the account data for the user account includes a history of the user's use of text types; The controller if the account data corresponding to the second identification information includes a second typing history using the second text type, determining the second text type based on the second typing history; 8. The server of claim 7, wherein if the second typing history is not included in the account data corresponding to the second identification information, the server determines the type of the second text based on the first typing history. [Claim 9] 1. A system comprising: an electronic device; and a server; The electronic device is When an audio signal is received via the user input interface unit, transmitting data including the audio signal to the server; outputting a result of performing an intention analysis on the speech signal received from the server; The server generating a plurality of first texts corresponding to the audio signals received from the electronic device; generating first identification information corresponding to the audio signal; Obtaining a user characteristic corresponding to the voice signal based on the first identification information; determining a second text corresponding to the voice signal from among the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user's characteristic among a history of text usage by characteristic stored in a database; transmitting a result of performing an intent analysis on the second text to the electronic device as a result of performing an intent analysis on the speech signal. [Claim 10] the server transmits at least one of the remaining first texts, excluding the second text, to the electronic device; The electronic device is outputting via a display a text object corresponding to at least one of the remainder of the plurality of first texts; The system according to claim 9, further comprising: requesting the server to perform an intention analysis on the text corresponding to the selected text object based on a user input for selecting the text object. [Claim 11] the database stores identification information corresponding to a user account and account data for the user account; The server If second identification information corresponding to the first identification information is stored in the database, obtaining characteristics of the user from account data corresponding to the second identification information; 10. The system of claim 9, wherein if the second identification information is not stored in the database or if the user characteristics are not included in the account data corresponding to the second identification information, the user characteristics are obtained from the first identification information. [Claim 12] The electronic device transmits a user list including at least one user identification information to the server; The server searching the database for third identification information corresponding to the user identification information included in the user list; 12. The system of claim 11, further comprising: comparing the first identification information with the third identification information to determine the second identification information. [Claim 13] The database stores a history of the use of text types by characteristics; The server determining the second text type based on a first type history in which the second text type was used in association with the user's characteristic among the history of text type usage according to the characteristic; The system of claim 11 , further comprising: generating a result of performing an intent analysis on the second text based on the determined type. [Claim 14] the account data for the user account includes a history of the user's use of text types; The server if the account data corresponding to the second identification information includes a second typing history using the second text type, determining the second text type based on the second typing history; 14. The system of claim 13, wherein if the second typing history is not included in the account data corresponding to the second identification information, the system determines the type of the second text based on the first typing history. [Claim 15] the server transmits at least one type corresponding to the second text to the electronic device; The electronic device is outputting via a display a type object corresponding to at least one type corresponding to the second text; The system according to claim 13, further comprising: requesting the server to perform an intention analysis for a type corresponding to the selected type object based on a user input for selecting the type object.

Claims

1. a server, a communication unit for communicating with the electronic device; A database that stores the history of text usage by characteristics; a controller; The controller generating a plurality of first texts corresponding to the audio signals received from the electronic device; generating first identification information corresponding to the audio signal; Obtaining a user characteristic corresponding to the voice signal based on the first identification information; determining a second text corresponding to the speech signal from among the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user's characteristic among the history of text usage by characteristic; and transmitting a result of performing an intention analysis on the second text to the electronic device.

2. 2. The server of claim 1, wherein the first identification information includes a feature vector for a voice print of the audio signal.

3. The server of claim 1 , wherein the user characteristics include at least one of age and gender.

4. the database stores identification information corresponding to a user account and account data for the user account; The controller If second identification information corresponding to the first identification information is stored in the database, obtaining characteristics of the user from account data corresponding to the second identification information; The server of claim 1, characterized in that if the second identification information is not stored in the database or if the account data corresponding to the second identification information does not include the user characteristics, the user characteristics are obtained from the first identification information.

5. The controller receiving a user list including at least one user identification from the electronic device; searching the database for third identification information corresponding to the user identification information included in the user list; 5. The server of claim 4, wherein the second identification information is determined by comparing the first identification information with the third identification information.

6. the account data for the user account stored in the database includes a history of the user's use of text; The controller If the account data corresponding to the second identification information includes a second text history using at least one of the plurality of first texts, determining the second text based on the second text history; 5. The server of claim 4, wherein if the second text history is not included in the account data corresponding to the second identification information, the second text is determined based on the first text history.

7. The database stores a history of the use of text types by characteristics; The controller determining the second text type based on a first type history in which the second text type was used in association with the user's characteristic among the history of using the text type according to the characteristic; The server of claim 4 , further comprising: generating a result of performing an intention analysis on the second text based on the determined type.

8. the account data for the user account includes a history of the user's use of text types; The controller if the account data corresponding to the second identification information includes a second typing history using the second text type, determining the second text type based on the second typing history; 8. The server of claim 7, wherein if the second typing history is not included in the account data corresponding to the second identification information, the server determines the type of the second text based on the first typing history.

9. 1. A system comprising: an electronic device; and a server; The electronic device is When an audio signal is received via the user input interface unit, transmitting data including the audio signal to the server; and outputting a result of performing an intention analysis on the voice signal received from the server. The server generating a plurality of first texts corresponding to the audio signals received from the electronic device; generating first identification information corresponding to the audio signal; Obtaining a user characteristic corresponding to the voice signal based on the first identification information; determining a second text corresponding to the speech signal from among the plurality of first texts based on a first text history in which each of the plurality of first texts was used in association with the user's characteristic among a history of text usage by characteristic stored in a database; transmitting a result of performing an intent analysis on the second text to the electronic device as a result of performing an intent analysis on the speech signal.

10. the server transmits at least one of the remaining first texts, excluding the second text, to the electronic device; The electronic device is outputting via a display a text object corresponding to at least one of the remainder of the plurality of first texts; The system of claim 9, further comprising: requesting the server to perform an intention analysis on the text corresponding to the selected text object based on a user input for selecting the text object.

11. the database stores identification information corresponding to a user account and account data for the user account; The server If second identification information corresponding to the first identification information is stored in the database, obtaining characteristics of the user from account data corresponding to the second identification information; 10. The system of claim 9, wherein if the second identification information is not stored in the database or if the user characteristics are not included in the account data corresponding to the second identification information, the user characteristics are obtained from the first identification information.

12. The electronic device transmits a user list including at least one user identification information to the server; The server searching the database for third identification information corresponding to the user identification information included in the user list; 12. The system of claim 11, further comprising: comparing the first identification information with the third identification information to determine the second identification information.

13. The database stores a history of the use of text types by characteristics; The server determining the second text type based on a first type history in which the second text type was used in association with the user's characteristic among the history of using the text type according to the characteristic; The system of claim 11 further comprising: generating a result of performing an intent analysis on the second text based on the determined type.

14. the account data for the user account includes a history of the user's use of text types; The server if the account data corresponding to the second identification information includes a second typing history using the second text type, determining the second text type based on the second typing history; 14. The system of claim 13, wherein if the second typing history is not included in the account data corresponding to the second identification information, the system determines the type of the second text based on the first typing history.

15. the server transmits at least one type corresponding to the second text to the electronic device; The electronic device is outputting via a display a type object corresponding to at least one type corresponding to the second text; The system of claim 13, further comprising: requesting the server to perform an intention analysis for a type corresponding to the selected type object based on a user input for selecting the type object.

Citation Information

Patent Citations

  • Voice processor

    JP1996054250A

  • Speech chat system, information processor, speech recognition method and program

    JP2008287210A

  • Content-retrieving device using voice recognition processing function, program, and method

    JP2010085522A

  • Generation apparatus, generation method and generation program

    JP2018194902A

  • server

    JP2023024713A