Electronic device and server
The system addresses the inconvenience and security issues of manual account entry by using voice recognition to authenticate and personalize user experiences in electronic devices, enhancing convenience and security.
Patent Information
- Application Number
- PCT/KR2024/010498
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-22
AI Technical Summary
Existing voice recognition technologies in electronic devices require manual entry of account information for user authentication, leading to inconvenience and security risks, especially when multiple users share a device.
An electronic device and server system that registers and identifies users based on their voice, allowing for automatic account login and personalized profile optimization using voice recognition technology.
Enables secure and convenient user authentication through voice recognition, eliminating the need for manual account entry and providing personalized responses based on user characteristics.
Smart Images

Figure KR2024010498_22012026_PF_FP_ABST
Abstract
Description
Electronic devices and servers
[0001] The present disclosure relates to electronic devices and servers, and more particularly, to electronic devices and servers utilizing voice recognition technology.
[0002] With recent technological advancements, research on voice recognition technology, which processes speech, is actively underway. In particular, research on voice recognition technology, which originated with smartphones, is now being conducted widely in various fields related to user convenience, such as home appliances used in homes and offices, as well as in vehicles.
[0003] Voice recognition technology is commonly used when a user controls an electronic device using their voice. For example, when a user utters a command to control an electronic device, the electronic device can directly recognize and process the user's voice and operate according to the corresponding command. Alternatively, the device can transmit the voice to a voice processing server and then operate according to the corresponding command received from the server.
[0004] Meanwhile, the services and functions provided through electronic devices are becoming increasingly diverse. Furthermore, users register accounts for various services and then log in with their registered accounts to access them. In this case, service providers utilize the user information managed for each account to provide optimal features or information tailored to the user.
[0005] Traditionally, when attempting to log in to a service, users would manually enter their account information, such as their account ID (identification number) and / or password. However, this presents a significant inconvenience, requiring users to manually enter their account information for each service. Furthermore, maintaining a logged-in status to eliminate the inconvenience of entering account information can lead to security issues, such as third parties accessing the user's account information. Furthermore, when multiple users share a single electronic device, this presents a problem: each user must enter their own account information to log in each time they use the service.
[0006] The present disclosure aims to solve the above-mentioned and other problems.
[0007] Another purpose is to provide an electronic device and server that can register identifying information about a user's voice to the user's account.
[0008] Another purpose is to provide an electronic device and server capable of identifying a user based on the user's voice.
[0009] Another object is to provide an electronic device and server capable of providing a profile image for a user account optimized for the user based on the user's speech history.
[0010] Another object is to provide an electronic device and server capable of providing a response in response to user characteristics identified from voice.
[0011] In order to achieve the above object, an electronic device according to an embodiment of the present disclosure includes a user input interface unit that transmits a signal corresponding to an input; a memory; and a control unit, wherein, when a voice signal is received through the user input interface unit, the control unit can identify a user account corresponding to the voice signal, add a keyword corresponding to the voice signal included in a result of performing intent analysis on the voice signal to data for the user account stored in the memory, and update a profile image for the user account based on the keyword included in the data for the user account.
[0012] In order to achieve the above object, a server according to one embodiment of the present disclosure includes a communication unit that communicates with an electronic device; a database; and a controller, wherein, when a voice signal is received through the communication unit, the controller can identify a user account corresponding to the voice signal, add a keyword corresponding to the voice signal included in a result of performing intent analysis on the voice signal to data for the user account stored in the database, and update a profile image for the user account based on the keyword included in the data for the user account.
[0013] The effects of the electronic device and server according to the present disclosure are described as follows.
[0014] According to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.
[0015] According to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.
[0016] According to at least one embodiment of the present disclosure, a profile image for a user account optimized for the user may be provided based on the user's speech history.
[0017] According to at least one embodiment of the present disclosure, a response can be provided in response to a characteristic of a user identified from voice.
[0018] Further scope of the applicability of the present disclosure will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present disclosure will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present disclosure, are given by way of example only.
[0019] FIG. 1 is a diagram illustrating a system according to one embodiment of the present disclosure.
[0020] Figure 2 is an internal block diagram of the electronic device of Figure 1.
[0021] Figure 3 is a drawing referenced in the description of the server of Figure 1.
[0022] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.
[0023] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0024] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an electronic device according to one embodiment of the present disclosure.
[0025] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0026] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.
[0027] FIGS. 9 to 13 are drawings for reference in explaining a process for registering identification information for a user's voice to a user's account according to one embodiment of the present disclosure.
[0028] FIG. 14 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0029] FIG. 15 is a flowchart of a method of operating a server according to one embodiment of the present disclosure.
[0030] FIGS. 16 to 23 are drawings for reference in explaining the operation of an electronic device according to embodiments of the present disclosure.
[0031] Hereinafter, the present disclosure will be described in detail with reference to the drawings. In the drawings, portions irrelevant to the description are omitted to clearly and concisely describe the present disclosure, and the same reference numerals are used for identical or extremely similar portions throughout the specification.
[0032] The suffixes "module" and "part" used in the following description are given solely for the convenience of writing this specification and do not impart any particularly significant meaning or role to the components themselves. Therefore, the terms "module" and "part" may be used interchangeably.
[0033] In this application, it should be understood that terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0034] Additionally, while terms such as "first" and "second" may be used in this specification to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another.
[0035] FIG. 1 is a diagram illustrating a system according to various embodiments of the present invention.
[0036] Referring to FIG. 1, the system (10) may include an electronic device (100) and / or a server (400).
[0037] The electronic device (100) can transmit and receive data to and from at least one server (400). For example, the electronic device (100) can transmit and receive data to and from at least one server (400) via a network (300) such as the Internet.
[0038] According to one embodiment, at least one server (400) may include a server that performs voice recognition, a server that processes data using a super-giant artificial intelligence model, a server that provides content, etc.
[0039] The electronic device (100) may include a video display device (100a), an air conditioner (100b), a refrigerator (100c), an air purifier (100d), a washing machine (100e), a vehicle (100f), etc. In the present disclosure, the electronic device (100) is described as an example of a video display device (100a), but the present invention is not limited thereto.
[0040] The image display device (100a) may be a device that processes and outputs an image. The image display device (100a) is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.
[0041] The video display device (100a) can receive a broadcast signal, process it, and output the processed broadcast image. When the video display device (100a) receives a broadcast signal, the video display device (100a) may correspond to a broadcast receiving device.
[0042] The video display device (100a) can receive broadcast signals wirelessly via an antenna, or can receive broadcast signals wired via a cable. For example, the video display device (100a) can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.
[0043] Figure 2 is an internal block diagram of the electronic device of Figure 1.
[0044] Referring to FIG. 2, the electronic device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).
[0045] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulator unit (120).
[0046] Meanwhile, unlike the drawing, the electronic device (100) may include only the broadcast receiving unit (105) and the external device interface unit (130) among the broadcast receiving unit (105), the external device interface unit (130), and the network interface unit (135). That is, the electronic device (100) may not include the network interface unit (135).
[0047] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.
[0048] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).
[0049] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.
[0050] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.
[0051] The demodulation unit (120) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (110).
[0052] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.
[0053] The stream signal output from the demodulation unit (120) can be input to the control unit (170). The control unit (170) can output an image through the display (180) and output an audio through the audio output unit (185) after performing demultiplexing, image / audio signal processing, etc.
[0054] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).
[0055] The external device interface unit (130) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.
[0056] In addition, the external device interface unit (130) can establish a communication network with various remote control devices (200) to receive a control signal related to the operation of the electronic device (100) from the remote control device (200) or transmit data related to the operation of the electronic device (100) to the remote control device (200).
[0057] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit can include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).
[0058] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, in mirroring mode, the external device interface unit (130) may receive device information, running application information, application images, etc. from the mobile terminal.
[0059] The external device interface unit (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.
[0060] The network interface unit (135) can provide an interface for connecting the electronic device (100) to a wired / wireless network including the Internet.
[0061] The network interface unit (135) may include a communication module (not shown) for connection to a wired / wireless network. For example, the network interface unit (135) may include a communication module for WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), etc.
[0062] The network interface unit (135) can transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.
[0063] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and information related thereto provided by a content provider or network provider via a network.
[0064] The network interface unit (135) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.
[0065] The network interface unit (135) can select and receive a desired application from among applications open to the public through a network.
[0066] The storage unit (140) may store programs for signal processing and control within the control unit (170), or may store processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored application programs upon request from the control unit (170).
[0067] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (170).
[0068] The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).
[0069] The storage unit (140) can store information about a specific broadcast channel through a channel memory function such as a channel map.
[0070] Although the storage unit (140) of FIG. 2 is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).
[0071] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.
[0072] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.
[0073] For example, a user input signal such as power on / off, channel selection, screen setting, etc. may be transmitted / received from a remote control device (200), a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. may be transmitted to the control unit (170), a user input signal input from a sensor unit (not shown) that senses a user's gesture may be transmitted to the control unit (170), or a signal from the control unit (170) may be transmitted to the sensor unit.
[0074] The input unit (160) may be provided on one side of the main body of the electronic device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.
[0075] The input unit (160) can receive various user commands related to the operation of the electronic device (100) and transmit a control signal corresponding to the input command to the control unit (170).
[0076] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.
[0077] The control unit (170) may include at least one processor, and may control the overall operation of the electronic device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.
[0078] The control unit (170) can demultiplex a stream input through the tuner unit (110), the demodulator unit (120), the external device interface unit (130), or the network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.
[0079] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).
[0080] The display (180) may include a display panel (not shown) having a plurality of pixels.
[0081] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display (180) may convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (170) to generate driving signals for the plurality of pixels.
[0082] The display (180) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display, etc., and may also be a 3D display. The 3D display (180) can be divided into a glasses-free type and a glasses type.
[0083] Meanwhile, the display (180) is configured as a touch screen and can be used as an input device in addition to an output device.
[0084] The audio output unit (185) receives a signal processed by the control unit (170) and outputs it as voice.
[0085] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (170) can also be input to an external output device through the external device interface unit (130).
[0086] The voice signal processed in the control unit (170) can be output as sound to the audio output unit (185). In addition, the voice signal processed in the control unit (170) can be input to an external output device through the external device interface unit (130).
[0087] Although not shown in FIG. 2, the control unit (170) may include a demultiplexing unit, an image processing unit, etc.
[0088] In addition, the control unit (170) can control the overall operation within the electronic device (100). For example, the control unit (170) can control the tuner unit (110) to select (tune) a broadcast corresponding to a channel selected by the user or a previously stored channel.
[0089] In addition, the control unit (170) can control the electronic device (100) by a user command or internal program input through the user input interface unit (150).
[0090] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a moving image, and may be a 2D image or a 3D image.
[0091] Meanwhile, the control unit (170) can cause a predetermined 2D object to be displayed within an image displayed on the display (180). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.
[0092] Meanwhile, the electronic device (100) may further include a camera (not shown). The camera can capture images of the user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the electronic device (100) above the display (180) or can be separately positioned. Image information captured by the camera can be input to the control unit (170).
[0093] The control unit (170) can recognize the user's location based on the image captured by the camera. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the electronic device (100). In addition, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.
[0094] The control unit (170) can detect the user's gesture based on an image captured from the camera unit, a signal detected from the sensor unit, or a combination thereof.
[0095] The power supply unit (190) can supply power to the entire electronic device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a system on chip (SOC), a display (180) for displaying images, and an audio output unit (185) for outputting audio.
[0096] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.
[0097] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (150) and display or output the same as voice on the remote control device (200).
[0098] Meanwhile, the electronic device (100) described above may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.
[0099] Meanwhile, the block diagram of the electronic device (100) illustrated in FIG. 2 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the electronic device (100) actually implemented.
[0100] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.
[0101] Figure 3 is a drawing referenced in the description of the server of Figure 1.
[0102] Referring to FIG. 3, the server (400) may include a relay server (410), an STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), a user identification server (440), and / or an account server (450). In the present disclosure, the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) are described as being distinct from each other, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), the user identification server (440), and the account server (450) may be configured as one server.
[0103] The relay server (410) can communicate with the electronic device (100). The relay server (410) can transfer data between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100). The relay server (410) can store at least a portion of the data transferred between the STT server (420), the NLP server (430), the user identification server (440), and the electronic device (100).
[0104] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the electronic device (100) via the relay server (410). The STT server (420) may also be referred to as an ASR (Automatic Speech Recognition) server.
[0105] The STT server (420) can improve the accuracy of speech-to-text conversion using a language model. The language model can refer to a model that can calculate the probability of a sentence or the probability of a subsequent word given previous words. For example, the language model can include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. In other words, the STT server (420) can use the language model to determine whether text data converted from speech data has been appropriately converted, thereby improving the accuracy of conversion into text data.
[0106] The NLP server (430) can receive text data. Based on the received text data, the NLP server (430) can perform intent analysis on the text data. The NLP server (430) can transmit intent analysis information indicating the results of the intent analysis to the electronic device (100) via the relay server (410).
[0107] According to one embodiment, the NLP server (430) may sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate intent analysis information. The morphological analysis step is a step of classifying text data corresponding to speech uttered by a user into morphemes, which are the smallest units having meaning, and determining which part of speech each classified morpheme has. The syntax analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and to determine what kind of relationship exists between each of the classified phrases. Through the syntax analysis step, the subject, object, and modifiers of the speech uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the speech uttered by the user using the results of the syntax analysis step. Specifically, the speech act analysis step is a step of determining the intent of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.
[0108] The user identification server (440) can receive voice data. Based on the voice data, the user identification server (440) can extract voice features. Here, the voice features may include the voice waveform, the voice frequency band, the voice power spectrum, etc. Extracting voice features will be described below with reference to FIGS. 4 and 5.
[0109] The user identification server (440) can obtain a feature vector of a voice from the features of the voice. The user identification server (440) can obtain a feature vector of a voice from the features of the voice based on a linear predictive coefficient, a cepstrum, a Mel Frequency Cepstral Coefficient (MFCC), and filter bank energy.
[0110] The user identification server (440) can determine the similarity between a plurality of feature vectors. The user identification server (440) can determine the similarity between a plurality of feature vectors using cosine similarity, Euclidean similarity, etc. In the present disclosure, the similarity between a first voice input and a second voice input is calculated based on cosine similarity, but is not limited thereto. For example, a first vector corresponding to a first text and a second vector corresponding to a second text may be generated. In this case, the cosine similarity between the first vector and the second vector may be calculated based on the following mathematical expression 1.
[0111]
[0112] Here, is the dot product of two vectors, and can mean the magnitude of two vectors. That is, cosine similarity can be calculated as the value obtained by dividing the inner product of two vectors by the product of the magnitudes of each vector. Cosine similarity can range from -1 to 1, and the closer it is to 1, the more similar the two vectors can be judged to be.
[0113] The user identification server (440) can determine whether the users who uttered the voice are the same user based on the similarity between multiple feature vectors. For example, if the similarity between the first feature vector corresponding to the first voice input and the second feature vector corresponding to the second voice input is above a predetermined standard, the user identification server (440) can determine that the user who uttered the first voice input and the user who uttered the second voice input are the same user.
[0114] According to one embodiment, the user identification server (440) may obtain a vector by processing a feature vector of a voice using an algorithm such as a GMM (Gaussian mixture model) supervector, i-vector, d-vector, x-vector, etc. The user identification server (440) may determine whether the user who uttered the voice is the same based on the similarity between a first vector processed corresponding to a first feature vector and a second vector processed corresponding to a second feature vector.
[0115] The user identification server (440) can store voice data. The user identification server (440) can store data regarding the voice print (hereinafter, "voice print information"). Here, the voice print information can include a voice feature vector and / or a vector obtained by processing the voice feature vector.
[0116] The user identification server (440) can store a database for voices. The database for voices can include unique identification information (hereinafter, device identification information) corresponding to an electronic device (100), unique identification information (hereinafter, user identification information) corresponding to a user account, voice data mapped to user identification information, voice print information mapped to user identification information, etc.
[0117] Device identification information, user identification information, voice data, voice print information, etc. included in the voice database may be stored in a user identification server (440) in association with each other. For example, at least one piece of device identification information, multiple voice data, and / or multiple voice print information may be mapped to the user identification information. In other words, it may be interpreted that the device identification information, voice data, voice print information, etc. are mapped to a user account and stored in the user identification server (440). In the present disclosure, an example in which multiple voice data and multiple voice print information are all mapped to the user identification information included in the voice database will be described.
[0118] The user identification server (440) can update the voiceprint information contained in the voice database based on the voice data contained in the voice database. For example, the user identification server (440) can generate voiceprint information corresponding to the voice data contained in the voice database using an algorithm different from the algorithm previously used. In this case, the user identification server (440) can change the voiceprint information contained in the voice database to the newly generated voiceprint information.
[0119] The account server (450) can manage data regarding user accounts. The account server (450) can manage user account IDs, passwords, user identification information, device identification information mapped to the user account, and whether or not to agree to terms and conditions related to various functions.
[0120] The account server (450) can store a database of user accounts. The database of user accounts may include the user account ID, password, user identification information, device identification information mapped to the user account, the registration date and time of the user account, whether or not the user has agreed to terms and conditions related to various functions, and the date and time of agreement to the terms and conditions.
[0121] The account server (450) can communicate with the electronic device (100). For example, the account server (450) can create and register a user account based on data from the electronic device (100). For example, the account server (450) can approve login to a user account based on an ID and password received from the electronic device (100).
[0122] FIG. 4 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.
[0123] Referring to FIG. 4, the server (400) may include a preprocessing unit (460), a controller (470), a communication unit (480), and / or a database (490).
[0124] The preprocessing unit (460) can preprocess voice received through the communication unit (480) or voice stored in the database (490).
[0125] The preprocessing unit (460) may be implemented as a separate chip from the controller (470) or as a chip included in the controller (470).
[0126] The preprocessing unit (460) can receive a voice signal (spoken by a user) and filter out noise signals from the voice signal before converting the received voice signal into text data.
[0127] When a preprocessing unit (460) is provided in the electronic device (100), it can recognize a trigger word for activating voice recognition of the electronic device (100). The preprocessing unit (460) converts the trigger word received through the user input interface unit (150) into text data, and if the converted text data corresponds to a previously stored trigger word, it can be determined that the trigger word has been recognized.
[0128] The preprocessing unit (460) can convert the noise-removed voice signal into a power spectrum.
[0129] Power spectrum can be a parameter that indicates which frequency components are included in the waveform of a temporally varying voice signal and at what magnitude.
[0130] The power spectrum shows the distribution of the squared amplitude values of the waveform of a voice signal according to frequency. This is explained with reference to Figure 5.
[0131] FIG. 5 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0132] Referring to FIG. 5, a voice signal (510) is illustrated. The voice signal (460) may be received from an external device or may be a signal pre-stored in memory (170).
[0133] The x-axis of the voice signal (510) represents time, and the y-axis can represent the amplitude size.
[0134] The power spectrum processing unit (463) can convert a voice signal (510) whose x-axis is the time axis into a power spectrum (520) whose x-axis is the frequency axis. The power spectrum processing unit (463) can convert the voice signal (510) into a power spectrum (520) using a fast Fourier transform (FFT). The x-axis of the power spectrum (520) represents frequency, and the y-axis represents the square value of the amplitude.
[0135] Referring again to FIG. 4, the functions of the preprocessing unit (460) and the controller (470) described in FIG. 4 can also be performed in the NLP server (430).
[0136] The preprocessing unit (460) may include a wave processing unit (461), a frequency processing unit (462), a power spectrum processing unit (463), a speech to text (STT) conversion unit (464), etc.
[0137] The wave processing unit (461) can extract the waveform of the voice.
[0138] The frequency processing unit (462) can extract the frequency band of the voice.
[0139] The power spectrum processing unit (463) can extract the power spectrum of the voice.
[0140] Power spectrum can be a parameter that indicates which frequency components are included in a given temporally varying waveform and at what magnitude.
[0141] The speech-to-text (STT) conversion unit (464) can convert speech into text. The speech-to-text conversion unit (464) can convert speech in a specific language into text in that language.
[0142] The controller (470) can control the overall operation of the server (400). The controller (470) can include a voice analysis unit (471), a text analysis unit (472), a feature clustering unit (473), a text mapping unit (474), and / or a voice synthesis unit (475).
[0143] The voice analysis unit (471) can extract voice characteristic information by using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed by the preprocessing unit (460). The voice characteristic information may include one or more of the speaker's gender information, the speaker's voice (or tone), pitch, the speaker's speech style, the speaker's speaking speed, and the speaker's emotion. In addition, the voice characteristic information may further include the speaker's voice tone.
[0144] The text analysis unit (472) can extract key phrases from the text converted by the voice-to-text conversion unit (464). If the text analysis unit (472) detects a difference in tone between phrases in the converted text, it can extract the phrases with different tones as key phrases. If the text analysis unit (472) changes the frequency band between phrases by more than a preset band, it can determine that the tone has changed. The text analysis unit (472) can also extract key words within the phrases of the converted text. Key words may be nouns present within the phrases, but this is merely an example.
[0145] The feature clustering unit (473) can classify the speaker's speech type using the characteristic information of the voice extracted from the voice analysis unit (471). The feature clustering unit (473) can classify the speaker's speech type by assigning weights to each of the type items constituting the characteristic information of the voice. The feature clustering unit (473) can classify the speaker's speech type using the attention technique of the deep learning model.
[0146] The text mapping unit (474) can translate a text converted into a first language into a text in a second language. The text mapping unit (474) can map a text translated into a second language to a text in the first language. The text mapping unit (474) can map the main expressions constituting the text in the first language to the corresponding expressions in the second language. The text mapping unit (474) can map the utterance types corresponding to the main expressions constituting the text in the first language to the expressions in the second language. This is to apply the classified utterance types to the expressions in the second language.
[0147] The voice synthesis unit (475) can generate a synthesized voice by applying the speech type and the speaker's tone classified by the feature clustering unit (473) to the main expression phrases of the text translated into a second language by the text mapping unit (474).
[0148] The controller (470) can determine the user's speech characteristics using one or more of the transmitted text data or power spectrum (520).
[0149] User speech characteristics may include the user's gender, the user's pitch, the user's timbre, the user's speech topic, the user's speech rate, the user's voice volume, etc.
[0150] The controller (470) can obtain the frequency of the voice signal (510) and the amplitude corresponding to the frequency by using the power spectrum (520).
[0151] The controller (470) can determine the gender of the user who has spoken using the frequency band of the power spectrum (470). For example, the controller (470) can determine the gender of the user as male if the frequency band of the power spectrum (520) is within a preset first frequency band range.
[0152] The controller (470) can determine the user's gender as female if the frequency band of the power spectrum (520) is within a preset second frequency band range. Here, the second frequency band range may be larger than the first frequency band range.
[0153] The controller (470) can determine the pitch of a sound by using the frequency band of the power spectrum (520). For example, the controller (470) can determine the pitch of a sound based on the amplitude within a specific frequency band range.
[0154] The controller (470) can determine the user's tone by using the frequency band of the power spectrum (520). For example, the controller (470) can determine a frequency band with an amplitude greater than a certain size among the frequency bands of the power spectrum (520) as the user's main vocal range, and determine the determined main vocal range as the user's tone.
[0155] The controller (470) can determine the user's speaking speed through the number of syllables spoken per unit time from the converted text data.
[0156] For the converted text data of the controller (470), the user's speech topic can be determined using the Bag-Of-Word Model technique.
[0157] The Bag-Of-Word Model technique extracts frequently used words based on their frequency within a sentence. Specifically, the Bag-Of-Word Model technique extracts unique words from a sentence, expresses the frequency of each extracted word as a vector, and determines the characteristics of the utterance topic. For example, if words like "running" and "physical strength" frequently appear in the controller (470) text data, the user's utterance topic can be classified as exercise.
[0158] The controller (470) can determine the topic of a user's speech from text data using a known text categorization technique. The controller (470) can extract keywords from the text data to determine the topic of the user's speech.
[0159] The controller (470) can determine the user's vocal volume by considering amplitude information across the entire frequency band. For example, the controller (470) can determine the user's vocal volume based on the average or weighted average of the amplitudes across each frequency band of the power spectrum.
[0160] The communication unit (480) can communicate with an external server via wired or wireless communication. The communication unit (480) can communicate with an electronic device (100) via wired or wireless communication.
[0161] The database (490) can store speech in a first language included in the content. The database (490) can store a synthesized speech in which the speech in the first language is converted into speech in a second language. The database (490) can store a first text corresponding to the speech in the first language and a second text in which the first text is translated into a second language. The database (490) can store various learning models required for speech recognition.
[0162] Meanwhile, the control unit (170) of the electronic device (100) illustrated in FIG. 2 may be equipped with the preprocessing unit (460) and the controller (470) illustrated in FIG. 4. That is, the control unit (170) of the electronic device (100) may perform the functions of the preprocessing unit (460) and the controller (470).
[0163] FIG. 6 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an electronic device according to one embodiment of the present disclosure.
[0164] That is, the voice recognition and synthesis process of FIG. 6 may be performed by the control unit (170) of the electronic device (100) without going through the server.
[0165] Referring to FIG. 6, the control unit (170) of the electronic device (100) may include an STT engine (610), an NLP engine (620), and a voice synthesis engine (630). Each engine may be either hardware or software.
[0166] The STT engine (610) can perform the function of the STT server (420) of FIG. 5. That is, the STT engine (610) can convert voice data into text data.
[0167] The NLP engine (620) can perform the function of the NLP server (430) of FIG. 5. That is, the NLP engine (620) can obtain intent analysis information indicating the speaker's intent from converted text data.
[0168] The speech synthesis engine (630) can perform the functions of a speech synthesis server. The speech synthesis engine (630) can search for syllables or words corresponding to given text data from a database, synthesize a combination of the searched syllables or words, and generate a synthesized speech.
[0169] The voice synthesis engine (630) may include a preprocessing engine (631) and a TTS engine (632).
[0170] The preprocessing engine (631) can preprocess text data before generating synthetic speech. Specifically, the preprocessing engine (631) performs tokenization, which divides the text data into tokens, which are meaningful units. After tokenization, the preprocessing engine (631) can perform a cleansing operation to remove unnecessary characters and symbols to remove noise. Thereafter, the preprocessing engine (631) can integrate word tokens with different expression methods to generate the same word token. Thereafter, the preprocessing engine (631) can remove meaningless word tokens (stopwords).
[0171] The TTS engine (632) can synthesize voice corresponding to preprocessed text data and generate a synthesized voice.
[0172] FIG. 7 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0173] Referring to FIG. 7, the electronic device (100) can determine, in operation S701, whether a user account is logged into the server (400). For example, a user can log into the server (400) with a user account by entering the ID and password of the user account.
[0174] According to one embodiment, when a user first logs into a server (400) using an electronic device (100), the electronic device (100) may include user identification information corresponding to the user account in a user list. For example, when three different user accounts log into the server (400) using the electronic device (100), the user list stored in the electronic device (100) may include three different user identification information.
[0175] In operation S702, the electronic device (100) can determine whether voice-related identification information (hereinafter, "Voice ID") is registered for a user account logged into the server (400). Here, the Voice ID may include voiceprint information stored in the user identification server (440). For example, the server (400) may transmit to the electronic device (100) whether a voice ID is registered for a user account logged into the server (400).
[0176] According to one embodiment, the server (400) can determine whether a voice ID is registered based on whether voice print information is mapped to user identification information, which is unique identification information corresponding to a user account logged into the server (400). In this case, if the voice ID corresponds to an unregistered user account, the number of voice print information mapped to the user identification information may be 0.
[0177] According to one embodiment, the server (400) may determine that a voice ID is registered if the number of voice prints mapped to the user identification information is two or more, a predetermined number, and may determine that the voice ID is not registered if the number is less than the predetermined number. For example, in the case of a user account with a registered voice ID, six different voice prints may be mapped to the user identification information. For example, in the case of a user account with an unregistered voice ID, five or fewer voice prints may be mapped to the user identification information.
[0178] According to one embodiment, a flag value indicating whether a voice ID is registered may be mapped to user identification information stored in the server (400). Here, the user identification information to which the flag value is mapped may be stored in the user identification server (440) and / or the account server (450). The server (400) may determine whether a voice ID is registered based on the flag value mapped to the user identification information. For example, if the voice ID is an unregistered user account, the flag value mapped to the user identification information may be 0, and if the voice ID is a registered user account, the flag value mapped to the user identification information may be 1.
[0179] In operation S703, if a voice ID has not been registered for a user account, the electronic device (100) may initiate a process for registering a voice ID. For example, when the electronic device (100) initiates a process for registering a voice ID, it may transmit data including device identification information, user identification information, a value indicating the start of voice ID registration, etc., to the server (400).
[0180] The electronic device (100) can output preset text in operation S704. The electronic device (100) can output any one of a plurality of preset texts. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) can output the preset text through the display (180).
[0181] According to one embodiment, the server (400) may transmit any one of a plurality of preset texts to the electronic device (100) in a preset order. At this time, the electronic device (100) may output the preset text received from the server (400).
[0182] The electronic device (100) can determine, in operation S705, whether a voice is input for a preset text. For example, the electronic device (100) can determine whether a voice is input through a microphone included in the input unit (160) within a preset time period. At this time, a voice signal corresponding to the voice input through the microphone can be transmitted to the control unit (170) through the user input interface unit (150). For example, the electronic device (100) can determine whether data including a voice signal corresponding to a voice spoken by a user is received from the remote control device (200) within a preset time period.
[0183] In operation S706, when a voice is input for a preset text, the electronic device (100) can transmit voice data containing a voice signal corresponding to the voice to the server (400). At this time, the electronic device (100) can transmit device identification information, user identification information, a language code indicating the type of language, etc., together with the voice data to the server (400).
[0184] The server (400) can convert a voice signal included in voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponds to a preset text. For example, the server (400) can determine whether the two texts correspond based on the similarity between the text converted from the voice signal and the preset text.
[0185] The server (400) can generate voiceprint information corresponding to the voice signal when the text converted from the voice signal corresponds to a preset text. The server (400) can map the voiceprint information generated for the preset text to user identification information and store it. The server (400) can map the voice data received for the preset text to user identification information and store it.
[0186] The electronic device (100), in operation S707, can determine whether the processing of the voice for the preset text was successful based on the response received from the server (400). For example, if the text converted from the voice signal and the preset text correspond to each other, the server (400) can notify the electronic device (100) of the success of the processing of the voice. For example, if the voiceprint information corresponding to the voice signal is generated, the server (400) can notify the electronic device (100) of the success of the processing of the voice.
[0187] Meanwhile, in operation S708, if voice input for a preset text is not performed, or if voice processing for a preset text fails, the electronic device (100) may determine whether to retry voice input. For example, the electronic device (100) may retry voice input based on a user input requesting retry voice input. In this case, the electronic device (100) may output the preset text again.
[0188] In operation S709, if the voice processing for a preset text is successful, the electronic device (100) can determine whether the processing for all texts is complete. For example, if the voice processing for a predetermined number of six preset texts is successful, the processing for all texts can be completed. Meanwhile, if the processing for five preset texts is complete, the electronic device (100) can output the last preset text.
[0189] The electronic device (100) may complete the process of registering a voice ID when processing of all texts is completed in operation S710. For example, if the electronic device (100) is a video display device (100a), the electronic device (100) may output a screen indicating completion of voice ID registration through the display (180). For example, the electronic device (100) may transmit data indicating completion of voice ID registration to the account server (450).
[0190] FIG. 8 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.
[0191] Referring to FIG. 8, the electronic device (100) can perform a login to the server (400) using a user account in operation S801.
[0192] The electronic device (100) can initiate a process of registering a voice ID in operation S802.
[0193] The electronic device (100) can output a first text among a plurality of preset texts in operation S803.
[0194] The electronic device (100) can receive a first voice for a first text in operation S804.
[0195] The electronic device (100) can transmit first voice data including a voice signal corresponding to the first voice to the server (400) in operation S805.
[0196] The server (400), in operation S806, can process the first voice for the first text based on the first voice data received from the electronic device (100). The server (400) can convert a voice signal corresponding to the first voice included in the first voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the first voice and the first text correspond to each other.
[0197] The server (400) may, in operation S807, notify the electronic device (100) of the completion of processing for the first voice. For example, the server (400) may notify the electronic device (100) of the success of processing for the first voice based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.
[0198] Meanwhile, the server (400) can generate first voiceprint information for the first voice based on the voice signal corresponding to the first voice, based on the correspondence between the text converted from the voice signal corresponding to the first voice and the first text.
[0199] The server (400) may, in operation S808, store first voice data and first voiceprint information for the first voice. The server (400) may store the first voice data and the first voiceprint information by mapping them to user identification information corresponding to a logged-in user account.
[0200] The electronic device (100) can sequentially output the second text to the fifth text. The electronic device (100) can sequentially receive the second voice to the fifth voice, which correspond to the second text to the fifth text, respectively. The electronic device (100) can sequentially transmit the second voice data to the fifth voice data, which correspond to the second voice to the fifth voice, respectively, to the server (400).
[0201] The server (400) can process the second to fifth voices based on the second to fifth voice data received from the electronic device (100), respectively. In addition, the server (400) can sequentially generate and store second to fifth voice information corresponding to the second to fifth voices, respectively.
[0202] The electronic device (100) can output a sixth text among a plurality of preset texts in operation S809.
[0203] The electronic device (100) can receive a sixth voice for a sixth text in operation S810.
[0204] The electronic device (100) can transmit sixth voice data including a voice signal corresponding to the sixth voice to the server (400) in operation S811.
[0205] The server (400) can process the sixth voice for the sixth text based on the sixth voice data received from the electronic device (100) in operation S812. The server (400) can convert a voice signal corresponding to the sixth voice included in the sixth voice data received from the electronic device (100) into text. The server (400) can determine whether the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.
[0206] The server (400) can, in operation S813, notify the electronic device (100) of the completion of processing for the sixth voice.
[0207] Meanwhile, the server (400) can generate sixth voice information for the sixth voice based on the voice signal corresponding to the sixth voice, if the text converted from the voice signal corresponding to the sixth voice and the sixth text correspond to each other.
[0208] The server (400) may store the sixth voice data and the sixth voiceprint information for the sixth voice in operation S814. The server (400) may store the sixth voice data and the sixth voiceprint information by mapping them to user identification information corresponding to the logged-in user account. At this time, the user identification information corresponding to the logged-in user account may be mapped with six different voice data and multiple voiceprint information.
[0209] The electronic device (100) may complete the process of registering a voice ID in operation S815. For example, the electronic device (100) may complete the process of registering a voice ID based on the completion of processing of a predetermined number of six different preset texts.
[0210] Referring to FIG. 9, if a user account is not logged in to the server (400), the electronic device (100) may output a login screen (900) related to logging in to the server (400) through the display (180). The login screen (900) may include an object (910) indicating a non-logged-in state, a login object (920) for executing login, etc. When a user selects the login object (920) using a pointer (205) corresponding to the remote control device (200), the electronic device (100) may output a screen for inputting an ID and password. At this time, the user may input the ID and password of the user account to log in to the server (400) with the user account.
[0211] Referring to FIG. 10, if a voice ID is not registered in a user account logged in to the server (400), the electronic device (100) can output a first account screen (1000) corresponding to the user account with the unregistered voice ID. The first account screen (1000) can include an object (1010) representing a logged-in user account, an object (1020) corresponding to the registration of a voice ID, etc. When a user selects an object (1020) corresponding to the registration of a voice ID using a pointer (205), the electronic device (100) can initiate a process of registering a voice ID.
[0212] Meanwhile, referring to FIG. 11, if a voice ID is already registered in a user account logged in to the server (400), the electronic device (100) can output a second account screen (1100) corresponding to the user account in which the voice ID is registered. The second account screen (1100) can include an object (1110) representing the logged-in user account, a re-registration object (1120) corresponding to re-registration of the voice ID, a deletion object (1130) corresponding to deletion of the voice ID, an activation object (1140) corresponding to use of a function related to the voice ID, etc. The user can select the activation object (1140) using the pointer (205) to activate or deactivate use of the function related to the voice ID.
[0213] Referring to FIG. 12, when an object (1020) corresponding to the registration of a voice ID is selected on the first account screen (1000), or when a re-registration object (1120) is selected on the second account screen (1100), the electronic device (100) can output a start screen (1200) for initiating the registration of a voice ID. When a user selects a start object (1210) using a pointer (205), the electronic device (100) can output a text screen for outputting preset text.
[0214] Referring to FIG. 13, the electronic device (100) can output a text screen (1300) that outputs any one of a plurality of preset texts. The text screen (1300) can include preset text (1301), a text sequence (1302), a termination object (1310) that terminates a process of registering a voice ID, an input object (1320) that receives voice, etc.
[0215] When the user selects the end object (1310) using the pointer (205), the process of registering a voice ID may be terminated. For example, when the process of registering a voice ID is terminated, all data stored in the server (400) while the process of registering a voice ID is in progress may be deleted.
[0216] When a user selects an input object (1320) using a pointer (205), the electronic device (100) can receive voice for the text.
[0217] According to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a text screen (1300) is displayed, the electronic device (100) can receive a voice for the text based on the user input of pressing the predetermined button received from the remote control device (200).
[0218] Meanwhile, according to one embodiment, when a user presses a predetermined button (e.g., a voice input button) included in a remote control device (200) while a process for registering a voice ID is in progress, the electronic device (100) may stop the process for registering a voice ID based on a user input for pressing the predetermined button received from the remote control device (200). At this time, the user input for pressing the predetermined button (e.g., a voice input button) included in the remote control device (200) may correspond to a user input for initiating voice recognition for a voice received through the remote control device (200). The electronic device (100) may perform an operation related to voice recognition on voice data including a voice signal received from the remote control device (200).
[0219] FIG. 14 is a flowchart of an operation method of an electronic device according to one embodiment of the present disclosure.
[0220] Referring to FIG. 14, the electronic device (100) can receive a voice in operation S1401. For example, the electronic device (100) can receive a voice through a microphone included in the input unit (160). At this time, a voice signal corresponding to the voice input through the microphone can be transmitted to the control unit (170) through the user input interface unit (150). For example, the electronic device (100) can receive data including a voice signal corresponding to a voice spoken by a user from a remote control device (200).
[0221] In operation S1402, the electronic device (100) may transmit voice data corresponding to the input voice to the server (400) that performs voice recognition. At this time, the electronic device (100) may transmit device identification information, a user list, a language code indicating the type of language, etc. to the server (400) together with the voice data. For example, the electronic device (100) may transmit voice data including a voice signal of a preset unit such as a syllable or a word to the server (400). That is, when a user utters a sentence, the electronic device (100) may transmit a voice signal of a preset unit to the server (400) while a voice input corresponding to the sentence or phrase is received from the remote control device (200).
[0222] The electronic device (100) may, in operation S1403, receive the result of processing the voice from the server (400). For example, the result of processing the voice may include text corresponding to the voice, intent analysis information resulting from performing natural language processing on the voice, user identification information corresponding to the voice, etc.
[0223] The server (400) can generate voiceprint information for a voice input into the electronic device (100) based on voice data received from the electronic device (100). The server (400) can search a database for voiceprint information (hereinafter, candidate voiceprint information) corresponding to user identification information included in a user list received from the electronic device (100). The server (400) can determine whether the candidate voiceprint information and the generated voiceprint information correspond to each other. The server (400) can determine user identification information, to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped, among the candidate voiceprint information, as user identification information corresponding to the voice input into the electronic device (100). Meanwhile, if there is no candidate voiceprint information corresponding to the generated voiceprint information, the server (400) can determine that there is no user identification information corresponding to the voice input into the electronic device (100).
[0224] The server (400) can transmit the result of processing voice data received from the electronic device (100) to the electronic device (100). For example, the server (400) can transmit to the electronic device (100) text converted from voice data received from the electronic device (100), intent analysis information indicating the result of performing intent analysis on the converted text, user identification information corresponding to the voice, and whether or not user identification information corresponding to the voice exists.
[0225] According to one embodiment, the electronic device (100) can convert a voice signal corresponding to a voice into text through the STT engine (810) included in the control unit (170). The electronic device (100) can perform intent analysis on the text corresponding to the voice input converted by the STT engine (810) through the NLP engine (820) included in the control unit (170), thereby obtaining intent analysis information.
[0226] The electronic device (100) can determine, in operation S1404, whether a user account corresponding to the voice exists. For example, the electronic device (100) can confirm the user account corresponding to the voice based on user identification information corresponding to the voice received from the server (400).
[0227] According to one embodiment, the memory (140) of the electronic device (100) can store voiceprint information mapped to user identification information. For example, the electronic device (100) can receive voiceprint information mapped to user identification information from the server (400) and store the voiceprint information in the memory (140). The electronic device (100) can generate voiceprint information for a voice signal. The electronic device (100) can search for candidate voiceprint information in the memory (140). The electronic device (100) can determine whether the candidate voiceprint information and the generated voiceprint information correspond to each other. The electronic device (100) can determine user identification information, to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped, among the candidate voiceprint information, as user identification information of a user account corresponding to the voice.
[0228] In operation S1405, if a user account corresponding to the voice exists, the electronic device (100) can determine whether the input voice is a voice requiring user identification based on the result of processing the voice.
[0229] According to one embodiment, the electronic device (100) can determine whether the input voice is a voice requiring user identification based on the type of command corresponding to the voice. For example, the electronic device (100) can determine that the input voice is not a voice requiring user identification based on whether the input voice corresponds to a command for adjusting the settings of the electronic device (100) (e.g., volume, screen brightness, etc.). For example, the electronic device (100) can determine that the input voice is a voice requiring user identification based on whether the input voice corresponds to a command for providing a predetermined function of the electronic device (100) (e.g., content provision function, search function, etc.).
[0230] In operation S1406, the electronic device (100) may add a keyword corresponding to the input voice to data (hereinafter, “account data”) for a user account corresponding to the voice stored in the memory (140) based on the result of processing the voice. Here, the account data may include the ID of the user account, the user’s name, age, gender, family relationship, a profile image for the user account, the user’s usage history, keywords, etc. For example, if the input voice is “Turn on the baseball broadcast,” the electronic device (100) may determine “baseball” as a keyword corresponding to the input voice. For example, if the input voice is “Search for news on YouTube,” the electronic device (100) may determine “YouTube” and “news” as keywords corresponding to the input voice.
[0231] The electronic device (100) may determine whether to update the profile image in operation S1407. For example, the electronic device (100) may determine to update the profile image when a user input for updating the profile image is received through the user input interface (150). For example, the electronic device (100) may determine to update the profile image when a predetermined period corresponding to the update of the profile image has arrived.
[0232] According to one embodiment, the electronic device (100) may determine a priority for each of the multiple keywords included in the account data. For example, if the number of times a first keyword is stored is greater than that of a second keyword, the priority of the first keyword may be higher than that of the second keyword. For example, the priority of a first keyword set by a user as an important keyword may be higher than that of a second keyword that is not set as an important keyword.
[0233] The electronic device (100) may determine a predetermined number of keywords from among a plurality of keywords included in the account data as a keyword group based on priorities. For example, the electronic device (100) may determine five keywords from among the plurality of keywords in descending order of priority as a first keyword group. In this case, the electronic device (100) may determine to update the profile image if the first keyword group and the second keyword group previously determined based on priorities are different. In other words, the electronic device (100) may determine to update the profile image if a keyword belonging to a keyword group among the plurality of keywords included in the account data is changed.
[0234] The electronic device (100) may generate a profile image when updating the profile image in operation S1408. The electronic device (100) may include the generated profile image in the account data stored in the memory (140).
[0235] According to one embodiment, the electronic device (100) may generate a profile image using a learning model learned through machine learning stored in the memory (140). For example, the electronic device (100) may generate at least one image corresponding to a keyword group included in the account data using the learning model stored in the memory (140). At this time, the electronic device (100) may generate at least one image based at least on the user's characteristics (e.g., age, gender, family relationship, etc.) included in the account data. Machine learning may refer to a method in which a computer learns through data without a human directly instructing the computer with logic, thereby enabling the computer to solve problems. Deep learning refers to an artificial intelligence technology in which a computer can learn on its own like a human by teaching a computer a human's way of thinking based on artificial neural networks (ANNs). An artificial neural network (ANN) may be implemented in the form of software or in the form of hardware such as a chip. For example, an artificial neural network (ANN) can include various types of algorithms, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), and a deep belief network (DBN).
[0236] According to one embodiment, the electronic device (100) can generate a profile image using a super-large artificial intelligence algorithm model.
[0237] For example, the electronic device (100) may transmit a first prompt for the generation of an image corresponding to a keyword group included in the account data to a server utilizing a large-scale artificial intelligence algorithm model. At this time, the electronic device (100) may generate a profile image based on at least one first image included in the response to the first prompt received from the server utilizing the large-scale artificial intelligence algorithm model.
[0238] For example, the electronic device (100) may transmit a second prompt for the generation of an image corresponding to the user's characteristics (e.g., age, gender, family relationships, etc.) and keyword groups (e.g., soccer, basketball, baseball, entertainment content, cooking) included in the account data to a server using a super-large artificial intelligence algorithm model. At this time, the electronic device (100) may generate a profile image based on at least one second image included in the response to the second prompt received from the server using the super-large artificial intelligence algorithm model.
[0239] According to one embodiment, the electronic device (100) can generate a profile image based on the currently set profile image for the user account. This allows for continuity between the profile image before and after the update.
[0240] For example, if the currently set profile image for a user account is an image directly set by the user using image data, the electronic device (100) may generate a profile image corresponding to the currently set profile image for the user account and the keyword group. On the other hand, if the currently set profile image for the user account is not an image set by the user using image data, the electronic device (100) may generate a profile image corresponding to the keyword group.
[0241] For example, when the electronic device (100) updates a profile image according to a user input for updating a profile image received through the user input interface unit (150), the electronic device (100) may generate a profile image corresponding to a keyword group. Meanwhile, when the electronic device (100) updates a profile image according to at least one of a change in a keyword belonging to a keyword group and the arrival of a predetermined period corresponding to an update of the profile image, the electronic device (100) may generate a profile image corresponding to a currently set profile image for the keyword group and the user account.
[0242] According to one embodiment, the electronic device (100) may generate a profile image based on a family relationship among the user's characteristics included in the account data. For example, if the family relationship of the first user is 'son', the electronic device (100) may generate the profile image of the first user based on at least one of the profile image of the second user whose family relationship is 'father', the profile image of the third user whose family relationship is 'mother', and the profile image of the fourth user whose family relationship is 'brother'. Through this, the profile images of family members may have a relationship with each other.
[0243] According to one embodiment, the electronic device (100) may generate a plurality of candidate images based on keywords included in account data. The electronic device (100) may output the plurality of candidate images through the display (180). When a user input for selecting one of the plurality of candidate images is received through the user input interface (150), the electronic device (100) may update the profile image based on the image selected according to the user input.
[0244] Meanwhile, at least some of the operations of the electronic device (100) described in FIG. 14 may be performed in the server (400). In this regard, a description will be given with reference to FIG. 15.
[0245] Figure 15 is a flowchart illustrating a server operation method according to one embodiment of the present disclosure. Any details that overlap with those described in Figure 14 will be omitted for detailed explanation.
[0246] Referring to FIG. 15, the server (400) can receive voice data including a voice signal from the electronic device (100) in operation S1501.
[0247] In operation S1502, the server (400) can perform intent analysis on the voice spoken by the user and obtain the results of the voice processing. For example, the server (400) can convert the voice signal included in the voice data into text. At this time, the server (400) can obtain intent analysis information, which is the result of performing intent analysis on the converted text.
[0248] In operation S1503, the server (400) can generate voice information for a voice input into the electronic device (100) based on voice data received from the electronic device (100).
[0249] The server (400) can determine, in operation S1504, whether a user account corresponding to the voice exists. For example, the server (400) can search for candidate voiceprint information in a database for voices. The server (400) can determine, among the candidate voiceprint information, the user identification information to which the candidate voiceprint information corresponding to the generated voiceprint information is mapped, as the user identification information corresponding to the voice input into the electronic device (100). The server (400) can confirm the user account corresponding to the voice based on the user identification information corresponding to the voice.
[0250] In operation S1505, if a user account corresponding to the voice exists, the server (400) can determine whether the input voice is a voice requiring user identification based on the result of processing the voice.
[0251] In operation S1506, the server (400) can add a keyword corresponding to the input voice to data for a user account corresponding to the voice stored in the database (490) based on the result of processing the voice.
[0252] The server (400) may transmit the result of processing the voice to the electronic device (100) in operation S1507. For example, the electronic device (100) may output a response corresponding to the text included in the result of processing the voice received from the server (400) through the display (180). For example, the electronic device (100) may perform an operation corresponding to a command included in the result of processing the voice.
[0253] The server (400) may determine whether to update the profile image in operation S1508. For example, the server (400) may determine to update the profile image when a user input for updating the profile image is received from the electronic device (100). For example, the server (400) may determine to update the profile image when a predetermined period corresponding to the update of the profile image has arrived. For example, the server (400) may determine to update the profile image when a keyword belonging to a keyword group among multiple keywords included in account data stored in the database (490) is changed.
[0254] The server (400) can create a profile image when updating a profile image in operation S1509.
[0255] For example, the server (400) may transmit a predetermined prompt for the generation of an image corresponding to a keyword included in account data stored in the database (490) to an external server utilizing a large-scale artificial intelligence algorithm model. In this case, the server (400) may generate a profile image based on at least one image included in a response to the predetermined prompt received from the external server utilizing the large-scale artificial intelligence algorithm model.
[0256] For example, when the server (400) updates a profile image according to a user input for updating a profile image received from the electronic device (100), the server (400) may generate an image corresponding to a keyword group. Meanwhile, when the electronic device (100) updates a profile image according to at least one of a change in a keyword belonging to a keyword group and the arrival of a predetermined period corresponding to an update of the profile image, the electronic device (100) may generate an image corresponding to a currently set profile image for the keyword group and the user account.
[0257] The server (400) can transmit the generated profile image to the electronic device (100) in operation S1510. Meanwhile, the server (400) can store the generated profile image in account data stored in the database (490).
[0258] The server (400) can generate multiple candidate images based on keywords included in the account data. The server (400) can transmit the multiple candidate images to the electronic device (100). When a user input for selecting one of the multiple candidate images is received from the electronic device (100), the server (400) can update the profile image based on the image selected according to the user input.
[0259] Referring to the reference numeral 1601 of FIG. 16, when the electronic device (100) receives a voice spoken by a user, the electronic device (100) may display text (1610) corresponding to the voice spoken by the user on a screen (1600) output through the display (180) based on a result of processing the voice. The electronic device (100) may determine, based on the result of processing the voice, that the voice spoken by the user corresponds to a command to execute an exercise application. Here, the exercise application may be an application linked to a user account. The electronic device (100) may determine, based on the fact that the input voice corresponds to a command to execute an application linked to a user account, that the input voice is a voice requiring user identification.
[0260] Meanwhile, the electronic device (100) can identify a user account corresponding to the voice based on the result of processing the voice spoken by the user. For example, the electronic device (100) can identify a user account corresponding to the voice based on user identification information corresponding to the voice received from the server (400).
[0261] Referring to the reference numeral 1602 of FIG. 16, the electronic device (100) can generate a response corresponding to the result of performing intent analysis on the voice spoken by the user. The electronic device (100) can output a response (1620) notifying that the exercise application is executed based on the fact that the voice spoken by the user corresponds to a command to execute the exercise application. At this time, the electronic device (100) can output a profile image (1630) currently set for the user account corresponding to the voice together with the response (1620) notifying that the exercise application is executed.
[0262] Meanwhile, the electronic device (100) can determine "exercise" as a keyword corresponding to the voice spoken by the user. The electronic device (100) can add "exercise," which is a keyword corresponding to the voice, to the account data stored in the memory (140).
[0263] Referring to FIG. 17, the electronic device (100) can output an account screen (1000, 1100) corresponding to a user account. An object (1010, 1110) representing a logged-in user account included in the account screen (1000, 1100) can include a profile image (1700) currently set for the user account and an object (1710) corresponding to the creation of the profile image.
[0264] When a user selects an object (1710) corresponding to the creation of a profile image using a pointer (205), the electronic device (100) may initiate a process for updating the profile image. For example, the electronic device (100) may generate the profile image using a learning model learned through machine learning stored in the memory (140), a server utilizing a super-large artificial intelligence algorithm model, etc. For example, the electronic device (100) may transmit a user input for updating the profile image to the server (400).
[0265] Referring to FIG. 18, the electronic device (100) may generate a plurality of candidate images (1810) based on keywords included in account data. At this time, the electronic device (100) may generate a plurality of candidate images (1810) by considering the user's characteristics (e.g., age, gender, family relationship, etc.) included in the account data together with the keywords included in the account data. For example, the electronic device (100) may generate a plurality of candidate images (1810) based on the user's age (30s), the user's gender (male), and 'soccer', 'basketball', 'baseball', 'game', and 'sports' belonging to the keyword group among the keywords included in the account data.
[0266] The electronic device (100) can output a screen (1800) including a plurality of candidate images (1810) through the display (180). When a user selects one of the plurality of candidate images (1810) using a pointer (205), the electronic device (100) can update the profile image based on the selected image.
[0267] Referring to FIG. 19, when a profile image is updated, an object (1010) representing a logged-in user account included in an account screen (1000) corresponding to the user account may include an updated profile image (1900).
[0268] Referring to the reference numeral 2001 of FIG. 20, when an electronic device (100) receives a voice spoken by a user, the electronic device (100) may output text (2010) corresponding to the voice spoken by the user based on the result of processing the voice. Based on the result of processing the voice, the electronic device (100) may determine that the voice spoken by the user corresponds to a command for performing a login.
[0269] Referring to the reference numeral 2002 of FIG. 20, the electronic device (100) can identify a user account corresponding to the voice based on the result of processing the voice spoken by the user. The electronic device (100) can output a response (2020) notifying that a login is performed with the user account corresponding to the voice based on the result of performing intent analysis on the voice spoken by the user. At this time, the electronic device (100) can output a currently set profile image (2030) for the user account corresponding to the voice together with the response (2020) notifying that a login is performed.
[0270] According to one embodiment, the electronic device (100) may generate a response based on the elapsed time since the profile image was updated. For example, if a predetermined period of time has not elapsed since the profile image was updated, the electronic device (100) may generate a response that includes the name of the user corresponding to the user account. For example, if a predetermined period of time has elapsed since the profile image was updated, the electronic device (100) may generate a response that does not include the name of the user corresponding to the user account. That is, if a predetermined period of time has elapsed since the profile image was updated, the electronic device (100) may output only the currently set profile image (2030) for the user account, thereby simply and intuitively indicating the user who made the voice call.
[0271] Referring to FIG. 21, the electronic device (100) can switch the user account logged into the server (400) from the currently logged-in user account to the user account of the user who uttered the voice. For example, the electronic device (100) can log out the currently logged-in user account and then log into the server (400) using the user account of the user who uttered the voice.
[0272] The electronic device (100) may output a user switching screen (2100) related to user switching while switching user accounts. The user switching screen (2100) may include a profile image (2110) corresponding to the currently logged-in user account, a profile image (2120) corresponding to the user account of the user who uttered the voice, an object (2130) indicating the switching of user accounts, etc.
[0273] Referring to FIG. 22, the electronic device (100) may generate a plurality of candidate images (1810) based on keywords included in account data. For example, when a user utters a voice requesting an update of a profile image, the electronic device (100) may generate a plurality of candidate images (2220) based on the results of intent analysis performed on the voice uttered by the user. The electronic device (100) may display the plurality of candidate images (2220) on a screen (1600) output through the display (180).
[0274] The electronic device (100) may output a notification (2210) indicating the creation of a profile image. For example, the notification (2210) indicating the creation of a profile image may include information about factors used in the creation of multiple candidate images (2220), such as the user's speech, age, and gender.
[0275] Referring to FIG. 23, the electronic device (100) can obtain characteristic information of a voice spoken by a user. For example, the electronic device (100) can determine the gender of the speaker, the tone of the speaker, the speech style, the speaking speed, the emotion of the speaker, etc. based on at least one of the voice waveform, the voice frequency band, and the voice power spectrum.
[0276] The electronic device (100) can determine the type of voice spoken by the user. For example, based on acquired voice characteristic information, the electronic device (100) can determine one of a plurality of preset types as the type of voice spoken by the user. The electronic device (100) can output a response (2310, 2320) indicating the creation of a profile image based on a screen pattern (2315, 2325) corresponding to the voice type. Through this, the electronic device (100) can provide a response optimized for the user's characteristics identified from the voice spoken by the user.
[0277] Referring to the reference numeral 2301 of FIG. 23, if the speaker's tone is determined to be 'warm', the electronic device (100) may determine the voice type as a first type corresponding to 'warm'. At this time, the electronic device (100) may output a response (2310) indicating the creation of a profile image based on a first screen pattern (2315) corresponding to the first type.
[0278] Meanwhile, referring to reference numeral 2302 of FIG. 23, if the speaker's tone is determined to be 'cheerful', the electronic device (100) may determine the voice type as a second type corresponding to 'cheerfulness'. At this time, the electronic device (100) may output a response (2320) indicating the creation of a profile image based on a second screen pattern (2325) corresponding to the second type.
[0279] As described above, according to at least one embodiment of the present disclosure, identification information for a user's voice can be registered in the user's account.
[0280] Additionally, according to at least one embodiment of the present disclosure, a user can be identified based on the user's voice.
[0281] Additionally, according to at least one embodiment of the present disclosure, a profile image for a user account optimized for the user may be provided based on the user's speech history.
[0282] Additionally, according to at least one embodiment of the present disclosure, a response can be provided in response to a characteristic of the user identified from the voice.
[0283] Referring to FIGS. 1 to 23, an electronic device (100) according to one aspect of the present disclosure includes a user input interface unit (150) that transmits a signal corresponding to an input; a memory (140); and a control unit (170). When a voice signal is received through the user input interface unit (150), the control unit (170) can identify a user account corresponding to the voice signal, add a keyword corresponding to the voice signal included in a result of performing intent analysis on the voice signal to data for the user account stored in the memory (140), and update a profile image for the user account based on the keyword included in the data for the user account.
[0284] In addition, according to one aspect of the present disclosure, the system further includes a network interface unit (135) that communicates with a first server (400) that performs voice recognition, and the memory (140) stores a user list including at least one user identification information corresponding to a user account that has a history of logging into the first server (400), and the control unit (170) transmits the user list together with data including the voice signal to the first server (400), and can confirm a user account corresponding to the voice signal based on the user identification information corresponding to the voice signal received from the first server (400).
[0285] In addition, according to one aspect of the present disclosure, a network interface unit (135) for communicating with a second server using a super-giant artificial intelligence model is further included, and the control unit (170) can transmit a predetermined prompt for the generation of an image corresponding to a keyword included in data for the user account to the second server, and update the profile image based on an image included in a response to the predetermined prompt received from the second server.
[0286] In addition, according to one aspect of the present disclosure, the control unit (170) determines a first keyword group from among a plurality of keywords included in data for the user account based on priorities, and when at least one of the cases where the first keyword group is different from the second keyword group determined based on the priorities immediately before and when a predetermined period has arrived, generates an image corresponding to the first keyword group, and updates the profile image based on the generated image.
[0287] In addition, according to one aspect of the present disclosure, when a user input for generating an image is received through the user input interface unit (150), the control unit (170) may generate a first image corresponding to the first keyword group, and when at least one of the cases where the first keyword group is different from the second keyword group and the predetermined period has arrived, the control unit (170) may generate a second image corresponding to a profile image currently set for the first keyword group and the user account.
[0288] In addition, according to one aspect of the present disclosure, the control unit (170) may acquire user characteristics including at least one of age and gender included in data for the user account, generate an image corresponding to a keyword included in the data for the user account and the user characteristics, and update the profile image based on the generated image.
[0289] In addition, according to one aspect of the present disclosure, a display (180) is further included, and the control unit (170) generates a plurality of candidate images based on keywords included in data for the user account, outputs the plurality of candidate images through the display (180), and when a user input for selecting one of the plurality of candidate images is received through the user input interface unit (150), the profile image can be updated based on the selected image.
[0290] In addition, according to one aspect of the present disclosure, the control unit (170) can generate a response corresponding to the result of performing intent analysis on the voice signal, and output the generated response and the profile image currently set for the user account through the display (180).
[0291] In addition, according to one aspect of the present disclosure, the control unit (170) may generate a first response including the name of the user corresponding to the user account if a predetermined period of time has not elapsed since the profile image was updated, and may generate a second response not including the name of the user if the predetermined period of time has elapsed since the profile image was updated.
[0292] In addition, according to one aspect of the present disclosure, the control unit (170) may determine the type of voice corresponding to the voice signal, and, if the type of the voice is a first type, output the generated response based on a first screen pattern corresponding to the first type, and, if the type of the voice is a second type, output the generated response based on a second screen pattern corresponding to the second type.
[0293] A server (400) according to one aspect of the present disclosure includes a communication unit (480) that communicates with an electronic device (100); a database (490); and a controller (470). When a voice signal is received through the communication unit (480), the controller (470) can identify a user account corresponding to the voice signal, add a keyword corresponding to the voice signal included in a result of performing intent analysis on the voice signal to data for the user account stored in the database (490), and update a profile image for the user account based on the keyword included in the data for the user account.
[0294] In addition, according to one aspect of the present disclosure, the controller (470) generates first identification information corresponding to the voice signal, and when a user list including at least one user identification information is received through the communication unit (480), searches the database (490) for second identification information corresponding to the user identification information included in the user list, and when the first identification information corresponds to the second identification information, determines a user account corresponding to the second identification information as the user account corresponding to the voice signal.
[0295] In addition, according to one aspect of the present disclosure, the controller (470) may transmit, through the communication unit (480), a predetermined prompt for the generation of an image corresponding to a keyword included in data for the user account to an external server using a super-giant artificial intelligence model, and update the profile image based on an image included in a response to the predetermined prompt received from the external server.
[0296] In addition, according to one aspect of the present disclosure, the controller (470) determines a first keyword group from among a plurality of keywords included in data for the user account based on priorities, and when at least one of the cases where the first keyword group is different from a second keyword group determined based on the priorities immediately before and when a predetermined period has arrived, generates an image corresponding to the first keyword group, and updates the profile image based on the generated image.
[0297] In addition, according to one aspect of the present disclosure, when a user input for generating an image is received through the communication unit (480), the controller (470) may generate a first image corresponding to the first keyword group, and when at least one of the cases where the first keyword group is different from the second keyword group and the predetermined period has arrived, the controller (470) may generate a second image corresponding to a profile image currently set for the first keyword group and the user account.
[0298] The attached drawings are only intended to facilitate understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure.
[0299] Meanwhile, the operating method of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices that store data that can be read by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across network-connected computer systems, so that the processor-readable code can be stored and executed in a distributed manner.
[0300] In addition, although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. User input interface section that transmits a signal corresponding to input; memory; and Including a control unit, The above control unit, When a voice signal is received through the user input interface section, the user account corresponding to the voice signal is checked, Adding a keyword corresponding to the voice signal included in the result of performing intent analysis on the voice signal to the data for the user account stored in the memory, An electronic device characterized in that it updates a profile image for the user account based on keywords included in data for the user account.
2. In paragraph 1, Further comprising a network interface unit for communicating with a first server performing voice recognition, The above memory stores a user list including at least one user identification information corresponding to a user account that has logged in to the first server, The above control unit, Transmitting the above user list together with data including the above voice signal to the first server, An electronic device characterized in that it verifies a user account corresponding to the voice signal based on user identification information corresponding to the voice signal received from the first server.
3. In paragraph 1, It further includes a network interface unit that communicates with a second server that uses a super-giant artificial intelligence model. The above control unit, Transmitting a predetermined prompt for the generation of an image corresponding to a keyword included in the data for the user account to the second server, An electronic device characterized in that the profile image is updated based on an image included in a response to the predetermined prompt received from the second server.
4. In paragraph 1, The above control unit, Based on the priority, a first keyword group is determined from among a plurality of keywords included in the data for the user account, If at least one of the cases where the first keyword group is different from the second keyword group determined based on the priority just before and a predetermined period has arrived, an image corresponding to the first keyword group is generated, An electronic device characterized in that it updates the profile image based on the generated image.
5. In paragraph 4, The above control unit, When a user input for generating an image is received through the user input interface section, a first image corresponding to the first keyword group is generated, An electronic device characterized in that, when at least one of the first keyword group is different from the second keyword group and the predetermined period has arrived, a second image corresponding to the profile image currently set for the first keyword group and the user account is generated.
6. In paragraph 1, The above control unit, Obtaining user characteristics including at least one of age and gender, which are included in the data for the user account; Generate an image corresponding to the keywords contained in the data for the user account and the characteristics of the user, An electronic device characterized in that it updates the profile image based on the generated image.
7. In paragraph 1, Including more displays, The above control unit, Generate multiple candidate images based on keywords contained in the data for the above user account, Outputting the above plurality of candidate images through the above display, An electronic device characterized in that, when a user input for selecting one of the plurality of candidate images is received through the user input interface unit, the profile image is updated based on the selected image.
8. In paragraph 1, Including more displays, The above control unit, Generate a response corresponding to the result of performing intent analysis on the above voice signal, An electronic device characterized in that it outputs the currently set profile image and the generated response for the user account through the display.
9. In paragraph 8, The above control unit, If a predetermined period of time has not passed since the above profile image was updated, a first response is generated that includes the name of the user corresponding to the above user account, An electronic device characterized in that, when the predetermined period of time has elapsed after the profile image is updated, a second response is generated that does not include the user's name.
10. In paragraph 8, The above control unit, Determine the type of voice corresponding to the above voice signal, If the type of the above voice is the first type, output the generated response based on the first screen pattern corresponding to the first type, An electronic device characterized in that, when the type of the voice is the second type, the generated response is output based on a second screen pattern corresponding to the second type.
11. Communication unit for communicating with electronic devices; database; and Includes a controller, The above controller, When a voice signal is received through the above communication unit, the user account corresponding to the voice signal is verified, Adding a keyword corresponding to the voice signal included in the result of performing intent analysis on the voice signal to the data for the user account stored in the database, A server characterized in that it updates a profile image for the user account based on keywords included in data for the user account.
12. In paragraph 11, The above controller, Generate first identification information corresponding to the above voice signal, When a user list including at least one user identification information is received through the above communication unit, second identification information corresponding to the user identification information included in the user list is searched for in the database, A server characterized in that, when the first identification information corresponds to the second identification information, the user account corresponding to the second identification information is determined as the user account corresponding to the voice signal.
13. In paragraph 11, The above controller, Through the above communication unit, a predetermined prompt for the generation of an image corresponding to a keyword included in the data for the user account is transmitted to an external server using a super-giant artificial intelligence model, A server characterized in that it updates the profile image based on an image included in a response to the predetermined prompt received from the external server.
14. In paragraph 11, The above controller, Based on the priority, a first keyword group is determined from among a plurality of keywords included in the data for the user account, If at least one of the cases where the first keyword group is different from the second keyword group determined based on the priority just before and a predetermined period has arrived, an image corresponding to the first keyword group is generated, A server characterized in that it updates the profile image based on the generated image.
15. In paragraph 14, The above controller, When a user input for generating an image is received through the above communication unit, a first image corresponding to the first keyword group is generated, A server characterized in that, if at least one of the first keyword group is different from the second keyword group and the predetermined period has arrived, a second image corresponding to the profile image currently set for the first keyword group and the user account is generated.
Citation Information
Patent Citations
Aerial work platforms and how to operate them to avoid accidents
KR1020260000946A
electronic apparatus, method for contolling mobile apparatus by electronic apparatus and computer-readable recording medium
KR102582332B1
Method, apparatus and computer program for providing automated calibration service for user's profile image
KR102608777B1
Communication module package
KR102723650B1
Method for relaying emotion display data and system for same
WO2013048225A2