Electronic device and operating method thereof

The electronic device simplifies voiceprint registration by generating utterance groups from voice data and automatically matching account profiles, addressing inconvenience and security risks in existing voiceprint registration processes.

WO2026010011A1PCT designated stage Publication Date: 2026-01-08LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/009512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-04
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing voiceprint registration processes in display devices are inconvenient as they require multiple utterances and can pose security risks if the user logged into the device does not match the speaker for registration.

Method used

An electronic device generates utterance groups based on collected voice data, matches voice commands with these groups, and automatically registers the account profile if a matching profile exists, eliminating the need for separate utterances and ensuring the user and speaker are the same.

Benefits of technology

This method simplifies the voiceprint registration process and reduces security risks by automatically matching the user and speaker, enhancing user convenience and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024009512_08012026_PF_FP_ABST
    Figure KR2024009512_08012026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment of the present disclosure may comprise: a memory; and a controller for generating one or more utterance groups by using voice data collected from one or more users, receiving a voice command, acquiring an utterance group matching the voice command from among the one or more utterance groups, and when an account profile matching utterance group information of the acquired utterance group exists among a plurality of account profiles stored in the memory, matching the account profile and the utterance group information and storing same in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of operation thereof

[0001] The present disclosure relates to an electronic device, and more particularly, to an electronic device capable of providing a voice print recognition service.

[0002] Recent display devices offer voice recognition services, which provide a variety of services through the user's spoken voice. One example of a voice recognition service is a voiceprint recognition service, which registers a voiceprint corresponding to the user's spoken voice to an account.

[0003] The display device displays multiple sentences for voiceprint registration, and the user utters voices corresponding to the multiple sentences. Voice features are then extracted from the user's spoken voices, and the voiceprint registration process is performed, matching the extracted voice features to an account.

[0004] However, the voice registration process may be inconvenient because the user must utter voices corresponding to multiple sentences.

[0005] Additionally, there is an issue with verifying that the user logged into the display device and the speaker for voiceprint registration are the same. If the user logged into the display device and the speaker for voiceprint registration are not the same, the speaker performing the voiceprint registration can use information about the logged-in account, potentially posing a security risk.

[0006] An object of the present disclosure may be to provide a display device that easily performs a voiceprint registration process without a separate voiceprint registration utterance process.

[0007] An object of the present disclosure may be to match a user of an account logged into a display device with a speaker for voice registration.

[0008] An object of the present disclosure may be to provide an automatic voiceprint registration service by generating a group of utterances capable of identifying a speaker based on collected voice data.

[0009] An electronic device according to one embodiment of the present disclosure may include a memory, a controller configured to generate one or more utterance groups using voice data collected from one or more users, receive a voice command, obtain an utterance group matching the voice command among the one or more utterance groups, and, if an account profile matching utterance group information of the obtained utterance group exists among a plurality of account profiles stored in the memory, match the account profile with the utterance group information and store the matched account profile in the memory.

[0010] According to one embodiment of the present disclosure, an electronic device may include a method of operating an electronic device, the method including: generating one or more utterance groups using voice data collected from one or more users; receiving a voice command; acquiring an utterance group matching the voice command among the one or more utterance groups; and, if an account profile matching utterance group information of the acquired utterance group exists among a plurality of account profiles, matching the account profile with the utterance group information and storing the matching account profile.

[0011] According to an embodiment of the present disclosure, the inconvenience of the voiceprint registration process can be eliminated as the voiceprint registration process is performed without a separate utterance process for voiceprint registration.

[0012] According to an embodiment of the present disclosure, the user of the account logged into the display device and the speaker for voice registration are matched, so that information errors between the speaker and the account logged into the display device can be reduced.

[0013] According to an embodiment of the present disclosure, a voiceprint registration service can be automatically implemented by generating a group of utterances capable of identifying the speaker based on collected voice data. This can simplify the voiceprint registration process.

[0014] Figure 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0015] Figure 2 is a block diagram of a remote control device according to one embodiment of the present invention.

[0016] Figure 3 shows an example of an actual configuration of a remote control device according to one embodiment of the present invention.

[0017] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0018] FIG. 5 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.

[0019] FIG. 6 is a drawing for explaining the configuration of a system according to one embodiment of the present disclosure.

[0020] FIG. 7 is a flowchart for explaining an operating method of a display device according to an embodiment of the present disclosure.

[0021] FIG. 8 is a diagram illustrating a process of generating one or more utterance groups using uttered speech according to an embodiment of the present disclosure.

[0022] FIG. 9 is a drawing illustrating an example of displaying a voice registration window according to one embodiment of the present disclosure.

[0023] FIG. 10 is a drawing illustrating an example of displaying a notification of completion of voice registration according to one embodiment of the present disclosure.

[0024] FIGS. 11A to 11D are drawings for explaining embodiments related to the present disclosure.

[0025] FIG. 12 is a sequence diagram for explaining an operation method of a system according to an embodiment of the present disclosure.

[0026] Figure 13 is a diagram illustrating the process of updating the utterance group and voiceprint recognition model after voiceprint registration is completed.

[0027] Hereinafter, embodiments related to the present invention will be described in more detail with reference to the drawings. The suffixes "module" and "part" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles.

[0028] A display device according to an embodiment of the present invention is, for example, an intelligent display device that adds computer-assisted functionality to its broadcast reception function. While faithfully performing the broadcast reception function, it can also be equipped with Internet functionality and other features, providing a more user-friendly interface, such as a manual input device, touch screen, or space remote control. Furthermore, with support for wired or wireless Internet functionality, it can connect to the Internet and a computer, enabling functions such as email, web browsing, banking, or gaming. A standardized, general-purpose operating system can be used for these various functions.

[0029] Accordingly, the display device described in the present invention can perform various user-friendly functions, for example, by allowing various applications to be freely added or deleted on a general-purpose operating system kernel. More specifically, the display device can be a network TV, HBB TV, smart TV, LED TV, OLED TV, etc., and in some cases, it can also be applied to smartphones.

[0030] FIG. 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0031] Referring to FIG. 1, the display device (100) may include a broadcast receiving unit (130), an external device interface (135), a memory (140), a user input interface (150), a controller (170), a wireless communication interface (173), a display (180), a speaker (185), and a power supply circuit (190).

[0032] The broadcast receiving unit (130) may include a tuner (131), a demodulator (132), and a network interface (133).

[0033] The tuner (131) can select a specific broadcast channel according to a channel selection command. The tuner (131) can receive a broadcast signal for the selected specific broadcast channel.

[0034] The demodulator (132) can separate the received broadcast signal into a video signal, an audio signal, and a data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into a form that can be output.

[0035] The external device interface (135) can receive an application or a list of applications within an adjacent external device and transmit it to the controller (170) or memory (140).

[0036] The external device interface (135) can provide a connection path between the display device (100) and the external device. The external device interface (135) can receive one or more of images and audio output from an external device connected wirelessly or wiredly to the display device (100) and transmit them to the controller (170). The external device interface (135) can include a plurality of external input terminals. The plurality of external input terminals can include an RGB terminal, one or more HDMI (High Definition Multimedia Interface) terminals, and a component terminal.

[0037] A video signal of an external device input through an external device interface (135) can be output through a display (180). A voice signal of an external device input through an external device interface (135) can be output through a speaker (185).

[0038] An external device that can be connected to the external device interface (135) may be any one of a set-top box, a Blu-ray player, a DVD player, a game console, a sound bar, a smartphone, a PC, a USB memory, and a home theater, but this is only an example.

[0039] The network interface (133) may provide an interface for connecting the display device (100) to a wired / wireless network including the Internet. The network interface (133) may transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.

[0040] Additionally, some content data stored in the display device (100) can be transmitted to a selected user or electronic device among other users or other electronic devices pre-registered in the display device (100).

[0041] The network interface (133) can access a predetermined web page through a connected network or another network linked to the connected network. That is, by accessing a predetermined web page through a network, data can be transmitted or received with the corresponding server.

[0042] In addition, the network interface (133) can receive content or data provided by a content provider or network operator. That is, the network interface (133) can receive content such as movies, advertisements, games, VOD, broadcast signals, etc. and information related thereto provided from a content provider or network provider via a network.

[0043] Additionally, the network interface (133) can receive firmware update information and update files provided by the network operator, and can transmit data to the Internet or content provider or network operator.

[0044] The network interface (133) can select and receive a desired application from among applications open to the public via a network.

[0045] The memory (140) stores a program for each signal processing and control within the controller (170), and can store signal-processed image, voice, or data signals.

[0046] In addition, the memory (140) may perform a function for temporary storage of video, audio, or data signals input from an external device interface (135) or a network interface (133), and may store information about a specific image through a channel memory function.

[0047] The memory (140) can store an application or a list of applications input from an external device interface (135) or a network interface (133).

[0048] The display device (100) can play content files (video files, still image files, music files, document files, application files, etc.) stored in the memory (140) and provide them to the user.

[0049] The user input interface (150) can transmit a signal input by the user to the controller (170) or transmit a signal from the controller (170) to the user. For example, the user input interface (150) can receive and process control signals such as power on / off, channel selection, and screen setting from the remote control device (200) according to various communication methods such as Bluetooth, Ultra Wideband (WB), ZigBee, Radio Frequency (RF) communication, or infrared (IR) communication, or process control signals from the controller (170) to be transmitted to the remote control device (200).

[0050] In addition, the user input interface (150) can transmit control signals input from local keys (not shown) such as a power key, channel key, volume key, and setting value to the controller (170).

[0051] An image signal processed by the controller (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, an image signal processed by the controller (170) can be input to an external output device through an external device interface (135).

[0052] The voice signal processed by the controller (170) can be output as audio to the speaker (185). In addition, the voice signal processed by the controller (170) can be input to an external output device through the external device interface (135).

[0053] In addition, the controller (170) can control the overall operation within the display device (100).

[0054] In addition, the controller (170) can control the display device (100) by a user command or internal program input through the user input interface (150), and can connect to a network to enable the user to download a desired application or application list into the display device (100).

[0055] The controller (170) enables the user-selected channel information, etc. to be output through a display (180) or speaker (185) together with processed video or audio signals.

[0056] In addition, the controller (170) allows a video signal or audio signal from an external device, for example, a camera or camcorder, input through the external device interface (135) to be output through the display (180) or speaker (185) in accordance with an external device video playback command received through the user input interface (150).

[0057] Meanwhile, the controller (170) can control the display (180) to display an image, for example, a broadcast image input through a tuner (131), an external input image input through an external device interface (135), an image input through a network interface, or an image stored in a memory (140) can be controlled to be displayed on the display (180). In this case, the image displayed on the display (180) can be a still image or a moving image, and can be a 2D image or a 3D image.

[0058] In addition, the controller (170) can control the playback of content stored in the display device (100), received broadcast content, or external input content input from outside, and the content can be in various forms such as broadcast video, external input video, audio file, still image, connected web screen, and document file.

[0059] The wireless communication interface (173) can communicate with an external device through wired or wireless communication. The wireless communication interface (173) can perform short-range communication with the external device. To this end, the wireless communication interface (173) can support short-range communication using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies. This wireless communication interface (173) can support wireless communication between the display device (100) and a wireless communication system, between the display device (100) and another display device (100), or between the display device (100) and a network in which the display device (100, or an external server) is located via a short-range wireless communication network (Wireless Area Network). The short-range wireless communication network can be a short-range wireless personal area network (Wireless Personal Area Network).

[0060] Here, the other display device (100) may be a wearable device (e.g., a smartwatch, smart glasses, a head-mounted display (HMD)) or a mobile terminal such as a smart phone that can exchange data with (or be linked to) the display device (100) according to the present invention. The wireless communication interface (173) may detect (or recognize) a wearable device capable of communication around the display device (100).

[0061] Furthermore, if the detected wearable device is a device certified to communicate with the display device (100) according to the present invention, the controller (170) can transmit at least a portion of the data processed in the display device (100) to the wearable device via the wireless communication interface (173). Accordingly, a user of the wearable device can utilize the data processed in the display device (100) via the wearable device.

[0062] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal processed by the controller (170) or a video signal, data signal, etc. received from an external device interface (135) into R, G, and B signals, respectively.

[0063] Meanwhile, since the display device (100) illustrated in FIG. 1 is merely an embodiment of the present invention, some of the illustrated components may be integrated, added, or omitted depending on the specifications of the display device (100) actually implemented.

[0064] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0065] According to another embodiment of the present invention, the display device (100) may receive and play back an image through a network interface (133) or an external device interface (135) without having a tuner (131) and a demodulator (132), unlike that shown in FIG. 1.

[0066] For example, the display device (100) may be implemented separately as an image processing device, such as a set-top box, for receiving broadcast signals or contents according to various network services, and a content playback device for playing contents input from the image processing device.

[0067] In this case, the operating method of the display device according to the embodiment of the present invention to be described below may be performed by any one of the display device (100) described with reference to FIG. 1, as well as an image processing device such as the separated set-top box, or a content playback device having a display (180) and a speaker (185).

[0068] Next, a remote control device according to an embodiment of the present invention will be described with reference to FIGS. 2 and 3.

[0069] FIG. 2 is a block diagram of a remote control device according to an embodiment of the present invention, and FIG. 3 shows an example of an actual configuration of a remote control device (200) according to an embodiment of the present invention.

[0070] First, referring to FIG. 2, the remote control device (200) may include a fingerprint recognition device (210), a wireless communication circuit (220), a user input interface (230), a sensor (240), an output interface (250), a power supply circuit (260), a memory (270), a controller (280), and a microphone (290).

[0071] Referring to FIG. 2, the wireless communication circuit (220) transmits and receives signals with any one of the display devices according to the embodiments of the present invention described above.

[0072] The remote control device (200) may be equipped with an RF circuit (221) capable of transmitting and receiving signals with the display device (100) according to RF communication standards, and an IR circuit (223) capable of transmitting and receiving signals with the display device (100) according to IR communication standards. In addition, the remote control device (200) may be equipped with a Bluetooth circuit (225) capable of transmitting and receiving signals with the display device (100) according to Bluetooth communication standards. In addition, the remote control device (200) may be equipped with an NFC circuit (227) capable of transmitting and receiving signals with the display device (100) according to NFC (Near Field Communication) communication standards, and a WLAN circuit (229) capable of transmitting and receiving signals with the display device (100) according to WLAN (Wireless LAN) communication standards.

[0073] In addition, the remote control device (200) transmits a signal containing information about the movement of the remote control device (200) to the display device (100) through a wireless communication circuit (220).

[0074] Meanwhile, the remote control device (200) can receive a signal transmitted by the display device (100) through the RF circuit (221), and, if necessary, can transmit commands for power on / off, channel change, volume change, etc. to the display device (100) through the IR circuit (223).

[0075] The user input interface (230) may be configured as a keypad, buttons, a touchpad, or a touch screen. The user can input commands related to the display device (100) to the remote control device (200) by operating the user input interface (230). If the user input interface (230) includes a hard key button, the user can input commands related to the display device (100) to the remote control device (200) by pushing the hard key button. This will be described with reference to FIG. 3.

[0076] Referring to FIG. 3, the remote control device (200) may include a plurality of buttons. The plurality of buttons may include a fingerprint recognition button (212), a power button (231), a home button (232), a live button (233), an external input button (234), a volume control button (235), a voice recognition button (236), a channel change button (237), a confirmation button (238), and a back button (239).

[0077] The fingerprint recognition button (212) may be a button for recognizing a user's fingerprint. In one embodiment, the fingerprint recognition button (212) may be capable of a push operation, and may receive a push operation and a fingerprint recognition operation.

[0078] The power button (231) may be a button for turning the power of the display device (100) on / off.

[0079] The home button (232) may be a button for moving to the home screen of the display device (100).

[0080] The live button (233) may be a button for displaying a real-time broadcast program.

[0081] The external input button (234) may be a button for receiving an external input connected to the display device (100).

[0082] The volume control button (235) may be a button for adjusting the size of the volume output by the display device (100).

[0083] The voice recognition button (236) may be a button for receiving a user's voice and recognizing the received voice.

[0084] The channel change button (237) may be a button for receiving a broadcast signal of a specific broadcast channel.

[0085] The confirmation button (238) may be a button for selecting a specific function, and the back button (239) may be a button for returning to the previous screen.

[0086] Let's explain Figure 2 again.

[0087] When the user input interface (230) has a touch screen, the user can input commands related to the display device (100) using the remote control device (200) by touching the soft keys of the touch screen. In addition, the user input interface (230) may have various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the present invention.

[0088] The sensor (240) may include a gyro sensor (241) or an acceleration sensor (243), and the gyro sensor (241) may sense information about the movement of the remote control device (200).

[0089] For example, the gyro sensor (241) can sense information about the operation of the remote control device (200) based on the x, y, and z axes, and the acceleration sensor (243) can sense information about the movement speed of the remote control device (200). Meanwhile, the remote control device (200) can further include a distance measuring sensor, so as to sense the distance to the display (180) of the display device (100).

[0090] The output interface (250) can output a video or audio signal corresponding to an operation of the user input interface (230) or a signal transmitted from the display device (100).

[0091] The user can recognize whether the output interface (250) is manipulating the user input interface (230) or controlling the display device (100).

[0092] For example, the output interface (250) may include an LED (251) that lights up when the user input interface (230) is operated or a signal is transmitted and received with the display device (100) via the wireless communication unit (225), a vibrator (253) that generates vibrations, a speaker (255) that outputs sound, or a display (257) that outputs images.

[0093] In addition, the power supply circuit (260) supplies power to the remote control device (200), and power waste can be reduced by stopping the power supply when the remote control device (200) does not move for a predetermined period of time.

[0094] The power supply circuit (260) can resume power supply when a predetermined key provided in the remote control device (200) is operated.

[0095] The memory (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).

[0096] When the remote control device (200) wirelessly transmits and receives signals through the display device (100) and the RF circuit (221), the remote control device (200) and the display device (100) transmit and receive signals through a predetermined frequency band.

[0097] The controller (280) of the remote control device (200) can store and reference information regarding the frequency band that can wirelessly transmit and receive signals with the display device (100) paired with the remote control device (200) in the memory (270).

[0098] The controller (280) controls all matters related to the control of the remote control device (200). The controller (280) can transmit a signal corresponding to a predetermined key operation of the user input interface (230) or a signal corresponding to the movement of the remote control device (200) sensed by the sensor (240) to the display device (100) via the wireless communication unit (225).

[0099] Additionally, the microphone (290) of the remote control device (200) can acquire voice.

[0100] A plurality of microphones (290) may be provided.

[0101] Next, Figure 4 is described.

[0102] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0103] Figure 4 (a) illustrates that a pointer (205) corresponding to a remote control device (200) is displayed on a display (180).

[0104] The user can move or rotate the remote control device (200) up and down, left and right. The pointer (205) displayed on the display (180) of the display device (100) corresponds to the movement of the remote control device (200). This remote control device (200) can be called a space remote control because, as shown in the drawing, the pointer (205) moves and is displayed according to the movement in 3D space.

[0105] Figure 4 (b) illustrates that when a user moves the remote control device (200) to the left, the pointer (205) displayed on the display (180) of the display device (100) also moves to the left correspondingly.

[0106] Information about the movement of the remote control device (200) detected through the sensor of the remote control device (200) is transmitted to the display device (100). The display device (100) can calculate the coordinates of the pointer (205) from the information about the movement of the remote control device (200). The display device (100) can display the pointer (205) to correspond to the calculated coordinates.

[0107] Figure 4 (c) illustrates a case where a user moves the remote control device (200) away from the display (180) while pressing a specific button within the remote control device (200). As a result, a selection area within the display (180) corresponding to the pointer (205) can be zoomed in and displayed in an enlarged manner.

[0108] Conversely, when the user moves the remote control device (200) closer to the display (180), the selection area within the display (180) corresponding to the pointer (205) may be zoomed out and displayed in a reduced size.

[0109] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.

[0110] Additionally, when a specific button within the remote control device (200) is pressed, recognition of up, down, left, and right movements may be excluded. That is, when the remote control device (200) moves away from or toward the display (180), up, down, left, and right movements may not be recognized, and only forward and backward movements may be recognized. When a specific button within the remote control device (200) is not pressed, only the pointer (205) moves in accordance with the up, down, left, and right movements of the remote control device (200).

[0111] Meanwhile, the movement speed or movement direction of the pointer (205) can correspond to the movement speed or movement direction of the remote control device (200).

[0112] Meanwhile, the pointer in this specification refers to an object displayed on the display (180) in response to the operation of the remote control device (200). Accordingly, objects of various shapes other than the arrow shape illustrated in the drawing can be used as the pointer (205). For example, the pointer may be a concept including a point, a cursor, a prompt, a thick outline, etc. In addition, the pointer (205) may be displayed corresponding to one point on the horizontal or vertical axis on the display (180), or may be displayed corresponding to multiple points such as lines or surfaces.

[0113] FIG. 5 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.

[0114] Referring to FIG. 5, the AI ​​server (500) may refer to a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network.

[0115] The AI ​​server (500) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network.

[0116] The AI ​​server (500) may be included as part of the configuration of the AI ​​device (100) and may perform at least part of the AI ​​processing together.

[0117] The AI ​​server (500) may include a communication unit (510), a memory (530), a learning processor (540), and a processor (560).

[0118] The communication unit (510) can transmit and receive data with an external device such as a display device (100).

[0119] The memory (530) may include a model storage unit (531).

[0120] The model storage unit (531) can store a model (or artificial neural network, 531a) that is being learned or has been learned through the learning processor (540).

[0121] A learning processor (540) can train an artificial neural network (531a) using learning data. The learning model can be used while mounted on the AI ​​server (500) of the artificial neural network, or can be mounted on an external device such as a display device (100).

[0122] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (530).

[0123] The processor (560) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.

[0124] FIG. 6 is a drawing for explaining the configuration of a system according to one embodiment of the present disclosure.

[0125] Referring to FIG. 6, the system (60) may include a display device (100), an AI server (500), a STT (Speech-To-Text) server (620), an NLP (Natural Language Processing) server (630), and a speaker recognition server (600).

[0126] Each component constituting the system (60) can perform wireless communication with each other. The wireless communication may be Internet communication.

[0127] The speaker recognition server (600) may be referred to as a speaker recognition device. The speaker recognition server (600) can identify a speaker by comparing the voice characteristics of a voice signal with pre-stored voice characteristics.

[0128] The AI ​​server (500) can transmit a voice signal transmitted from the display device (100) to the speaker recognition server (600) or the STT server (620).

[0129] The AI ​​server (500) may be a relay server that relays communication between the display device (100) and an external server.

[0130] The AI ​​server (500) may be omitted. In this case, the display device (100) may directly transmit a voice signal to the speaker recognition server (600) or the STT server (620).

[0131] The STT (Speech-To-Text) server (620) can convert a voice signal into text data.

[0132] The STT server (620) can transmit the converted text data to the NLP server (630).

[0133] The NLP (Natural Language Processing) server (630) can obtain intent analysis results based on text data received from the STT server (620) using a natural language processing (NLP) engine.

[0134] The NLP server (630) can transmit the acquired intent analysis results to the display device (100) via the AI ​​server (500).

[0135] The speaker database (601) may be included in the speaker recognition server (600) or may be provided separately from the speaker recognition server (600).

[0136] The speaker database (601) can store a speaker profile corresponding to each speaker. The speaker profile may include one or more of the speaker's gender, age, nickname, payment information, email address, and phone number. The speaker profile may be referred to as an account profile.

[0137] The speaker database (601) can store embedding vectors corresponding to each of a plurality of registered speakers and speaker identification information matched to each embedding vector.

[0138] A speaker profile may include an embedding vector representing the voice features of a registered speaker.

[0139] Hereinafter, the speaker registration process may be a process of extracting features of a specific speaker's speech using speech recognition technology, matching the extracted features to speaker identification information, and storing them. The speaker registration process may be a voiceprint registration process.

[0140] The speaker identification process may be a process of identifying whether the pre-registered speaker and the currently speaking speaker match.

[0141] The speaker registration process and speaker identification process may be processes for providing personalized voice recognition services.

[0142] FIG. 7 is a flowchart for explaining an operating method of a display device according to an embodiment of the present disclosure.

[0143] The display device (100) below can also be applied to an electronic device equipped with a display (180). The electronic device can be any one of an air conditioner, a refrigerator, a robot vacuum cleaner, or a vehicle.

[0144] The display device (100) may also be referred to as an electronic device (100).

[0145] The controller (170) of the display device (100) can generate one or more speech groups based on speech data corresponding to the speaker's voice (S701).

[0146] In one embodiment, the controller (170) can receive speech from one or more speakers from a remote control device (200).

[0147] In another embodiment, the controller (170) can receive voices spoken by one or more speakers through a microphone (not shown) provided in the display device (100).

[0148] One or more speakers may utter a voice at different times. The voice may be a voice command uttered to utilize the voice recognition service of the display (100).

[0149] The controller (170) can generate one or more utterance groups based on a voice data set corresponding to the received voices.

[0150] The controller (170) can obtain an embedding vector (voice feature vector) representing voice features from voice data. The controller (170) can extract the embedding vector from the voice data using the MFCC (Mel-Frequency Cepstral Coefficients) technique.

[0151] The MFCC technique may be a technique that converts voice data into a frequency-based power spectrum and performs an inverse Fourier transform on the power spectrum to extract an embedding vector.

[0152] The controller (170) can cluster a plurality of embedding vectors and generate one or more utterance groups based on the clustering results.

[0153] The controller (170) can create one or more utterance groups by clustering multiple embedding vectors using the K-means clustering technique or the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) technique.

[0154] The controller (170) can store an embedding vector corresponding to speech data in the memory (140). The stored embedding vector can be used to generate a group of utterances.

[0155] The controller (170) can store a PCM (Pulse Code Modulation) voice file corresponding to the voice spoken by the speaker in the memory (140). The voice data may be a PCM voice file. The controller (170) can create one or more utterance groups based on the PCM voice file.

[0156] FIG. 8 is a diagram illustrating a process of generating one or more utterance groups using uttered speech according to an embodiment of the present disclosure.

[0157] Referring to Figure 8, the first speaker (A) is 35 years old and female. The second speaker (B) is 42 years old and male. The third speaker (C) is 13 years old and male.

[0158] Voice characteristics may vary depending on the speaker's age and gender. In the present disclosure, information about the voices uttered by the speaker while using a voice recognition service is collected in advance, and one or more utterance groups can be created using the collected information.

[0159] The controller (170) can extract an embedding vector from each of the voices (801) spoken by the first speaker (A) and cluster the extracted embedding vectors to create a first utterance group (810).

[0160] The controller (170) can extract an embedding vector from each of the voices (803) spoken by the second speaker (A) and cluster the extracted embedding vectors to create a second utterance group (830).

[0161] The controller (170) can extract an embedding vector from each of the voices (805) spoken by the third speaker (C) and cluster the extracted embedding vectors to create a third utterance group (850).

[0162] Each of the first speech group (810), the second speech group (830), and the third speech group (850) may be a group clustered according to age and gender, which are items for distinguishing voice features.

[0163] Again, Figure 7 is explained.

[0164] When one or more utterance groups are created, the controller (170) can display a voice registration window on the display (180) (S703).

[0165] In one embodiment, the controller (170) may display a voice registration window for voice registration on the display (180) when one or more utterance groups are created.

[0166] In another embodiment, the controller (170) may recognize the voice of a speaker who has not registered his / her voiceprint, and if one or more utterance groups are created, a voiceprint registration window may be displayed on the display (180).

[0167] In another embodiment, when the speaker performs voiceprint registration through a voiceprint registration application installed on the display device (100), the display of the voiceprint registration window may be omitted.

[0168] FIG. 9 is a drawing illustrating an example of displaying a voice registration window according to one embodiment of the present disclosure.

[0169] Referring to FIG. 9, the controller (170) of the display device (100) can display a voiceprint registration window (900) for voiceprint registration on the display (180). The voiceprint registration window (900) can be referred to as a voiceprint registration notification.

[0170] The controller (170) may display a voiceprint registration window (900) on the display (180) when one or more utterance groups are created. The controller (170) may display a voiceprint registration window (900) when the voice of a non-registered voiceprint speaker is recognized and one or more utterance groups are created. The voiceprint registration window (900) may be displayed in the form of a pop-up window.

[0171] Again, Figure 7 is explained.

[0172] The controller (170) can receive a voice command and obtain a utterance group matching the received voice command among one or more utterance groups (S707).

[0173] The controller (170) can receive a voice command agreeing to voice registration, and can extract a group of utterances matching the voice command from one or more groups of utterances according to the received voice command.

[0174] The voice command for consenting to voice registration may be a command indicating consent or affirmation, such as "yes." The controller (170) may convert the voice command into text data, and through intent analysis of the converted text data, may recognize that the voice command is a command indicating consent.

[0175] The controller (170) can extract a voice feature vector representing a voice characteristic from a voice command. The controller (170) can obtain an utterance group that matches the extracted voice feature vector from one or more utterance groups. The controller (170) can obtain an utterance group with the closest distance between voice feature vectors from one or more utterance groups as the matching utterance group.

[0176] Referring to FIG. 9, the speaker utters the voice command <Yes> to agree to voice registration while the voice registration window (900) is displayed on the display (180).

[0177] The controller (170) can compare the acquired utterance group information with multiple account profiles (S709).

[0178] Utterance group information may be an item identifying an utterance group. Utterance group information may be items that serve as criteria for clustering utterance groups. Utterance group information may include one or more of: age (or age range), gender, a first voice feature vector identifying age, or a second voice feature vector identifying gender.

[0179] An account profile may be a profile representing a user account logged into the display device (100). The account profile may include one or more of the account's nickname (or title), age, gender, language, payment information (e.g., credit card information), preferred genre, or preferred content.

[0180] The account profile may be information pre-registered in the account of the display device (100).

[0181] The memory (140) can store multiple account profiles corresponding to each of multiple user accounts.

[0182] The controller (170) can determine whether an account profile matching the speech group information exists among multiple account profiles. For example, if the speech group information indicates an age range of 30s and a gender of female, the controller (170) can extract an account profile matching this.

[0183] If the controller (170) determines that there is an account profile matching the utterance group information among the multiple account profiles (S711), the controller can complete voice registration (S713).

[0184] Completing voice registration may be a process of matching voice features corresponding to the acquired utterance group with the profile of the logged-in account and storing them in memory (140).

[0185] In one embodiment, if the controller (170) determines that there is an account profile matching the utterance group information among the multiple account profiles, the controller (170) may display a voice registration completion notification indicating that voice registration is complete on the display (180).

[0186] In another embodiment, the controller (170) may complete voice registration if the account profile corresponding to the account currently logged into the display device (100) and the utterance group information corresponding to the acquired utterance group match.

[0187] This is because, if voice registration is performed even when the user of the account currently logged into the display device (100) and the speaker do not match, the speaker may use the information of the logged-in account, which may cause a security problem.

[0188] FIG. 10 is a drawing illustrating an example of displaying a notification of completion of voice registration according to one embodiment of the present disclosure.

[0189] Referring to FIG. 10, the controller (170) of the display device (100) can display a voice registration completion notification (1000) on the display (180) if there is an account profile matching the acquired speech group information.

[0190] Speakers can automatically perform voiceprint registration by simply uttering "yes," as shown in Figure 9, without the hassle of having to utter multiple sentences. This significantly improves the convenience of voiceprint registration.

[0191] Again, Figure 7 is explained.

[0192] If the controller (170) determines that there is no account profile matching the ignition group information among the multiple account profiles, it can display a notification for new account registration on the display (180) (S715).

[0193] If the controller (170) determines that there is no account profile matching the ignition group information among the multiple account profiles, it may display a notification for new account registration on the display (180).

[0194] A notification for a new account sign-up may be a notification to match the newly entered account profile to the speaker's speaking group information.

[0195] FIGS. 11A to 11D are drawings for explaining embodiments related to the present disclosure.

[0196] In FIGS. 11a to 11d, it is assumed that multiple ignition groups (810, 830, 850) are created in advance, as in FIG. 8.

[0197] Figure 11a is an example of an embodiment in which voice registration is automatically completed when the profile of a logged-in account and the speech group information match.

[0198] Referring to FIG. 11a, the display device (100) can display a login screen (1100) on the display (180). The login screen (1100) can include first and second account items (1110, 1120) and an account addition item (1130).

[0199] Each of the first and second account items (1110, 1120) may be an item representing a pre-registered account for the service of the display device (100). Each of the first and second account items (1110, 1120) may include one or more of an account nickname or an account icon identifying the account.

[0200] The Add Account item (1130) may be an item for creating a new account.

[0201] When the first account item (1110) is selected, the display device (100) can log in with the first account corresponding to the first account item (1110). Upon performing the log in, the display device (100) can display a first account icon (1101) on the display (180) indicating that the user is logged in with the first account item (1110).

[0202] The first account icon (1101) may be displayed until logging out. The first account icon (1101) may be displayed at the top of the screen of the display (180), but this is only an example.

[0203] In one embodiment, the display device (100) may display a voice registration window (1140) on the display (180) to inquire about voice registration when one or more utterance groups are created.

[0204] In another embodiment, the display device (100) may display a voice registration window (1140) on the display (180) to inquire about voice registration when the user of the first account is an unregistered voice registration user and one or more speech groups are created.

[0205] The speaker (K) utters the consent voice command “Yes.” The display device (100) can receive the consent voice command and extract a voice feature vector representing voice characteristics from the consent voice command.

[0206] The display device (100) can determine whether there is a utterance group that matches the extracted voice feature vector among one or more utterance groups.

[0207] In one embodiment, the display device (100) can extract a group of utterances from a voice feature vector using an artificial neural network-based matching model stored in a memory (140).

[0208] The matching model may be an artificial intelligence model trained through supervised learning. The training data set for training the matching model may include training speech feature vectors and utterance group labels that identify utterance groups.

[0209] The matching model can be trained to minimize a loss function representing the difference between the target feature vector inferred from the learning speech feature vector and the utterance group label.

[0210] The display device (100) can output a voice registration completion notification (1150) when there is a utterance group that matches the extracted voice feature vector among one or more utterance groups, and the utterance group information corresponding to the matching utterance group matches the first account profile of the first account item (1110).

[0211] In one embodiment, a voiceprint recognition service may be provided through voiceprint registration. The voiceprint recognition service may be a service that stores the voiceprint of a speaker (K), compares the voiceprint with the characteristics of the voice spoken by the speaker (K), and performs a function of responding to the voice only when the characteristics of the voice and the voiceprint match.

[0212] Voice registration may be the process of matching a voiceprint with an account's profile.

[0213] Figure 11b is an example of an embodiment that outputs a notification indicating that voice registration is not possible when the speech group information corresponding to the voice of the speaker (K) does not match the logged-in account.

[0214] Referring to Figure 11b, the speaker (K) can log in through a second account entry (1120). The user of the actual second account entry (1120) and the speaker (K) may be different individuals. For example, the second account entry (1120) may correspond to the father's account, and the speaker (K) may be the son.

[0215] The display device (100) may display a second account icon (1103) on the display (180) indicating that the user is logged in as a second account item (1120) according to the login operation.

[0216] In one embodiment, the display device (100) may display a voice registration window (1140) on the display (180) to inquire about voice registration when one or more utterance groups are created.

[0217] In another embodiment, the display device (100) may display a voice registration window (1140) on the display (180) to inquire about voice registration when the user of the second account is an unregistered voice registration user and one or more speech groups have been created.

[0218] The speaker (K) utters the consent voice command “Yes.” The display device (100) can receive the consent voice command and extract a voice feature vector representing voice characteristics from the consent voice command.

[0219] The display device (100) can determine whether there is a utterance group that matches the extracted voice feature vector among one or more utterance groups.

[0220] The display device (100) can determine whether there is an utterance group that matches the extracted voice feature vector among one or more utterance groups, and whether utterance group information corresponding to the matching utterance group matches the second account profile of the second account item (1120).

[0221] The display device (100) can output a voice registration failure notification (1160) when the speech group information corresponding to the voice of the speaker (K) does not match the second account profile of the second account item (1120).

[0222] The voiceprint registration failure notification (1160) may be a notification indicating that the speaker and the logged-in account do not match, and that voiceprint registration is only possible through the speaker's own account.

[0223] If voice registration is completed even when the second account profile of the second account item (1120) and the speaker's (K) speech group information do not match, a security issue arises in that personal information (payment information) corresponding to the second account item (1120) can be used as the speaker's (K) voice.

[0224] In an embodiment of the present disclosure, if the speech group information matching the speaker's voice does not match the logged-in account profile, voiceprint registration may not be permitted to prevent misuse of personal information.

[0225] In another embodiment, if the speech group information corresponding to the voice of the speaker (K) does not match each of the account profiles stored in the memory (140), the display device (100) may display a new account subscription notification (1170) on the display (180) to induce new account subscription, as in FIG. 11c.

[0226] A new subscription notification (1170) may be displayed if there is no account profile matching the speaker's (K) speaking group information.

[0227] When the display device (100) receives a voice command of consent from the speaker (K) while displaying a new account registration notification (1170), the display device (100) can display a new registration window (1180) on the display (180).

[0228] The new sign-up window (1180) may be a window for entering an account profile including one or more of the account's nickname, age, gender, or payment information.

[0229] When an account profile is entered through a new subscription window (1180), the display device (100) can match the account profile and the speaker's (K) speech group information and store them in the memory (140).

[0230] FIG. 11d is an embodiment of displaying a selection guidance notification (1190) that guides selection of one of the multiple account profiles when there are multiple account profiles matching the speech group information corresponding to the voice of the speaker (K).

[0231] Referring to FIG. 11d, the display device (100) can display a selection prompt notification (1190) on the display (180) that prompts selection of one of a plurality of account profiles when the speaker's speech group information matches the first account item (1110) and the second account item (1120).

[0232] The display device (100) can identify each of the first account item (1110) and the second account item (1120) matching the acquired utterance group through a highlight box.

[0233] The display device (100) can match and store the utterance group information and the account profile of the selected item when either one of the first account item (1110) and the second account item (1120) is selected.

[0234] After that, the display device (100) can display a voice registration completion notification (1191) indicating that voice registration has been completed on the display (180).

[0235] FIG. 12 is a sequence diagram for explaining an operation method of a system according to an embodiment of the present disclosure.

[0236] A system according to one embodiment of the present disclosure may include a display device (100) and an AI server (500). A system according to another embodiment of the present disclosure may include a display device (100) and a speaker recognition server (600).

[0237] That is, the server that is linked to the display device (100) in FIG. 12 may be either an AI server (500) or a speaker recognition server (600).

[0238] In the following, it is assumed that the server that is linked to the display device (100) is an AI server (500), but the operations performed by the AI ​​server (500) may be performed instead by the speaker recognition server (600).

[0239] Referring to FIG. 12, the controller (170) of the display device (100) can collect voice data (S1201) and transmit the collected voice data to the AI ​​server (500) via the network interface (133) (S1203).

[0240] The controller (170) can receive voices spoken by one or more speakers from a remote control device (200) or voices spoken by one or more speakers through a microphone (not shown) provided in the display device (100).

[0241] The processor (560) of the AI ​​server (500) can generate one or more speech groups based on the received voice data (S1205).

[0242] The processor (560) can generate one or more utterance groups based on a voice data set corresponding to the received voices.

[0243] The processor (560) can obtain an embedding vector (voice feature vector) representing voice features from voice data. The controller (170) can extract the embedding vector from the voice data using the MFCC (Mel-Frequency Cepstral Coefficients) technique.

[0244] The processor (560) can cluster a plurality of embedding vectors and generate one or more utterance groups based on the clustering results.

[0245] The processor (560) can cluster a plurality of embedding vectors using the K-means clustering technique or the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) technique to create one or more utterance groups.

[0246] The description of the process of creating a utterance group is replaced with the example of Fig. 8.

[0247] The processor (560) of the AI ​​server (500) can transmit information about one or more speech groups generated through the communication unit (510) to the display device (100) (S1207).

[0248] Utterance group information may be an item identifying an utterance group. Utterance group information may be items that serve as criteria for clustering utterance groups. Utterance group information may include one or more of age (or age range), gender, a voice feature vector identifying age, or a voice feature vector identifying gender.

[0249] The processor (560) can transmit utterance group information of each of one or more utterance groups to the display device (100) via the communication unit (510).

[0250] The controller (170) of the display device (100) can display a voice registration window on the display (180) upon receiving information about one or more speech groups (S1209).

[0251] In one embodiment, when the controller (170) receives information about one or more utterance groups, it may display a voice registration window for voice registration on the display (180).

[0252] In another embodiment, the controller (170) may recognize the voice of a speaker who has not registered his / her voiceprint and, if information about one or more utterance groups is received, display a voiceprint registration window on the display (180).

[0253] In another embodiment, when the speaker performs voiceprint registration through a voiceprint registration application installed on the display device (100), the display of the voiceprint registration window may be omitted.

[0254] The controller (170) of the display device (100) can receive a voice command (S1211) and transmit the received voice command to the AI ​​server (500) via the network interface (133) (S1213).

[0255] The controller (170) can receive a voice command agreeing to register the voice signature and transmit the received voice command to the AI ​​server (500) via the network interface (133).

[0256] The processor (560) of the AI ​​server (500) can obtain a utterance group that matches the received voice command among one or more utterance groups (S1215).

[0257] The processor (560) can extract a voice feature vector representing a voice characteristic from a voice command. The processor (560) can obtain an utterance group that matches the extracted voice feature vector from one or more utterance groups. The processor (560) can obtain an utterance group with the closest distance between voice feature vectors from one or more utterance groups as the matching utterance group.

[0258] The processor (560) of the AI ​​server (500) can compare the acquired utterance group information with multiple account profiles (S1217).

[0259] The processor (560) may store multiple account profiles in advance in the memory (530).

[0260] If the processor (560) of the AI ​​server (500) determines that there is an account profile matching the utterance group information among the multiple account profiles (S1219), the processor can match the account profile and the utterance group information and store them in the memory (530) (S1221).

[0261] The processor (560) can complete voice registration by matching the account profile and the speech group information and storing them in the memory (530).

[0262] After that, the processor (560) of the AI ​​server (500) can transmit a message indicating completion of the voice registration to the display device (100) through the communication unit (510) (S1223).

[0263] The controller (170) of the display device (100) can display a voice registration completion notification on the display (180) based on a message indicating the completion of voice registration (S1225).

[0264] The notification of completion of voice registration is as in the embodiment of Fig. 11a.

[0265] If the processor (560) of the AI ​​server (500) determines that there is no account profile matching the speech group information among the multiple account profiles (S1219), it can transmit a message indicating the completion of voice registration to the display device (100) through the communication unit (510) (S1227).

[0266] The controller (170) of the display device (100) can display a new account registration notification on the display (180) based on a message indicating completion of non-registration of the voice character (S1229).

[0267] The new account sign-up notification is as in the embodiment of Fig. 11c.

[0268] Figure 13 is a diagram illustrating the process of updating the utterance group and voiceprint recognition model after voiceprint registration is completed.

[0269] FIG. 13 may be operations performed after step S713 of FIG. 7.

[0270] The controller (170) of the display device (100) can collect the speaker's voice data after the voice registration is completed (S1301).

[0271] The controller (170) can continuously collect voice data of the speaker through a microphone provided in a remote control device (200) or a display device (100).

[0272] The controller (170) can update the speech group and voice recognition model based on the collected voice data (S1303).

[0273] The controller (170) can extract speech features from the speaker's speech data and create a new utterance group corresponding to the extracted speech features. The controller (170) can compare the utterance group corresponding to the existing speaker with the new utterance group. The controller (170) can measure the similarity between the utterance group corresponding to the existing speaker and the new utterance group, and if the similarity is above a certain similarity, the new utterance group can be merged with the existing utterance group.

[0274] The controller (170) can measure the similarity between a utterance group corresponding to an existing speaker and a new utterance group, and if the similarity is less than a certain similarity, the new utterance group can be added separately from the existing utterance group.

[0275] A voiceprint recognition model may be a model that recognizes voiceprints by comparing voice features of voice data with previously stored voice features. The voiceprint recognition model may be stored in memory (140).

[0276] The voice recognition model can be an artificial neural network-based model trained using a Recurrent Neural Network (RNN).

[0277] The controller (170) can update a voice recognition model using the voice features of the collected voice data. The controller (170) can retrain the voice recognition model using the collected voice features. Through this process, the voice recognition model can be updated, and the voice recognition rate for the speaker can be improved.

[0278] An electronic device (100) according to an embodiment of the present disclosure may include a memory (140), a controller (170) that generates one or more speech groups using voice data collected from one or more users, receives a voice command, obtains a speech group matching the voice command among the one or more speech groups, and, if an account profile matching speech group information of the obtained speech group exists among a plurality of account profiles stored in the memory, matches the account profile with the speech group information and stores the matched account profile in the memory.

[0279] The above controller (170) can use the collected voice data to create one or more speech groups according to age and gender.

[0280] The above utterance group information may include at least one of the age, the gender, a first voice feature vector corresponding to the age, or a second voice feature vector corresponding to the gender.

[0281] The electronic device (100) may further include a display (180), and the controller (170)

[0282] When the voice of an unregistered speaker is recognized and one or more utterance groups are created, a voice registration notification for voice registration can be displayed on the display (180).

[0283] The electronic device (100) may further include a display (180), and the controller (170) may display a voice registration completion notification indicating that voice registration has been completed on the display (180).

[0284] The electronic device (100) may further include a display (180), and the controller (170) may display a voice registration failure notification on the display (180) when the account profile of the account logged into the electronic device and the speech group information do not match.

[0285] The electronic device (100) may further include a display (180), and the controller (170) may display a new account sign-up notification for new account sign-up on the display (180) when the account profile of the account logged into the electronic device (100) and the utterance group information do not match.

[0286] The above account profile may include age and gender corresponding to the above account.

[0287] According to one embodiment of the present disclosure, the above-described method can be implemented as processor-readable code on a medium in which a program is recorded. Examples of processor-readable media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices.

[0288] The display device described above is not limited to the configuration and method of the embodiments described above, and the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.

Claims

In electronic devices, memory; A controller that generates one or more speech groups using voice data collected from one or more users, receives a voice command, obtains a speech group matching the voice command among the one or more speech groups, and, if an account profile matching the speech group information of the obtained speech group exists among a plurality of account profiles stored in the memory, matches the account profile and the speech group information and stores the matched account profile in the memory. Electronic devices. In the first paragraph, The above controller Using the collected voice data, one or more speech groups are created based on age and gender. Electronic devices. In the second paragraph, The above utterance group information is At least one of the age, the gender, the first voice feature vector corresponding to the age, or the second voice feature vector corresponding to the gender Electronic devices. In the first paragraph, Including more displays, The above controller When the voice of an unregistered speaker is recognized and one or more utterance groups are created, a voice registration notification for voice registration is displayed on the display. Electronic devices. In the first paragraph, Including more displays, The above controller A voice registration completion notification is displayed on the display indicating that voice registration has been completed. Electronic devices. In the first paragraph, Including more displays, The above controller If the account profile of the account logged in with the above electronic device and the above utterance group information do not match, a voice registration failure notification is displayed on the above display. Electronic devices. In the first paragraph, Including more displays, The above controller If the account profile of the account logged in with the above electronic device does not match the above ignition group information, a new account registration notification for new account registration is displayed on the above display. Electronic devices. In the first paragraph, The above account profile is Including the age and gender corresponding to the above account Electronic devices. In a method of operating an electronic device, A step of generating one or more groups of utterances using voice data collected from one or more users; Step of receiving a voice command; A step of obtaining an utterance group matching the voice command among the one or more utterance groups; and If there is an account profile that matches the utterance group information of the acquired utterance group among multiple account profiles, a step of matching and storing the account profile and the utterance group information is included. How electronic devices work. In paragraph 9, The above generating steps are A step of generating one or more speech groups according to age and gender using the collected voice data. How electronic devices work. In Article 10, The above utterance group information is At least one of the age, the gender, the first voice feature vector corresponding to the age, or the second voice feature vector corresponding to the gender How electronic devices work. In paragraph 9, If the voice of an unregistered speaker is recognized and one or more utterance groups are generated, the step of displaying a voice registration notification for voice registration is further included. How electronic devices work. In paragraph 9, Further comprising the step of displaying a voice registration completion notification indicating that the voice registration has been completed. How electronic devices work. In the first paragraph, If the account profile of the account logged in with the electronic device and the speech group information do not match, the step of displaying a notification of voice registration failure is further included. How electronic devices work. In paragraph 9, If the account profile of the account logged in with the electronic device does not match the ignition group information, the method further includes a step of displaying a new account registration notification for new account registration. How electronic devices work.

Citation Information

Patent Citations

  • Method and apparatus for outputting information

    JP2019091419A

  • Method for adding account, terminal, server, and computer storage medium

    KR1020170139650A

  • Smart movement system for the elder

    KR1020210010779A

  • Tire vulcanization device capable of two-step demolding and tire vulcanization method using the same

    KR102463551B1

  • User account matching based on a natural language utterance

    US20210224367A1