Display device and operation method thereof
By receiving and displaying personalized prompts through a display device, and collecting and transmitting user voice data, the problem of low voice recognition rate in existing technologies is solved, and the accuracy of recognizing the latest content and dialect words is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-05
- Publication Date
- 2026-04-07
AI Technical Summary
The prompts provided during the current speaker registration process lack personalization, resulting in low speech recognition rates, especially for the latest content and dialect words.
The display device receives personalized prompts from the AI server via a network interface, including the names of the latest content and dialect words. It displays these prompts and collects the user's voice data, which is then transmitted to the speaker recognition server and the STT server for feature matching and model training.
It improves speaker recognition and speech recognition rates, especially for the latest content with insufficient data and the accuracy of recognizing dialect words.
Smart Images

Figure CN121816613A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a display device that collects a user's spoken voice using statements for speaker registration. Background Technology
[0002] Speaker recognition services refer to services that use speech recognition technology to identify and distinguish specific speakers. Speaker recognition services can be used in various application areas, including security, education, voice commands, and automation systems.
[0003] Speaker recognition services can analyze the characteristics of a speaker's speech, store each speaker's unique speech patterns, and identify the speaker based on the stored speech patterns.
[0004] For speaker recognition services, a speaker registration process is required to register the speaker's voice.
[0005] Prompts are provided during the speaker registration process to improve the accuracy of speech recognition.
[0006] The speaker utters a prompt, and the voice recognition device, such as a mobile device or TV, identifies the speaker by recognizing the uttered prompt.
[0007] However, the prompts provided during the previous speaker registration process only offered trigger instructions or common words, resulting in monotonous speech recognition that did not take into account the user's vocal characteristics (dialect, tone of voice).
[0008] In addition, there is a problem with the low speech recognition rate for words such as the names of the latest content. Summary of the Invention
[0009] The problem the invention aims to solve
[0010] The purpose of this invention is to compensate for speaker recognition rate and speech recognition rate by using prompts during the user's speaker registration process.
[0011] The purpose of this invention is to extract the speech features of prompts during the speaker registration process and provide personalized services to the speaker.
[0012] means for solving problems
[0013] The display device of this invention includes: a display; a network interface for communicating with one or more servers; and a controller that obtains a speaker registration statement including a prompt for speaker registration, displays the generated speaker registration statement on the display when a request for speaker registration is received, obtains a voice command corresponding to the displayed speaker registration statement, and transmits voice data corresponding to the obtained voice command to the one or more servers through the network interface; the prompt may include one of the following: the name of the latest content and a word representing a regional dialect.
[0014] The operation method of the display device according to an embodiment of the present invention includes: obtaining a speaker registration statement including a prompt for speaker registration; displaying the generated speaker registration statement upon receiving a request for speaker registration; obtaining a voice instruction corresponding to the displayed speaker registration statement; and transmitting voice data corresponding to the obtained voice instruction to one or more servers; wherein the prompt may include one of the name of the latest content and a word representing a regional dialect.
[0015] Invention Effects
[0016] According to embodiments of the present invention, by collecting speech data on the names of recent content with limited data, words that conform to dialects, and words with low recognition rates, it is possible to improve speaker recognition rate and the conversion accuracy of the STT model.
[0017] According to embodiments of the present invention, by extracting features of the user's voice, it is possible to provide an optimized voice recognition service to an individual. Attached Figure Description
[0018] Figure 1 The block diagram illustrates the configuration of a display device according to an embodiment of the present invention.
[0019] Figure 2 This is a block diagram of a remote control device according to an embodiment of the present invention.
[0020] Figure 3 This illustrates an actual configuration example of a remote control device according to an embodiment of the present invention.
[0021] Figure 4 An example of applying a remote control device according to an embodiment of the present invention is shown.
[0022] Figure 5 An artificial intelligence (AI) server according to an embodiment of the present invention is shown.
[0023] Figure 6This is a ladder diagram illustrating the action method of an AI system according to an embodiment of the present invention.
[0024] Figure 7 This is a ladder diagram illustrating the action method of an AI system according to an embodiment of the present invention.
[0025] Figures 8a to 8c This is a diagram illustrating examples of various forms of speaker registration statements according to embodiments of the present invention.
[0026] Figure 9 This is a diagram illustrating the process of collecting voice data for prompts according to an embodiment of the present invention.
[0027] Figure 10 This is a flowchart illustrating the operation method of a display device according to an embodiment of the present invention. Detailed Implementation
[0028] Hereinafter, embodiments relating to the present invention will be described in more detail with reference to the accompanying drawings. The suffixes “module” and “part” used in the following description are assigned or used interchangeably for ease of writing only, and do not inherently have a distinguishing meaning or function.
[0029] The display device of this invention is, for example, an intelligent display device that adds computer support to the broadcast receiving function, or adds internet functionality while remaining faithful to the broadcast receiving function. It can have more convenient interfaces such as a handwriting input device, a touchscreen, or a spatial remote control. Furthermore, with support for wired or wireless internet access, it can connect to the internet and a computer, and perform functions such as email, browser, banking, or games. For these various functions, a standardized general-purpose operating system (OS) can be used.
[0030] Therefore, in the display device described in this invention, various applications can be freely added or removed, for example, on a general-purpose OS kernel, thereby enabling the execution of various user-friendly functions. More specifically, the display device may be, for example, a network television (TV), a hybrid broadcast broadband television (HBBTV), a smart TV, an LED TV, an OLED TV, etc., and may also be applied to smartphones, depending on the circumstances.
[0031] Figure 1 The block diagram illustrates the configuration of a display device according to an embodiment of the present invention.
[0032] Reference Figure 1The display device 100 may include a broadcast receiver 130, an external device interface 135, a memory 140, a user input interface 150, a controller 170, a wireless communication interface 173, a display 180, a speaker 185, and a power supply circuit 190.
[0033] The broadcast receiver 130 may include a tuner 131, a demodulator 132, and a network interface 133.
[0034] Tuner 131 can select a specific broadcast channel according to a channel selection command. Tuner 131 can receive broadcast signals about the selected specific broadcast channel.
[0035] The demodulator 132 can separate the received broadcast signal into video signal, audio signal, and data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into an output form.
[0036] The external device interface 135 can receive applications or application directories from adjacent external devices and pass them to the controller 170 or the memory 140.
[0037] External device interface 135 provides a connection path between display device 100 and external devices. External device interface 135 can receive one or more video or audio signals output from an external device connected to display device 100 wirelessly or via a wired connection and transmit them to controller 170. External device interface 135 may include a plurality of external input terminals. The plurality of external input terminals may include RGB terminals, one or more High Definition Multimedia Interface (HDMI) terminals, and component video terminals.
[0038] Image signals from external devices input via external device interface 135 can be output via display 180. Audio signals from external devices input via external device interface 135 can be output via speaker 185.
[0039] The external device that can be connected to the external device interface 135 can be any of the following: set-top box, Blu-ray player, DVD player, game console, soundbar, smartphone, PC, USB storage device, home theater system, but this is just one example.
[0040] Network interface 133 can provide an interface for connecting the display device 100 to a wired / wireless network, including the Internet. Network interface 133 can send or receive data with other users or other electronic devices via the accessed network or another network linked to the accessed network.
[0041] Additionally, a portion of the content data stored on the display device 100 can be sent to other users pre-registered with the display device 100 or to selected users or selected electronic devices in other electronic devices.
[0042] Network interface 133 can access designated web pages through the connected network or another network connected to the connected network. That is, it can access designated web pages through the network and send or receive data with the corresponding server.
[0043] Furthermore, network interface 133 can receive content or data provided by content providers or network operators. That is, network interface 133 can receive content such as movies, advertisements, games, video on demand (VOD), and broadcast signals, as well as related information, provided by content providers or network providers via the network.
[0044] In addition, network interface 133 can receive firmware update information and update files provided by network operators, and can send data to the Internet, content providers, or network operators.
[0045] Network interface 133 can select and receive desired applications from publicly available (open) applications via the network.
[0046] The memory 140 can store programs for various signal processing and control within the controller 170, and store processed image, voice, or data signals.
[0047] In addition, the memory 140 can also perform the function of temporarily storing image, voice or data signals input from the external device interface 135 or the network interface 133, and can also store information related to a specified image through the channel memory function.
[0048] The memory 140 can store applications or application directories input from the external device interface 135 or the network interface 133.
[0049] The display device 100 can play content files (video files, still image files, music files, text files, application files, etc.) stored in the memory 140 and provide them to the user.
[0050] User input interface 150 can transmit user-input signals to controller 170, or transmit signals from controller 170 to user. For example, user input interface 150 can receive and process control signals such as power on / off, channel selection, and screen settings from remote control device 200 via various communication methods such as Bluetooth, UltraWideband, ZigBee, radio frequency (RF) communication, or infrared (IR) communication, or send control signals from controller 170 to remote control device 200.
[0051] In addition, the user input interface 150 can transmit control signals input from local keys (not shown) such as the power button, channel button, volume button, and setting value to the controller 170.
[0052] The image signal processed in the controller 170 is input to the display 180, and can be displayed as an image corresponding to the image signal. In addition, the image signal processed in the controller 170 can be input to an external output device through the external device interface 135.
[0053] The voice signal processed in controller 170 can be output via speaker 185. Alternatively, the voice signal processed in controller 170 can be input to an external output device via external device interface 135.
[0054] In addition, the controller 170 can control the overall operation within the display device 100.
[0055] In addition, the controller 170 can control the display device 100 using user commands or internal programs input through the user input interface 150, and can download the user's required applications or application directories to the display device 100 by accessing the network.
[0056] The controller 170 enables the user-selected channel information, along with the processed video or audio signals, to be output via the display 180 or the speaker 185.
[0057] Additionally, the controller 170 can output image or audio signals from external devices, such as cameras or camcorders, input through the external device interface 135 via the display 180 or speaker 185, based on the external device image playback command received through the user input interface 150.
[0058] On the other hand, the controller 170 can control the display 180 to display images, such as broadcast images input via the tuner 131, externally input images input via the external device interface 135, images input via the network interface, or images stored in the memory 140. In this case, the images displayed on the display 180 can be still images or moving images, and can be 2D images or 3D images.
[0059] In addition, the controller 170 can control the playback of content stored in the display device 100, or received broadcast content, or externally input content, which can be various forms such as broadcast images, externally input images, audio files, still images, accessed web pages, and text files.
[0060] The wireless communication interface 173 can communicate with external devices via wired or wireless communication. The wireless communication interface 173 can perform short-range communication with external devices. For this purpose, the wireless communication interface 173 can utilize at least one of the following technologies to support short-range communication: Bluetooth™, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wideband (UWB), ZigBee, Near Field Communication (NFC), Wi-Fi, Wi-Fi Direct, and Wireless Universal Serial Bus. Such a wireless communication interface 173 can support wireless communication between the display device 100 and the wireless communication system, between the display device 100 and other display devices, or between the network where the display device 100 (or an external server) resides, via wireless area networks. Short-range wireless communication networks can be called wireless personal area networks (PANs).
[0061] Here, other display devices 100 may be mobile terminals capable of exchanging (or being linked with) data with the display device 100 of the present invention, such as wearable devices (e.g., smartwatches, smart glasses, head-mounted displays, smartphones, etc.). The wireless communication interface 173 can detect (or identify) wearable devices capable of communication around the display device 100.
[0062] Furthermore, if the detected wearable device is certified to communicate with the display device 100 of the present invention, the controller 170 can transmit at least a portion of the data processed in the display device 100 to the wearable device via the wireless communication interface 173. Therefore, the user of the wearable device can utilize the data processed in the display device 100 through the wearable device.
[0063] The display 180 can convert the image signals, data signals, OSD signals processed in the controller 170 or the image signals and data signals received from the external device interface 135 into R, G, and B signals respectively to generate drive signals.
[0064] on the other hand, Figure 1 The display device 100 shown is merely one embodiment of the present invention. Some of the constituent elements shown may be integrated, added, or omitted according to the specifications of the actual implemented display device 100.
[0065] That is, as needed, two or more constituent elements can be combined into one constituent element, or one constituent element can be subdivided into two or more constituent elements. Furthermore, the functions performed in each module are only for illustrating embodiments of the present invention, and their specific actions or devices do not limit the scope of the present invention.
[0066] According to another embodiment of the present invention, the display device 100 and Figure 1 Unlike the one shown, it does not have tuner 131 and demodulator 132, and can receive and play images via network interface 133 or external device interface 135.
[0067] For example, the display device 100 can be separated into an image processing device, such as a set-top box, for receiving broadcast signals or content from various network services, and a content playback device for playing content input from the image processing device.
[0068] In this case, the operation method of the display device of the embodiments of the present invention described below can be derived not only from, but also from, as referred to Figure 1The operation can be performed by the display device 100 described herein, or by any of the separate set-top box or other image processing device, or by a content playback device having a display 180 and an audio output unit 185.
[0069] Next, refer to Figures 2 to 3 This invention describes a remote control device according to an embodiment of the present invention.
[0070] Figure 2 This is a block diagram of a remote control device according to an embodiment of the present invention. Figure 3 This illustrates an actual configuration example of a remote control device 200 according to an embodiment of the present invention.
[0071] First, refer to Figure 2 The remote control device 200 may include a fingerprint reader 210, a wireless communication circuit 220, a user input interface 230, a sensor 240, an output interface 250, a power supply circuit 260, a memory 270, a controller 280, and a microphone 290.
[0072] Reference Figure 2 The wireless communication circuit 220 transmits and receives signals from any one of the display devices in the aforementioned embodiments of the present invention.
[0073] The remote control device 200 may have an RF circuit 221 capable of transmitting and receiving signals with the display device 100 according to an RF communication standard, and an IR circuit 223 capable of transmitting and receiving signals with the display device 100 according to an IR communication standard. Additionally, the remote control device 200 may have a Bluetooth circuit 225 capable of transmitting and receiving signals with the display device 100 according to a Bluetooth communication standard. Furthermore, the remote control device 200 may have an NFC circuit 227 capable of transmitting and receiving signals with the display device 100 according to a Near Field Communication (NFC) standard, and a WLAN circuit 229 capable of transmitting and receiving signals with the display device 100 according to a Wireless Local Area Network (WLAN) standard.
[0074] In addition, the remote control device 200 transmits signals including information related to the movement of the remote control device 200 to the display device 100 via the wireless communication circuit 220.
[0075] On the other hand, the remote control device 200 can receive signals transmitted by the display device 100 through the RF circuit 221, and can transmit commands to the display device 100 regarding power on / off, channel change, volume change, etc., through the IR circuit 223 as needed.
[0076] The user input interface 230 may consist of a key group, buttons, a touchpad, or a touchscreen. The user can input commands related to the display device 100 to the remote control device 200 by operating the user input interface 230. If the user input interface 230 has physical buttons, the user can input commands related to the display device 100 to the remote control device 200 by pressing the physical buttons. For this, refer to... Figure 3 Please provide an explanation.
[0077] Reference Figure 3 The remote control device 200 may include a plurality of buttons. The plurality of buttons may include a fingerprint recognition button 212, a power button 231, a home button 232, a live button 233, an external input button 234, a volume control button 235, a voice recognition button 236, a channel change button 237, an confirmation button 238, and a back button 239.
[0078] The fingerprint recognition button 212 can be a button used to recognize a user's fingerprint. As one embodiment, since the fingerprint recognition button 212 can be pressed, it can also receive both pressing and fingerprint recognition actions.
[0079] The power button 231 can be a button used to turn the power to / from the display device 100 on or off.
[0080] The home button 232 can be a button used to move to the home screen of the display device 100.
[0081] The live broadcast button 233 can be used to display live broadcast programs.
[0082] External input key 234 can be a key for receiving external input connected to display device 100.
[0083] The volume control key 235 can be used to adjust the volume output by the display device 100.
[0084] The voice recognition key 236 can be a key used to receive and recognize the user's voice.
[0085] Channel change key 237 can be a key used to receive broadcast signals from a specific broadcast channel.
[0086] The confirmation key 238 can be used to select a specific function, and the back key 239 can be used to return to the previous screen.
[0087] To reiterate Figure 2 .
[0088] When the user input interface 230 has a touchscreen, the user can input commands related to the display device 100 to the remote control device 200 by touching the soft keys on the touchscreen. Alternatively, the user input interface 230 may have various input devices that the user can operate, such as scroll keys or joysticks; this embodiment does not limit the scope of the invention.
[0089] Sensor 240 may have a gyroscope sensor 241 or an accelerometer sensor 243. The gyroscope sensor 241 can sense motion-related information of the remote control device 200.
[0090] For example, the gyroscope sensor 241 can sense motion-related information of the remote control device 200 with reference to the x, y, and z axes, and the accelerometer sensor 243 can sense information such as the moving speed of the remote control device 200. On the other hand, since the remote control device 200 can also have a distance measurement sensor, it can sense the distance between itself and the display 180 of the display device 100.
[0091] The output interface 250 can output image or voice signals that correspond to the operation of the user input interface 230 or to the signals transmitted from the display device 100.
[0092] Users can identify whether the user input interface 230 is being operated or whether the display device 100 is being controlled through the output interface 250.
[0093] For example, the output interface 250 may have an LED 251 that is lit when the user input interface 230 is operated or when the wireless communication unit 225 sends and receives signals with the display device 100, or a vibrator 253 that generates vibration, or a speaker 255 that outputs sound, or a display 257 that outputs images.
[0094] In addition, the power supply circuit 260 supplies power to the remote control device 200, and can stop the power supply if the remote control device 200 does not move during a specified period of time, thereby reducing power waste.
[0095] The power supply circuit 260 can restore power when a designated button on the remote control device 200 is operated.
[0096] The memory 270 can store various programs, application data, etc. required for the control or operation of the remote control device 200.
[0097] When the remote control device 200 wirelessly transmits and receives signals with the display device 100 via the RF circuit 221, the remote control device 200 and the display device 100 transmit and receive signals via a specified frequency band.
[0098] The controller 280 of the remote control device 200 can store and refer to information in the memory 270, such as information about the frequency band that enables wireless transmission and reception of signals with the display device 100 paired with the remote control device 200.
[0099] The controller 280 controls all matters related to the control of the remote control device 200. The controller 280 can transmit signals corresponding to specified key operations of the user input interface 230 or signals corresponding to the movement of the remote control device 200 sensed in the sensor 240 to the display device 100 via the wireless communication unit 225.
[0100] In addition, the microphone 290 of the remote control device 200 can acquire voice.
[0101] Microphone 290 can have a plurality of microphones.
[0102] Next, let's explain. Figure 4 .
[0103] Figure 4 An example of applying a remote control device according to an embodiment of the present invention is shown.
[0104] Figure 4 Example (a) shows a pointer 205 corresponding to the remote control device 200 displayed on the display 180.
[0105] Users can move or rotate the remote control device 200 up and down, left and right. The pointer 205 displayed on the display 180 of the display device 100 corresponds to the movement of the remote control device 200. As shown in the figure, in such a remote control device 200, the pointer 205 moves with the movement of the remote control device 200 in 3D space, and therefore can be named a spatial remote control.
[0106] Figure 4 As illustrated in (b), if the user moves the remote control device 200 to the left, the pointer 205 displayed on the display 180 of the display device 100 also moves to the left accordingly.
[0107] The movement-related information of the remote control device 200 detected by the sensors of the remote control device 200 is transmitted to the display device 100. The display device 100 can calculate the coordinates of the pointer 205 based on the movement-related information of the remote control device 200. The display device 100 can display the pointer 205 in a manner corresponding to the calculated coordinates.
[0108] Figure 4Example (c) illustrates a situation where, with a specific button within the remote control device 200 pressed, the user moves the remote control device 200 away from the display 180. As a result, the selection area within the display 180 corresponding to the pointer 205 can be magnified and expanded.
[0109] Conversely, when the user moves the remote control device 200 closer to the display 180, the selection area within the display 180 corresponding to the pointer 205 can be reduced and displayed in a smaller size.
[0110] On the other hand, it is also possible to select a smaller area when the remote control device 200 is far away from the display 180, and to select a larger area when the remote control device 200 is close to the display 180.
[0111] Furthermore, when a specific button within the remote control device 200 is pressed, the recognition of up / down and left / right movements can be excluded. That is, when the remote control device 200 is moved away from or closer to the display 180, up / down and left / right movements are not recognized, and only forward / backward movements are recognized. When the specific button within the remote control device 200 is not pressed, only the pointer 205 moves with the up / down and left / right movements of the remote control device 200.
[0112] On the other hand, the moving speed or direction of the pointer 205 can correspond to the moving speed or direction of the remote control device 200.
[0113] On the other hand, the pointer in this specification refers to an object displayed on the display 180 corresponding to the action of the remote control device 200. Therefore, the pointer 205 can be an object of various shapes, in addition to the arrow shape shown in the figure. For example, it can be a concept including a point, cursor, prompt, thick outline, etc. Moreover, the pointer 205 can not only be displayed corresponding to a single position on the horizontal and vertical axes of the display 180, but also corresponding to multiple positions such as a line or surface.
[0114] Figure 5 An artificial intelligence (AI) server according to an embodiment of the present invention is shown.
[0115] Reference Figure 5 AI server 500 can refer to a device that uses machine learning algorithms to train artificial neural networks or uses pre-trained artificial neural networks.
[0116] AI Server 500 can consist of multiple servers to perform distributed processing and can be defined as a 5G network.
[0117] AI server 500 may also be included in AI device 100 as part of it, together performing at least a part of AI processing.
[0118] AI server 500 may include communication unit 510, memory 530, learning processor 540 and processor 560.
[0119] The communication unit 510 can send and receive data with external devices such as the display device 100.
[0120] The memory 530 may include a model storage unit 531.
[0121] The model storage unit 531 can use the learning processor 540 to store the training or already trained model (or artificial neural network) 531a.
[0122] The learning processor 540 can use training data to train the artificial neural network 531a. The trained model can be used in the AI server 500 mounted on the artificial neural network, or it can be used in an external device such as the display device 100.
[0123] The training model can be implemented using hardware, software, or a combination of both. If part or all of the training model is implemented in software, one or more instructions constituting the training model can be stored in memory 530.
[0124] The processor 560 can use a trained model to infer result values from new input data and generate responses or control instructions based on the inferred result values.
[0125] Figure 6 This is a diagram illustrating the configuration of an AI system according to an embodiment of the present invention.
[0126] Reference Figure 6 The AI system 60 may include an AI server 500, a display device 100, a speaker recognition server 610, and a speech-to-text (STT) server 630.
[0127] In one embodiment, the speaker recognition server 610 and the STT server 630 may also be composed of a single server.
[0128] In another embodiment, AI server 500 may include speaker recognition server 610 and STT server 630.
[0129] The components of an AI system can communicate with each other via the internet.
[0130] AI Server 500 can generate more than one prompt for speaker registration.
[0131] The AI server 500 can transmit one or more generated prompts to the display device 100 via the communication unit 510.
[0132] The display device 100 can display the speaker registration statement on the display 180 based on one or more received prompts.
[0133] The display device 100 can receive voice commands corresponding to the speaker's registered statement.
[0134] The display device 100 can transmit the voice data and speaker identification information corresponding to the obtained voice command to the speaker recognition server 610 via the network interface 133.
[0135] The speaker recognition server 610 can match the features of the received voice data with the speaker identification information and store them in the speaker database.
[0136] The display device 100 can transmit voice data corresponding to the obtained voice command to the STT server 630 via the network interface 133.
[0137] The STT server 630 can convert received voice data into text data and update the STT model based on the converted text data.
[0138] Figure 7 This is a ladder diagram illustrating the action method of an AI system according to an embodiment of the present invention.
[0139] The processor 560 of the AI server 500 can generate one or more prompts for speaker registration (S701).
[0140] Speaker registration can be the process of extracting the features of the speech produced by the speaker and matching the extracted speech features with the speaker's identification information.
[0141] If the speaker registration process is completed, the display device 100 extracts the features of the speaker's voice. The speech recognition function can only be performed if there is speaker identification information that matches the extracted features.
[0142] For speaker registration, a speaker registration statement is required, which includes a prompt that will be displayed on display device 100.
[0143] For this purpose, processor 560 can generate more than one prompt for speaker registration.
[0144] In one embodiment, the processor 560 may generate prompts based on data stored in a regional dialect database. The regional dialect database may store words representing dialects used for a language in multiple regions.
[0145] For example, a regional dialect database may include multiple words identified as dialects in a first region and multiple words identified as dialects in a second region.
[0146] The processor 560 can obtain standard words that can be used as dialect pronunciations in the area corresponding to the location information of the display device 100 as prompts.
[0147] The regional dialect database can be included in AI Server 500 or set up separately from AI Server 500.
[0148] In another embodiment, the processor 560 may generate a prompt based on data stored in a recent content database. The recent content database may store information about recent content. This information may include one or more of the content's name, abbreviation, and aliases.
[0149] The latest content database can periodically receive information about content from content providers.
[0150] Processor 560 can extract the name of the latest content from the latest content database and use the extracted name of the latest content as a prompt.
[0151] In another embodiment, the processor 560 can obtain the prompt based on the word recognition rate. The processor 560 can obtain a word as a prompt if the word recognition rate of a specific word is lower than a preset recognition rate.
[0152] Word recognition rate can be expressed as the percentage of a specific word that is spoken that is recognized. Specifically, it can be represented as the ratio of the number of successful recognitions to the total number of times the word is spoken.
[0153] The processor 560 can obtain words with a recognition rate lower than a preset recognition rate from a plurality of words stored in the database as prompts.
[0154] In another embodiment, the processor 560 can obtain prompts from the names of the latest content and words in the dialect representing the region whose word recognition rate is lower than a preset recognition rate.
[0155] The processor 560 of the AI server 500 can transmit one or more generated prompts to the display device 100 via the communication unit 510 (S703).
[0156] In one embodiment, the processor 560 can transmit one or more generated prompts for speaker registration to the display device 100 via the communication unit 510.
[0157] In another embodiment, the processor 560 can generate one or more speaker registration statements including one or more prompts, and can transmit the generated speaker registration statements to the display device 100 via the communication unit 510.
[0158] The speaker registration statement can be a statement that includes a wake word for activating the voice recognition function of the display device 100 and a generated prompt.
[0159] The controller 170 of the display device 100 can display the speaker registration statement on the display 180 based on one or more received prompts (S705).
[0160] The controller 170 can generate one or more speaker registration statements, including one or more received prompts and wake words.
[0161] When a request to register a speaker via voice is received, the controller 170 may display the speaker registration statement on the display 180.
[0162] The controller 170 can receive requests for speaker registration via voice by executing an application that provides voice recognition services.
[0163] The controller 170 of the display device 100 can obtain voice commands corresponding to the speaker's registered statement (S707).
[0164] In one embodiment, the controller 170 can obtain voice commands corresponding to the speaker's registered statement from the remote control device 200 via the user input interface 150. The user speaks the speaker's registered statement through the remote control device 200, and the remote control device 200 can transmit the voice command spoken by the user to the display device 100.
[0165] In another embodiment, the controller 170 may receive voice commands corresponding to the speaker's registered statement via a microphone (not shown) on the display device 200.
[0166] The controller 170 of the display device 100 can send the voice data and speaker identification information corresponding to the obtained voice command to the speaker recognition server 610 (S709) via the network interface 133.
[0167] Speaker-identifying information can include information that allows the speaker to be identified.
[0168] Speaker identification information may include the speaker's name, account, email address, and other information that can identify the speaker.
[0169] The speaker recognition server 610 can match the features of the received speech data with the speaker identification information and store them in the speaker database (S711).
[0170] The speaker recognition server 610 can extract features from speech data. These features may include one or more of the following: pitch, timbre, gender, volume, and frequency band.
[0171] In one embodiment, the speaker recognition server 610 may use Mel-Frequency Cepstral Coefficient (MFCC) technology to extract features from the speech data.
[0172] MFCC technology can transform speech data into Mel frequency scale in the frequency domain, thereby extracting the frequency characteristics of speech.
[0173] In another embodiment, the speaker recognition server 610 may utilize a deep learning-based speech feature extraction model to extract features from the speech data.
[0174] Speech feature extraction models can be models that use neural network architectures such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to extract features from speech data.
[0175] The speaker recognition server 610 can match the features of the extracted speech data with speaker identification information and store them in the speaker database.
[0176] A speaker database can store speaker identification information for each of a plurality of speakers and a set of speech features (including a plurality of speech features) that match the speaker identification information.
[0177] After the speaker registration process, when the speaker speaks, the speaker recognition server 610 can compare the characteristics of the spoken voice signal with the characteristics stored in the speaker database to determine whether the speaker is the same.
[0178] If one or more features of the spoken voice signal match one or more features stored in the database, the speaker recognition server 610 can determine that the speaker is the same.
[0179] In particular, the speaker recognition server 610 can extract and store features of the speech data corresponding to the prompts representing dialects, thereby improving the recognition accuracy of specific speakers.
[0180] On the other hand, the controller 170 of the display device 100 can transmit voice data corresponding to the obtained voice command to the STT server 630 (S713) via the network interface 133.
[0181] The STT server 630 can convert the received voice data into text data (S715) and update the STT model based on the converted text data (S717).
[0182] The STT server 630 can use the converted text data as training data to update the STT model. The STT model can be a model that converts speech data into text data.
[0183] That is, for words containing dialects in the speaker's registered statement, the names of the latest content, or prompts with low word recognition rates, it may be a situation where not much data has been collected.
[0184] The STT server 630 can retrain the STT model using the speaker's speech data in response to prompts collected by the display device 100 during the speaker registration process.
[0185] This can greatly improve the conversion accuracy of the STT model.
[0186] Figures 8a to 8c This is a diagram illustrating examples of speaker registration statements in various forms according to embodiments of the present invention.
[0187] The display device 100 can display a speaker registration screen 800 on the display 180 according to a request from a speaker to register by speaking.
[0188] Reference Figure 8a The display device 100 can display the first speaker registration statement 810, including the first prompt 811, on the speaker registration screen 800.
[0189] The first prompt, 811, can be the name of the latest content.
[0190] The user can issue voice commands to the remote control device 200 corresponding to the first speaker registration statement 810, which includes the first prompt 811.
[0191] The display device 100 can receive voice commands from the remote control device 200 and transmit the voice data corresponding to the received voice commands to the speaker recognition server 610 and the STT server 630.
[0192] Thus, according to embodiments of the present invention, speech data on the names of recent content with limited data can be collected to improve speaker recognition rate and the conversion accuracy of the STT model.
[0193] Reference Figure 8b The display device 100 can display a second speaker registration statement 830, including a second prompt 831, on the speaker registration screen 800.
[0194] The second prompt 831 can be a standard word that can be pronounced in a dialect. That is, the second prompt 831 can be a word pronounced differently in different regions of a country. The second prompt 831 can be a word that can be recognized as a dialect in the regional information of the display device 100.
[0195] The user can issue voice commands to the remote control device 200 corresponding to the second speaker registration statement 830, which includes the second prompt 831.
[0196] The display device 100 can receive voice commands from the remote control device 200 and transmit the voice data corresponding to the received voice commands to the speaker recognition server 610 and the STT server 630.
[0197] Thus, according to embodiments of the present invention, by collecting speech data of words corresponding to dialects, it can be used to improve speaker recognition rate and the conversion accuracy of STT model.
[0198] Reference Figure 8c The display device 100 can display a third speaker registration statement 850, including a third prompt 851, on the speaker registration screen 800.
[0199] The third prompt, 851, can be a word whose recognition rate is lower than the preset recognition rate.
[0200] The user can issue voice commands to the remote control device 200 corresponding to the third speaker registration statement 850, which includes the third prompt 851.
[0201] The display device 100 can receive voice commands from the remote control device 200 and transmit the voice data corresponding to the received voice commands to the speaker recognition server 610 and the STT server 630.
[0202] Thus, according to embodiments of the present invention, speech data for words with low word recognition rates can be collected to improve speaker recognition rates and the conversion accuracy of the STT model.
[0203] In another embodiment, the display device 100 may sequentially display a first speaker registration statement 810, a second speaker registration statement 830, and a third speaker registration statement 850.
[0204] Therefore, voice data for the first prompt 811, the second prompt 831, and the third prompt 851 can be collected at once.
[0205] Figure 9 This is a diagram illustrating the process of collecting voice data for prompts according to an embodiment of the present invention.
[0206] AI server 500 can receive the name of the latest content from the latest content database (DB) 910.
[0207] AI Server 500 can obtain the name of the latest content as a prompt for speaker registration.
[0208] AI server 500 can receive words corresponding to the dialect of the region corresponding to the location information from regional dialect DB930 based on the location information of display device 100.
[0209] AI server 500 can obtain dialect words from the region corresponding to the location information of display device 100 as prompts for speaker registration.
[0210] The AI server 500 can obtain words with a recognition rate lower than a pre-set recognition rate from its own database as prompts for speaker registration.
[0211] AI server 500 can transmit one or more generated prompts to display device 100.
[0212] The display device 100 can use one or more generated prompts to generate speaker registration statements.
[0213] like Figures 8a to 8c As shown, when a request to register a speaker via voice is received, the display device 100 may display a first speaker registration statement 810, a second speaker registration statement 830, or a third speaker registration statement 850 on the display 180.
[0214] The display device 100 can obtain the voice data spoken by the user in response to the speaker's registered statement, and transmit the obtained voice data and speaker identification information to the speaker recognition server 610 and the STT server 630.
[0215] The speaker recognition server 610 can extract the speech signal corresponding to the prompt from the received speech data and obtain the features of the extracted speech signal.
[0216] The speaker recognition server 610 can match the features of the acquired speech signal with the speaker identification information and store them in the speaker DB950.
[0217] The Speaker DB950 can match and store the speaker identification information of each of a plurality of speakers with the speaker's speech features.
[0218] After a speaker registers, the speaker recognition server 610 can extract the features of the speaker's voice commands received from the display device 100, compare the extracted features with the stored features, and perform speaker recognition.
[0219] The speaker recognition server 610 can transmit the speaker recognition results to the display device 100, and the display device 100 can output the speaker recognition results.
[0220] The STT server 630 can convert voice data received from the display device 100 into text data. The STT server 630 can extract prompt data corresponding to the prompt from the converted text data.
[0221] The STT server 630 can use cue data collected during speaker registration to retrain the STT model 970. The STT model 970 requires a large amount of data to improve text conversion accuracy.
[0222] If the STT model 970 is retrained using cue data collected during the speaker registration process, the text conversion accuracy of the STT model 970 can be greatly improved.
[0223] Figure 10 This is a flowchart illustrating the operation method of a display device according to an embodiment of the present invention.
[0224] Reference Figure 10 The controller 170 of the display device 100 can receive one or more prompts for speaker registration (S1001).
[0225] In one embodiment, the controller 170 may receive more than one prompt from the AI server 500.
[0226] In another embodiment, controller 170 may receive the name of the latest content from the content provider server and use the received name as a prompt.
[0227] In another embodiment, controller 170 may receive the name of the latest content from the latest content DB910.
[0228] In another embodiment, the controller 170 can receive words from the regional dialect DB930 that represent the dialect of the region corresponding to the location information of the display device 100, and obtain the received words as prompts.
[0229] The controller 170 of the display device 100 can generate a speaker registration statement including the generated prompt (S1003).
[0230] The controller 170 can generate multiple speaker registration statements, each including a prompt.
[0231] As another example, controller 170 can generate speaker registration statements that include multiple prompts.
[0232] The controller 170 of the display device 100 displays a speaker registration statement on the display 180 according to the speaker registration request (S1005).
[0233] When a speaker registration request is received, the controller 170 can display the generated speaker registration statement on the display 180.
[0234] The controller 170 of the display device 100 can receive voice commands spoken by the user that correspond to the speaker's registered statement (S1007).
[0235] In one embodiment, the controller 170 can receive voice commands spoken by a user that correspond to a speaker's registered statement from the remote control device 200.
[0236] In another embodiment, the controller 170 may receive voice commands spoken by a user corresponding to a speaker's registered statement via a microphone (not shown) on the display device 100.
[0237] The controller 170 of the display device 100 can transmit voice data corresponding to the voice command to the speaker recognition server 610 and the STT server 630 via a network interface (S1009).
[0238] The controller 170 can send voice data and speaker identification information to the speaker recognition server 610.
[0239] The speaker recognition server 610 can extract the speaker's speech features from the received speech data, match the extracted speech features with the speaker identification information, and store them in the speaker DB950.
[0240] The STT server 630 can use the received voice data for training the STT model.
[0241] Thus, according to embodiments of the present invention, by collecting speech data on the names of recent content with limited data, words corresponding to dialects, and words with low recognition rates, it is possible to improve speaker recognition rate and the conversion accuracy of the STT model.
[0242] According to an embodiment of the present invention, the aforementioned method can be implemented on a medium for recording programs in processor-readable code. Examples of processor-readable media include read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, optical data storage device, etc.
Claims
1. A display device, wherein, include: monitor; A network interface for communicating with more than one server; as well as The controller obtains a speaker registration statement including a prompt for speaker registration, and upon receiving a request for speaker registration, displays the generated speaker registration statement on the display, obtains a voice command corresponding to the displayed speaker registration statement, and transmits voice data corresponding to the obtained voice command to one or more servers through the network interface. The notification includes the name of the latest content and one of the following: a word representing the local dialect.
2. The display device according to claim 1, wherein, The prompts are for words whose recognition rate is lower than a preset recognition rate.
3. The display device according to claim 1, wherein, The controller receives the name of the latest content from the servers included in the one or more servers.
4. The display device according to claim 1, wherein, The controller obtains the location information of the display device and retrieves the words representing the dialect of the region corresponding to the location information from the regional dialect database as the prompt.
5. The display device according to claim 1, wherein, The controller sequentially displays a plurality of speaker registration statements, each with a different prompt, on the display.
6. The display device according to claim 1, wherein, The controller transmits the voice data to the speaker recognition server and the speech-to-text server through the network interface.
7. The display device according to claim 6, wherein, The controller transmits speaker identification information and the voice data to the speaker recognition server.
8. A method for operating a display device, wherein, include: The steps to obtain a speaker registration statement that includes prompts for speaker registration; The step of displaying the generated speaker registration statement upon receiving a request for speaker registration; The steps to obtain the voice command corresponding to the displayed speaker registration statement; as well as The step of transmitting the voice data corresponding to the obtained voice command to one or more servers; The notification includes the name of the latest content and one of the following: a word representing the local dialect.
9. The method of operating the display device according to claim 8, wherein, The prompts are for words whose recognition rate is lower than a preset recognition rate.
10. The method of operating the display device according to claim 8, wherein, It also includes the step of receiving the name of the latest content from the servers included in the one or more servers.
11. The method of operating the display device according to claim 8, wherein, Also includes: The step of obtaining the position information of the display device; as well as The step of obtaining the prompt from a regional dialect database by using the words representing the dialect of the region corresponding to the location information as the prompt.
12. The method of operating the display device according to claim 8, wherein, The display steps include: The steps involve sequentially displaying multiple speaker registration statements, each with a different prompt.
13. The method of operating the display device according to claim 8, wherein, The transmission steps include: The step of transmitting the voice data to the speaker recognition server and the speech-to-text server.
14. The method of operating the display device according to claim 13, wherein, The transmission steps include: The step of transmitting the speaker identification information and the voice data to the speaker recognition server.