Display device and operating method thereof

The display device enhances voice recognition by separating and prioritizing user voice data in noisy environments, ensuring accurate command execution and reducing power usage.

WO2026029218A1PCT designated stage Publication Date: 2026-02-05LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/011066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing voice recognition technologies in display devices struggle to accurately distinguish user voices from multiple voices in noisy environments, leading to repeated recognition attempts and unintended command execution due to ambient sounds and interference from other speakers.

Method used

A display device equipped with microphones, a controller, and a display that separates user voice data from ambient noise, selects relevant voice data based on user information, gain values, and location, and displays results accordingly, minimizing power consumption by stopping acquisition of unselected data.

Benefits of technology

Improves voice recognition quality by accurately identifying the intended user's voice command even in noisy conditions, reducing missed results and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024011066_05022026_PF_FP_ABST
    Figure KR2024011066_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A display device and an operating method thereof according to an embodiment of the present disclosure are to provide a result desired by a user by distinguishing a user's voice during voice recognition in a situation in which several voices exist. When there are a plurality of pieces of voice data separated from an audio signal, and voice data to be displayed as result information is selected from among the plurality of pieces of voice data, information acquisition of unselected voice data may be stopped.
Need to check novelty before this filing date? Find Prior Art

Description

Display device and method of operation thereof

[0001] The present disclosure relates to a display device and an operating method thereof for processing multiple voices mixed during voice recognition.

[0002] The input interface for conventional display devices was mainly a remote control, and users performed functions by pressing various input buttons provided on the remote control.

[0003] With the rapid development of digital technology, not only display devices but also content are supporting high specifications, high definition, and various functions. However, performing these functions using only the buttons on the existing remote control was not only difficult, but also very inconvenient.

[0004] This led to the adoption of input technology through voice recognition in display devices.

[0005] However, although there have been many advancements in technology, existing voice recognition technology still has difficulty accurately recognizing users' voices in various situations.

[0006] In particular, when using voice recognition, there was a problem in which the voice recognition process was repeated and retried when an unintended sound source was recognized due to interference from the surroundings.

[0007] Alternatively, there was a problem where the user's intended result of giving a voice command was different from that of the intended user due to the voice of another person.

[0008] The present disclosure provides a display device and an operating method thereof that distinguishes a user's voice during voice recognition in a situation where multiple voices are present and provides a result desired by the user.

[0009] The present disclosure seeks to minimize degradation in the quality of speech recognition even when ambient sounds, such as the voices of other people other than the user's speech, are mixed in.

[0010] A display device according to an embodiment of the present disclosure is intended to extract voice data of a user who intends to issue an actual voice command by distinguishing the voice through voice recognition, separating voice data included in an audio signal, and then providing a corresponding result.

[0011] A display device according to an embodiment of the present disclosure includes at least one microphone, a controller for acquiring user information and result information corresponding to at least one voice data separated from a signal acquired through the microphone, and a display for displaying the result information, wherein when a plurality of voice data are separated from a signal, the controller can stop acquiring information of unselected voice data when voice data to be displayed as result information is selected from among the plurality of voice data.

[0012] The controller can turn off the microphone when the speech data to be displayed as result information is finished speaking.

[0013] The controller can select voice data corresponding to a pre-registered user from among multiple voice data based on user information as voice data to be displayed as result information.

[0014] When a plurality of pieces of voice data to be displayed as result information are selected, the controller can display a plurality of pieces of result information corresponding to each of the selected pieces of voice data on the display.

[0015] When there are multiple microphones, the controller can obtain location information of multiple users corresponding to multiple voice data, and display each of the multiple result information in an area corresponding to the location of each user based on the location information.

[0016] The controller can obtain a gain value for each of a plurality of voice data and display a plurality of result information differently based on the gain value.

[0017] The controller obtains a first gain value of first voice data among a plurality of voice data and a second gain value of second voice data among a plurality of voice data, and if the first gain value is greater than the second gain value, the result information of the first voice data can be displayed larger than the result information of the second voice data.

[0018] The controller can display the result information of the first voice data and the result information of the second voice data in the same size if the difference between the first gain value and the second gain value is less than a preset threshold value.

[0019] The controller can display only the result information of the voice data having a large gain value if the difference between the first gain value and the second gain value is greater than a preset threshold value.

[0020] When the controller acquires new second voice data while displaying result information of the first voice data, the controller can display the result information corresponding to the second voice data differently based on user information corresponding to the second voice data.

[0021] If the user corresponding to the first voice data and the user corresponding to the second voice data are the same, the controller can obtain and display the second result information corresponding to the second voice data within the first result information of the first voice data.

[0022] If the user corresponding to the first voice data and the user corresponding to the second voice data are different, the controller displays the second result information and the third result information together, and the third result information may be result information obtained corresponding to the second voice data regardless of the first result information.

[0023] When the controller acquires multiple voice data from a signal acquired through a microphone, it can display result information corresponding to voice data corresponding to a user who uttered the activation word, and when the controller acquires multiple voice data from a signal received from a remote control device, it can display result information corresponding to voice data corresponding to a user with the largest gain value.

[0024] The controller can select voice data containing a preset keyword from among multiple voice data as voice data to be displayed as result information.

[0025] A method of operating a display device according to an embodiment of the present disclosure includes a step of separating at least one voice data from a signal acquired through a microphone, a step of acquiring user information and result information corresponding to the separated voice data, and a step of displaying the result information. In a case where there are multiple voice data separated from a signal, if voice data to be displayed as result information is selected from among the multiple voice data, the method may further include a step of stopping acquisition of information of unselected voice data.

[0026] According to an embodiment of the present disclosure, even if multiple voices are mixed, it is possible to provide a result for the voice of a user who actually gave a voice command, thereby improving the quality of voice recognition.

[0027] According to an embodiment of the present disclosure, when multiple voices are mixed during voice recognition, by providing results for one or more voices, the problem of missing the desired result of a user who has issued a voice command is minimized, thereby improving user satisfaction.

[0028] According to an embodiment of the present disclosure, when voice data to be provided as result information is selected from among a plurality of voice data, there is an effect of minimizing the resulting cost and power consumption by stopping information acquisition of voice data that is not selected.

[0029] FIG. 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0030] Figure 2 is a block diagram of a remote control device according to one embodiment of the present invention.

[0031] Figure 3 shows an example of an actual configuration of a remote control device (200) according to one embodiment of the present invention.

[0032] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0033] FIG. 5 is a drawing showing various examples of display devices to which embodiments of the present disclosure can be applied.

[0034] FIG. 6 is a control block diagram illustrating a method for a display device according to an embodiment of the present disclosure to receive and process an audio signal.

[0035] FIG. 7 is a flowchart illustrating an operation method of a display device according to a first embodiment of the present disclosure.

[0036] FIGS. 8 and 9 are diagrams for explaining a method of selecting voice data to be displayed as result information according to the first embodiment of the present disclosure.

[0037] FIGS. 10 and 11 are diagrams for explaining a method of selecting voice data to be displayed as result information according to a second embodiment of the present disclosure.

[0038] FIG. 12 is a diagram for explaining a method for selecting voice data to be displayed as result information according to a third embodiment of the present disclosure.

[0039] FIG. 13 is a diagram for explaining a method of selecting voice data to be displayed as result information according to the fourth embodiment of the present disclosure.

[0040] FIG. 14 is a drawing for explaining the operation of a display device according to an embodiment of the present disclosure after selecting voice data to be displayed as result information.

[0041] FIG. 15 is a flowchart illustrating an operation method of a display device according to a second embodiment of the present disclosure.

[0042] FIGS. 16 and 17 are diagrams for explaining a method of displaying result information of a plurality of voice data according to the first embodiment of the present disclosure.

[0043] Figures 18 and 19 are exemplary drawings showing a case where the speech of the first and second voice data ends at different times.

[0044] FIG. 20 is a diagram for explaining a method for displaying result information of a plurality of voice data according to a second embodiment of the present disclosure.

[0045] FIG. 21 is a first exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0046] FIG. 22 is a second exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0047] FIG. 23 is a third exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0048] FIG. 24 is an exemplary diagram showing a method for providing results when a display device according to an embodiment of the present disclosure recognizes continuous speech.

[0049] Hereinafter, embodiments related to the present invention will be described in more detail with reference to the drawings. The suffixes "module" and "part" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles.

[0050] A display device according to an embodiment of the present invention is, for example, an intelligent display device that adds computer-assisted functionality to its broadcast reception function. While faithfully performing the broadcast reception function, it can also be equipped with Internet functionality and other features, providing a more user-friendly interface, such as a manual input device, touch screen, or space remote control. Furthermore, with support for wired or wireless Internet functionality, it can connect to the Internet and a computer, enabling functions such as email, web browsing, banking, or gaming. A standardized, general-purpose operating system can be used for these various functions.

[0051] Accordingly, the display device described in the present invention can perform various user-friendly functions, for example, by allowing various applications to be freely added or deleted on a general-purpose operating system kernel. More specifically, the display device can be a network TV, HBB TV, smart TV, LED TV, OLED TV, etc., and in some cases, it can also be applied to smartphones.

[0052] FIG. 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0053] Referring to FIG. 1, the display device (100) may include a broadcast receiving unit (130), an external device interface (135), a memory (140), a user input interface (150), a controller (170), a wireless communication interface (173), a microphone (175), a display (180), a speaker (185), and a power supply circuit (190).

[0054] The broadcast receiving unit (130) may include a tuner (131), a demodulator (132), and a network interface (133).

[0055] The tuner (131) can select a specific broadcast channel according to a channel selection command. The tuner (131) can receive a broadcast signal for the selected specific broadcast channel.

[0056] The demodulator (132) can separate the received broadcast signal into a video signal, an audio signal, and a data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into a form that can be output.

[0057] The external device interface (135) can receive an application or a list of applications within an adjacent external device and transmit it to the controller (170) or memory (140).

[0058] The external device interface (135) can provide a connection path between the display device (100) and the external device. The external device interface (135) can receive one or more of images and audio output from an external device connected wirelessly or wiredly to the display device (100) and transmit them to the controller (170). The external device interface (135) can include a plurality of external input terminals. The plurality of external input terminals can include an RGB terminal, one or more HDMI (High Definition Multimedia Interface) terminals, and a component terminal.

[0059] A video signal of an external device input through an external device interface (135) can be output through a display (180). A voice signal of an external device input through an external device interface (135) can be output through a speaker (185).

[0060] An external device that can be connected to the external device interface (135) may be any one of a set-top box, a Blu-ray player, a DVD player, a game console, a sound bar, a smartphone, a PC, a USB memory, and a home theater, but this is only an example.

[0061] The network interface (133) may provide an interface for connecting the display device (100) to a wired / wireless network including the Internet. The network interface (133) may transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.

[0062] Additionally, some content data stored in the display device (100) can be transmitted to a selected user or electronic device among other users or other electronic devices pre-registered in the display device (100).

[0063] The network interface (133) can access a predetermined web page through a connected network or another network linked to the connected network. That is, by accessing a predetermined web page through a network, data can be transmitted or received with the corresponding server.

[0064] In addition, the network interface (133) can receive content or data provided by a content provider or network operator. That is, the network interface (133) can receive content such as movies, advertisements, games, VOD, broadcast signals, etc. and information related thereto provided from a content provider or network provider via a network.

[0065] Additionally, the network interface (133) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.

[0066] The network interface (133) can select and receive a desired application from among applications open to the public via a network.

[0067] The memory (140) stores a program for each signal processing and control within the controller (170), and can store signal-processed image, voice, or data signals.

[0068] In addition, the memory (140) may perform a function for temporary storage of video, audio, or data signals input from an external device interface (135) or a network interface (133), and may store information about a specific image through a channel memory function.

[0069] The memory (140) can store an application or a list of applications input from an external device interface (135) or a network interface (133).

[0070] The display device (100) can play content files (video files, still image files, music files, document files, application files, etc.) stored in the memory (140) and provide them to the user.

[0071] The user input interface (150) can transmit a signal input by the user to the controller (170) or transmit a signal from the controller (170) to the user. For example, the user input interface (150) can receive and process control signals such as power on / off, channel selection, and screen setting from the remote control device (200) according to various communication methods such as Bluetooth, Ultra Wideband (WB), ZigBee, Radio Frequency (RF) communication, or infrared (IR) communication, or process control signals from the controller (170) to be transmitted to the remote control device (200).

[0072] In addition, the user input interface (150) can transmit control signals input from local keys (not shown) such as a power key, channel key, volume key, and setting value to the controller (170).

[0073] An image signal processed by the controller (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, an image signal processed by the controller (170) can be input to an external output device through an external device interface (135).

[0074] The voice signal processed by the controller (170) can be output as audio to the speaker (185). In addition, the voice signal processed by the controller (170) can be input to an external output device through the external device interface (135).

[0075] In addition, the controller (170) can control the overall operation within the display device (100).

[0076] In addition, the controller (170) can control the display device (100) by a user command or internal program input through the user input interface (150), and can connect to a network to enable the user to download a desired application or application list into the display device (100).

[0077] The controller (170) enables the user-selected channel information, etc. to be output through a display (180) or speaker (185) together with processed video or audio signals.

[0078] In addition, the controller (170) allows a video signal or audio signal from an external device, for example, a camera or camcorder, input through the external device interface (135) to be output through the display (180) or speaker (185) in accordance with an external device video playback command received through the user input interface (150).

[0079] Meanwhile, the controller (170) can control the display (180) to display an image, for example, a broadcast image input through a tuner (131), an external input image input through an external device interface (135), an image input through a network interface, or an image stored in a memory (140) can be controlled to be displayed on the display (180). In this case, the image displayed on the display (180) can be a still image or a moving image, and can be a 2D image or a 3D image.

[0080] In addition, the controller (170) can control the playback of content stored in the display device (100), received broadcast content, or external input content input from outside, and the content can be in various forms such as broadcast video, external input video, audio file, still image, connected web screen, and document file.

[0081] The wireless communication interface (173) can communicate with an external device through wired or wireless communication. The wireless communication interface (173) can perform short-range communication with the external device. To this end, the wireless communication interface (173) can support short-range communication using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies. This wireless communication interface (173) can support wireless communication between the display device (100) and a wireless communication system, between the display device (100) and another display device (100), or between the display device (100) and a network in which the display device (100, or an external server) is located via a short-range wireless communication network (Wireless Area Network). The short-range wireless communication network can be a short-range wireless personal area network (Wireless Personal Area Network).

[0082] Here, the other display device (100) may be a wearable device (e.g., a smartwatch, smart glasses, a head-mounted display (HMD)) or a mobile terminal such as a smart phone that can exchange data with (or be linked to) the display device (100) according to the present invention. The wireless communication interface (173) may detect (or recognize) a wearable device capable of communication around the display device (100).

[0083] Furthermore, if the detected wearable device is a device certified to communicate with the display device (100) according to the present invention, the controller (170) can transmit at least a portion of the data processed in the display device (100) to the wearable device via the wireless communication interface (173). Accordingly, a user of the wearable device can utilize the data processed in the display device (100) via the wearable device.

[0084] The microphone (175) can acquire audio. The microphone (175) can acquire audio around the display device (100).

[0085] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal processed by the controller (170) or a video signal, data signal, etc. received from an external device interface (135) into R, G, and B signals, respectively.

[0086] Meanwhile, since the display device (100) illustrated in FIG. 1 is merely an embodiment of the present invention, some of the illustrated components may be integrated, added, or omitted depending on the specifications of the display device (100) actually implemented.

[0087] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0088] According to another embodiment of the present invention, the display device (100) may receive and play back an image through a network interface (133) or an external device interface (135) without having a tuner (131) and a demodulator (132), unlike that shown in FIG. 1.

[0089] For example, the display device (100) may be implemented separately as an image processing device, such as a set-top box, for receiving broadcast signals or contents according to various network services, and a content playback device for playing contents input from the image processing device.

[0090] In this case, the operating method of the display device according to the embodiment of the present invention described below may be performed by any one of the display device (100) described with reference to FIG. 1, as well as an image processing device such as the separated set-top box, or a content playback device having a display (180) and an audio output unit (185).

[0091] The speaker (185) receives a signal processed by the controller (170) and outputs it as voice.

[0092] The power supply circuit (190) supplies power to the entire display device (100). In particular, it can supply power to a controller (170) that can be implemented in the form of a system on chip (SOC), a display (180) for displaying images, and a speaker (185) for audio output.

[0093] Specifically, the power supply circuit (190) may include a converter that converts AC power into DC power and a dc / dc converter that converts the level of the DC power.

[0094] Next, a remote control device according to an embodiment of the present invention will be described with reference to FIGS. 2 and 3.

[0095] FIG. 2 is a block diagram of a remote control device according to an embodiment of the present invention, and FIG. 3 shows an example of an actual configuration of a remote control device (200) according to an embodiment of the present invention.

[0096] First, referring to FIG. 2, the remote control device (200) may include a fingerprint recognition device (210), a wireless communication circuit (220), a user input interface (230), a sensor (240), an output interface (250), a power supply circuit (260), a memory (270), a controller (280), and a microphone (290).

[0097] Referring to FIG. 2, the wireless communication circuit (220) transmits and receives signals with any one of the display devices according to the embodiments of the present invention described above.

[0098] The remote control device (200) may be equipped with an RF circuit (221) capable of transmitting and receiving signals with the display device (100) according to RF communication standards, and an IR circuit (223) capable of transmitting and receiving signals with the display device (100) according to IR communication standards. In addition, the remote control device (200) may be equipped with a Bluetooth circuit (225) capable of transmitting and receiving signals with the display device (100) according to Bluetooth communication standards. In addition, the remote control device (200) may be equipped with an NFC circuit (227) capable of transmitting and receiving signals with the display device (100) according to NFC (Near Field Communication) communication standards, and a WLAN circuit (229) capable of transmitting and receiving signals with the display device (100) according to WLAN (Wireless LAN) communication standards.

[0099] In addition, the remote control device (200) transmits a signal containing information about the movement of the remote control device (200) to the display device (100) through a wireless communication circuit (220).

[0100] Meanwhile, the remote control device (200) can receive a signal transmitted by the display device (100) through the RF circuit (221), and, if necessary, can transmit commands for power on / off, channel change, volume change, etc. to the display device (100) through the IR circuit (223).

[0101] The user input interface (230) may be configured as a keypad, buttons, a touchpad, or a touch screen. The user can input commands related to the display device (100) to the remote control device (200) by operating the user input interface (230). If the user input interface (230) includes a hard key button, the user can input commands related to the display device (100) to the remote control device (200) by pushing the hard key button. This will be described with reference to FIG. 3.

[0102] Referring to FIG. 3, the remote control device (200) may include a plurality of buttons. The plurality of buttons may include a fingerprint recognition button (212), a power button (231), a home button (232), a live button (233), an external input button (234), a volume control button (235), a voice recognition button (236), a channel change button (237), a confirmation button (238), and a back button (239).

[0103] The fingerprint recognition button (212) may be a button for recognizing a user's fingerprint. In one embodiment, the fingerprint recognition button (212) may be capable of a push operation, and may receive a push operation and a fingerprint recognition operation.

[0104] The power button (231) may be a button for turning the power of the display device (100) on / off.

[0105] The home button (232) may be a button for moving to the home screen of the display device (100).

[0106] The live button (233) may be a button for displaying a real-time broadcast program.

[0107] The external input button (234) may be a button for receiving an external input connected to the display device (100).

[0108] The volume control button (235) may be a button for adjusting the size of the volume output by the display device (100).

[0109] The voice recognition button (236) may be a button for receiving a user's voice and recognizing the received voice.

[0110] The channel change button (237) may be a button for receiving a broadcast signal of a specific broadcast channel.

[0111] The confirmation button (238) may be a button for selecting a specific function, and the back button (239) may be a button for returning to the previous screen.

[0112] Let's explain Figure 2 again.

[0113] When the user input interface (230) has a touch screen, the user can input commands related to the display device (100) using the remote control device (200) by touching the soft keys of the touch screen. In addition, the user input interface (230) may have various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the present invention.

[0114] The sensor (240) may include a gyro sensor (241) or an acceleration sensor (243), and the gyro sensor (241) may sense information about the movement of the remote control device (200).

[0115] For example, the gyro sensor (241) can sense information about the operation of the remote control device (200) based on the x, y, and z axes, and the acceleration sensor (243) can sense information about the movement speed of the remote control device (200). Meanwhile, the remote control device (200) can further include a distance measuring sensor, so as to sense the distance to the display (180) of the display device (100).

[0116] The output interface (250) can output a video or audio signal corresponding to an operation of the user input interface (230) or a signal transmitted from the display device (100).

[0117] The user can recognize whether the output interface (250) is manipulating the user input interface (230) or controlling the display device (100).

[0118] For example, the output interface (250) may include an LED (251) that lights up when the user input interface (230) is operated or a signal is transmitted and received with the display device (100) via the wireless communication unit (225), a vibrator (253) that generates vibrations, a speaker (255) that outputs sound, or a display (257) that outputs images.

[0119] In addition, the power supply circuit (260) supplies power to the remote control device (200), and power waste can be reduced by stopping the power supply when the remote control device (200) does not move for a predetermined period of time.

[0120] The power supply circuit (260) can resume power supply when a predetermined key provided in the remote control device (200) is operated.

[0121] The memory (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).

[0122] When the remote control device (200) wirelessly transmits and receives signals through the display device (100) and the RF circuit (221), the remote control device (200) and the display device (100) transmit and receive signals through a predetermined frequency band.

[0123] The controller (280) of the remote control device (200) can store and reference information regarding the frequency band that can wirelessly transmit and receive signals with the display device (100) paired with the remote control device (200) in the memory (270).

[0124] The controller (280) controls all matters related to the control of the remote control device (200). The controller (280) can transmit a signal corresponding to a predetermined key operation of the user input interface (230) or a signal corresponding to the movement of the remote control device (200) sensed by the sensor (240) to the display device (100) via the wireless communication unit (225).

[0125] Additionally, the microphone (290) of the remote control device (200) can acquire voice.

[0126] A plurality of microphones (290) may be provided.

[0127] Next, Figure 4 is described.

[0128] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0129] Figure 4 (a) illustrates that a pointer (205) corresponding to a remote control device (200) is displayed on a display (180).

[0130] The user can move or rotate the remote control device (200) up and down, left and right. The pointer (205) displayed on the display (180) of the display device (100) corresponds to the movement of the remote control device (200). This remote control device (200) can be called a space remote control because, as shown in the drawing, the pointer (205) moves and is displayed according to the movement in 3D space.

[0131] Figure 4 (b) illustrates that when a user moves the remote control device (200) to the left, the pointer (205) displayed on the display (180) of the display device (100) also moves to the left correspondingly.

[0132] Information about the movement of the remote control device (200) detected through the sensor of the remote control device (200) is transmitted to the display device (100). The display device (100) can calculate the coordinates of the pointer (205) from the information about the movement of the remote control device (200). The display device (100) can display the pointer (205) to correspond to the calculated coordinates.

[0133] Figure 4 (c) illustrates a case where a user moves the remote control device (200) away from the display (180) while pressing a specific button within the remote control device (200). As a result, a selection area within the display (180) corresponding to the pointer (205) can be zoomed in and displayed in an enlarged manner.

[0134] Conversely, when the user moves the remote control device (200) closer to the display (180), the selection area within the display (180) corresponding to the pointer (205) may be zoomed out and displayed in a reduced size.

[0135] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.

[0136] Additionally, when a specific button within the remote control device (200) is pressed, recognition of up, down, left, and right movements may be excluded. That is, when the remote control device (200) moves away from or toward the display (180), up, down, left, and right movements may not be recognized, and only forward and backward movements may be recognized. When a specific button within the remote control device (200) is not pressed, only the pointer (205) moves in accordance with the up, down, left, and right movements of the remote control device (200).

[0137] Meanwhile, the movement speed or movement direction of the pointer (205) can correspond to the movement speed or movement direction of the remote control device (200).

[0138] Meanwhile, the pointer in this specification refers to an object displayed on the display (180) in response to the operation of the remote control device (200). Accordingly, objects of various shapes other than the arrow shape illustrated in the drawing can be used as the pointer (205). For example, the pointer may be a concept including a point, a cursor, a prompt, a thick outline, etc. In addition, the pointer (205) may be displayed corresponding to one point on the horizontal or vertical axis on the display (180), or may be displayed corresponding to multiple points such as lines or surfaces.

[0139] Meanwhile, although the display device is described as a TV in FIGS. 1 to 4, the display device according to the embodiment of the present disclosure may correspond to various devices equipped with a display (180).

[0140] FIG. 5 is a drawing showing various examples of display devices to which embodiments of the present disclosure can be applied.

[0141] In Fig. 5 (a), a smartphone is shown, (b) a TV, (c) a monitor, (d) a washing machine, (e) a refrigerator, (f) an air conditioner, and (g) a car. The display device according to an embodiment of the present disclosure may include various devices as shown in Figs. 5 (a) to (g). That is, the display device according to an embodiment of the present disclosure may include various devices having a display (180).

[0142] In addition, the display device may include various devices capable of receiving and processing audio signals from a microphone (175) provided therein or a microphone provided in a remote control device (200) or a smartphone, etc. In addition, in the present disclosure, at least one of the display device and the remote control device (200) may be provided with multiple microphones (175). When there are multiple microphones (175), the direction of speech of voice data can be recognized.

[0143] FIG. 6 is a control block diagram illustrating a method for a display device according to an embodiment of the present disclosure to receive and process an audio signal.

[0144] The display device (100) may include an audio signal processing module for receiving and processing audio signals.

[0145] The audio signal processing module may be included in the controller (170) of Fig. 1. In this case, the operations described below are operations performed by the controller (170).

[0146] The audio signal processing module may include at least one of a speaker recognition module (1701), an STT processing module (1703), and an NLP processing module (1705).

[0147] The speaker recognition module (1701) can recognize a user who has spoken voice data included in an audio signal. Here, the audio signal can be input through a microphone (175) or received from an external source such as a network interface (133). The speaker recognition module (1701) can recognize a user who has spoken voice data by using voice characteristics (e.g., voiceprints) of each user stored in a memory (140) or the like. For example, the speaker recognition module (1701) can recognize a user by comparing voice data included in an audio signal with previously stored voiceprints and extracting a voiceprint that matches or is most similar to the voiceprint.

[0148] The speaker recognition module (1701) can extract at least one voice data from an audio signal. If the audio signal includes multiple voice data, the speaker recognition module (1701) can separate the multiple voice data.

[0149] The speaker recognition module (1701) can recognize a user corresponding to each of a plurality of separated voice data. For example, the speaker recognition module (1701) can separate first voice data and second voice data from an audio signal and recognize a first user corresponding to the first voice data and a second user corresponding to the second voice data.

[0150] In addition, the speaker recognition module (1701) can transmit at least one voice data separated from the audio signal to the STT processing module (1703). At this time, the speaker recognition module (1703) can also transmit recognized user information together with the voice data.

[0151] The STT processing module (1703) can process at least one voice data separated by the speaker recognition module (1701) through STT (Speech to Text). When multiple voice data are input, the STT processing module (1703) can process each of the multiple voice data through STT.

[0152] Meanwhile, the speaker recognition module (1701) may be included in the STT processing module (1703). In this case, the STT processing module (1703) can receive an audio signal, separate voice data from the audio signal, and perform both user recognition and STT processing.

[0153] The STT processing module (1703) can transmit the STT processing result to the NLP processing module (1305).

[0154] The NLP processing module (1305) can receive STT processed text and perform NLP (Natural Language Processing).

[0155] NLP obtains results based on the intent analysis and analysis of the input text.

[0156] Meanwhile, in Fig. 6, the speaker recognition module (1701), the STT processing module (1703), and the NLP processing module (1705) are described separately, but this is merely for convenience of explanation. In other words, at least two or more of these may be integrated into a single module, or they may be divided into three or more modules.

[0157] Additionally, at least one of the speaker recognition module (1701), the STT processing module (1703), and the NLP processing module (1705) may perform signal processing through communication with an external server (not shown). For example, the speaker recognition module (1701) may recognize a user through communication with a voiceprint database (not shown). Alternatively, the STT processing module (1703) may perform STT processing through communication with an STT server (not shown). Alternatively, the NLP processing module (1705) may perform NLP through communication with an NLP server (not shown).

[0158] That is, the display device (100) according to the embodiment of the present disclosure can separate at least one voice data from an audio signal using various methods including the above-described method, and perform user recognition, STT processing, and NLP processing on the separated voice data.

[0159] Meanwhile, when acquiring audio signals, ambient noise, such as the voices of others, may prevent the user from receiving the desired results. Therefore, the present disclosure aims to more accurately provide the user with the desired results, regardless of ambient noise. Specifically, the present disclosure aims to more accurately recognize users giving voice commands and improve the accuracy of providing results based on voice commands.

[0160] FIG. 7 is a flowchart illustrating an operation method of a display device according to a first embodiment of the present disclosure.

[0161] The controller (170) can obtain an audio signal (S10).

[0162] The controller (170) can obtain an audio signal from a microphone (175) or a user input interface (150).

[0163] The microphone (175) can acquire an audio signal by converting ambient sounds into electrical signals in response to a microphone-on command. The microphone (175) can recognize preset voice commands, etc. as microphone-on commands.

[0164] The user input interface (150) can receive an audio signal from the remote control device (200). The remote control device (200) can obtain an audio signal by converting ambient sounds into electrical signals through the microphone (290) in response to a microphone-on command. The remote control device (200) can recognize the input of a specific button, such as a voice recognition button (236), as a microphone-on command.

[0165] The controller (170) can obtain an audio signal from at least one of a microphone (175) and a user input interface (150).

[0166] The controller (170) can separate voice data from an audio signal (S20).

[0167] An audio signal may contain only one user's voice. Alternatively, the audio signal may contain multiple users' voices. Accordingly, the controller (170) can separate voice data corresponding to each user's voice from the audio signal.

[0168] The controller (170) can obtain information on each separated voice data (S30).

[0169] That is, the controller (170) can obtain information corresponding to each voice data included in the audio signal. Here, the information can include at least one of user information and result information. The controller (170) can analyze each voice data to obtain at least one of user information and result information.

[0170] The controller (170) can obtain user information corresponding to the voice data by comparing the voice print of the voice data with a pre-stored voice print. For example, if the controller (170) separates the first and second voice data from the audio signal, the controller (170) can obtain first user information corresponding to the first voice data and second user information corresponding to the second voice data.

[0171] The controller (170) can obtain result information by processing voice data through STT and NLP.

[0172] As soon as the controller (170) starts acquiring audio signals, it can analyze voice data in real time and acquire corresponding information.

[0173] The controller (170) can determine whether there are multiple voice data separated from the audio signal (S40).

[0174] If there is one voice data, the controller (170) can provide a result corresponding to one voice data (S50).

[0175] If there are multiple voice data, the controller (170) can select voice data to be displayed as result information among the multiple voice data (S1100).

[0176] For example, when the controller (170) separates first and second voice data from an audio signal, it can select voice data to be displayed as result information among the first and second voice data. There may be various methods for selecting voice data to be displayed as result information.

[0177] Next, a method for selecting voice data to be displayed as result information will be described with reference to examples of FIGS. 8 to 13.

[0178] First, FIGS. 8 and 9 are diagrams for explaining a method of selecting voice data to be displayed as result information according to the first embodiment of the present disclosure.

[0179] The controller (170) can determine whether an audio signal has been acquired from a microphone (175) provided in the display device (100) (S1111). Alternatively, the controller (170) can also determine whether an audio signal has been acquired from a remote control device (200) via a user input interface (150).

[0180] When the controller (170) obtains an audio signal from the remote control device (200), it can select voice data based on the gain value (S1113).

[0181] For example, when the controller (170) obtains an audio signal from the remote control device (200), it can obtain the gain value of each voice data. The controller (170) can select the voice data with the largest gain value and display the corresponding result information. That is, the controller (170) can select the voice data with the largest gain value as the voice data to be displayed as the result information.

[0182] When the controller (170) obtains an audio signal from a microphone (175), it can select voice data based on the trigger word (S1115).

[0183] For example, when the controller (170) obtains an audio signal from a microphone (175), it can obtain whether each voice data includes a preset trigger word. The controller (170) can select voice data including the trigger word as voice data to be displayed as result information.

[0184] Referring to FIG. 9, the controller (170) can separate the first voice data, “Hi LG, find a movie,” and the second voice data, “find a song,” from the audio signal. In the case of (a) of FIG. 9, the controller (170) can select the first voice data including the preset activation word, “Hi LG,” based on the audio signal obtained from the microphone (175), and display “movie search results” as the corresponding result information. In the case of FIG. 9 (b), the controller (170) can select the second voice data having a large gain value based on the audio signal obtained from the remote control device (200), and display “song search results” as the corresponding result information.

[0185] According to the first embodiment, there is an advantage in that the accuracy of providing results for voice commands is increased by selecting voice data indicated as result information in a different manner depending on the acquisition path of the audio signal.

[0186] Next, FIGS. 10 and 11 are diagrams for explaining a method of selecting voice data to be displayed as result information according to the second embodiment of the present disclosure.

[0187] The controller (170) can detect a preset keyword from each voice data (S1121).

[0188] The controller (170) can obtain whether or not a preset keyword is included in each voice data.

[0189] The controller (170) can select voice data from which a keyword is detected as voice data to be displayed as result information (S1123).

[0190] Referring to the example of Fig. 11, the controller (170) can separate the first audio data, “Find a movie,” and the second audio data, “That song was exciting,” from the audio signal. The controller (170) can detect preset keywords from each of the first and second audio data. If the controller (170) detects a keyword only in the first audio data, it can select the first audio data as the audio data to be displayed as result information. Accordingly, the controller (170) can display the movie search results as result information of the first audio data.

[0191] Meanwhile, keywords can vary. For example, keywords such as "Find Me," "Search Me," and "Play Me" may be set, but these are merely examples. Furthermore, keywords may vary depending on the type of display device. For example, if the display device is a car, keywords may be set to "Guide Me," "Find Me Directions," or "Find Me Routes." If the display device is a refrigerator, keywords may be set to "Recommend Me," "Find Me Recipe," or "Find Me Ingredients."

[0192] FIG. 12 is a diagram for explaining a method for selecting voice data to be displayed as result information according to a third embodiment of the present disclosure.

[0193] The controller (170) can obtain user information corresponding to each voice data (S1131).

[0194] User information may include the number of times voice data is selected, i.e., the frequency of voice commands. That is, the display device (100) may count the number of times each user has issued a voice command and store the count as user information.

[0195] The controller (170) can select voice data of a user with a high voice command frequency as voice data to be displayed as result information (S1133).

[0196] For example, if the frequency of the user's voice command corresponding to the first voice data is higher than the frequency of the user's voice command corresponding to the second voice data, the controller (170) may select the first voice data as the voice data to be displayed as result information.

[0197] According to the third embodiment, by selecting voice data of a user who frequently uses voice commands, there is an advantage of increasing the accuracy of providing results for voice commands.

[0198] FIG. 13 is a diagram for explaining a method of selecting voice data to be displayed as result information according to the fourth embodiment of the present disclosure.

[0199] The controller (170) can select voice data based on user information corresponding to the voice data. For example, the display device (100) may have previously registered and stored user information. The registered user information may include a voiceprint. The controller (170) can compare each of a plurality of voice data with a voiceprint stored as user information and extract voice data that matches the pre-stored voiceprint. The controller (170) can recognize a user whose voice data matches the pre-stored voiceprint as a pre-registered user. If a pre-stored user exists, the controller (170) can select the voice data of the pre-registered user as the voice data to be displayed as result information.

[0200] Referring to FIG. 12, the controller (170) can separate the first voice data, “Find entertainment shows featuring Yoo Jae-suk,” and the second voice data, “What should I eat for dinner tonight,” from the audio signal. The controller (170) can compare the first voice data and the second voice data with registered user information to obtain user information for each voice data. The controller (170) may obtain first user (user 1) information as user information corresponding to the first voice data, and may not obtain user information corresponding to the second voice data. In this case, the controller (170) may select the first voice data corresponding to the pre-registered first user as voice data to be displayed as result information.

[0201] That is, according to the fourth embodiment, the controller (170) can select voice data of a pre-registered user as voice data to be displayed as result information based on user information of each voice data.

[0202] Meanwhile, both the first and second voice data may be recognized as speech from a registered user. In this case, the controller (170) may select the voice data of the user of the currently logged-in account as the voice data to be displayed as result information.

[0203] The first to fourth embodiments described above are merely examples. That is, the controller (170) can select voice data to be displayed as result information in various ways.

[0204] Again, Figure 7 is explained.

[0205] The controller (170) can display the result information of the selected voice data and stop acquiring information of the unselected voice data (S1200).

[0206] The controller (170) can turn off the microphone (175) when the speaking of the selected voice data is finished (S1300).

[0207] FIG. 14 is a drawing for explaining the operation of a display device according to an embodiment of the present disclosure after selecting voice data to be displayed as result information.

[0208] Referring to the example of Fig. 14, the controller (170) can separate the first voice data, “Find a variety show featuring Yoo Jae-seok,” from the audio signal and the second voice data, “Isn’t Kian84 more interesting these days than Yoo Jae-seok?”

[0209] The controller (170) may select the first voice data among the first voice data and the second voice data as the voice data to be displayed as result information. In this case, the controller (170) may stop acquiring information about the second voice data. That is, the controller (170) may transmit the second voice data to an external server for information acquisition, and when the voice data to be displayed as result information is selected, the controller may stop transmitting the second voice data to the external server.

[0210] For example, the controller (170) may select the first voice data as the voice data to be displayed as result information at time t1. Accordingly, the controller (170) may stop transmitting the second voice data after transmitting "the fortress" to an external server. Thereafter, the controller (170) may turn off the microphone (175) when the end of the speech of the selected first voice data is recognized.

[0211] In this way, by stopping the acquisition of information from unselected voice data, there is an advantage in minimizing the power and cost consumed for acquiring unnecessary information.

[0212] Meanwhile, the display device (100) can also select multiple pieces of voice data to be displayed as result information.

[0213] FIG. 15 is a flowchart illustrating an operation method of a display device according to a second embodiment of the present disclosure.

[0214] The controller (170) can obtain an audio signal (S10), separate voice data from the audio signal (S20), and obtain information on each of the separated voice data (S30).

[0215] The controller (170) determines whether there are multiple voice data separated from the audio signal (S40), and if there is only one voice data, it can provide a result corresponding to one voice data (S50).

[0216] Steps S10 to S50 are the same as those described in Fig. 7, so redundant descriptions will be omitted.

[0217] If there are multiple voice data separated from an audio signal, the controller (170) can display result information for each of the multiple voice data (S2100).

[0218] The controller (170) may display result information corresponding to all of the plurality of voice data separated from the audio signal. Alternatively, the controller (170) may display result information corresponding to a preset number (e.g., two) of voice data among the plurality of voice data separated from the audio signal.

[0219] Meanwhile, there may be various ways to display the result information of multiple voice data.

[0220] FIGS. 16 and 17 are diagrams for explaining a method of displaying result information of a plurality of voice data according to the first embodiment of the present disclosure.

[0221] The controller (170) can obtain the first voice data, "Find a variety show featuring Yoo Jae-suk", and the second voice data, "What should I eat for dinner tonight?". The controller (170) can recognize the first and second voice data as speeches of registered users, the first and second users, respectively. Accordingly, the controller (170) can perform STT processing and NLP on the first and second voice data, respectively, and display corresponding result information.

[0222] The result information may include STT processed text and NLP results.

[0223] Referring to FIG. 17, the controller (170) may first display text obtained by STT processing the first voice data upon acquisition of the first voice data. The controller (170) may then display text corresponding to the first voice data, and when second voice data is acquired simultaneously, the controller (170) may simultaneously display text obtained by STT processing the second voice data. Thereafter, when the first and second voice data are finished speaking, the controller (170) may acquire results based on NLP and display them together.

[0224] At this time, the controller (170) can determine the position at which the result information of the plurality of voice data is displayed according to the user's position. Specifically, if the display device (100) is equipped with a plurality of microphones (175), the controller (170) can obtain the position information of the plurality of users corresponding to the plurality of voice data. The controller (170) can display each of the plurality of result information in an area corresponding to the position of each user based on the user's position information. For example, if the position of the first user is on the left and the position of the second user is on the right based on the direction in which the screen is viewed, the controller (170) can display the result information of the first voice data on the left and display the result information of the second voice data on the right.

[0225] Meanwhile, adjusting the display position of the result information according to the user's position may vary depending on the type of display device (100). For example, if the display device (100) is a car, the controller (170) may always display the result information of the voice data on the left side (towards the driver).

[0226] In the examples of FIGS. 16 and 17, the speech of the first and second voice data ends at the same time, but the speech of the first and second voice data may end at different times.

[0227] Figures 18 and 19 are exemplary drawings showing a case where the speech of the first and second voice data ends at different times.

[0228] The controller (170) can obtain the first voice data, “Find a variety show that Yoo Jae-seok appears in” and the second voice data, “Isn’t Kian84 more interesting these days than Yoo Jae-seok?”

[0229] The controller (170) may first display text obtained by STT processing the first voice data upon acquisition of the first voice data. The controller (170) may display text corresponding to the first voice data, and when second voice data is acquired together, the controller (170) may display text obtained by STT processing the second voice data together. Thereafter, if the speech of the first voice data is terminated first, the controller (170) may perform NLP on only the first voice data and display the result accordingly first. In this case, if the speech of the second voice data is not terminated, the text obtained by STT processing the second voice data may continue to be displayed. In addition, if the speech of the second voice data is terminated, the controller (170) may display the result of NLP processing the second voice data together.

[0230] Meanwhile, in the examples of FIGS. 17 and 19, the result information of the first and second voice data is displayed in the same size. Depending on the embodiment, the controller (170) may display the result information of the first and second voice data in different sizes.

[0231] FIG. 20 is a diagram for explaining a method for displaying result information of a plurality of voice data according to a second embodiment of the present disclosure.

[0232] The controller (170) can obtain an audio signal (S10), separate voice data from the audio signal (S20), and obtain information on each of the separated voice data (S30).

[0233] The controller (170) determines whether there are multiple voice data separated from the audio signal (S40), and if there is only one voice data, it can provide a result corresponding to one voice data (S50).

[0234] Steps S10 to S50 are the same as those described in Fig. 7, so redundant descriptions will be omitted.

[0235] The controller (170) can obtain the gain value of each of the plurality of voice data (S3100).

[0236] For example, the controller (170) can obtain a first gain value of first voice data among a plurality of voice data and a second gain value of second voice data among a plurality of voice data.

[0237] The controller (170) can display result information of multiple voice data based on the gain value (S3200).

[0238] The gain value may be a numerical value representing the size of the voice data. The controller (170) may calculate the gain value of the voice data in various ways. For example, the controller (170) may calculate the gain value based on the maximum decibel of each voice data. In other words, the controller (170) may calculate a larger gain value as the maximum decibel increases.

[0239] The controller (170) can display result information of multiple voice data in various ways based on the gain value. Hereinafter, various methods for displaying result information of multiple voice data based on the gain value will be described.

[0240] FIG. 21 is a first exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0241] According to the first embodiment, the controller (170) can adjust the size of the result information to be displayed based on the gain value of the voice data. Specifically, the controller (170) can display the result information in a larger size as the gain value of the voice data increases, and can display the result information in a smaller size as the gain value of the voice data decreases.

[0242] Referring to the example of Fig. 21, the controller (170) can obtain the gain values ​​of the first and second voice data as 7 and 3, respectively. In this case, the controller (170) can display the first result information of the first voice data larger than the second result information of the second voice data.

[0243] The larger the gain value, the closer it is to the user's voice data that attempted to give a voice command, so by displaying the corresponding result information in a larger size, there is an advantage in that it can provide the user with the information he needs in a larger size.

[0244] Additionally, the controller (170) may adjust the position of the result information to be displayed based on the gain value. For example, the controller (170) may display the result information of voice data with a large gain value among multiple voice data in the center, and display the result information of voice data with a small gain value on the side.

[0245] In addition, the controller (170) can display various result information depending on the gain value.

[0246] According to another embodiment, the controller (170) can display the result information of voice data having a large gain value among a plurality of voice data on the full screen, and display the result information of voice data having a small gain value as a PIP (Picture In Picture).

[0247] According to another embodiment, the controller (170) can selectably display result information of a plurality of voice data as tabs, and at the same time, arrange the tabs in order of increasing gain value.

[0248] Meanwhile, the controller (170) can adjust the size of the result information based on the gain value, and can also display it in the same size if the difference between the gain values ​​is small.

[0249] FIG. 22 is a second exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0250] The controller (170) can obtain the first and second gain values ​​corresponding to the first and second voice data, respectively, as 5 and 3. The controller (170) can calculate the difference between the gain values ​​of the first and second voice data, respectively. If the difference between the gain values ​​is less than a preset first threshold value (e.g., 3), the controller (170) can display the result information of the first voice data and the result information of the second voice data in the same size.

[0251] In this way, when the difference between the threshold values ​​is small, it is difficult to clearly distinguish which voice is the main speaker's speech, so there is an advantage in minimizing the problem of the main speaker's results being displayed small by providing them at the same size.

[0252] If the difference between gain values ​​is large, the controller (170) may not display the result information of voice data with a small gain value.

[0253] FIG. 23 is a third exemplary drawing showing a method for a display device according to an embodiment of the present disclosure to display result information based on a gain value of voice data.

[0254] The controller (170) can obtain the first and second gain values ​​corresponding to the first and second voice data as 7 and 1, respectively. The controller (170) can calculate the difference between the gain values ​​of the first and second voice data, respectively. If the difference between the gain values ​​is greater than or equal to a preset second threshold value (e.g., 5), the controller (170) can not display the result information of the voice data having a small gain value. In other words, if the difference between the gain values ​​is greater than or equal to a preset second threshold value (e.g., 5), the controller (170) can only display the result information of the voice data having a large gain value.

[0255] In this way, when the difference between the threshold values ​​is large, there is an advantage in that only necessary information can be provided by not providing the result information of voice data estimated to be ambient noise.

[0256] Meanwhile, in the case of continuous speech in which the display device (100) receives an audio signal within a predetermined time after providing the result of the voice data, the result information can be provided differently based on user information.

[0257] FIG. 24 is an exemplary diagram showing a method for providing results when a display device according to an embodiment of the present disclosure recognizes continuous speech.

[0258] The controller (170) can acquire new second voice data while displaying the result information of the first voice data. Referring to the example of FIG. 23, after acquiring the first voice data, "Hi LG, find a movie for me," the controller (170) can recognize the first voice data as the first user's speech and display the corresponding result information.

[0259] The controller (170) may acquire second voice data while displaying result information corresponding to the first voice data. In this case, the controller (170) may provide different result information corresponding to the second voice data based on user information corresponding to the second voice data.

[0260] First, refer to the example of (a) in Fig. 23. The controller (170) can obtain the second voice data, “Show only romance,” while displaying the result information corresponding to the first voice data. If the second voice data is recognized as the first user’s speech, the controller (170) can obtain and display the result information corresponding to the second voice data within the result information of the first voice data. That is, if the second voice data is recognized as the same user’s speech corresponding to the currently displayed result information, the controller (170) can obtain and display the result information corresponding to the second voice data within the result information of the first voice data.

[0261] Meanwhile, the controller (170) can obtain the second voice data, “Find me a drama,” while displaying the result information corresponding to the first voice data. The controller (170) can obtain the second voice data, “Find me a drama,” while displaying the result information corresponding to the first voice data. If the second voice data is recognized as speech of a second user, the controller (170) can obtain and display the result information corresponding to the second voice data regardless of the result information of the first voice data. That is, if the second voice data is recognized as speech of a user different from the user corresponding to the result information being displayed, the controller (170) can obtain and display the result information corresponding to the second voice data regardless of the result information of the first voice data.

[0262] And, at this time, the controller (170) may obtain result information corresponding to the second voice data within the result information of the first voice data and display them together.

[0263] In summary, when the controller (170) acquires second voice data, which is the speech of the second user, while displaying first result information corresponding to first voice data, which is the speech of the first user, it may display second result information, which acquires the result information of the second voice data within the first result information, and third result information, which acquires the result information of the second voice data regardless of the first result information.

[0264] As described above, recognizing continuous speech has the advantage of providing different result information depending on whether the user is the same or not, thereby providing each user with the desired result more accurately.

[0265] According to one embodiment of the present disclosure, the above-described method can be implemented as processor-readable code on a medium in which a program is recorded. Examples of processor-readable media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices.

[0266] The display device described above is not limited to the configuration and method of the embodiments described above, and the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.

Claims

1. At least one microphone; A controller that obtains user information and result information corresponding to at least one voice data separated from a signal obtained through the microphone; and Includes a display that displays the above result information, The above controller When there are multiple voice data separated from the above signal, if voice data to be displayed as result information is selected from among the multiple voice data, information acquisition of the voice data that is not selected is stopped. Display device.

2. In claim 1, The above controller When the speech of the voice data to be displayed as the above result information is finished, the microphone is turned off. Display device.

3. In claim 1, The above controller Based on the user information, the voice data corresponding to the registered user among the plurality of voice data is selected as the voice data to be displayed as the result information. Display device.

4. In claim 1, The above controller When multiple voice data to be displayed as the above result information is selected, multiple result information corresponding to each of the multiple selected voice data is displayed on the display. Display device.

5. In claim 4, The above controller In case there are multiple microphones, location information of multiple users corresponding to the multiple voice data is obtained, and each of the multiple result information is displayed in an area corresponding to the location of each user based on the location information. Display device.

6. In claim 4, The above controller Obtaining a gain value for each of the plurality of voice data, and displaying the plurality of result information differently based on the gain value. Display device.

7. In claim 6, The above controller Obtaining a first gain value of first voice data among the plurality of voice data and a second gain value of second voice data among the plurality of voice data, If the first gain value is greater than the second gain value, the result information of the first voice data is displayed to be greater than the result information of the second voice data. Display device.

8. In claim 7, The above controller If the difference between the first gain value and the second gain value is less than a preset threshold value, the result information of the first voice data and the result information of the second voice data are displayed in the same size. Display device.

9. In claim 7, The above controller If the difference between the first gain value and the second gain value is greater than a preset threshold value, only the result information of the voice data having a large gain value is displayed. Display device.

10. In claim 1, The above controller When new second voice data is acquired while displaying the result information of the first voice data, the result information corresponding to the second voice data is displayed differently based on the user information corresponding to the second voice data. Display device.

11. In claim 10, The above controller If the user corresponding to the first voice data and the user corresponding to the second voice data are the same, the second result information corresponding to the second voice data is obtained and displayed within the first result information of the first voice data. Display device.

12. In claim 10, The above controller If the user corresponding to the first voice data and the user corresponding to the second voice data are different, the second result information and the third result information are displayed together. The third result information is result information obtained in response to the second voice data regardless of the first result information. Display device.

13. In claim 1, The above controller When multiple voice data are acquired from the signal acquired through the above microphone, result information corresponding to the voice data corresponding to the user who uttered the activation word is displayed, When multiple voice data are acquired from a signal received from a remote control device, the result information corresponding to the voice data corresponding to the user with the largest gain value is displayed. Display device.

14. In claim 1, The above controller Selecting voice data containing a preset keyword among the above plurality of voice data as the voice data to be displayed as the result information Display device.

15. A step of separating at least one voice data from a signal acquired through a microphone; A step of obtaining user information and result information corresponding to separated voice data; and Including a step of displaying the above result information, In the case where there are multiple voice data separated from the signal, if voice data to be displayed as result information is selected from among the multiple voice data, the method further includes a step of stopping information acquisition of the voice data that is not selected. How the display device operates.

Citation Information

Patent Citations

  • Device and method for recognizing voice

    KR1020140072727A

  • BICM reception device and method corresponding to 64-symbol mapping and low density parity check codeword with 16200 length, 2 / 15 rate

    KR1020220032041A

  • Method of manufacturing optical laminate, and optical laminate

    KR1020240157559A

  • Apparatus and method for l1 andl2 triggered mobility procedure for inter central unit situation

    KR1020250125602A

  • KR20230066797A