Image display apparatus and system comprising same

The video display device and system address the accuracy issues in conventional voice recognition systems by incorporating user feedback on voice recognition results, improving recognition accuracy and system reliability.

WO2025110258A1PCT designated stage expired Publication Date: 2025-05-30LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/018700
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional voice recognition systems in video display devices often struggle with accuracy, requiring users to repeatedly speak commands until the desired recognition result is achieved, leading to inconvenience and reduced system reliability.

Method used

A video display device and system that request feedback on the result of voice recognition, determining the similarity between consecutive voice inputs, and outputting a screen for user feedback when the similarity meets a predetermined standard, allowing the system to perform operations based on user corrections.

Benefits of technology

This approach improves the accuracy of speech recognition by allowing user feedback on voice recognition results, reducing the need for repeated voice inputs and enhancing system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023018700_30052025_PF_FP_ABST
    Figure KR2023018700_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image display apparatus and a system comprising same. The image display apparatus according to one embodiment of the present disclosure comprises: a display; a user input interface unit for receiving a user input; and a control unit, wherein the control unit can determine, if a first voice input is received through the user input interface unit, whether the first voice input was received within a predetermined time after a second voice input had been received immediately before, determine the similarity between the first voice input and the second voice input if the first voice input is received within the predetermined time after the second voice input was received, output, if the similarity is greater than or equal to a predetermined reference, through the display, a screen for requesting feedback about the result of performing voice recognition on the first voice input, and perform an operation corresponding to a feedback input received through the user input interface unit.
Need to check novelty before this filing date? Find Prior Art

Description

Video display device and system including the same

[0001] The present disclosure relates to a video display device and a system including the same.

[0002] A video display device is a device that displays images for the user to view. For example, a video display device may include a television (TV), monitor, or notebook computer equipped with a liquid crystal display (LCD) using liquid crystals or an organic light-emitting diode (OLED) display using organic light-emitting diodes (OLED).

[0003] With the advancement of technology, various services that apply voice recognition technology are being developed and provided in various fields, and various technologies are being developed to apply voice recognition technology to video display devices such as TVs.

[0004] Systems using conventional speech recognition technology perform speech recognition on each input from a user and perform actions based on the results of the speech recognition. When the actual speech produced by the user differs from the results of the speech recognition performed on that speech, the user typically re-speaks the same or similar speech to achieve the desired function or action. However, if the results of the speech recognition for the user's actual speech repeatedly differ from the actual speech produced by the user, the user must repeatedly re-speak until the desired speech recognition result is achieved, resulting in inconvenience and lowering the reliability of the system.

[0005] The present disclosure aims to solve the above-mentioned and other problems.

[0006] Another object is to provide a video display device and a system including the same that can request feedback on the result of performing speech recognition, taking into account the user's speech intention.

[0007] Another object is to provide a video display device and a system including the same that can accurately perform a function or action desired by a user based on the user's feedback on the result of performing voice recognition.

[0008] Another object is to provide a video display device and a system including the same that can improve the accuracy of speech recognition based on user feedback on the result of performing speech recognition.

[0009] In order to achieve the above object, an image display device according to one embodiment of the present disclosure includes: a display; a user input interface unit for receiving a user input; and a control unit, wherein, when a first voice input is received through the user input interface unit, the control unit determines whether the first voice input was received within a predetermined time after a second voice input was received immediately before; when the first voice input is received within the predetermined time after the second voice input is received, the control unit determines a similarity between the first voice input and the second voice input; and when the similarity is equal to or greater than a predetermined standard, outputs a screen requesting feedback on a result of performing voice recognition on the first voice input through the display; and performs an operation corresponding to a feedback input received through the user input interface unit.

[0010] In order to achieve the above object, a system according to one embodiment of the present disclosure includes a video display device and a first server that performs voice recognition on a voice input received from the video display device, wherein the first server, when a first voice input is received from the video display device, determines whether the first voice input was received within a predetermined time after a second voice input was received immediately before, and, when the first voice input was received within the predetermined time after the second voice input was received, determines a similarity between the first voice input and the second voice input, and, when the similarity is equal to or higher than a predetermined standard, transmits a command corresponding to the calculated similarity to the video display device, and the video display device outputs a screen requesting feedback on a result of performing voice recognition on the first voice input based on the command received from the first server, and performs an operation corresponding to the received feedback input.

[0011] The effects of the video display device and the system including the same according to the present disclosure are described as follows.

[0012] According to at least one embodiment of the present disclosure, feedback on the result of performing speech recognition can be requested by taking into account the user's speech intention, thereby preventing the user from repeatedly speaking.

[0013] According to at least one embodiment of the present disclosure, a user's desired function or action can be accurately performed based on the user's feedback on the result of performing voice recognition.

[0014] According to at least one embodiment of the present disclosure, the accuracy of speech recognition can be improved based on user feedback on the result of performing speech recognition.

[0015] Further scope of the applicability of the present disclosure will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present disclosure will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present disclosure, are given by way of example only.

[0016] FIG. 1 is a diagram illustrating an image display system according to one embodiment of the present disclosure.

[0017] Figure 2 is an internal block diagram of the image display device of Figure 1.

[0018] Figure 3 is an internal block diagram of the control unit of Figure 2.

[0019] FIG. 4a is a drawing illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.

[0020] Figure 5 is a drawing referenced in the description of the first server of Figure 1.

[0021] FIG. 6 is a flowchart of an operation method of an image display device according to one embodiment of the present disclosure.

[0022] FIG. 7 is a flowchart of a method of operating a system according to one embodiment of the present disclosure.

[0023] FIGS. 8 to 11 are drawings for reference in explaining the operation of a video display device according to one embodiment of the present disclosure.

[0024] FIG. 12 is a flowchart of a method of operating a system according to another embodiment of the present disclosure.

[0025] FIG. 13 and FIG. 14 are drawings for reference in explaining the operation of an image display device according to another embodiment of the present disclosure.

[0026] FIG. 15 and FIG. 16 are flowcharts of a method of operating a system according to various embodiments of the present disclosure.

[0027] Hereinafter, the present disclosure will be described in detail with reference to the drawings. In the drawings, portions irrelevant to the description are omitted to clearly and concisely describe the present disclosure, and the same reference numerals are used for identical or extremely similar portions throughout the specification.

[0028] The suffixes "module" and "part" used in the following description are given solely for the convenience of writing this specification and do not impart any particularly significant meaning or role to the components themselves. Therefore, the terms "module" and "part" may be used interchangeably.

[0029] In this application, it should be understood that terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0030] Additionally, while terms such as "first" and "second" may be used in this specification to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another.

[0031] FIG. 1 is a diagram illustrating an image display system according to various embodiments of the present invention.

[0032] Referring to FIG. 1, the image display system (10) may include an image display device (100) and / or a remote control device (200).

[0033] The image display device (100) may be a device that processes and outputs an image. The image display device (100) is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.

[0034] The video display device (100) can receive a broadcast signal, process the signal, and output the processed broadcast image. When the video display device (100) receives a broadcast signal, the video display device (100) may correspond to a broadcast receiving device.

[0035] The video display device (100) can receive broadcast signals wirelessly via an antenna, or can receive broadcast signals wired via a cable. For example, the video display device (100) can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.

[0036] The remote control device (200) can be connected to the image display device (100) by wire and / or wirelessly, and can provide various control signals to the image display device (100). At this time, the remote control device (200) can include a device that establishes a wired or wireless network with the image display device (100), and transmits various control signals to the image display device (100) through the established network, or receives signals related to various operations processed in the image display device (100) from the image display device (100).

[0037] For example, various input devices such as a mouse, keyboard, space remote control, trackball, joystick, etc. can be used as the remote control device (200). The remote control device (200) can be referred to as an external device, and it is to be noted in advance that external devices and remote control devices can be used interchangeably as needed.

[0038] The video display device (100) can be connected to only a single remote control device (200) or can be connected to two or more remote control devices (200) simultaneously, and can change objects displayed on the screen or adjust the status of the screen based on control signals provided from each remote control device (200).

[0039] Meanwhile, the video display system (10) may further include at least one server (300). The video display device (100) may transmit and receive data with at least one server (300). For example, the video display device (100) may transmit and receive data with at least one server (300) via a network such as the Internet.

[0040] According to one embodiment, at least one server (300) may include a first server (400) that performs voice recognition, a second server (500) that processes data using a super-giant artificial intelligence model (hereinafter, super-giant AI), a third server (600) that provides content, etc.

[0041] Figure 2 is an internal block diagram of the image display device of Figure 1.

[0042] Referring to FIG. 2, the video display device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).

[0043] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulator unit (120).

[0044] Meanwhile, unlike the drawing, the image display device (100) may include only the broadcast reception unit (105) and the external device interface unit (130) among the broadcast reception unit (105), the external device interface unit (130), and the network interface unit (135). That is, the image display device (100) may not include the network interface unit (135).

[0045] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.

[0046] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).

[0047] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.

[0048] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.

[0049] The demodulation unit (120) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (110).

[0050] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.

[0051] The stream signal output from the demodulation unit (120) can be input to the control unit (170). The control unit (170) can output an image through the display (180) and output an audio through the audio output unit (185) after performing demultiplexing, image / audio signal processing, etc.

[0052] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).

[0053] The external device interface unit (130) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.

[0054] In addition, the external device interface unit (130) can establish a communication network with various remote control devices (200) as illustrated in FIG. 1, and receive a control signal related to the operation of the image display device (100) from the remote control device (200) or transmit data related to the operation of the image display device (100) to the remote control device (200).

[0055] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit can include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).

[0056] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, in mirroring mode, the external device interface unit (130) may receive device information, running application information, application images, etc. from the mobile terminal.

[0057] The external device interface unit (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.

[0058] The network interface unit (135) can provide an interface for connecting the video display device (100) to a wired / wireless network including the Internet.

[0059] The network interface unit (135) may include a communication module (not shown) for connection to a wired / wireless network. For example, the network interface unit (135) may include a communication module for WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), etc.

[0060] The network interface unit (135) can transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.

[0061] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and information related thereto provided by a content provider or network provider via a network.

[0062] The network interface unit (135) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.

[0063] The network interface unit (135) can select and receive a desired application from among applications open to the public through a network.

[0064] The storage unit (140) may store programs for signal processing and control within the control unit (170), or may store processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored application programs upon request from the control unit (170).

[0065] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (170).

[0066] The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).

[0067] The storage unit (140) can store information about a specific broadcast channel through a channel memory function such as a channel map.

[0068] Although the storage unit (140) of FIG. 2 is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).

[0069] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.

[0070] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.

[0071] For example, a user input signal such as power on / off, channel selection, screen setting, etc. may be transmitted / received from a remote control device (200), a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. may be transmitted to the control unit (170), a user input signal input from a sensor unit (not shown) that senses a user's gesture may be transmitted to the control unit (170), or a signal from the control unit (170) may be transmitted to the sensor unit.

[0072] The input unit (160) may be provided on one side of the main body of the video display device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.

[0073] The input unit (160) can receive various user commands related to the operation of the video display device (100) and transmit a control signal corresponding to the input command to the control unit (170).

[0074] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.

[0075] The control unit (170) may include at least one processor, and may control the overall operation of the image display device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.

[0076] The control unit (170) can demultiplex a stream input through the tuner unit (110), the demodulator unit (120), the external device interface unit (130), or the network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.

[0077] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).

[0078] The display (180) may include a display panel (not shown) having a plurality of pixels.

[0079] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display (180) may convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (170) to generate driving signals for the plurality of pixels.

[0080] The display (180) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display, etc., and may also be a 3D display. The 3D display (180) can be divided into a glasses-free type and a glasses type.

[0081] Meanwhile, the display (180) is configured as a touch screen and can be used as an input device in addition to an output device.

[0082] The audio output unit (185) receives a signal processed by the control unit (170) and outputs it as voice.

[0083] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (170) can also be input to an external output device through the external device interface unit (130).

[0084] The voice signal processed in the control unit (170) can be output as sound to the audio output unit (185). In addition, the voice signal processed in the control unit (170) can be input to an external output device through the external device interface unit (130).

[0085] Although not shown in FIG. 2, the control unit (170) may include a demultiplexing unit, an image processing unit, etc. This will be described later with reference to FIG. 3.

[0086] In addition, the control unit (170) can control the overall operation within the video display device (100). For example, the control unit (170) can control the tuner unit (110) to select (tune) a broadcast corresponding to a channel selected by the user or a previously stored channel.

[0087] In addition, the control unit (170) can control the image display device (100) by a user command or internal program input through the user input interface unit (150).

[0088] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a moving image, and may be a 2D image or a 3D image.

[0089] Meanwhile, the control unit (170) can cause a predetermined 2D object to be displayed within an image displayed on the display (180). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.

[0090] Meanwhile, the image display device (100) may further include a camera (not shown). The camera can capture images of a user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the image display device (100) above the display (180) or can be separately positioned. Image information captured by the camera can be input to the control unit (170).

[0091] The control unit (170) can recognize the user's location based on the image captured by the camera. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the image display device (100). In addition, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.

[0092] The control unit (170) can detect the user's gesture based on an image captured from the camera unit, a signal detected from the sensor unit, or a combination thereof.

[0093] The power supply unit (190) can supply power to the entire image display device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a system on chip (SOC), a display (180) for image display, and an audio output unit (185) for audio output.

[0094] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.

[0095] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (150) and display or output the same as voice on the remote control device (200).

[0096] Meanwhile, the above-described video display device (100) may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.

[0097] Meanwhile, the block diagram of the image display device (100) illustrated in FIG. 2 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the image display device (100) actually implemented.

[0098] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0099] Figure 3 is an internal block diagram of the control unit of Figure 2.

[0100] Referring to FIG. 3, a control unit (170) according to one embodiment of the present invention may include a demultiplexer (310), an image processing unit (320), a processor (330), an OSD generation unit (340), a mixer (345), a frame rate conversion unit (350), and / or a formatter (360). In addition, an audio processing unit (not shown) and a data processing unit (not shown) may be further included.

[0101] The demultiplexer (310) can demultiplex an input stream. For example, when MPEG-2 TS is input, it can be demultiplexed to separate it into video, audio, and data signals, respectively. Here, the stream signal input to the demultiplexer (310) may be a stream signal output from the tuner (110), the demodulator (120), or the external device interface (130).

[0102] The image processing unit (320) can perform image processing of a demultiplexed image signal. To this end, the image processing unit (320) may be equipped with an image decoder (325) and a scaler (335).

[0103] The video decoder (325) can decode a demultiplexed video signal, and the scaler (335) can perform scaling so that the resolution of the decoded video signal can be output on the display (180).

[0104] The video decoder (325) may include decoders of various standards. For example, it may include an MPEG-2, H.264 decoder, a 3D video decoder for color images and depth images, a decoder for multi-view images, etc.

[0105] The processor (330) can control the overall operation within the video display device (100) or the control unit (170). For example, the processor (330) can control the tuner (110) to select (tune) a broadcast corresponding to a channel selected by the user or a pre-stored channel.

[0106] In addition, the processor (330) can control the image display device (100) by a user command or internal program input through the user input interface unit (150).

[0107] Additionally, the processor (330) can perform data transmission control with the network interface unit (135) or the external device interface unit (130).

[0108] Additionally, the processor (330) can control the operation of the demultiplexing unit (310), the image processing unit (320), the OSD generation unit (340), etc. within the control unit (170).

[0109] The OSD generation unit (340) can generate OSD signals based on user input or on its own. For example, based on a user input signal input through the input unit (160), it can generate signals for displaying various information in the form of graphics or text on the screen of the display (180).

[0110] The generated OSD signal may include various data such as the user interface screen of the video display device (100), various menu screens, widgets, icons, etc. In addition, the generated OSD signal may include a 2D object or a 3D object.

[0111] Additionally, the OSD generation unit (340) can generate a pointer that can be displayed on the display (180) based on a pointing signal input from the remote control device (200).

[0112] The OSD generation unit (340) may include a pointing signal processing unit (not shown) that generates a pointer. It is also possible for the pointing signal processing unit (not shown) to be provided separately rather than within the OSD generation unit (240).

[0113] The mixer (345) can mix the OSD signal generated by the OSD generation unit (340) and the decoded image signal processed by the image processing unit (320). The mixed image signal can be provided to the frame rate conversion unit (350).

[0114] The frame rate converter (FRC) (350) can convert the frame rate of an input video. Meanwhile, the frame rate converter (350) can also output the video as is without a separate frame rate conversion.

[0115] The formatter (360) can arrange left-eye image frames and right-eye image frames of a frame rate-converted 3D image. In addition, it can output a synchronization signal (Vsync) for opening the left-eye glasses and right-eye glasses of a 3D viewing device (not shown).

[0116] Meanwhile, the formatter (360) can change the format of the input video signal into a video signal for display on the display (180) and output it.

[0117] Additionally, the formatter (360) can change the format of a 3D video signal. For example, it can change the format to any one of various 3D formats, such as a side-by-side format, a top-down format, a frame sequential format, an interlaced format, and a checker box format.

[0118] Meanwhile, the formatter (360) can also convert a 2D image signal into a 3D image signal. For example, according to a 3D image generation algorithm, an edge or a selectable object can be detected within a 2D image signal, and an object or a selectable object according to the detected edge can be separated and generated as a 3D image signal. At this time, the generated 3D image signal can be separated and aligned into a left-eye image signal (L) and a right-eye image signal (R), as described above.

[0119] Meanwhile, although not shown in the drawing, a 3D processor (not shown) for 3D effect signal processing may be further placed after the formatter (360). This 3D processor may process brightness, tint, and color adjustments of the image signal to improve the 3D effect. For example, it may perform signal processing to make the image clear at a close distance and blurry at a long distance. Meanwhile, the function of this 3D processor may be incorporated into the formatter (360) or incorporated into the image processing unit (320).

[0120] Meanwhile, the audio processing unit (not shown) within the control unit (170) can perform audio processing of the demultiplexed audio signal. For this purpose, the audio processing unit (not shown) can be equipped with various decoders.

[0121] Additionally, the audio processing unit (not shown) within the control unit (170) can process bass, treble, volume control, etc.

[0122] A data processing unit (not shown) within the control unit (170) can perform data processing of a demultiplexed data signal. For example, if the demultiplexed data signal is an encoded data signal, it can be decoded. The encoded data signal may be electronic program guide information (EPG) information that includes broadcast information such as the start time and end time of a broadcast program broadcast on each channel.

[0123] Meanwhile, the block diagram of the control unit (170) illustrated in FIG. 3 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the control unit (170) actually implemented.

[0124] In particular, the frame rate conversion unit (350) and the formatter (360) are not provided within the control unit (170), but may be provided separately, or may be provided separately as one module.

[0125] FIG. 4a is a drawing illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.

[0126] Referring to FIG. 4a, it can be confirmed that a pointer (205) corresponding to a remote control device (200) is displayed on the display (180) of the video display device (100).

[0127] Referring to (a) of Fig. 4a, the user can move or rotate the remote control device (200) up and down, left and right, forward and backward. At this time, the pointer (205) displayed on the display (180) of the image display device (100) can be displayed in response to the movement of the remote control device (200). Since the pointer (205) of this remote control device (200) moves and is displayed according to the movement in 3D space, as shown in the drawing, it can be called a space remote control or a 3D pointing device.

[0128] Referring to (b) of FIG. 4a, when the user moves the remote control device (200) to the left, it can be confirmed that the pointer (205) displayed on the display (180) of the image display device (100) also moves to the left in response to the movement of the remote control device (200).

[0129] Information about the movement of the remote control device (200) detected through the sensor of the remote control device (200) can be transmitted to the image display device (100). The image display device (100) can calculate the coordinates of the pointer (205) from the information about the movement of the remote control device (200). The image display device (100) can display the pointer (205) to correspond to the calculated coordinates.

[0130] Referring to (c) of FIG. 4a, a user can move the remote control device (200) away from the display (180) while pressing a specific button provided on the remote control device (200). As a result, a selection area within the display (180) corresponding to the pointer (205) can be zoomed in and displayed in an enlarged manner. Conversely, when a user moves the remote control device (200) closer to the display (180) while pressing a specific button provided on the remote control device (200), a selection area within the display (180) corresponding to the pointer (205) can be zoomed out and displayed in a reduced manner.

[0131] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.

[0132] Meanwhile, when the user presses a specific button within the remote control device (200), recognition of up, down, left, and right movements may be excluded. That is, when the remote control device (200) moves away from or toward the display (180), up, down, left, and right movements are not recognized, and only forward and backward movements may be recognized. When the user does not press a specific button within the remote control device (200), only up, down, left, and right movements of the remote control device (200) may be recognized, and accordingly, only the pointer (205) may be moved.

[0133] Meanwhile, the movement speed or movement direction of the pointer (205) can correspond to the movement speed or movement direction of the remote control device (200).

[0134] Referring to FIG. 4b, the remote control device (200) may include a wireless communication unit (220), a user input unit (230), a sensor unit (240), an output unit (250), a power supply unit (260), a storage unit (270), and / or a control unit (280).

[0135] The wireless communication unit (220) can transmit and receive signals with the image display device (100).

[0136] In this embodiment, the remote control device (200) may be equipped with an RF module (221) capable of transmitting and receiving signals with the image display device (100) according to RF (Radio frequency) communication standards. In addition, the remote control device (200) may be equipped with an IR module (223) capable of transmitting and receiving signals with the image display device (100) according to IR (Infrared radiation) communication standards.

[0137] The remote control device (200) can transmit a signal including information about the movement of the remote control device (200) to the image display device (100) through the RF module (221). The remote control device (200) can receive the signal transmitted by the image display device (100) through the RF module (221).

[0138] The remote control device (200) can transmit commands for power on / off, channel change, volume change, etc. to the video display device (100) through the IR module (223).

[0139] The user input unit (230) may be composed of a keypad, buttons, a touch pad, a touch screen, etc. The user can input commands related to the video display device (100) to the remote control device (200) by operating the user input unit (230).

[0140] When the user input unit (230) has a hard key button, the user can input a command related to the video display device (100) to the remote control device (200) through a push operation of the hard key button.

[0141] When the user input unit (230) has a touch screen, the user can input commands related to the video display device (100) using the remote control device (200) by touching the soft keys of the touch screen.

[0142] Meanwhile, the user input unit (230) may be equipped with various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the present invention.

[0143] The user input unit (230) may be equipped with a microphone. The user may speak into the microphone provided in the user input unit (230). At this time, the microphone provided in the user input unit (230) may receive the voice spoken by the user.

[0144] The sensor unit (240) may be equipped with a gyro sensor (241) or an acceleration sensor (243). The gyro sensor (241) can sense the movement of the remote control device (200).

[0145] The gyro sensor (241) can sense information about the operation of the remote control device (200) based on the x, y, and z axes. The acceleration sensor (243) can sense information about the movement speed of the remote control device (200). Meanwhile, the sensor unit (240) may further include a distance measuring sensor capable of sensing the distance from the display (180).

[0146] The output unit (250) can output an image or sound corresponding to the operation of the user input unit (230) or to a signal transmitted from the image display device (100). Through the output unit (250), the user can recognize whether the user input unit (230) is being operated or whether the image display device (100) is being controlled.

[0147] The output unit (250) may include an LED module (251) including at least one light-emitting element (e.g., an LED (Light Emitting Diode)), a vibration module (253) that generates vibration, a sound output module (255) that outputs sound, and / or a display module (257) that outputs an image.

[0148] The power supply unit (260) can supply power to each component provided in the remote control device (200). The power supply unit (260) can include at least one battery (not shown).

[0149] The power supply unit (260) can prevent unnecessary power consumption by stopping the power supply to each component provided in the remote control device (200) when movement of the remote control device (200) is not detected for a predetermined period of time through the sensor unit (240).

[0150] The power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when a predetermined event occurs. For example, the power supply unit (260) can resume power supply to each component when a predetermined key equipped in the remote control device (200) is operated. For example, the power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when movement of the remote control device (200) is detected through the sensor unit (240).

[0151] The storage unit (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).

[0152] When the remote control device (200) wirelessly transmits and receives signals through the image display device (100) and the RF module (221), the remote control device (200) and the image display device (100) can transmit and receive signals through a predetermined frequency band. The control unit (280) of the remote control device (200) can store and reference information about the frequency band, etc., through which signals can be wirelessly transmitted and received between the remote control device (200) and the image display device (100) paired therewith in the storage unit (270).

[0153] The control unit (280) may include at least one processor, and may control the overall operation of the remote control device (200) using the processor included therein.

[0154] The control unit (280) can transmit a control signal corresponding to a predetermined key operation of the user input unit (230) or a control signal corresponding to the movement of the remote control device (200) sensed by the sensor unit (240) to the image display device (100) via the wireless communication unit (220).

[0155] The user input interface unit (150) of the video display device (100) may be equipped with a wireless communication unit (151) capable of transmitting and receiving signals wirelessly with a remote control device (200), and a coordinate value calculation unit (155) capable of calculating the coordinate value of a pointer corresponding to the operation of the remote control device (200).

[0156] The user input interface unit (150) can wirelessly transmit and receive signals with the remote control device (200) via the RF module (152). In addition, the user input interface unit (150) can receive signals transmitted by the remote control device (200) according to the IR communication standard via the IR module (153).

[0157] The coordinate value calculation unit (155) can calculate the coordinate values ​​(x, y) of the pointer (205) to be displayed on the display (170) by correcting hand tremors or errors from a signal corresponding to the operation of the remote control device (200) received through the wireless communication unit (151).

[0158] A transmission signal of a remote control device (200) input to a video display device (100) through a user input interface unit (150) can be transmitted to a control unit (170) of the video display device (100). The control unit (170) of the video display device (100) can check information about the operation and key operation of the remote control device (200) from the signal transmitted from the remote control device (200) and control the video display device (100) in response thereto.

[0159] As another example, the remote control device (200) can calculate the pointer coordinate values ​​corresponding to the operation and output them to the user input interface unit (150) of the image display device (100). In this case, the user input interface unit (150) of the image display device (100) can transmit information about the received pointer coordinate values ​​to the control unit (170) without a separate hand shake or error correction process.

[0160] In addition, as another example, the coordinate value calculation unit (155) may be provided inside the control unit (170) rather than the user input interface unit (150), unlike in the drawing.

[0161] Figure 5 is a drawing referenced in the description of the first server of Figure 1.

[0162] Referring to FIG. 5, the first server (400) may include a relay server (410), an STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), an AI (Artificial Intelligence) server (440), and / or a database (450). In the present disclosure, the relay server (410), the STT server (420), the NLP server (430), and the AI ​​server (440) are described as being distinct from each other, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), and the AI ​​server (440) may be configured as one server.

[0163] The relay server (410) can communicate with the video display device (100). The relay server (410) can transfer data between the STT server (420), the NLP server (430), and the video display device (100). The relay server (410) can store at least a portion of the data transferred between the STT server (420), the NLP server (430), and the video display device (100).

[0164] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the video display device (100) via the relay server (410). The STT server (420) may also be referred to as an ASR (Automatic Speech Recognition) server.

[0165] The STT server (420) can improve the accuracy of speech-to-text conversion using a language model. The language model can refer to a model that can calculate the probability of a sentence or the probability of a subsequent word given previous words. For example, the language model can include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. In other words, the STT server (420) can use the language model to determine whether text data converted from speech data has been appropriately converted, thereby improving the accuracy of conversion into text data.

[0166] The NLP server (430) can receive text data. Based on the received text data, the NLP server (430) can perform intent analysis on the text data. The NLP server (430) can transmit intent analysis information indicating the results of the intent analysis to the video display device (100) via the relay server (410).

[0167] According to one embodiment, the NLP server (430) may sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate intent analysis information. The morphological analysis step is a step of classifying text data corresponding to speech uttered by a user into morphemes, which are the smallest units having meaning, and determining which part of speech each classified morpheme has. The syntax analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and to determine what kind of relationship exists between each of the classified phrases. Through the syntax analysis step, the subject, object, and modifiers of the speech uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the speech uttered by the user using the results of the syntax analysis step. Specifically, the speech act analysis step is a step of determining the intent of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.

[0168] The AI ​​server (440) may transmit a response according to a request from the NLP server (430) to the relay server (410) or the NLP server (430). For example, when a content search request is received from the NLP server (430), the AI ​​server (440) may transmit data for at least one content corresponding to the received content search request to the NLP server (430). For example, when a recommendation request for a keyword associated with a predetermined keyword (hereinafter, “associated keyword”) is received from the NLP server (430), the AI ​​server (440) may transmit data for at least one associated keyword corresponding to the predetermined keyword to the NLP server (430).

[0169] In one embodiment, the AI ​​server (440) may process data using a super-large AI. For example, the AI ​​server (440) may provide a response to a prompt received from the NLP server (430).

[0170] In one embodiment, the AI ​​server (440) may include multiple AI agent servers. Here, the multiple AI agent servers may each correspond to a type of request from the NLP server (430). For example, the AI ​​server (440) may utilize a first AI agent server when a text-related request is received from the NLP server (430), and may utilize a second AI agent server when a video-related request is received.

[0171] The database (450) can store data used to generate responses from the AI ​​server (440). The database (450) can include multiple sub-databases (451 to 453). For example, the database (450) can include a first database (451) that stores text such as syllables and words, a second database (452) that stores voices, a third database (453) that stores various learning models, etc.

[0172] FIG. 6 is a flowchart of an operation method of an image display device according to one embodiment of the present disclosure.

[0173] Referring to FIG. 6, the image display device (100) can receive a first voice input in operation S610. For example, when a user's voice is input into the remote control device (200) through a microphone included in the remote control device (200), the remote control device (200) can transmit data regarding the input user's voice to the image display device (100). At this time, the image display device (100) can receive a voice input including data regarding the user's voice from the remote control device (200).

[0174] According to one embodiment, the video display device (100) can receive a voice input from the remote control device (200) when the voice input receiving function (hereinafter, voice recognition function) is activated. For example, when a user presses a specific button included in the remote control device (200), the remote control device (200) can transmit a start command to the video display device (100). Meanwhile, when the voice recognition function of the video display device (100) is activated, the remote control device (200) can activate a microphone included in the user input unit (230).

[0175] Meanwhile, when a voice input is received, the video display device (100) can transmit the voice input to the first server (400). In addition, the video display device (100) can receive text corresponding to the voice input from the first server (400). For example, when the video display device (100) transmits voice data to the first server (400), the STT server (420) can convert the voice data into text data and transmit it to the video display device (100).

[0176] According to one embodiment, the video display device (100) can transmit voice data in preset units such as syllables and words to the first server (400). That is, when a user utters a sentence, the video display device (100) can transmit voice data in preset units to the first server (400) and receive text data in preset units from the first server (400) while a voice input corresponding to the sentence or phrase is received from the remote control device (200).

[0177] According to one embodiment, the video display device (100) can output text corresponding to a voice input through the display (180). For example, the video display device (100) can output text corresponding to a preset unit of text data received from the first server (400) through the display (180) while receiving a voice input corresponding to a sentence or phrase from the remote control device (200).

[0178] The video display device (100) can determine, in operation S620, whether the first voice input was received within a predetermined time after the second voice input was received immediately before. For example, the video display device (100) can determine whether the first voice input was received before a predetermined time of 15 seconds elapses from the time the second voice input was received from the remote control device (200).

[0179] In operation S630, if it is determined that the first voice input has been received within a predetermined time after the second voice input has been received, the video display device (100) can determine the similarity between the first voice input and the second voice input. That is, if voice inputs with a high degree of similarity are continuously received within a predetermined time, the video display device (100) can determine that the user has spoken again within the predetermined time because the voice actually spoken by the user and the result of voice recognition are different from each other.

[0180] For example, the image display device (100) can store a text (hereinafter, referred to as the second text) corresponding to the result of performing voice recognition on the second voice input. At this time, the image display device (100) can determine the similarity between the text (hereinafter, referred to as the first text) corresponding to the result of performing voice recognition on the first voice input and the second text if the first voice input is received within a predetermined time. Meanwhile, the image display device (100) can delete the second text if the first voice input is not received within a predetermined time after the second voice input is received.

[0181] According to one embodiment, the similarity between the first voice input and the second voice input can be calculated using a text-based similarity calculation method such as cosine similarity, Jaccard coefficient, correlation coefficient, and Hamming distance. In the present disclosure, the similarity between the first voice input and the second voice input is calculated based on cosine similarity, but is not limited thereto. For example, a first vector corresponding to a first text and a second vector corresponding to a second text can be generated. In this case, the cosine similarity between the first vector and the second vector can be calculated based on the following mathematical expression 1.

[0182]

[0183] Here, is the dot product of two vectors, and can mean the magnitude of two vectors. That is, cosine similarity can be calculated as the value obtained by dividing the inner product of two vectors by the product of the magnitudes of each vector. Cosine similarity can range from -1 to 1, and the closer it is to 1, the more similar the two vectors can be judged to be.

[0184] According to one embodiment, the similarity between the first voice input and the second voice input may be calculated by the first server (400). For example, if the image display device (100) determines that the first voice input has been received within a predetermined time after the second voice input has been received, the image display device (100) may transmit the first text and the second text to the first server (400). At this time, the first server (400) may transmit the result of calculating the similarity between the first text and the second text to the image display device (100). The image display device (100) may determine the similarity between the first voice input and the second voice input based on the similarity between the first text and the second text received from the first server (400).

[0185] The video display device (100) can determine, in operation S640, whether the similarity between the first voice input and the second voice input is greater than a predetermined threshold. For example, the video display device (100) can determine whether the cosine similarity between the first voice input and the second voice input is greater than a preset threshold.

[0186] In operation S650, the video display device (100) may provide the user with the result of performing voice recognition on the first voice input if the similarity between the first voice input and the second voice input is above a predetermined standard. For example, the video display device (100) may output a screen including the first text through the display (180).

[0187] According to one embodiment, the video display device (100) may output a screen requesting feedback on the result of performing voice recognition on a first voice input. For example, the screen requesting feedback may include a first text, a message requesting confirmation of whether the first text corresponds to the voice spoken by the user, an object used for correcting the result of performing voice recognition, a group of candidates for text corresponding to the first voice input (hereinafter, "candidate text"), etc.

[0188] The video display device (100), in operation S660, can receive user feedback (hereinafter, feedback input) regarding the result of performing voice recognition on the first voice input. For example, the video display device (100) can receive a feedback input from the remote control device (200) that determines whether the first text corresponds to the voice spoken by the user. For example, the video display device (100) can receive a feedback input from the remote control device (200) that corrects the first text, such as deleting or modifying the first text. For example, the video display device (100) can receive a feedback input from the remote control device (200) that selects one of a plurality of candidate texts.

[0189] The video display device (100) can perform an operation corresponding to user feedback in operation S670. For example, when an application providing a streaming service is running, the video display device (100) can search for content based on text (hereinafter, "final text") determined to correspond to the first voice input based on the feedback input. For example, when an application using a search engine is running, the video display device (100) can output search results for the final text.

[0190] According to one embodiment, the image display device (100) may transmit result data corresponding to the feedback input to the first server (400). For example, the result data may include a first text, a second text, a final text, etc. The first server (400) may store the result data received from the image display device (100). The first server (400) may train a language model based on the result data.

[0191] Meanwhile, in operation S680, if a second voice input is received and a predetermined amount of time has passed before a first voice input is received, or if the similarity between the first voice input and the second voice input is below a predetermined standard, the video display device (100) may perform an operation corresponding to the first voice input. For example, if an application providing a streaming service is running, the video display device (100) may search for content corresponding to the first text.

[0192] Figure 7 is a flowchart illustrating a method of operating a system according to one embodiment of the present disclosure. Any details that overlap with those described in Figure 6 will be omitted for detailed explanation.

[0193] Referring to FIG. 7, the video display device (100) can receive voice input in operation S701.

[0194] The video display device (100) can transmit the received voice input to the first server (400) in operation S702.

[0195] The first server (400) can, in operation S703, convert voice data corresponding to voice input received from the video display device (100) into text data.

[0196] The first server (400) can transmit text data corresponding to voice input to the video display device (100) in operation S704.

[0197] The video display device (100), in operation S705, can determine the similarity between the first voice input currently received and the second voice input received immediately before, if a voice input is received again within a predetermined time after the voice input was received immediately before.

[0198] In operation S706, the video display device (100) can provide the user with the result of performing voice recognition on the first voice input when the similarity between the first voice input and the second voice input is greater than a predetermined standard.

[0199] The video display device (100) can receive user feedback on the result of performing voice recognition on the first voice input in operation S707.

[0200] The video display device (100) can perform an operation corresponding to the user's feedback in operation S708.

[0201] The video display device (100) can transmit result data corresponding to the feedback input to the first server (400) in operation S709.

[0202] The first server (400) can store result data received from the image display device (100) in operation S710.

[0203] Referring to FIG. 8, when a user (20) first utters the second voice input (801), "Manchester City Cornaldo," the remote control device (200) can transmit the second voice input received through the microphone to the video display device (100). In the present disclosure, "Back Cornaldo" is explained as an example of a player belonging to the youth team of a professional soccer club called "Manchester City."

[0204] The video display device (100) can transmit the second voice input (801) received from the remote control device (200) to the first server (400). At this time, since the name 'Back Ronaldo' is not well-known to the public as a player belonging to the youth team of 'Manchester City', the first server (400) can generate 'Manchester City Ronaldo' as the second text (810) by considering the number of syllables of the second voice input (801) and the relevance to 'Manchester City'. The video display device (100) can output the second text (810) received from the first server (400) through the display (180).

[0205] The video display device (100) can perform an action corresponding to the second text (810). For example, when an application providing a streaming service is running, the video display device (100) can search for content corresponding to 'Manchester City Ronaldo'.

[0206] Referring to FIG. 9, when a user (20) utters the first voice input (901) “Manchester City Back Cornaldo,” the remote control device (200) can transmit the first voice input (901) received through the microphone to the image display device (100). The image display device (100) can transmit the first voice input (901) received from the remote control device (200) to the first server (400). At this time, the first server (400) can generate “Manchester City McDonald’s” as the first text (910) by considering the number of syllables of the first voice input (901), the relevance to “Manchester City,” etc. The image display device (100) can output the first text (910) received from the first server (400) through the display (180).

[0207] At this time, if the first voice input (901) is input within a predetermined time of 15 seconds after the second voice input (801) is input, the video display device (100) can determine the similarity between the first text (910) and the second text (810). At this time, the similarity between 'Manchester City Ronaldo' and 'Manchester City McDonald's' may be above a predetermined standard.

[0208] Referring to FIG. 10, the video display device (100) can output a first screen requesting feedback on the result of performing voice recognition on the first voice input (901) when the similarity between the first text (910) and the second text (810) is a predetermined standard.

[0209] The first screen requesting feedback may include an object (1000) on which text corresponding to the feedback input is displayed. In this case, the object (1000) on which the text is displayed may display a first text (910).

[0210] The first screen requesting feedback may include a virtual keyboard (1005) used for correcting the results of voice recognition. The virtual keyboard (1005) may be a software component that allows input of characters without requiring physical keys. For example, the virtual keyboard (1005) may include a plurality of keys, each of which represents a character. The user may select a key included in the virtual keyboard (1005) based on the position of the pointer (205) displayed on the first screen using a remote control device (200). At this time, the image display device (100) may generate or modify text included in the object (1000) on which text is displayed based on the input of selecting a key included in the virtual keyboard (1005) received from the remote control device (200).

[0211] The first screen requesting feedback may further include a message (1020) requesting confirmation of whether the first text (910) corresponds to the voice spoken by the user.

[0212] Referring to FIG. 11, when a user completes input of text (1110) using a virtual keyboard (1005), the text (1110) input by the user can be determined as the final text corresponding to the first voice input (901).

[0213] At this time, the video display device (100) may display a message (1120) indicating that input for the final text (1110) has been completed on the first screen requesting feedback. In addition, the video display device (100) may perform an operation corresponding to the final text (1110). For example, when an application providing a streaming service is running, the video display device (100) may search for content corresponding to 'Manchester City back Conaldo'.

[0214] Figure 12 is a flowchart illustrating a method of operating a system according to another embodiment of the present disclosure. Any details that overlap with those described in Figures 6 to 11 will be omitted for brevity.

[0215] Referring to FIG. 12, the video display device (100) can receive voice input in operation S1201.

[0216] The video display device (100) can transmit the received voice input to the first server (400) in operation S1202.

[0217] The first server (400) can, in operation S1203, convert voice data corresponding to voice input received from the video display device (100) into text data.

[0218] The first server (400) can transmit text data corresponding to voice input to the video display device (100) in operation S1204.

[0219] In operation S1205, if a voice input is received again within a predetermined time after a voice input was received immediately before, the video display device (100) can determine the similarity between the first voice input currently received and the second voice input received immediately before.

[0220] In operation S1206, if the similarity between the first voice input and the second voice input is greater than or equal to a predetermined standard, the video display device (100) may request the first server (400) to transmit multiple candidate texts for the first voice input. For example, the video display device (100) may transmit voice data for the first voice input and / or first text for the first voice input to the first server (400). The video display device (100) may also transmit second text for the second voice input to the first server (400) together with the voice data and / or first text.

[0221] According to one embodiment, when the first server (400) stores text corresponding to the result of performing voice recognition on a voice input received from the image display device (100), the image display device (100) may transmit only a request for transmission of candidate text for the first voice input to the first server (400).

[0222] The first server (400) may, in operation S1207, generate multiple candidate texts for the first voice input. For example, the first server (400) may generate multiple candidate texts based on the waveforms of syllables included in the first voice input, the relationship between words, the probability of the next word appearing given the previous word, etc.

[0223] In one embodiment, the first server (400) may generate n candidate texts in descending order of priority according to an n-best method. For example, the first server (400) may generate four candidate texts with high priority based on the results of performing voice recognition on the first voice input.

[0224] The first server (400) can transmit multiple candidate texts to the video display device (100) in operation S1208.

[0225] The video display device (100) may, in operation S1209, provide the user with the results of performing voice recognition on the first voice input. For example, the video display device (100) may output a screen displaying multiple candidate texts received from the first server (400) through the display (180).

[0226] The video display device (100) may, in operation S1210, receive user feedback regarding the results of performing voice recognition on the first voice input. For example, the video display device (100) may receive an input from the remote control device (200) to select one of a plurality of candidate texts.

[0227] The video display device (100) can perform an operation corresponding to the user's feedback in operation S1211.

[0228] The video display device (100) can transmit result data corresponding to the feedback input to the first server (400) in operation S1212.

[0229] The first server (400) can store the result data received from the image display device (100) in operation S1213.

[0230] Referring to FIG. 13, in a case where a user first utters a second voice input (801) of “Manchester City Conaldo” and then utters a first voice input (901) of “Manchester City Back Conaldo” within a predetermined time, if the similarity between the first text (910) and the second text (810) is above a predetermined standard, the video display device (100) can request the first server (400) to transmit a plurality of candidate texts for the first voice input.

[0231] The video display device (100) can output a second screen requesting feedback on the result of performing voice recognition on the first voice input (901).

[0232] The second screen requesting feedback may include objects (1301 to 1304) corresponding to a plurality of candidate texts received from the first server (400). The user may select one of the objects (1301 to 1304) corresponding to the plurality of candidate texts based on the position of the pointer (205) displayed on the second screen using the remote control device (200).

[0233] The second screen requesting feedback may further include a message (1300) requesting confirmation of the voice spoken by the user.

[0234] Referring to FIG. 14, when a user selects a first object (1301) among objects (1301 to 1304) corresponding to multiple candidate texts, the candidate text corresponding to the first object (1301) can be determined as the final text corresponding to the first voice input (901).

[0235] At this time, the video display device (100) may display a message (1410) indicating that the input of the final text has been completed on a second screen requesting feedback. Furthermore, the video display device (100) may perform an action corresponding to the final text. For example, if an application using a search engine is running, the video display device (100) may output search results for "Manchester City back Conaldo."

[0236] Figures 15 and 16 are flowcharts illustrating the operation method of the system according to various embodiments of the present disclosure. Detailed descriptions of content overlapping with those described in Figures 6 to 14 will be omitted.

[0237] Referring to FIG. 15, the video display device (100) can receive voice input in operation S1501.

[0238] The video display device (100) can transmit the received voice input to the first server (400) in operation S1502.

[0239] The first server (400) can, in operation S1503, convert voice data corresponding to voice input received from the video display device (100) into text data.

[0240] The first server (400) can transmit text data corresponding to voice input to the video display device (100) in operation S1504.

[0241] In operation S1505, if a voice input is received again within a predetermined time after a voice input was received immediately before, the video display device (100) can determine the similarity between the first voice input currently received and the second voice input received immediately before.

[0242] In operation S1506, if the similarity between the first voice input and the second voice input is greater than a predetermined standard, the video display device (100) may transmit a prompt based on the voice input to the second server (500). The second server (500) may be a server that processes data using ultra-large AI. For example, the video display device (100) may generate a prompt requesting the transmission of candidate text for the first voice input and transmit it to the second server (500).

[0243] According to one embodiment, the video display device (100) may generate a prompt requesting the transmission of candidate text for a first voice input using a template for generating a prompt stored in the memory (140). For example, the video display device (100) may generate a prompt by inserting the first text into a predetermined template.

[0244] According to one embodiment, the video display device (100) may generate a prompt requesting the transmission of candidate text for the first voice input based on at least one of a currently running application, currently output content, a currently selected broadcast channel, and a current point in time at which the prompt is generated. For example, the video display device (100) may generate a prompt comprising content indicating that the first voice input was received while specific content was being output through the display (180) at a specific point in time. In this case, the second server (500) may process the request for the first voice input using information about the specific content being output at the specific point in time.

[0245] According to one embodiment, the video display device (100) may generate a prompt based on the user's usage history stored in the memory (140). For example, the video display device (100) may determine that the user's preferred content genres are sports, movies, and games, based on the user's usage history stored in the memory (140). In this case, the video display device (100) may generate a prompt comprising content indicating that a first voice input has been received from a user who prefers a specific genre of content. In this case, the second server (500) may process a request regarding the first voice input using the specific genre preferred by the user.

[0246] The second server (500) may, in operation S1507, generate a response to a prompt received from the video display device (100). For example, the second server (500) may generate a response to a prompt requesting the transmission of candidate text for the first voice input using a large-scale AI. In this case, the second server (500) may generate candidate text for at least one first voice input.

[0247] The second server (500) may, in operation S1508, transmit a response to the prompt to the video display device (100). The response to the prompt may include at least one candidate text.

[0248] The video display device (100) may, in operation S1509, provide the user with the results of performing voice recognition on the first voice input. For example, the video display device (100) may output a second screen displaying candidate text included in a response to a prompt received from the second server (500) through the display (180).

[0249] The video display device (100) may, in operation S1510, receive user feedback regarding the results of performing voice recognition on the first voice input. For example, the video display device (100) may receive an input from the remote control device (200) to select one of the candidate texts.

[0250] The video display device (100) can perform an operation corresponding to the user's feedback in operation S1511.

[0251] The video display device (100) can transmit result data corresponding to the feedback input to the first server (400) in operation S1512.

[0252] The first server (400) can store the result data received from the image display device (100) in operation S1513.

[0253] Referring to FIG. 16, the video display device (100) can receive voice input in operation S1601.

[0254] The video display device (100) can transmit the received voice input to the first server (400) in operation S1602.

[0255] The first server (400) can, in operation S1603, convert voice data corresponding to voice input received from the video display device (100) into text data.

[0256] The first server (400) can transmit text data corresponding to voice input to the video display device (100) in operation S1604.

[0257] In operation S1605, if a voice input is received again within a predetermined time after a voice input was received immediately before, the first server (400) can determine the similarity between the first voice input currently received and the second voice input received immediately before.

[0258] In operation S1606, if the similarity between the first voice input and the second voice input is greater than a predetermined standard, the first server (400) may transmit a prompt based on the voice input to the second server (500). For example, the first server (400) may generate a prompt requesting the transmission of candidate text for the first voice input and transmit it to the second server (500).

[0259] The second server (500) can generate a response to a prompt received from the video display device (100) in operation S1607.

[0260] The second server (500) can, in operation S1608, transmit a response to the prompt to the first server (400).

[0261] In operation S1609, the first server (400) may transmit the result of performing voice recognition on the first voice input to the image display device (100) based on a response to a prompt received from the second server (500). For example, the result of performing voice recognition on the first voice input may include candidate text for the first voice input. Meanwhile, the result of performing voice recognition on the first voice input may include a command corresponding to the calculated similarity.

[0262] The video display device (100) may, in operation S1610, provide the user with the results of performing voice recognition on the first voice input. For example, the video display device (100) may output a second screen displaying candidate text received from the first server (400) through the display (180).

[0263] The video display device (100) may, in operation S1611, receive user feedback regarding the results of performing voice recognition on the first voice input. For example, the video display device (100) may receive an input from the remote control device (200) to select one of the candidate texts.

[0264] The video display device (100) can perform an operation corresponding to the user's feedback in operation S1612.

[0265] The video display device (100) can transmit result data corresponding to the feedback input to the first server (400) in operation S1613.

[0266] The first server (400) can store the result data received from the image display device (100) in operation S1614.

[0267] Meanwhile, in this drawing, the first server (400) transmits a prompt to the second server (500), and the second server (500) transmits the result of performing voice recognition on the first voice input to the video display device (100) based on a response to the prompt, but is not limited thereto.

[0268] For example, if the similarity between the first voice input and the second voice input is determined to be above a predetermined standard, the first server (400) may transmit a command corresponding to the calculated similarity to the video display device (100). At this time, the video display device (100) may output a first screen requesting feedback including a virtual keyboard (1005) based on the command received from the first server (400).

[0269] For example, if the similarity between the first voice input and the second voice input is determined to be above a predetermined standard, the first server (400) may generate a plurality of candidate texts for the first voice input and transmit them to the video display device (100). At this time, the video display device (100) may output a second screen requesting feedback including objects (1301 to 1304) corresponding to the plurality of candidate texts based on the plurality of candidate texts received from the first server (400).

[0270] As described above, according to at least one embodiment of the present disclosure, feedback on the result of performing speech recognition can be requested by taking into account the user's speech intention, thereby preventing the user from repeatedly speaking.

[0271] Additionally, according to at least one embodiment of the present disclosure, a function or action desired by the user can be accurately performed based on the user's feedback on the result of performing voice recognition.

[0272] Additionally, according to at least one embodiment of the present disclosure, the accuracy of speech recognition can be improved based on user feedback on the result of performing speech recognition.

[0273] Referring to FIGS. 1 to 16, an image display device (100) according to one aspect of the present disclosure includes a display (180); a user input interface unit (150) for receiving a user input; and a control unit (170). When a first voice input is received through the user input interface unit (150), the control unit (170) determines whether the first voice input has been received within a predetermined time after a second voice input has been received immediately before, determines the similarity between the first voice input and the second voice input when the first voice input has been received within the predetermined time after the second voice input has been received, and outputs a screen requesting feedback on the result of performing voice recognition on the first voice input through the display (180) when the similarity is equal to or greater than a predetermined standard, and performs an operation corresponding to a feedback input received through the user input interface unit (150).

[0274] In addition, according to one aspect of the present disclosure, the control unit (170) may perform an operation corresponding to the first voice input when the second voice input is received and the predetermined time has elapsed, or when the similarity is less than the predetermined standard.

[0275] In addition, according to one aspect of the present disclosure, the system further includes a network interface unit (135) that communicates with at least one server, and the control unit (170) transmits result data corresponding to the feedback input to a first server (400) that performs voice recognition on a voice input received from the image display device based on the feedback input being received through the user input interface unit (150), and the result data may include a final text determined to correspond to the first voice input based on the feedback input.

[0276] Additionally, according to one aspect of the present disclosure, the result data may further include at least one of a first text corresponding to a result of performing the voice recognition on the first voice input and a second text corresponding to a result of performing the voice recognition on the second voice input.

[0277] In addition, according to one aspect of the present disclosure, the control unit (170) may receive, from the first server (400), a plurality of candidate texts according to an n-best method corresponding to the result of performing the voice recognition on the first voice input, output a screen requesting the feedback including the plurality of candidate texts, and determine a text corresponding to the first voice input based on an input for selecting one of the plurality of candidate texts received through the user input interface unit (150).

[0278] In addition, according to one aspect of the present disclosure, the screen requesting the feedback includes a virtual keyboard (1005), and the control unit (170) can determine text corresponding to the first voice input based on an input of selecting a key included in the virtual keyboard (1005) received through the user input interface unit (150).

[0279] Additionally, according to one aspect of the present disclosure, the screen requesting feedback may include a message requesting correction of the result of performing the voice recognition on the first voice input.

[0280] In addition, according to one aspect of the present disclosure, the system further includes a network interface unit (135) that communicates with at least one server, and the control unit (170) transmits a predetermined prompt based on the first voice input to a second server (500) that uses a super-giant artificial intelligence model, and outputs a screen requesting the feedback, which includes at least one candidate text included in a response to the predetermined prompt received from the second server (500), through the display (180).

[0281] In addition, according to one aspect of the present disclosure, the image display device further includes a memory (140) that stores a usage history, and the control unit (170) can obtain information on content preferred by the user from the usage history and generate the predetermined prompt based on the information on the content preferred by the user.

[0282] In addition, according to one aspect of the present disclosure, the control unit (170) may generate the predetermined prompt based on at least one of a country corresponding to the user, an application currently running on the video display device, content currently output through the video display device, a broadcast channel currently selected on the video display device, and a current time point at which the prompt is generated.

[0283] A system (10) according to one aspect of the present disclosure includes an image display device (100) and a first server (400) that performs voice recognition on a voice input received from the image display device (100), wherein the first server (400) determines, when a first voice input is received from the image display device (100), whether the first voice input was received within a predetermined time after a second voice input was received immediately before, and determines, when the first voice input is received within the predetermined time after the second voice input is received, a similarity between the first voice input and the second voice input, and, when the similarity is equal to or greater than a predetermined standard, transmits a command corresponding to the calculated similarity to the image display device (100), and the image display device (100) outputs a screen requesting feedback on a result of performing voice recognition on the first voice input based on the command received by the first server (400), and performs an operation corresponding to the received feedback input.

[0284] In addition, according to one aspect of the present disclosure, the image display device (100) transmits result data corresponding to the feedback input to the first server (400) based on the feedback input being received, and the first server (400) learns a language model used for the voice recognition based on the result data, and the result data may include a final text determined to correspond to the first voice input based on the feedback input.

[0285] In addition, according to one aspect of the present disclosure, when the similarity is above a predetermined standard, the first server (400) transmits a plurality of candidate texts according to the n-best method corresponding to the result of performing the voice recognition on the first voice input to the image display device (100), and the image display device (100) outputs a screen requesting the feedback including the plurality of candidate texts, and determines the text corresponding to the first voice input based on an input for selecting one of the plurality of candidate texts.

[0286] In addition, according to one aspect of the present disclosure, the screen requesting the feedback includes a virtual keyboard, and the image display device (100) can generate text corresponding to the first voice input based on an input of selecting a key included in the virtual keyboard, and perform an operation corresponding to the generated text.

[0287] In addition, according to one aspect of the present disclosure, a second server (500) using a super-giant artificial intelligence model is further included, wherein the first server (400) transmits a predetermined prompt based on the first voice input to the second server (500), and transmits at least one candidate text included in a response to the predetermined prompt received from the second server (500) to the image display device (100), and the image display device (100) outputs a screen requesting the feedback including the plurality of candidate texts, and determines a text corresponding to the first voice input based on an input for selecting one of the plurality of candidate texts.

[0288] The attached drawings are only intended to facilitate understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure.

[0289] Meanwhile, the operating method of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices that store data that can be read by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across network-connected computer systems, so that the processor-readable code can be stored and executed in a distributed manner.

[0290] In addition, although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. Display; A user input interface section for receiving user input; and Including a control unit, The above control unit, When a first voice input is received through the user input interface section, it is determined whether the first voice input was received within a predetermined time after the second voice input was received immediately before. If the first voice input is received within the predetermined time after the second voice input is received, the similarity between the first voice input and the second voice input is determined, If the above similarity is above a predetermined standard, a screen requesting feedback on the result of performing voice recognition on the first voice input is output through the display, An image display device characterized by performing an action corresponding to a feedback input received through the user input interface unit.

2. In paragraph 1, The above control unit, A video display device characterized in that when the second voice input is received and the first voice input is received after the predetermined time has elapsed or when the similarity is below the predetermined standard, an operation corresponding to the first voice input is performed.

3. In paragraph 1, Further comprising a network interface section communicating with at least one server, The above control unit, Based on the feedback input being received through the user input interface section, the result data corresponding to the feedback input is transmitted to the first server that performs voice recognition for the voice input received from the image display device, A video display device, characterized in that the result data includes a final text determined to correspond to the first voice input according to the feedback input.

4. In paragraph 3, An image display device, characterized in that the result data further includes at least one of a first text corresponding to a result of performing the voice recognition on the first voice input and a second text corresponding to a result of performing the voice recognition on the second voice input.

5. In paragraph 3, The above control unit, Receive from the first server a plurality of candidate texts according to the n-best method corresponding to the result of performing the speech recognition for the first speech input, Output a screen requesting the feedback including the above multiple candidate texts, An image display device characterized in that the text corresponding to the first voice input is determined based on an input for selecting one of the plurality of candidate texts received through the user input interface unit.

6. In paragraph 1, The screen requesting the above feedback includes a virtual keyboard, The above control unit, A video display device characterized in that the text corresponding to the first voice input is determined based on an input for selecting a key included in the virtual keyboard received through the user input interface unit.

7. In paragraph 6, The screen requesting the above feedback is: A display device characterized by including a message requesting correction of the result of performing the voice recognition for the first voice input.

8. In paragraph 1, Further comprising a network interface section communicating with at least one server, The above control unit, A predetermined prompt based on the first voice input is transmitted to a second server using a super-giant artificial intelligence model, A display device characterized in that it outputs a screen requesting feedback, including at least one candidate text included in a response to the predetermined prompt received from the second server, through the display.

9. In paragraph 8, Further comprising a memory for storing usage history for the above video display device, The above control unit, Obtain information about the user's preferred content from the above usage history, A video display device characterized in that it generates the predetermined prompt based on information about content preferred by the user.

10. In paragraph 8, The above control unit, A video display device characterized in that the predetermined prompt is generated based on at least one of a country corresponding to the user, an application currently running on the video display device, content currently output through the video display device, a broadcast channel currently selected on the video display device, and a current point in time when the prompt is generated.

11. A system including a video display device and a first server that performs voice recognition on voice input received from the video display device, The above first server, When a first voice input is received from the above video display device, it is determined whether the first voice input was received within a predetermined time after the second voice input was received immediately before, If the first voice input is received within the predetermined time after the second voice input is received, the similarity between the first voice input and the second voice input is determined, If the above similarity is greater than a predetermined standard, a command corresponding to the calculated similarity is transmitted to the image display device, The above video display device, Based on the command received from the first server, a screen is output requesting feedback on the result of performing voice recognition on the first voice input, A system characterized by performing an action responsive to received feedback input.

12. In paragraph 11, The above video display device, Based on the reception of the above feedback input, result data corresponding to the feedback input is transmitted to the first server, The above first server, Based on the above result data, the language model used for the speech recognition is learned, A system characterized in that the result data includes a final text determined to correspond to the first voice input according to the feedback input.

13. In paragraph 11, The above first server, If the above similarity is greater than a predetermined standard, multiple candidate texts corresponding to the result of performing the voice recognition for the first voice input according to the n-best method are transmitted to the video display device, The above video display device, Output a screen requesting the feedback including the above multiple candidate texts, A system characterized in that a text corresponding to the first speech input is determined based on an input selecting one of the plurality of candidate texts.

14. In paragraph 11, The screen requesting the above feedback includes a virtual keyboard, The above video display device, Generate text corresponding to the first voice input based on an input of selecting a key included in the virtual keyboard, A system characterized by performing an action corresponding to the generated text.

15. In paragraph 11, It further includes a second server that uses a super-giant artificial intelligence model, The above first server, Transmitting a predetermined prompt based on the first voice input to the second server, Transmitting at least one candidate text included in a response to the given prompt received from the second server to the video display device, The above video display device, Output a screen requesting the feedback including the above multiple candidate texts, A system characterized in that a text corresponding to the first speech input is determined based on an input selecting one of the plurality of candidate texts.

Citation Information

Patent Citations

  • Method, device, and program for speech recognition

    JP2003316386A

  • Arbitration between voice-enabled devices

    KR1020180039135A

  • Method and apparatus for speech recognition correction

    KR1020180054362A

  • Method for outputting speech recognition results based on determination of sameness and appratus using the same

    KR102153220B1

  • Calibration device for oral scanner

    KR102551332B1