Image display device and system comprising same
The device addresses the time gap in voice recognition by initiating real-time content data acquisition and processing, ensuring accurate and timely responses to user inputs, enhancing speech recognition and output summaries.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional video display devices experience a time gap between user voice input and content data acquisition, leading to discrepancies in the output content, making it difficult to provide the desired voice recognition result accurately.
The device initiates content data acquisition and processes voice recognition in real-time, utilizing a control unit to receive voice signals, process them based on content data, and output responses through the display, optionally involving a server for enhanced processing.
Enables accurate and timely acquisition of content data in response to user voice utterances, improving speech recognition accuracy and providing a summary of output content during voice recognition activation.
Smart Images

Figure KR2024018021_21052026_PF_FP_ABST
Abstract
Description
Image display device and system including the same
[0001] The present disclosure relates to an image display device and a system including the same.
[0002] A video display device is a device equipped with the function of displaying images that a user can view. For example, a video display device may include a TV, monitor, notebook computer, etc., equipped with a Liquid Crystal Display (LCD) using liquid crystals or an OLED display using organic light-emitting diodes (OLEDs).
[0003] With the advancement of technology, various services applying voice recognition technology are being developed and provided in diverse fields, and various technological developments are underway to apply voice recognition technology to video display devices such as TVs.
[0004] Generally, a system using speech recognition technology performs intent analysis on the voice spoken by a user when the user speaks, and processes a command corresponding to the voice spoken by the user based on the result of the intent analysis. In this case, if the user speaks a voice related to content output through a video display device, the system confirms that the voice spoken by the user is related to the content based on the result of the intent analysis, and acquires data regarding the content to process a command corresponding to the voice spoken by the user.
[0005] According to conventional methods, there is a time gap between when the user speaks and when the system acquires data regarding the content. If a video display device continues to output content while the user's spoken voice is being processed, the video or audio of the content output through the display device differs between the two times, presenting a problem in that it is difficult for the system to provide the voice recognition result desired by the user.
[0006] The present disclosure aims to solve the aforementioned problems and other problems.
[0007] Another objective is to provide a video display device capable of accurately acquiring data about content in response to a user's voice utterance, and a system including the same.
[0008] Another objective is to provide a video display device and a system including the same that can improve the accuracy of the result of performing speech recognition on a voice spoken by a user.
[0009] Another objective is to provide a video display device capable of providing a summary of the content of output while a voice recognition function is activated, and a system including the same.
[0010] An image display device according to one embodiment of the present disclosure for achieving the above objective comprises: a display; a user input interface unit that transmits a signal corresponding to a user input; and a control unit. When a voice recognition function is activated while outputting content through the display, the control unit initiates the acquisition of content data corresponding to the content, receives a voice signal through the user input interface unit, acquires a voice recognition result by processing the voice signal based on the content data, and outputs a response corresponding to the voice recognition result through the display.
[0011] A system according to one embodiment of the present disclosure for achieving the above objective comprises an image display device and a server, wherein the image display device, when a voice recognition function is activated while outputting content through a display, initiates the acquisition of content data corresponding to the content, transmits the content data to the server, transmits a voice signal received through a user input interface unit to the server, receives a voice recognition result processed from the voice signal from the server, outputs a response corresponding to the voice recognition result through the display, and the server, based on the content data, generates a voice recognition result processed from the voice signal received from the image display device and transmits the voice recognition result to the image display device.
[0012] The effects of the image display device and the system including the same according to the present disclosure are described as follows.
[0013] According to at least one embodiment of the present disclosure, data regarding content can be accurately obtained in response to a user's voice utterance.
[0014] According to at least one embodiment of the present disclosure, the accuracy of the result of performing speech recognition on a voice spoken by a user can be improved.
[0015] According to at least one embodiment of the present disclosure, a summary of the content output while the voice recognition function is activated can be provided.
[0016] Further scopes of the applicability of the present disclosure will become apparent from the following detailed description. However, since various changes and modifications within the spirit and scope of the present disclosure are clearly understood by those skilled in the art, specific embodiments, such as the detailed description and preferred embodiments of the present disclosure, should be understood as being given merely as examples.
[0017] FIG. 1 is a drawing illustrating an image display system according to one embodiment of the present disclosure.
[0018] Figure 2 is an internal block diagram of the image display device of Figure 1.
[0019] Figure 3 is an internal block diagram of the control unit of Figure 2.
[0020] FIG. 4a is a diagram illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.
[0021] FIG. 5 is a drawing referenced in the description of the first server of FIG. 1.
[0022] FIG. 6 is a block diagram illustrating the configuration of a first server according to one embodiment of the present disclosure.
[0023] FIG. 7 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0024] FIG. 8 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.
[0025] FIG. 9 is a flowchart of a method of operation of an image display device according to one embodiment of the present disclosure.
[0026] FIGS. 10a to 10c are flowcharts of a method of operation of a system according to various embodiments of the present disclosure.
[0027] FIGS. 11 to 19 are drawings referenced in the description relating to the processing of a user's voice input according to embodiments of the present disclosure.
[0028] FIGS. 20a and FIGS. 20b are flowcharts of a method of operation of an image display device according to another embodiment of the present disclosure.
[0029] FIGS. 21 to 26 are drawings referenced in the description relating to the processing of a user's voice input according to embodiments of the present disclosure.
[0030] FIGS. 27 to 29 are flowcharts of a method of operation of a system according to various embodiments of the present disclosure.
[0031] The present disclosure is described in detail below with reference to the drawings. In order to clearly and concisely explain the present disclosure, parts unrelated to the description have been omitted from the drawings, and the same reference numerals are used throughout the specification for identical or extremely similar parts.
[0032] The suffixes "module" and "part" for components used in the following description are assigned solely for the ease of drafting this specification and do not inherently confer any particularly significant meaning or role. Accordingly, the terms "module" and "part" may be used interchangeably.
[0033] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0034] Additionally, in this specification, terms such as first, second, etc. may be used to describe various elements, but these elements are not limited by these terms. These terms are used only to distinguish one element from another.
[0035] FIG. 1 is a drawing illustrating an image display system according to various embodiments of the present invention.
[0036] Referring to FIG. 1, the image display system (10) may include an image display device (100) and / or a remote control device (200).
[0037] The image display device (100) may be a device that processes and outputs an image. The image display device (100) is not specifically limited to any device capable of outputting a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.
[0038] The video display device (100) can receive a broadcast signal, process the signal, and output the processed broadcast video. When the video display device (100) receives a broadcast signal, the video display device (100) may correspond to a broadcast receiving device.
[0039] The video display device (100) may receive a broadcast signal wirelessly through an antenna or receive a broadcast signal wired through a cable. For example, the video display device (100) may receive a terrestrial broadcast signal, a satellite broadcast signal, a cable broadcast signal, an IPTV (Internet Protocol Television) broadcast signal, etc.
[0040] The remote control device (200) is connected to the image display device (100) via a wired and / or wireless connection and can provide various control signals to the image display device (100). At this time, the remote control device (200) may include a device that establishes a wired or wireless network with the image display device (100) and transmits various control signals to the image display device (100) through the established network, or receives signals related to various operations processed by the image display device (100) from the image display device (100).
[0041] For example, various input devices such as a mouse, keyboard, spatial remote control, trackball, and joystick may be used as the remote control device (200). The remote control device (200) may be referred to as an external device, and it should be noted in advance that the external device and the remote control device may be used interchangeably as needed below.
[0042] The video display device (100) may be connected to only a single remote control device (200) or simultaneously connected to two or more remote control devices (200), and may change the object displayed on the screen or adjust the state of the screen based on the control signal provided from each remote control device (200).
[0043] Meanwhile, the video display system (10) may further include at least one server (300). The video display device (100) may transmit and receive data to and from at least one server (300). For example, the video display device (100) may transmit and receive data to and from at least one server (300) via a network such as the Internet.
[0044] According to one embodiment, at least one server (300) may include a first server (400) that performs voice recognition, a second server (500) that processes data using a super-giant artificial intelligence model (hereinafter, super-giant AI), a third server (600) that provides content, etc.
[0045] Figure 2 is an internal block diagram of the image display device of Figure 1.
[0046] Referring to FIG. 2, the video display device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).
[0047] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulating unit (120).
[0048] Meanwhile, unlike the drawing, the video display device (100) may include only the broadcast receiving unit (105) and the external device interface unit (130) among the broadcast receiving unit (105), the external device interface unit (130), and the network interface unit (135). That is, the video display device (100) may not include the network interface unit (135).
[0049] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among the broadcast signals received through an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.
[0050] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).
[0051] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through a channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.
[0052] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive multiple channels of broadcast signals. Alternatively, a single tuner that simultaneously receives multiple channels of broadcast signals is also possible.
[0053] The demodulator (120) can receive the digital IF signal (DIF) converted by the tuner (110) and perform a demodulator operation.
[0054] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.
[0055] The stream signal output from the demodulation unit (120) can be input to the control unit (170). After performing demultiplexing, video / audio signal processing, etc., the control unit (170) can output video through the display (180) and output audio through the audio output unit (185).
[0056] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).
[0057] The external device interface section (130) can be connected wirelessly or via wired connection to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game console, camera, camcorder, computer (laptop), set-top box, etc., and can also perform input / output operations with the external devices.
[0058] Additionally, the external device interface unit (130) can establish a communication network with various remote control devices (200) such as those shown in FIG. 1, receive control signals related to the operation of the image display device (100) from the remote control device (200), or transmit data related to the operation of the image display device (100) to the remote control device (200).
[0059] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit may include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).
[0060] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, the external device interface unit (130) can receive device information, information on an application being executed, an application image, etc. from a mobile terminal in mirroring mode.
[0061] The external device interface section (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.
[0062] The network interface section (135) can provide an interface for connecting the video display device (100) to a wired / wireless network including the internet.
[0063] The network interface section (135) may include a communication module (not shown) for connection with a wired / wireless network. For example, the network interface section (135) may include a communication module for a WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), etc.
[0064] The network interface unit (135) can transmit or receive data with other users or other electronic devices through the connected network or another network linked to the connected network.
[0065] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and related information provided by a content provider or network provider through a network.
[0066] The network interface unit (135) can receive firmware update information and update files provided by the network operator and can transmit data to the internet or a content provider or network operator.
[0067] The network interface unit (135) can select and receive a desired application among the applications that are open to the public through the network.
[0068] The storage unit (140) may store programs for each signal processing and control within the control unit (170), and may also store signal-processed video, audio, or data signals. For example, the storage unit (140) may store applications designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored applications upon request from the control unit (170).
[0069] The program, etc. stored in the storage unit (140) is not specifically limited as long as it can be executed by the control unit (170).
[0070] The storage unit (140) may also perform the function of temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).
[0071] The storage unit (140) can store information regarding a predetermined broadcast channel through a channel memory function such as a channel map.
[0072] Although the storage unit (140) of FIG. 2 is illustrated in an embodiment in which it is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).
[0073] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.
[0074] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.
[0075] For example, user input signals such as power on / off, channel selection, and screen settings can be transmitted / received from a remote control device (200), user input signals input from local keys (not shown) such as a power key, channel key, volume key, and setting value can be transmitted to a control unit (170), user input signals input from a sensor unit (not shown) that senses a user's gesture can be transmitted to a control unit (170), or signals from a control unit (170) can be transmitted to a sensor unit.
[0076] The input unit (160) may be provided on one side of the main body of the image display device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.
[0077] The input unit (160) can receive various user commands related to the operation of the image display device (100) and can transmit a control signal corresponding to the input command to the control unit (170).
[0078] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.
[0079] The control unit (170) may include at least one processor and can control the overall operation of the image display device (100) using the included processor. Here, the processor may be a general processor such as a CPU (central processing unit). Of course, the processor may be a dedicated device such as an ASIC or a processor based on other hardware.
[0080] The control unit (170) can demultiplex a stream input through the tuner unit (110), demodulator unit (120), external device interface unit (130), or network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.
[0081] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).
[0082] The display (180) may include a display panel (not shown) having a plurality of pixels.
[0083] A plurality of pixels provided in the display panel may have RGB subpixels. Alternatively, a plurality of pixels provided in the display panel may have RGBW subpixels. The display (180) can convert an image signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) to generate a driving signal for a plurality of pixels.
[0084] The display (180) can be a PDP (Plasma Display Panel), LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), flexible display, etc., and can also be a 3D display. The 3D display (180) can be classified into a glasses-free type and a glasses type.
[0085] Meanwhile, the display (180) can be configured as a touch screen and used as an input device in addition to an output device.
[0086] The audio output unit (185) receives a voice-processed signal from the control unit (170) and outputs it as voice.
[0087] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. Additionally, the image signal processed by the control unit (170) can be input to an external output device through the external device interface unit (130).
[0088] The voice signal processed by the control unit (170) can be sound-outputted to the audio output unit (185). Additionally, the voice signal processed by the control unit (170) can be input to an external output device through the external device interface unit (130).
[0089] Although not illustrated in FIG. 2, the control unit (170) may include a demultiplexer, an image processing unit, etc. This will be described later with reference to FIG. 3.
[0090] In addition, the control unit (170) can control the overall operation within the video display device (100). For example, the control unit (170) can control the tuner unit (110) to select (Tuning) a broadcast corresponding to a channel selected by the user or a previously stored channel.
[0091] Additionally, the control unit (170) can control the image display device (100) by means of a user command or an internal program input through the user input interface unit (150).
[0092] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a video, and may be a 2D image or a 3D image.
[0093] Meanwhile, the control unit (170) can make a predetermined 2D object appear within the image displayed on the display (180). For example, the object may be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.
[0094] Meanwhile, the image display device (100) may further include a shooting unit (not shown). The shooting unit can photograph the user. The shooting unit may be implemented with one camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the shooting unit may be embedded in the image display device (100) on the upper part of the display (180) or may be placed separately. Image information captured by the shooting unit may be input to the control unit (170).
[0095] The control unit (170) can recognize the user's location based on the image captured by the shooting unit. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the image display device (100). Additionally, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.
[0096] The control unit (170) can detect a user's gesture based on each of the images captured by the shooting unit or the signals detected by the sensor unit, or a combination thereof.
[0097] The power supply unit (190) can supply power throughout the image display device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a System On Chip (SOC), a display (180) for image display, and an audio output unit (185) for audio output.
[0098] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a DC / DC converter (not shown) that converts the level of DC power.
[0099] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) may use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. Additionally, the remote control device (200) may receive video, audio, or data signals output from the user input interface unit (150) and display or output audio from the remote control device (200).
[0100] Meanwhile, the above-described video display device (100) may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.
[0101] Meanwhile, the block diagram of the image display device (100) shown in FIG. 2 is merely a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted according to the specifications of the image display device (100) actually implemented.
[0102] That is, as needed, two or more components may be combined into a single component, or a single component may be subdivided into two or more components. In addition, the functions performed in each block are intended to explain embodiments of the present invention, and the specific operations or devices do not limit the scope of the present invention.
[0103] Figure 3 is an internal block diagram of the control unit of Figure 2.
[0104] Referring to FIG. 3, a control unit (170) according to one embodiment of the present invention may include a demultiplexer (310), an image processing unit (320), a processor (330), an OSD generation unit (340), a mixer (345), a frame rate conversion unit (350), and / or a formatter (360). Additionally, it may further include an audio processing unit (not shown) and a data processing unit (not shown).
[0105] The demultiplexer (310) can demultiplex the input stream. For example, if an MPEG-2 TS is input, it can be demultiplexed to separate it into video, audio, and data signals. Here, the stream signal input to the demultiplexer (310) may be a stream signal output from the tuner (110), the demodulator (120), or the external device interface (130).
[0106] The image processing unit (320) can perform image processing of the demultiplexed image signal. To this end, the image processing unit (320) may be equipped with an image decoder (325) and a scaler (335).
[0107] The video decoder (325) can decode the demultiplexed video signal, and the scaler (335) can perform scaling so that the resolution of the decoded video signal can be output on the display (180).
[0108] The image decoder (325) may be equipped with decoders of various specifications. For example, it may be equipped with MPEG-2, H,264 decoders, 3D image decoders for color images and depth images, decoders for multiple viewpoint images, etc.
[0109] The processor (330) can control the overall operation within the video display device (100) or within the control unit (170). For example, the processor (330) can control the tuner (110) to tune a broadcast corresponding to a channel selected by the user or a pre-stored channel.
[0110] Additionally, the processor (330) can control the image display device (100) by means of a user command or an internal program input through the user input interface unit (150).
[0111] Additionally, the processor (330) can perform data transmission control with the network interface unit (135) or the external device interface unit (130).
[0112] Additionally, the processor (330) can control the operation of the demultiplexer (310), image processing unit (320), OSD generation unit (340), etc. within the control unit (170).
[0113] The OSD generation unit (340) can generate an OSD signal based on user input or independently. For example, based on a user input signal input through the input unit (160), it can generate a signal to display various information as graphics or text on the screen of the display (180).
[0114] The generated OSD signal may include various data such as a user interface screen of the video display device (100), various menu screens, widgets, and icons. Additionally, the generated OSD signal may include 2D objects or 3D objects.
[0115] Additionally, the OSD generation unit (340) can generate a pointer that can be displayed on the display (180) based on a pointing signal input from the remote control device (200).
[0116] The OSD generation unit (340) may include a pointing signal processing unit (not shown) that generates a pointer. It is also possible for the pointing signal processing unit (not shown) to be provided separately rather than being provided within the OSD generation unit (240).
[0117] The mixer (345) can mix the OSD signal generated by the OSD generation unit (340) and the decoded video signal processed by the video processing unit (320). The mixed video signal can be provided to the frame rate conversion unit (350).
[0118] The frame rate converter (FRC) (350) can convert the frame rate of the input video. Meanwhile, the frame rate converter (350) can also output the video as is without separate frame rate conversion.
[0119] The formatter (360) can arrange the left eye image frame and the right eye image frame of the frame rate-converted 3D image. And, it can output a synchronization signal (Vsync) for opening the left eye glass and the right eye glass of the 3D viewing device (not shown).
[0120] Meanwhile, the formatter (360) can convert the format of the input video signal into a video signal for display on the display (180) and output it.
[0121] Additionally, the formatter (360) can change the format of the 3D video signal. For example, it can be changed to any one of various 3D formats such as Side by Side format, Top / Down format, Frame Sequential format, Interlaced format, Checker Box format.
[0122] Meanwhile, the formatter (360) may convert a 2D video signal into a 3D video signal. For example, according to a 3D video generation algorithm, an edge or selectable object within the 2D video signal may be detected, and the object or selectable object corresponding to the detected edge may be separated and generated into a 3D video signal. At this time, the generated 3D video signal may be separated and aligned into a left eye video signal (L) and a right eye video signal (R), as described above.
[0123] Meanwhile, although not shown in the drawing, it is also possible to place a 3D processor (not shown) for 3D effect signal processing after the formatter (360). This 3D processor can process the brightness, tint, and color adjustment of the video signal to improve the 3D effect. For example, it can perform signal processing to make the near distance sharp and the far distance blurry. Meanwhile, the function of this 3D processor can be merged into the formatter (360) or merged within the image processing unit (320).
[0124] Meanwhile, an audio processing unit (not shown) within the control unit (170) can perform voice processing of a demultiplexed voice signal. To this end, the audio processing unit (not shown) may be equipped with various decoders.
[0125] In addition, the audio processing unit (not shown) within the control unit (170) can process bass, treble, volume control, etc.
[0126] A data processing unit (not shown) within the control unit (170) can perform data processing of a demultiplexed data signal. For example, if the demultiplexed data signal is an encoded data signal, it can be decoded. The encoded data signal may be Electronic Program Guide information containing broadcast information such as the start time and end time of a broadcast program aired on each channel.
[0127] Meanwhile, the block diagram of the control unit (170) shown in FIG. 3 is merely a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted according to the specifications of the control unit (170) actually implemented.
[0128] In particular, the frame rate converter (350) and the formatter (360) are not provided within the control unit (170), but are provided separately, or may be provided separately as a single module.
[0129] FIG. 4a is a diagram illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.
[0130] Referring to FIG. 4a, it can be seen that a pointer (205) corresponding to a remote control device (200) is displayed on the display (180) of the image display device (100).
[0131] Referring to (a) of FIG. 4a, the user can move or rotate the remote control device (200) up and down, left and right, and forward and backward. At this time, the pointer (205) displayed on the display (180) of the image display device (100) can be displayed in response to the movement of the remote control device (200). Since the pointer (205) of the remote control device (200) moves and is displayed according to the movement in 3D space as shown in the drawing, it can be named a spatial remote control or a 3D pointing device.
[0132] Referring to (b) of FIG. 4a, when the user moves the remote control device (200) to the left, it can be seen that the pointer (205) displayed on the display (180) of the image display device (100) also moves to the left in response to the movement of the remote control device (200).
[0133] Information regarding the movement of the remote control device (200) detected through the sensor of the remote control device (200) can be transmitted to the image display device (100). The image display device (100) can calculate the coordinates of the pointer (205) from the information regarding the movement of the remote control device (200). The image display device (100) can display the pointer (205) to correspond to the calculated coordinates.
[0134] Referring to (c) of FIG. 4a, while pressing a specific button provided on the remote control device (200), the user can move the remote control device (200) away from the display (180). By doing so, the selected area within the display (180) corresponding to the pointer (205) can be zoomed in and enlarged. Conversely, while pressing a specific button provided on the remote control device (200), if the user moves the remote control device (200) closer to the display (180), the selected area within the display (180) corresponding to the pointer (205) can be zoomed out and reduced.
[0135] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.
[0136] Meanwhile, when the user presses a specific button within the remote control device (200), recognition of up-down and left-right movement may be excluded. That is, when the remote control device (200) moves away from or closer to the display (180), up-down, left-right movement is not recognized, and only forward-backward movement may be recognized. When the user does not press a specific button within the remote control device (200), only up-down, left-right movement of the remote control device (200) can be recognized, and accordingly, only the pointer (205) can move.
[0137] Meanwhile, the movement speed or direction of movement of the pointer (205) can correspond to the movement speed or direction of movement of the remote control device (200).
[0138] Referring to FIG. 4b, the remote control device (200) may include a wireless communication unit (220), a user input unit (230), a sensor unit (240), an output unit (250), a power supply unit (260), a storage unit (270), and / or a control unit (280).
[0139] The wireless communication unit (220) can transmit and receive signals to and from the image display device (100).
[0140] In this embodiment, the remote control device (200) may be equipped with an RF module (221) capable of transmitting and receiving signals to and from the image display device (100) according to an RF (Radio frequency) communication standard. Additionally, the remote control device (200) may be equipped with an IR module (223) capable of transmitting and receiving signals to and from the image display device (100) according to an IR (Infrared radiation) communication standard.
[0141] The remote control device (200) can transmit a signal containing information regarding the movement of the remote control device (200), etc., to the video display device (100) through the RF module (221). The remote control device (200) can receive the signal transmitted by the video display device (100) through the RF module (221).
[0142] The remote control device (200) can transmit commands regarding power on / off, channel change, volume change, etc. to the video display device (100) through the IR module (223).
[0143] The user input unit (230) may be composed of a keypad, buttons, a touchpad, a touch screen, etc. The user can input commands related to the image display device (100) to the remote control device (200) by operating the user input unit (230).
[0144] If the user input unit (230) is equipped with a hard key button, the user can input commands related to the image display device (100) to the remote control device (200) through a push operation of the hard key button.
[0145] If the user input unit (230) is equipped with a touchscreen, the user can input commands related to the image display device (100) to the remote control device (200) by touching the soft keys on the touchscreen.
[0146] Meanwhile, the user input unit (230) may be equipped with various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the rights of the present invention.
[0147] The user input unit (230) may be equipped with a microphone. The user may speak into the microphone equipped in the user input unit (230). At this time, the microphone equipped in the user input unit (230) may receive the voice spoken by the user.
[0148] The sensor unit (240) may be equipped with a gyroscope sensor (241) or an accelerometer sensor (243). The gyroscope sensor (241) can sense the movement of the remote control device (200).
[0149] The gyroscope sensor (241) can sense information regarding the operation of the remote control device (200) based on the x, y, and z axes. The accelerometer sensor (243) can sense information regarding the movement speed of the remote control device (200), etc. Meanwhile, the sensor unit (240) may further be equipped with a distance measuring sensor capable of sensing the distance to the display (180).
[0150] The output unit (250) can output video or sound corresponding to the operation of the user input unit (230) or the signal transmitted from the video display device (100). Through the output unit (250), the user can recognize whether the user input unit (230) is operated or whether the video display device (100) is controlled.
[0151] The output unit (250) may include an LED module (251) comprising at least one light-emitting element (e.g., LED (Light Emitting Diode)), a vibration module (253) that generates vibration, a sound output module (255) that outputs sound, and / or a display module (257) that outputs an image.
[0152] The power supply unit (260) can supply power to each component equipped in the remote control device (200). The power supply unit (260) may include at least one battery (not shown).
[0153] The power supply unit (260) can prevent unnecessary power consumption by stopping the power supply to each component equipped in the remote control device (200) when the movement of the remote control device (200) is not detected for a predetermined period of time through the sensor unit (240).
[0154] The power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when a predetermined event occurs. For example, the power supply unit (260) can resume power supply to each component when a predetermined key equipped in the remote control device (200) is operated. For example, the power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when movement of the remote control device (200) is detected through the sensor unit (240).
[0155] The storage unit (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).
[0156] When the remote control device (200) wirelessly transmits and receives signals through the image display device (100) and the RF module (221), the remote control device (200) and the image display device (100) can transmit and receive signals through a predetermined frequency band. The control unit (280) of the remote control device (200) can store and refer to information regarding the frequency band, etc., for wirelessly transmitting and receiving signals with the image display device (100) paired with the remote control device (200) in the storage unit (270).
[0157] The control unit (280) may include at least one processor, and can control the overall operation of the remote control device (200) using the processor included therein.
[0158] The control unit (280) can transmit a control signal corresponding to a predetermined key operation of the user input unit (230) or a control signal corresponding to the movement of the remote control device (200) sensed by the sensor unit (240) to the image display device (100) through the wireless communication unit (220).
[0159] The user input interface unit (150) of the image display device (100) may be equipped with a wireless communication unit (151) capable of wirelessly transmitting and receiving signals with a remote control device (200), and a coordinate value calculation unit (155) capable of calculating coordinate values of a pointer corresponding to the operation of the remote control device (200).
[0160] The user input interface unit (150) can wirelessly transmit and receive signals to and from the remote control device (200) through the RF module (152). It can also receive signals transmitted by the remote control device (200) according to the IR communication standard through the IR module (153).
[0161] The coordinate value calculation unit (155) can calculate the coordinate values (x,y) of a pointer (205) to be displayed on a display (170) by correcting hand tremor or error from a signal corresponding to the operation of a remote control device (200) received through a wireless communication unit (151).
[0162] A transmission signal from a remote control device (200) input to a video display device (100) through a user input interface unit (150) can be transmitted to a control unit (170) of the video display device (100). The control unit (170) of the video display device (100) can check information regarding the operation and key operation of the remote control device (200) from the signal transmitted from the remote control device (200), and control the video display device (100) in response.
[0163] As another example, the remote control device (200) can calculate pointer coordinate values corresponding to the operation and output them to the user input interface unit (150) of the image display device (100). In this case, the user input interface unit (150) of the image display device (100) can transmit information regarding the received pointer coordinate values to the control unit (170) without a separate hand tremor or error correction process.
[0164] In addition, as another example, the coordinate value calculation unit (155) may be provided inside the control unit (170) rather than the user input interface unit (150), unlike in the drawing.
[0165] FIG. 5 is a drawing referenced in the description of the first server of FIG. 1.
[0166] Referring to FIG. 5, the first server (400) may include a relay server (410), a STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), an AI (Artificial Intelligence) server (440) and / or a database (450). In this disclosure, the relay server (410), the STT server (420), the NLP server (430), and the AI server (440) are described as being distinct from one another, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), and the AI server (440) may be configured as a single server.
[0167] The relay server (410) can communicate with the video display device (100). The relay server (410) can transmit data between the STT server (420), the NLP server (430), and the video display device (100). The relay server (410) can store at least a portion of the data transmitted between the STT server (420), the NLP server (430), and the video display device (100).
[0168] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the video display device (100) via the relay server (410). The STT server (420) may also be named an ASR (Automatic Speech Recognition) server.
[0169] The STT server (420) can improve the accuracy of speech-to-text conversion by using a language model. The language model may refer to a model capable of calculating the probability of a sentence or calculating the probability of the next word appearing given previous words. For example, the language model may include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. That is, the STT server (420) can determine whether the text data converted from speech data has been appropriately converted by using a language model, and thereby improve the accuracy of the conversion to text data.
[0170] The NLP server (430) can receive text data. The NLP server (430) can perform intent analysis on the text data based on the received text data. The NLP server (430) can transmit intent analysis information indicating the result of the intent analysis to the video display device (100) via the relay server (410).
[0171] According to one embodiment, an NLP server (430) can generate intent analysis information by sequentially performing a morphological analysis step, a syntactic analysis step, a speech act analysis step, a dialogue processing step, etc., on text data. The morphological analysis step is a step of classifying text data corresponding to the voice uttered by the user into morpheme units, which are the smallest units with meaning, and determining what part of speech each classified morpheme has. The syntactic analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and determining what relationship exists between each classified phrase. Through the syntactic analysis step, the subject, object, and modifiers of the voice uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the voice uttered by the user using the results of the syntactic analysis step. Specifically, the speech act analysis step is a step of determining the intent of the sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The dialogue processing stage is a stage that uses the results of the speech act analysis stage to determine whether to answer, respond to, or ask a question inquiring for additional information in response to the user's utterance.
[0172] The AI server (440) can transmit a response to a request from the NLP server (430) to the relay server (410) or the NLP server (430). For example, when the AI server (440) receives a request to search for content from the NLP server (430), it can transmit data for at least one piece of content corresponding to the received request to search for content to the NLP server (430). For example, when the AI server (440) receives a request to recommend a keyword (hereinafter referred to as an associated keyword) associated with a predetermined keyword from the NLP server (430), it can transmit data for at least one associated keyword corresponding to the predetermined keyword to the NLP server (430).
[0173] According to one embodiment, the AI server (440) can process data using a super-large AI. For example, the AI server (440) can provide a response to a prompt received from the NLP server (430).
[0174] According to one embodiment, the AI server (440) may include a plurality of AI agent servers. Here, the plurality of AI agent servers may each correspond to the type of request of the NLP server (430). For example, the AI server (440) may use a first AI agent server when a text-related request is received from the NLP server (430), and use a second AI agent server when a video-related request is received.
[0175] The database (450) can store data used to generate a response for the AI server (440). The database (450) may include a plurality of sub-databases (451 to 453). For example, the database (450) may include a first database (451) that stores text such as syllables and words, a second database (452) that stores voice, a third database (453) that stores various learning models, etc.
[0176] FIG. 6 is a block diagram illustrating the configuration of a first server according to one embodiment of the present disclosure.
[0177] Referring to FIG. 6, the first server (400) may include a preprocessing unit (460), a controller (470), a communication unit (480) and / or a database (490).
[0178] The preprocessing unit (460) can preprocess voice received through the communication unit (480) or voice stored in the database (490).
[0179] The preprocessing unit (460) may be implemented as a separate chip from the controller (470) or as a chip included in the controller (470).
[0180] The preprocessing unit (460) receives a voice signal (spoken by the user) and can filter out noise signals from the voice signal before converting the received voice signal into text data.
[0181] When a preprocessing unit (460) is provided in the image display device (100), it can recognize a starter word to activate voice recognition of the image display device (100). The preprocessing unit (460) converts the starter word received through the user input interface unit (150) into text data, and if the converted text data is text data corresponding to a previously stored starter word, it can determine that the starter word has been recognized.
[0182] The preprocessing unit (460) can convert the noise-removed voice signal into a power spectrum.
[0183] The power spectrum can be a parameter indicating what frequency components are included in the waveform of a time-varying speech signal and in what magnitude.
[0184] The power spectrum shows the distribution of squared amplitude values according to the frequency of the waveform of the voice signal. This will be explained with reference to Fig. 7.
[0185] FIG. 7 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0186] Referring to FIG. 7, a voice signal (710) is shown. The voice signal (460) may be received from an external device or may be a signal stored in memory (170) in advance.
[0187] The x-axis of the voice signal (710) can represent time, and the y-axis can represent the amplitude.
[0188] The power spectrum processing unit (463) can convert a voice signal (710) with the x-axis being the time axis into a power spectrum (720) with the x-axis being the frequency axis.
[0189] The power spectrum processing unit (463) can convert the voice signal (710) into a power spectrum (720) using a Fast Fourier Transform (FFT).
[0190] The x-axis of the power spectrum (720) represents frequency, and the y-axis represents the square of the amplitude.
[0191] Figure 6 is explained again.
[0192] The functions of the preprocessing unit (460) and controller (470) described in FIG. 6 can also be performed on the NLP server (430).
[0193] The preprocessing unit (460) may include a wave processing unit (461), a frequency processing unit (462), a power spectrum processing unit (463), a speech-to-text (STT) conversion unit (464), etc.
[0194] The wave processing unit (461) can extract the waveform of the voice.
[0195] The frequency processing unit (462) can extract the frequency band of the voice.
[0196] The power spectrum processing unit (463) can extract the power spectrum of the voice.
[0197] A power spectrum can be a parameter that indicates what frequency components are included in a waveform and in what magnitude when a waveform that varies over time is given.
[0198] The speech-to-text (STT) conversion unit (464) can convert speech into text.
[0199] The voice-to-text conversion unit (464) can convert the voice of a specific language into text of that language.
[0200] The controller (470) can control the overall operation of the first server (400).
[0201] The controller (470) may include a voice analysis unit (471), a text analysis unit (472), a feature clustering unit (473), a text mapping unit (474) and / or a voice synthesis unit (475).
[0202] The voice analysis unit (471) can extract voice characteristic information using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed in the preprocessing unit (460).
[0203] Information on vocal characteristics may include one or more of the following: the speaker's gender, the speaker's voice (or timbre, tone), pitch, the speaker's manner of speaking, the speaker's speech rate, and the speaker's emotion.
[0204] In addition, voice characteristic information may also include the speaker's timbre.
[0205] The text analysis unit (472) can extract key expression phrases from the text converted by the voice-to-text conversion unit (464).
[0206] When the text analysis unit (472) detects that the tone between phrases in the converted text is different, it can extract the phrase with a different tone as the main expression phrase.
[0207] The text analysis unit (472) can determine that the tone has changed if the frequency band between phrases changes beyond a preset band.
[0208] The text analysis unit (472) may extract key words within the phrases of the converted text. Key words may be nouns existing within the phrases, but this is merely an example.
[0209] The feature clustering unit (473) can classify the speaker's speech type using the voice characteristic information extracted from the voice analysis unit (471).
[0210] The feature clustering unit (473) can classify the speaker's speech type by assigning weights to each of the type items constituting the characteristic information of the voice.
[0211] The feature clustering unit (473) can classify the speaker's speech type using the attention technique of a deep learning model.
[0212] The text mapping unit (474) can translate text converted into a first language into text in a second language.
[0213] The text mapping unit (474) can map text translated into a second language to text in the first language.
[0214] The text mapping unit (474) can map major expression phrases constituting the text of the first language to corresponding phrases of the second language.
[0215] The text mapping unit (474) can map utterance types corresponding to major expression phrases constituting the text of the first language to phrases of the second language. This is to apply the classified utterance types to the phrases of the second language.
[0216] The voice synthesis unit (475) can generate synthesized voice by applying the utterance type classified in the feature clustering unit (473) and the speaker's timbre to the main expression phrases of the text translated into a second language in the text mapping unit (474).
[0217] The controller (470) can determine the user's speech characteristics using one or more of the transmitted text data or power spectrum (720).
[0218] The characteristics of a user's speech may include the user's gender, pitch, timbre, topic of speech, speed of speech, and volume.
[0219] The controller (470) can obtain the frequency of the voice signal (710) and the amplitude corresponding to the frequency by using the power spectrum (720).
[0220] The controller (470) can determine the gender of the user who spoke the voice by using the frequency band of the power spectrum (470).
[0221] For example, the controller (470) can determine the user's gender as male if the frequency band of the power spectrum (720) is within a preset first frequency band range.
[0222] The controller (470) can determine the user's gender as female if the frequency band of the power spectrum (720) is within a preset second frequency band range. Here, the second frequency band range may be larger than the first frequency band range.
[0223] The controller (470) can determine the pitch of the voice by using the frequency band of the power spectrum (720).
[0224] For example, the controller (470) can determine the pitch of the sound according to the amplitude within a specific frequency band range.
[0225] The controller (470) can determine the user's tone by using the frequency bands of the power spectrum (720). For example, the controller (470) can determine a frequency band among the frequency bands of the power spectrum (720) in which the amplitude is greater than a certain size as the user's main tone range, and determine the determined main tone range as the user's tone.
[0226] The controller (470) can determine the user's speech rate from the converted text data through the number of syllables spoken per unit time.
[0227] For the converted text data, the controller (470) can determine the topic of the user's utterance using the Bag-Of-Word Model technique.
[0228] The Bag-Of-Word Model is a technique that extracts frequently used words based on their frequency within a sentence. Specifically, the Bag-Of-Word Model extracts unique words from a sentence and represents the frequency of each extracted word as a vector to determine the characteristics of the utterance topic.
[0229] For example, if words such as <running>, <physical strength>, etc. appear frequently in the text data of the controller (470), the topic of the user's speech can be classified as exercise.
[0230] The controller (470) can determine the topic of the user's utterance from the text data by using a known text categorization technique. The controller (470) can determine the topic of the user's utterance by extracting keywords from the text data.
[0231] The controller (470) can determine the user's volume by considering amplitude information across the entire frequency band.
[0232] For example, the controller (470) can determine the user's volume based on the average or weighted average of the amplitudes in each frequency band of the power spectrum.
[0233] The communication unit (480) can communicate with an external server via wired or wireless means. The communication unit (480) can communicate with the image display device (100) via wired or wireless means.
[0234] The database (490) can store the voice of the first language included in the content.
[0235] The database (490) can store synthesized speech in which the speech of the first language is converted into the speech of the second language.
[0236] The database (490) can store a first text corresponding to the voice of the first language and a second text in which the first text is translated into the second language.
[0237] The database (490) may store various learning models required for speech recognition.
[0238] Meanwhile, the control unit (170) of the image display device (100) shown in FIG. 2 may be equipped with the preprocessing unit (460) and controller (470) shown in FIG. 6.
[0239] That is, the control unit (170) of the image display device (100) may perform the functions of the preprocessing unit (460) and the controller (470).
[0240] FIG. 8 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.
[0241] That is, the voice recognition and synthesis process of Fig. 8 may be performed by the control unit (170) of the image display device (100) without going through a server.
[0242] Referring to FIG. 8, the control unit (170) of the image display device (100) may include an STT engine (810), an NLP engine (820), and a speech synthesis engine (830).
[0243] Each engine can be either hardware or software.
[0244] The STT engine (810) can perform the function of the STT server (420) of FIG. 5. That is, the STT engine (810) can convert voice data into text data.
[0245] The NLP engine (820) can perform the function of the NLP server (430) of FIG. 5. That is, the NLP engine (820) can obtain intent analysis information indicating the speaker's intent from the converted text data.
[0246] The speech synthesis engine (830) can perform the function of a speech synthesis server.
[0247] The speech synthesis engine (830) can search for syllables or words corresponding to given text data from a database and synthesize combinations of the searched syllables or words to generate synthesized speech.
[0248] The speech synthesis engine (830) may include a preprocessing engine (831) and a TTS engine (832).
[0249] The preprocessing engine (831) can preprocess text data before generating synthesized speech.
[0250] Specifically, the preprocessing engine (831) performs tokenization, which divides text data into meaningful units called tokens.
[0251] After performing tokenization, the preprocessing engine (831) can perform a cleansing operation to remove unnecessary characters and symbols to remove noise.
[0252] After that, the preprocessing engine (831) can combine word tokens with different representation methods to generate the same word token.
[0253] After that, the preprocessing engine (831) can remove meaningless word tokens (stopwords).
[0254] The TTS engine (832) can synthesize speech corresponding to the preprocessed text data and generate synthesized speech.
[0255] FIG. 9 is a flowchart of a method of operation of an image display device according to one embodiment of the present disclosure.
[0256] Referring to FIG. 9, the image display device (100) can activate a voice recognition function in operation S901. For example, the image display device (100) can activate a voice recognition function when a signal corresponding to an input of pressing a predetermined button (e.g., a voice input button) is received from a remote control device (200). For example, the image display device (100) can activate a voice recognition function when a trigger word that activates the voice recognition function is input through a microphone included in the input unit (160).
[0257] The video display device (100) may, in response to the activation of the voice recognition function while outputting content through the display (180) in the S902 operation, begin acquiring data corresponding to the content (hereinafter, content data). The video display device (100) may store the acquired content data in the memory (140).
[0258] Content data may include data related to content output through the display (180). Content data may include video data for the video of the content, audio data for the audio of the content, caption data for the closed caption of the content, metadata for the content, data for the broadcast channel providing the content, data for the time when the content is provided, data for the currently running application, etc.
[0259] For example, the video display device (100) can acquire a predetermined number of video frames per second in response to the activation of the voice recognition function. For example, the video display device (100) can acquire images captured from a screen output through a predetermined number of displays (180) per second in response to the activation of the voice recognition function. For example, the video display device (100) can acquire audio of content in response to the activation of the voice recognition function.
[0260] According to one embodiment, the video display device (100) can acquire data regarding the time at which content data is acquired. For example, the video display device (100) can acquire data regarding the time at which each of the video frames of the content is acquired. For example, the video display device (100) can acquire data regarding the time at which each part of the audio of the content is acquired.
[0261] The video display device (100) can receive a voice input corresponding to a voice spoken by a user in operation S903. For example, the video display device (100) can receive a voice signal corresponding to the voice input from a remote control device (200). For example, the video display device (100) can receive a voice signal corresponding to the voice input through a microphone included in the input unit (160).
[0262] When a voice input is received, the video display device (100) can transmit a voice signal corresponding to the voice input to the server (400). The video display device (100) can receive text corresponding to the voice input from the server (400). For example, when the video display device (100) transmits a voice signal to the server (400), the STT server (420) of the server (400) can convert the voice signal into text and transmit it to the video display device (100).
[0263] Meanwhile, the video display device (100) may also convert a voice signal corresponding to a voice input into text through an STT engine (810) included in the control unit (170).
[0264] According to one embodiment, the image display device (100) can transmit a voice signal of a preset unit, such as a syllable or a word, to a server (400). That is, when a user speaks a sentence, the image display device (100) can transmit a voice signal of a preset unit to the server (400) while a voice input corresponding to the sentence or phrase is received from a remote control device (200), and can receive text of a preset unit from the server (400).
[0265] According to one embodiment, the image display device (100) can output text corresponding to voice input through a display (180). For example, while a voice signal corresponding to a sentence or phrase is received from a remote control device (200), the image display device (100) can output text of a preset unit received from a server (400) through the display (180).
[0266] Meanwhile, the video display device (100) may output text of a preset unit converted through the STT engine (810) included in the control unit (170) through the display (180).
[0267] The video display device (100) can determine whether the reception of voice input is completed in operation S904. For example, the video display device (100) can determine that the reception of voice input is completed when the reception of a signal corresponding to the input of pressing a predetermined button (e.g., voice input button) from the remote control device (200) is terminated. For example, the video display device (100) can determine that the reception of voice input is completed when a voice signal is not received through the microphone of the input unit (160) for a predetermined period of time or longer.
[0268] The image display device (100) can obtain the result of performing an intention analysis on the voice input (hereinafter, the intention analysis result) when the reception of the voice input is completed in operation S905.
[0269] For example, the video display device (100) may request the server (400) to perform intent analysis for the voice input. The NLP server (430) of the server (400) may perform intent analysis on the text corresponding to the voice input converted by the STT server (420). The server (400) may transmit the intent analysis results to the video display device (100). Here, the intent analysis results may include keywords included in the voice input, sentence components of the keywords, the intent of the sentence, commands corresponding to the voice input, etc.
[0270] For example, the video display device (100) can obtain an intent analysis result by performing an intent analysis on text corresponding to a voice input converted from an STT engine (810) through an NLP engine (820) included in a control unit (170).
[0271] The video display device (100) can determine whether the voice input is related to the content based on the result of the intent analysis in operation S906. For example, if the voice input is related to an object or place included in the image output through the display (180), the video display device (100) can determine that the voice input is related to the content. For example, if the voice input is related to audio output through a speaker included in the audio output unit (185), the video display device (100) can determine that the voice input is related to the content. For example, if the voice input is related to the content of the content, the video display device (100) can determine that the voice input is related to the content.
[0272] According to one embodiment, the video display device (100) can acquire content data from a start time corresponding to the activation of a voice recognition function until a preset end time. For example, the video display device (100) can acquire content data until the time when the reception of voice input is completed. For example, the video display device (100) can acquire content data until the time when an intention analysis result is acquired. For example, the video display device (100) can acquire content data until the time when it determines whether the voice input is related to the content.
[0273] The video display device (100) may decide to use content data when the voice input is related to content in operation S907.
[0274] According to one embodiment, the image display device (100) can determine the data corresponding to the keyword among the content data acquired from the start time to the end time based on the keyword included in the voice input.
[0275] For example, the video display device (100) can determine data corresponding to the keyword by comparing the time at which content data is acquired with the time at which the keyword is spoken. At this time, the video display device (100) can determine the content data acquired in a predetermined time interval corresponding to the time at which the keyword is spoken as the data corresponding to the keyword.
[0276] For example, the video display device (100) can obtain a result of extracting an object from a video frame included in the content data when the voice input is related to the video of the content. At this time, the video display device (100) can determine the video frame containing the object corresponding to the keyword included in the voice input as the data corresponding to the keyword.
[0277] According to one embodiment, the video display device (100) may use content data even when the voice input is not related to the content. For example, the video display device (100) may use content data to obtain a recommendation query corresponding to the content.
[0278] The image display device (100) can obtain a result of processing voice input in response to the intention analysis result in operation S908 (hereinafter, voice recognition result).
[0279] For example, the video display device (100) may request the server (400) to transmit the voice recognition result based on the intent analysis result. At this time, if the video display device (100) determines that content data will be used, it may transmit the content data to the server (400). The server (400) may generate a voice recognition result based on the intent analysis result and the content data and transmit it to the video display device (100).
[0280] For example, the video display device (100) can generate a voice recognition result based on the intention analysis result and content data.
[0281] The video display device (100) can output a response corresponding to the voice recognition result in the S909 operation. For example, the video display device (100) can output a screen corresponding to the voice recognition result through the display (180). For example, the video display device (100) can output audio corresponding to the voice recognition result through a speaker included in the input unit (160).
[0282] Meanwhile, the video display device (100) can perform a predetermined operation when the voice input corresponds to the performance of a predetermined operation.
[0283] According to one embodiment, the video display device (100) may obtain at least one recommendation query corresponding to the content. The video display device (100) may output the recommendation query along with a response corresponding to the voice recognition result through the display (180). For example, if the voice input is not related to the content, the video display device (100) may transmit content data to the server (400) and receive a recommendation query corresponding to the content from the server (400). For example, if the voice input is not related to the content, the video display device (100) may generate a recommendation query corresponding to the content.
[0284] According to one embodiment of the present disclosure, when the image display device (100) performs all of the acquisition and storage of content data and the generation of voice recognition results based on content data, the server (400) can accurately provide the result of voice recognition for the voice spoken by the user even if it only provides services for processing voice input and analyzing the intent of voice input. In addition, when the image display device (100) additionally performs processing of voice input and analysis of the intent of voice input, the result of voice recognition for the voice spoken by the user can be accurately provided regardless of communication with the server (400) through a network.
[0285] Meanwhile, at least some of the operations of the image display device (100) described in FIG. 9 may be performed on a server (400). In this regard, the following explanation will be provided with reference to FIG. 10a to FIG. 10c.
[0286] FIGS. 10a to 10c are flowcharts of a method of operation of a system according to various embodiments of the present disclosure. Detailed descriptions of content that overlaps with the content described in FIG. 9 are omitted. Detailed descriptions of content that overlaps with FIGS. 10a to 10c are omitted. Meanwhile, FIGS. 10a to 10c describe that a server (400) performs processing and intent analysis of voice input, but is not limited thereto.
[0287] Referring to FIG. 10a, the image display device (100) can activate a voice recognition function in operation S1001.
[0288] The video display device (100) can begin acquiring content data in response to the activation of the voice recognition function in operation S1002.
[0289] The video display device (100) can transmit the acquired content data to the server (400) in operation S1003. The video display device (100) can transmit data regarding the time at which the content data was acquired to the server (400) along with the content data.
[0290] The server (400) can store content data received from the video display device (100) in operation S1004. The server (400) can store data regarding the time at which the content data received from the video display device (100) was acquired.
[0291] The video display device (100) can receive voice input corresponding to the voice spoken by the user in operation S1005.
[0292] The video display device (100) can transmit a voice signal corresponding to the voice input to the server (400) in operation S1006.
[0293] The video display device (100) may request the server (400) to analyze the intent of the voice input in operation S1007. For example, when the reception of the voice input is completed, the video display device (100) may request the server (400) to analyze the intent of the voice input.
[0294] The video display device (100) can terminate the acquisition of content data in response to the completion of receiving voice input in operation S1008.
[0295] The server (400) can perform intent analysis on a voice signal received from the video display device (100) in operation S1009. Meanwhile, the server (400) may also perform intent analysis on a voice input even if a voice signal is not received from the video display device (100) for a predetermined period of time or longer, or if a request for intent analysis is not received from the video display device (100).
[0296] The server (400) can process content data based on the intent analysis result in the S1010 operation. For example, the server (400) can process content data based on the intent analysis result if the voice input is related to content. For example, if the voice input is not related to content, the server (400) can generate a recommendation query corresponding to the content based on the content data.
[0297] The server (400) can transmit a voice recognition result to the video display device (100) in operation S1011. For example, if the voice input is related to content, the server (400) can transmit a voice recognition result generated based on the intent analysis result and the result of processing content data to the video display device (100). For example, if the voice input is not related to content, the server (400) can transmit a voice recognition result generated based on the intent analysis result to the video display device (100) along with a recommendation query.
[0298] The image display device (100) can output a response corresponding to the voice recognition result in the S1012 operation.
[0299] According to one embodiment of the present disclosure, when a video display device (100) transmits content data to a server (400) and the server (400) performs storage of content data, generation of voice recognition results based on content data, etc., storage of content data in memory (140) in the video display device (100) can be omitted whenever the voice recognition function is used. In addition, the video display device (100) does not need to separately secure storage space in memory (140) for storage of content data.
[0300] Referring to FIG. 10b, the image display device (100) can activate a voice recognition function in operation S1021.
[0301] The video display device (100) can begin acquiring content data in response to the activation of the voice recognition function in operation S1022.
[0302] The video display device (100) can store the acquired content data in the memory (140) during the S1023 operation. The video display device (100) can store data regarding the time at which the content data was acquired in the memory (140).
[0303] The video display device (100) can receive voice input corresponding to the voice spoken by the user in operation S1024.
[0304] The video display device (100) can transmit a voice signal corresponding to the voice input to the server (400) in operation S1025.
[0305] The video display device (100) can request the server (400) to analyze the intent of the voice input in operation S1026.
[0306] The video display device (100) can terminate the acquisition of content data in response to the completion of receiving voice input in operation S1027.
[0307] The server (400) can perform an intention analysis on the voice signal received from the video display device (100) in operation S1028.
[0308] In operation S1029, the server (400) can transmit an intent analysis result corresponding to the voice input to the video display device (100). The video display device (100) may terminate the acquisition of content data in response to receiving the intent analysis result from the server (400).
[0309] The video display device (100) can determine whether the voice input is related to the content based on the result of the intention analysis in operation S1230.
[0310] In operation S1231, the video display device (100) can transmit at least some of the content data stored in memory (140) to the server (400) when the voice input is related to content. At this time, the video display device (100) can transmit data corresponding to the keyword included in the voice input among the content data stored in memory (140) to the server (400). For example, the video display device (100) can determine the data to be transmitted to the server (400) among the content data stored in memory (140) based on the time when the keyword included in the voice input was spoken.
[0311] The server (400) can process content data received from the video display device (100) in operation S1232. For example, the server (400) can generate a speech recognition result based on the intent analysis result and the content data.
[0312] The server (400) can transmit the voice recognition result to the image display device (100) in operation S1033.
[0313] The video display device (100) can output a response corresponding to the voice recognition result in the S1034 operation.
[0314] Meanwhile, the video display device (100) may omit the transmission of content data to the server (400) when the voice input is not related to the content. For example, the video display device (100) may perform a predetermined operation when the voice input corresponds to the performance of a predetermined operation.
[0315] According to one embodiment of the present disclosure, when a video display device (100) acquires and stores content data and transmits the content data to a server (400) based on an intent analysis result, the amount of data transmitted between the video display device (100) and the server (400) can be minimized. In addition, the load on the server (400) generated during the storage and processing of content data can be minimized.
[0316] Referring to FIG. 10c, the image display device (100) can activate a voice recognition function in operation S1041.
[0317] The video display device (100) may instruct the server (400) to acquire content data in response to the activation of the voice recognition function in operation S1042. For example, the video display device (100) may transmit to the server (400) a URL (Uniform Resource Locator) corresponding to the content, metadata for the content, information about an account required to access the content, etc.
[0318] The server (400) may begin acquiring content data in operation S1043. The server (400) may store the acquired content data.
[0319] The video display device (100) can receive voice input corresponding to the voice spoken by the user in operation S1044.
[0320] The video display device (100) can transmit a voice signal corresponding to the voice input to the server (400) in operation S1045.
[0321] The video display device (100) can request the server (400) to analyze the intent of the voice input in operation S1046.
[0322] The server (400) can terminate the acquisition of content data in response to the request for intent analysis in operation S1047.
[0323] The server (400) can perform intent analysis on a voice signal received from the video display device (100) in operation S1048. Meanwhile, the server (400) may perform intent analysis on the voice input even if no voice signal is received from the video display device (100) for a predetermined period of time or longer, or if no request for intent analysis is received from the video display device (100). At this time, the server (400) may terminate the acquisition of content data.
[0324] The server (400) can process content data based on the intent analysis result in operation S1049.
[0325] The server (400) can transmit the voice recognition result to the image display device (100) in the S1050 operation.
[0326] The video display device (100) can output a response corresponding to the voice recognition result in the S1050 operation.
[0327] According to one embodiment of the present disclosure, when the server (400) performs all of the following—acquiring and storing content data, generating voice recognition results based on content data, etc.—the load on the video display device (100) that occurs during the storage and processing of content data can be minimized. In addition, regardless of the performance of the video display device (100), the result of performing voice recognition on the voice spoken by the user can be accurately provided. Furthermore, the amount of data transmitted between the video display device (100) and the server (400) can be minimized.
[0328] Referring to FIG. 11, the image display device (100) can output a first content screen (1100) through the display (180). The image display device (100) can activate a voice recognition function while outputting content through the display (180).
[0329] The video display device (100) can receive voice input for a voice spoken by a user when the voice recognition function is activated. The video display device (100) can output text (1110) corresponding to the voice spoken by the user.
[0330] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a search for clothing worn by a male figure displayed on the screen. Additionally, the video display device (100) can determine that the voice input is related to the video of the content.
[0331] The video display device (100) may request the server (400) to transmit the results of a search for clothing worn by a male person displayed on the screen. At this time, the video display device (100) may transmit video data among the acquired content data to the server (400). The video display device (100) may also transmit metadata about the content to the server (400) along with the video data.
[0332] Referring to FIG. 12, the image display device (100) can acquire content data (1200) from a start time (1201) corresponding to the activation of the voice recognition function until a preset end time (1202).
[0333] The video display device (100) can identify the point in time (1203) when the keyword 'man' included in the voice input is spoken. The video display device (100) can determine the video frame corresponding to the point in time (1203) when 'man' is spoken among the video frames included in the content data (1200) as the data corresponding to the keyword. For example, the video display device (100) can determine the video frame acquired at the point in time (1203) when 'man' is spoken among the content data (1200) as the data to be transmitted to the server (400). For example, the video display device (100) can determine the video frame (1210) acquired from the start time (1210) to the point in time (1203) when 'man' is spoken among the content data (1200) as the data to be transmitted to the server (400).
[0334] Meanwhile, the video display device (100) can obtain a result of extracting an object from a video frame included in the content data (1200) when the voice input is related to the video of the content. For example, the control unit (170) of the video display device (100) can extract an object from a video frame. At this time, the video display device (100) can determine, based on the result of extracting the object, a video frame containing a male person among the video frames included in the content data (1200) as data corresponding to the keyword.
[0335] Referring to FIG. 13, the image display device (100) can output objects (1310, 1320) for search results regarding clothing worn by a male person displayed on the screen through the display (180) as a response corresponding to the voice recognition result. In correspondence with two male people displayed on the first content screen (1100), objects (1310, 1320) for search results regarding two types of clothing can be output through the display (180).
[0336] At this time, even when a second content screen (1300) different from the first content screen (1100) is displayed, the video display device (100) can provide the user with a response related to the video displayed at the time the user speaks.
[0337] When a user selects one of the objects (1310, 1320) for the search results using a pointer (205), the image display device (100) can output a screen (e.g., a web page) of clothing corresponding to the selected object.
[0338] Referring to FIG. 14, the image display device (100) can output a third content screen (1400) through a display (180). The image display device (100) can receive voice input for voice spoken by a user. The image display device (100) can output text (1410) corresponding to the voice spoken by the user.
[0339] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a search for music included in the audio of the content. Additionally, the video display device (100) can determine that the voice input is related to the audio of the content.
[0340] The video display device (100) may request the server (400) to transmit the results of searching for music included in the audio of the content. At this time, the video display device (100) may transmit audio data among the acquired content data to the server (400). The video display device (100) may also transmit metadata about the content to the server (400) along with the audio data.
[0341] According to one embodiment, the video display device (100) can determine the time when the keyword 'music' included in the voice input is spoken. The video display device (100) can determine the audio portion corresponding to the time when 'music' is spoken among the audio portions of the content as data corresponding to the keyword. For example, the video display device (100) can determine the audio portions obtained in a predetermined time interval including the time when 'music' is spoken among the content data as data corresponding to the keyword.
[0342] Referring to FIG. 15, the image display device (100) can output an object (1510) for the result of searching for music included in the audio of the content through the display (180) as a response corresponding to the voice recognition result.
[0343] At this time, even if audio related to a fourth content screen (1500) is output that is different from the audio related to a third content screen (1400) output at the time the user speaks, the video display device (100) can provide the user with a response related to the audio output at the time the user speaks.
[0344] Referring to FIG. 16, the image display device (100) can output a fifth content screen (1600) through a display (180). The image display device (100) can receive voice input for a voice spoken by a user. The image display device (100) can output text (1610) corresponding to the voice spoken by the user.
[0345] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a search for the current weather. Additionally, the video display device (100) can determine that the voice input is not related to content.
[0346] The video display device (100) may request the server (400) to transmit the results of searching for the current weather. Meanwhile, the video display device (100) may transmit content data to the server (400) to obtain a recommendation query corresponding to the content. For example, the video display device (100) may transmit video data, metadata, etc. to the server (400).
[0347] Referring to FIG. 17, the image display device (100) can output an object (1710) for the result of searching for the current weather through the display (180) as a response corresponding to the voice recognition result.
[0348] The video display device (100) can output an object (1720) for a recommendation query received from the server (400) through the display (180). The recommendation query received from the server (400) may include a query related to voice input spoken by the user, a query related to the content of the content, a query related to a video displayed on the screen, etc.
[0349] Referring to FIG. 18, the image display device (100) can receive voice input for voice spoken by a user. The image display device (100) can output text (1810) corresponding to the voice spoken by the user.
[0350] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a summary of the relationships between the characters. Additionally, the video display device (100) can determine that the voice input is related to the content of the content.
[0351] The video display device (100) may request the server (400) to transmit a summary of the relationships between the characters. At this time, the video display device (100) may transmit video data and / or subtitle data among the acquired content data to the server (400). The video display device (100) may also transmit metadata about the content to the server (400) along with the video data and / or subtitle data.
[0352] The server (400) can generate a summary of the relationships between characters based on content data received from the video display device (100). For example, the server (400) can identify the character corresponding to the voice spoken by the user based on video data and / or subtitle data. At this time, the server (400) can receive the relationships between characters from a second server (500) that processes data using a super-large AI, by using a prompt requesting the relationships between characters.
[0353] Referring to FIG. 19, the image display device (100) can output an object (1910) regarding a summary of the relationships between characters through a display (180) as a response corresponding to the voice recognition result.
[0354] FIGS. 20a and FIG. 20b are flowcharts of a method of operation of an image display device according to another embodiment of the present disclosure. Detailed descriptions of contents that overlap with those described in FIG. 9 will be omitted.
[0355] Referring to FIG. 20a, the image display device (100) can activate a voice recognition function in operation S2001.
[0356] The video display device (100) can start acquiring content data in response to the activation of the voice recognition function while outputting content through the display (180) in the S2002 operation.
[0357] The video display device (100) can receive voice input corresponding to the voice spoken by the user in the S2003 operation.
[0358] The video display device (100) can determine whether the reception of voice input is completed in the S2004 operation.
[0359] The image display device (100) can obtain an intention analysis result for the voice input when the reception of the voice input is completed in operation S2005.
[0360] The video display device (100) can determine whether the voice input is related to the content based on the result of the intention analysis in operation S2006.
[0361] The video display device (100) may decide to use content data when the voice input is related to content in operation S2007.
[0362] The image display device (100) can obtain a voice recognition result that processes the voice input in response to the intention analysis result in operation S2008.
[0363] The video display device (100) can output a response corresponding to the voice recognition result in the S2009 operation.
[0364] The video display device (100) can determine whether the voice recognition function is terminated during the operation S2010. For example, the video display device (100) can determine that the voice recognition function is terminated when the output of the response corresponding to the voice recognition result is terminated and only the content is output.
[0365] Meanwhile, the video display device (100) may determine that the voice recognition function has not been terminated while a response corresponding to the voice recognition result is being output. The video display device (100) may determine that the voice recognition function has not been terminated if additional user input related to voice input is received while a response corresponding to the voice recognition result is being output. For example, the video display device (100) may determine that the voice recognition function has not been terminated if voice input is received while a response corresponding to the voice recognition result is being output.
[0366] Referring to FIG. 20b, the image display device (100) can perform an operation based on user input while the voice recognition function is not terminated during operation S2011. For example, if additional voice input is received while the voice recognition function is not terminated, the image display device (100) can output a response corresponding to the voice recognition result for the additionally received voice input.
[0367] The video display device (100) can determine whether a predetermined period has arrived in operation S2012. Here, the predetermined period may mean a period for generating a summary of the content.
[0368] In operation S2013, the video display device (100) can obtain a summary of the content when a predetermined period arrives. For example, the video display device (100) can obtain content data for a predetermined time corresponding to the predetermined period. At this time, when a predetermined period arrives as the predetermined time elapses, the video display device (100) can obtain a summary corresponding to the content data obtained during the predetermined time.
[0369] According to one embodiment, the video display device (100) may request the server (400) to transmit a partial summary. For example, the video display device (100) may transmit voice data and / or subtitle data among the content data acquired during a predetermined time to the server (400). The video display device (100) may transmit video data acquired during a predetermined time to the server (400) together with voice data and / or subtitle data.
[0370] The server (400) can generate a partial summary using a Large Language Model (LLM). For example, the server (400) can acquire text data corresponding to voice data and / or subtitle data received from a video display device (100). At this time, the server (400) can generate a summary that briefly explains the text data using the Large Language Model (LLM).
[0371] According to one embodiment, the image display device (100) can generate a summary corresponding to content data acquired over a predetermined period of time. For example, the image display device (100) can generate a summary that briefly explains text data corresponding to content data using a large language model (LLM). The image display device (100) can determine whether the speech recognition function is terminated during operation S2014. The image display device (100) can acquire a partial summary according to a predetermined period while the speech recognition function is not terminated.
[0372] The video display device (100) can obtain a summary of the content when the voice recognition function is terminated in operation S2015. For example, if the voice recognition function is terminated before a predetermined period has arrived, the video display device (100) can obtain a summary corresponding to the content data obtained for a period shorter than a predetermined time.
[0373] The video display device (100) can output a summary of the content in operation S2016. For example, the video display device (100) can output a full summary including a partial summary through the display (180).
[0374] Referring to FIG. 21, the image display device (100) can output a content screen (2100) through a display (180). The image display device (100) can activate a voice recognition function while outputting content through the display (180).
[0375] The video display device (100) can receive voice input for a voice spoken by a user when the voice recognition function is activated. The video display device (100) can output text (2110) corresponding to the voice spoken by the user.
[0376] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a search for a male person displayed on the screen. Additionally, the video display device (100) can determine that the voice input is related to the video of the content.
[0377] The video display device (100) may request the server (400) to transmit the results of a search for a male person displayed on the screen. At this time, the video display device (100) may transmit video data among the acquired content data to the server (400). The video display device (100) may also transmit metadata about the content to the server (400) along with the video data.
[0378] Referring to FIG. 22, the image display device (100) can output an object (2210) for the result of searching for a male person displayed on the screen as a response corresponding to the voice recognition result through the display (180).
[0379] Referring to FIG. 23, the image display device (100) can receive voice input for a voice spoken by a user while the voice recognition function is not terminated. The image display device (100) can output text (2310) corresponding to the voice spoken by the user.
[0380] The video display device (100) can determine, based on the result of intent analysis of the voice input, that the voice spoken by the user corresponds to a search for a female person displayed on the screen. Additionally, the video display device (100) can determine that the voice input is related to the video of the content.
[0381] The video display device (100) may request the server (400) to transmit the results of a search for a female person displayed on the screen. At this time, the video display device (100) may transmit video data among the acquired content data to the server (400). The video display device (100) may also transmit metadata about the content to the server (400) along with the video data.
[0382] Referring to FIG. 24, the image display device (100) can output an object (2410) for the result of searching for a female person displayed on the screen as a response corresponding to the voice recognition result through the display (180).
[0383] Referring to FIG. 25, the image display device (100) can acquire content data (2500) from a start time (2501) corresponding to the activation of the voice recognition function until a time (2502) when the voice recognition function is terminated.
[0384] Referring to reference numeral 2510, the image display device (100) can obtain, according to a predetermined period, a first partial summary corresponding to the first content data (2511), a second partial summary corresponding to the second content data (2512), a third partial summary corresponding to the third content data (2513), and a fourth partial summary corresponding to the fourth content data (2514).
[0385] Referring to reference numeral 2520, the image display device (100) can obtain a result of searching for a male person displayed on the screen based on data (2521) corresponding to the keyword 'male anchor' included in the first voice input. Additionally, the image display device (100) can obtain a result of searching for a female person displayed on the screen based on data (2522) corresponding to the keyword 'female anchor' included in the second voice input.
[0386] Referring to FIG. 26, the image display device (100) can output an object (2610) for a summary when the voice recognition function is terminated. The object (2610) for a summary may include a full summary including a first part summary to a fourth part summary.
[0387] FIGS. 27 to 29 are flowcharts of a method of operation of a system according to various embodiments of the present disclosure. Detailed descriptions of content that overlaps with previously described content will be omitted. Detailed descriptions of content that overlaps with FIGS. 27 to 29 will be omitted. Meanwhile, FIGS. 27 to 29 describe that a server (400) performs processing and intent analysis of voice input, but is not limited thereto.
[0388] Referring to FIG. 27, the image display device (100) can activate a voice recognition function in operation S2701.
[0389] The video display device (100) can begin acquiring content data in response to the activation of the voice recognition function in operation S2702.
[0390] The video display device (100) can transmit the acquired content data to the server (400) in the operation S2703. The video display device (100) can transmit data regarding the time at which the content data was acquired to the server (400) along with the content data.
[0391] The server (400) can store content data received from the video display device (100) in operation S2704. The server (400) can store data regarding the time at which the content data received from the video display device (100) was acquired.
[0392] The video display device (100) can receive a first voice input corresponding to the voice spoken by the user in operation S2705.
[0393] The video display device (100) can transmit a first voice signal corresponding to a first voice input to the server (400) in operation S2706.
[0394] The video display device (100) can request the server (400) to analyze the intent of the first voice input in operation S2707.
[0395] The server (400) can perform an intention analysis on the first voice signal received from the video display device (100) in operation S2708.
[0396] The server (400) can process content data based on the result of intent analysis of the first voice signal in operation S2709.
[0397] The server (400) can transmit a first voice recognition result corresponding to a first voice input to a video display device (100) in operation S2710.
[0398] The image display device (100) can output a response to the first voice input based on the first voice recognition result in the operation S2711.
[0399] Meanwhile, the server (400) can generate a first summary in operation S2712. For example, the first summary may correspond to content data stored for a predetermined period corresponding to a predetermined cycle from the start time corresponding to the activation of the voice recognition function.
[0400] Meanwhile, the video display device (100) may request the server (400) to generate a first summary. At this time, the server (400) may generate the first summary in accordance with the request received from the video display device (100).
[0401] The image display device (100) can receive a second voice input corresponding to the voice spoken by the user in operation S2713.
[0402] The video display device (100) can transmit a second voice signal corresponding to the second voice input to the server (400) in operation S2714.
[0403] The video display device (100) may request the server (400) to analyze the intent of the second voice input in operation S2715.
[0404] The server (400) can perform an intention analysis on the second voice signal received from the video display device (100) in operation S2716.
[0405] The server (400) can process content data based on the result of intent analysis of the second voice signal in operation S2717.
[0406] The server (400) can transmit a second voice recognition result corresponding to the second voice input to the image display device (100) in operation S2718.
[0407] The image display device (100) can output a response corresponding to the second voice input based on the second voice recognition result in the operation of S2719.
[0408] Meanwhile, the server (400) can generate a second summary in operation S2720. For example, the second summary may correspond to content data stored for a predetermined period corresponding to a predetermined period from the last point in time when content data corresponding to the first summary was stored.
[0409] Meanwhile, the video display device (100) may request the server (400) to generate a second summary. At this time, the server (400) may generate the second summary in accordance with the request received from the video display device (100).
[0410] The video display device (100) can terminate the voice recognition function in operation S2721.
[0411] The video display device (100) can terminate the acquisition of content data in response to the termination of the voice recognition function in operation S2722.
[0412] The video display device (100) may request the transmission of a summary of the content in operation S2723.
[0413] The server (400) can generate a third summary in operation S2724. For example, the second summary may correspond to content data stored from the last point in time when content data corresponding to the second summary was stored.
[0414] The server (400) can transmit a summary of the content to the video display device (100) in operation S2725. For example, the server (400) can transmit a first summary, a second summary, and a third summary to the video display device (100) as a summary of the content.
[0415] Meanwhile, the server (400) may transmit multiple partial summaries to the video display device (100) respectively. For example, the server (400) may transmit the generated partial summary to the video display device (100) in response to the generation of a partial summary. At this time, the server (400) may generate a final partial summary and transmit it to the video display device (100) based on a request corresponding to the termination of a voice recognition function received from the video display device (100).
[0416] The video display device (100) can output a summary of the content in operation S2726.
[0417] Referring to FIG. 28, the image display device (100) can activate a voice recognition function in operation S2801.
[0418] The video display device (100) can begin acquiring content data in response to the activation of the voice recognition function in operation S2802.
[0419] The video display device (100) can store the acquired content data in the memory (140) during the S2803 operation. The video display device (100) can store data regarding the time at which the content data was acquired in the memory (140).
[0420] The video display device (100) can receive a first voice input corresponding to the voice spoken by the user in the S2804 operation.
[0421] The video display device (100) can transmit a first voice signal corresponding to a first voice input to the server (400) in operation S2805.
[0422] The video display device (100) can request the server (400) to analyze the intent of the first voice input in operation S2806.
[0423] The server (400) can perform an intention analysis on the first voice signal received from the video display device (100) in operation S2807.
[0424] The server (400) can transmit the result of the intention analysis for the first voice signal to the video display device (100) in operation S2808.
[0425] The video display device (100) can determine whether the first voice input is related to the content based on the result of the intention analysis of the first voice signal in operation S2809.
[0426] In operation S2810, the video display device (100) can transmit at least a portion of the content data stored in the memory (140) to the server (400) when the first voice input is related to content. At this time, the video display device (100) can transmit to the server (400) data corresponding to the keyword included in the first voice input among the content data stored in the memory (140).
[0427] The server (400) can process content data received from the video display device (100) in the S2811 operation. For example, the server (400) can generate a first voice recognition result corresponding to the first voice input based on the intent analysis result for the first voice signal and the content data.
[0428] The server (400) can transmit the first voice recognition result to the image display device (100) in operation S2812.
[0429] The image display device (100) can output a response corresponding to the first voice input based on the first voice recognition result in the S2813 operation.
[0430] The image display device (100) can generate a first summary in operation S2814.
[0431] The video display device (100) can receive a second voice input corresponding to the voice spoken by the user in operation S2815.
[0432] The video display device (100) can transmit a second voice signal corresponding to the second voice input to the server (400) in operation S2816.
[0433] The video display device (100) may request the server (400) to analyze the intent of the second voice input in operation S2817.
[0434] The server (400) can perform an intention analysis on the second voice signal received from the video display device (100) in operation S2818.
[0435] The server (400) can transmit the result of the intention analysis for the second voice signal to the image display device (100) in operation S2819.
[0436] The video display device (100) can determine whether the second voice input is related to the content based on the result of the intent analysis of the second voice signal in the operation S2820. If the second voice input is not related to the content, the video display device (100) may omit the transmission of content data to the server (400).
[0437] The image display device (100) can output a response corresponding to the second voice input in the S2821 operation. For example, the image display device (100) can perform a predetermined operation if the second voice input corresponds to the performance of a predetermined operation that is not related to content.
[0438] The image display device (100) can generate a second summary in operation S2822.
[0439] The video display device (100) can terminate the voice recognition function in operation S2823.
[0440] The video display device (100) can terminate the acquisition of content data in response to the termination of the voice recognition function in operation S2824.
[0441] The image display device (100) can generate a third summary in operation S2825.
[0442] The video display device (100) can output a summary of the content in operation S2826. For example, the video display device (100) can output a first summary, a second summary, and a third summary as a summary of the content.
[0443] Meanwhile, the server (400) may generate a summary of the content and transmit it to the video display device (100). In this case, the video display device (100) transmits content data used to generate the summary of the content to the server (400), and the server (400) may generate a summary of the content based on the content data received from the video display device (100).
[0444] Referring to FIG. 29, the image display device (100) can activate a voice recognition function in operation S2901.
[0445] The video display device (100) can instruct the server (400) to acquire content data in response to the activation of the voice recognition function in operation S2902.
[0446] The server (400) may initiate the acquisition of content data in operation S2903. The server (400) may store the acquired content data.
[0447] The video display device (100) can receive a first voice input corresponding to the voice spoken by the user in operation S2904.
[0448] The video display device (100) can transmit a first voice signal corresponding to a first voice input to the server (400) in operation S2905.
[0449] The video display device (100) can request the server (400) to analyze the intent of the first voice input in operation S2906.
[0450] The server (400) can perform an intention analysis on the first voice signal received from the video display device (100) in operation S2907.
[0451] The server (400) can process content data based on the result of intent analysis of the first voice signal in operation S2908.
[0452] The server (400) can transmit a first voice recognition result corresponding to a first voice input to a video display device (100) in operation S2909.
[0453] The image display device (100) can output a response corresponding to the first voice recognition result in the operation of S2910.
[0454] Meanwhile, the server (400) can generate a first summary in operation S2911.
[0455] The video display device (100) can receive a second voice input corresponding to the voice spoken by the user in operation S2912.
[0456] The video display device (100) can transmit a second voice signal corresponding to the second voice input to the server (400) in operation S2913.
[0457] The video display device (100) may request the server (400) to analyze the intent of the second voice input in operation S2914.
[0458] The server (400) can perform an intention analysis on the second voice signal received from the video display device (100) in operation S2915.
[0459] The server (400) can process content data based on the result of intent analysis of the second voice signal in operation S2916.
[0460] The server (400) can transmit a second voice recognition result corresponding to the second voice input to the image display device (100) in operation S2917.
[0461] The image display device (100) can output a response corresponding to the second voice recognition result in operation S2918.
[0462] Meanwhile, the server (400) can generate a second summary in operation S2919.
[0463] The video display device (100) can terminate the voice recognition function in operation S2920.
[0464] The video display device (100) can instruct the server (400) to terminate the acquisition of content data in response to the termination of the voice recognition function in operation S2921.
[0465] The server (400) can terminate the acquisition of content data in operation S2922.
[0466] The server (400) can generate a third summary in the S2923 operation.
[0467] The server (400) can transmit a summary of the content to the video display device (100) in operation S2924. For example, the server (400) can transmit a first summary, a second summary, and a third summary to the video display device (100) as a summary of the content.
[0468] The video display device (100) can output a summary of the content in operation S2925.
[0469] As described above, according to at least one embodiment of the present disclosure, data regarding content can be accurately obtained in response to a user's voice utterance.
[0470] In addition, according to at least one embodiment of the present disclosure, the accuracy of the result of performing speech recognition on a voice spoken by a user can be improved.
[0471] In addition, according to at least one embodiment of the present disclosure, a summary of the content output while the voice recognition function is activated can be provided.
[0472] Referring to FIGS. 1 to 29, an image display device (100) according to one aspect of the present disclosure comprises: a display (180); a user input interface unit (150) that transmits a signal corresponding to a user input; and a control unit (170). When the control unit (170) activates a voice recognition function while outputting content through the display (180), it initiates the acquisition of content data corresponding to the content, receives a voice signal through the user input interface unit (150), acquires a voice recognition result that processes the voice signal based on the content data, and outputs a response corresponding to the voice recognition result through the display (180).
[0473] Additionally, according to one aspect of the present disclosure, the control unit (170) acquires the content data from a start time corresponding to the activation of the voice recognition function until a preset end time, and the end time may correspond to any one of the time when the reception of the voice signal is completed, the time when an intention analysis result is obtained by performing an intention analysis on the voice signal, and the time when it is determined whether the voice signal is related to the content.
[0474] Additionally, according to one aspect of the present disclosure, the control unit (170) obtains an intent analysis result by performing intent analysis on a voice signal received through the user input interface unit (150), determines whether the voice signal is related to the content based on the intent analysis result, and if the voice signal is related to the content, obtains a first voice recognition result corresponding to the content data and the intent analysis result, and if the voice signal is not related to the content, obtains a second voice recognition result corresponding to the intent analysis result.
[0475] Additionally, according to one aspect of the present disclosure, the control unit (170) can determine data corresponding to a keyword included in the voice signal among the content data, and obtain the first voice recognition result based on the data corresponding to the keyword.
[0476] Additionally, according to one aspect of the present disclosure, the control unit (170) may acquire data for a first time point at which the content data was acquired, and compare the first time point with a second time point at which the keyword was uttered to determine data corresponding to the keyword.
[0477] Additionally, according to one aspect of the present disclosure, the control unit (170) may, when the voice signal is related to the image of the content, obtain a result of extracting an object from an image frame included in the content data, and determine data corresponding to the keyword based on an image frame containing an object corresponding to the keyword.
[0478] Additionally, according to one aspect of the present disclosure, the control unit (170) may obtain at least one recommendation query corresponding to the content data when the voice signal is not related to the content, and output the recommendation query together with the response through the display (180).
[0479] Additionally, according to one aspect of the present disclosure, the control unit (170) may obtain a summary of the content based on the content data, and when the voice recognition function is terminated, output the summary of the content through the display (180).
[0480] Additionally, according to one aspect of the present disclosure, a summary of the content may be text generated in correspondence with the content data through a Large Language Model (LLM).
[0481] Additionally, according to one aspect of the present disclosure, the control unit (170) acquires the content data for a predetermined time corresponding to a predetermined period while the voice recognition function is activated, and when the predetermined period arrives, acquires a partial summary corresponding to the content data acquired during the predetermined time, and the summary for the content may include the partial summary acquired according to the predetermined period.
[0482] A system (10) according to one aspect of the present disclosure includes an image display device (100) and a server (400). When the image display device (100) activates a voice recognition function while outputting content through a display (180), it initiates the acquisition of content data corresponding to the content, transmits the content data to the server (400), transmits a voice signal received through a user input interface unit (150) to the server (400), receives a voice recognition result processed from the voice signal from the server (400), outputs a response corresponding to the voice recognition result through the display (180), and the server (400) generates a voice recognition result processed from the voice signal received from the image display device (100) based on the content data, and transmits the voice recognition result to the image display device (100).
[0483] Additionally, according to one aspect of the present disclosure, the server (400) may generate an intent analysis result by performing intent analysis on the voice signal, determine whether the voice signal is related to the content based on the intent analysis result, and if the voice signal is related to the content, generate a first voice recognition result corresponding to the content data and the intent analysis result, and if the voice signal is not related to the content, generate a second voice recognition result corresponding to the intent analysis result.
[0484] Additionally, according to one aspect of the present disclosure, the server (400) can determine data corresponding to a keyword included in the voice signal among the content data, and generate the first voice recognition result based on the data corresponding to the keyword.
[0485] Additionally, according to one aspect of the present disclosure, the server (400) generates at least one recommendation query corresponding to the content data when the voice signal is not related to the content, transmits the recommendation query together with the voice recognition result to the image display device (100), and the image display device (100) can output the recommendation query received from the server (400) together with the response through the display (180).
[0486] Additionally, according to one aspect of the present disclosure, the server (400) generates a summary of the content based on the content data and transmits it to the image display device (100), and the image display device (100) can output the summary of the content through the display (180) when the voice recognition function is terminated.
[0487] The attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification, and the technical concept disclosed in this specification is not limited by the attached drawings and should be understood to include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this disclosure.
[0488] Meanwhile, the method of operation of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices in which data that can be read by a processor is stored. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc., and also include implementation in the form of a carrier wave, such as transmission over the Internet. Furthermore, the processor-readable recording medium may be distributed across networked computer systems, so that processor-readable code can be stored and executed in a distributed manner.
[0489] Furthermore, although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. Display; A user input interface unit that transmits a signal corresponding to user input; and It includes a control unit, and The above control unit is, When a voice recognition function is activated while content is being output through the above display, the acquisition of content data corresponding to the content is initiated, and A voice signal is received through the above user input interface unit, and Based on the above content data, a speech recognition result is obtained by processing the above voice signal, and A video display device characterized by outputting a response corresponding to the voice recognition result through the above display.
2. In Paragraph 1, The above control unit is, From the start time corresponding to the activation of the above voice recognition function until a preset end time, the above content data is acquired, and The above termination point is, A video display device characterized by corresponding to any one of the following: a time when the reception of the voice signal is completed, a time when an intention analysis result is obtained by performing an intention analysis on the voice signal, and a time when it is determined whether the voice signal is related to the content.
3. In Paragraph 1, The above control unit is, Obtaining an intent analysis result by performing intent analysis on a voice signal received through the above user input interface unit, and Based on the above intent analysis results, determine whether the voice signal is related to the content, and If the above voice signal is related to the above content, a first voice recognition result corresponding to the above content data and the above intent analysis result is obtained, and A video display device characterized by obtaining a second voice recognition result corresponding to the intent analysis result when the above voice signal is not related to the above content.
4. In Paragraph 3, The above control unit is, Among the above content data, data corresponding to keywords included in the voice signal is determined, and An image display device characterized by obtaining the first voice recognition result based on data corresponding to the above keyword.
5. In Paragraph 4, The above control unit is, Acquiring data for the first point in time when the above content data was acquired, and An image display device characterized by determining data corresponding to the keyword by comparing the first point in time and the second point in time when the keyword is spoken.
6. In Paragraph 4, The above control unit is, If the above voice signal is related to the video of the above content, a result of extracting an object from the video frame included in the above content data is obtained, and An image display device characterized by determining data corresponding to the keyword based on an image frame containing an object corresponding to the keyword.
7. In Paragraph 3, The above control unit is, If the above voice signal is not related to the above content, at least one recommendation query corresponding to the above content data is obtained, and A video display device characterized by outputting the above-mentioned recommendation query along with the above-mentioned response through the above-mentioned display.
8. In Paragraph 3, The above control unit is, Based on the above content data, obtain a summary of the above content, and A video display device characterized by outputting a summary of the content through the display when the above voice recognition function is terminated.
9. In Paragraph 8, A video display device characterized in that the summary of the above content is text generated in correspondence with the content data through a Large Language Model (LLM).
10. In Paragraph 8, The above control unit is, With the above voice recognition function activated, the above content data is acquired for a predetermined time corresponding to a predetermined period, and When the above predetermined period arrives, a partial summary corresponding to the content data acquired during the above predetermined time is obtained, and A video display device characterized in that the summary of the above content includes the above partial summary obtained according to the above predetermined period.
11. In a system including a video display device and a server, The above image display device is, When a voice recognition function is activated while content is being output through a display, the acquisition of content data corresponding to the content is initiated, and Transmit the above content data to the above server, and The voice signal received through the user input interface unit is transmitted to the server, and Receiving a voice recognition result processed from the above voice signal from the server, and Through the above display, a response corresponding to the voice recognition result is output, and The above server is, Based on the above content data, a voice recognition result is generated by processing the voice signal received from the above video display device, and A system characterized by transmitting the above voice recognition result to the above image display device.
12. In Paragraph 11, The above server is, Generate an intent analysis result by performing intent analysis on the above voice signal, and Based on the above intent analysis results, determine whether the voice signal is related to the content, and If the above voice signal is related to the above content, a first voice recognition result corresponding to the above content data and the above intent analysis result is generated, and A system characterized by generating a second speech recognition result corresponding to the intent analysis result when the above voice signal is not related to the above content.
13. In Paragraph 12, The above server is, Among the above content data, data corresponding to keywords included in the voice signal is determined, and A system characterized by generating the first speech recognition result based on data corresponding to the above keywords.
14. In Paragraph 12, The above server is, If the above voice signal is not related to the above content, at least one recommendation query corresponding to the above content data is generated, and The above recommendation query is transmitted to the above video display device together with the above voice recognition result, and The above image display device is, A system characterized by outputting the recommendation query received from the above server along with the response through the above display.
15. In Paragraph 12, The above server is, Based on the above content data, a summary of the above content is generated and transmitted to the video display device, and The above image display device is, A system characterized by outputting a summary of the content through the display when the above voice recognition function is terminated.