System including image display device and operation method thereof
The video display device system addresses the lack of context-aware summaries by using a language model to generate tailored summaries and recommendations, improving user engagement.
Patent Information
- Application Number
- PCT/KR2024/003558
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-09-25
AI Technical Summary
Conventional methods for summarizing video content fail to consider user preferences or context, leading to ineffective content summaries.
A video display device system that generates summaries using a language model to process text, audio, and subtitles, dividing content into segments for tailored summaries and recommending optimal viewing methods.
Enables context-aware summaries and personalized content recommendations, enhancing user engagement and content understanding.
Smart Images

Figure KR2024003558_25092025_PF_FP_ABST
Abstract
Description
System including a video display device and its operating method
[0001] The present disclosure relates to a system including a video display device and an operating method thereof.
[0002] A video display device is a device that displays images for the user to view. For example, a video display device may include a television (TV), monitor, or notebook computer equipped with a liquid crystal display (LCD) using liquid crystals or an organic light-emitting diode (OLED) display using organic light-emitting diodes (OLED).
[0003] With the recent growth of the multimedia content industry, countless diverse content is now available to users. It's practically impossible for users to watch each and every video to fully understand the content, so there's a growing need to summarize the content and provide users with it.
[0004] Traditionally, content providers or service providers typically edited video content to create a summary video or summarize the content and provide it to users. Meanwhile, technologies that analyze video content and summarize its content are also being developed. However, conventional methods simply combine segments of the video to create a summary without considering the overall context, making it difficult to reflect user preferences or tendencies.
[0005] The present disclosure aims to solve the above-mentioned and other problems.
[0006] Another object is a system including a video display device capable of generating a summary of content based on the entire text corresponding to the video, audio, subtitles, synopsis, etc. of the content, and a method of operating the same.
[0007] Another object is to provide a system including a video display device capable of generating a summary of content using a language model, and an operating method thereof.
[0008] Another object is to provide a system including a video display device capable of generating various summaries for each of a plurality of segments constituting content, and a method of operating the same.
[0009] Another object is to provide a system including a video display device capable of recommending an optimal viewing method for content and a method of operating the same.
[0010] In order to achieve the above object, an operating method of a system including a video display device according to one embodiment of the present disclosure may include an operation of generating an entire text corresponding to a predetermined content from data about the predetermined content; an operation of dividing the entire text into a plurality of segments; an operation of generating at least one summary for each of the plurality of segments through a language model (LM); and an operation of outputting information related to viewing of the predetermined content through the video display device based on the summary for each of the plurality of segments.
[0011] In order to achieve the above object, a system including a video display device according to one embodiment of the present disclosure includes a control unit that processes data for a predetermined content, and the control unit generates an entire text corresponding to the predetermined content from the data for the predetermined content, divides the entire text into a plurality of segments, generates at least one summary for each of the plurality of segments through a language model (LM), and based on the summary for each of the plurality of segments, outputs information related to viewing of the predetermined content through the video display device.
[0012] The effects of a system including a video display device according to the present disclosure and its operating method are described as follows.
[0013] According to at least one embodiment of the present disclosure, a summary of content can be generated based on full text corresponding to video, audio, subtitles, synopsis, etc. of the content.
[0014] According to at least one embodiment of the present disclosure, a summary of content can be generated using a language model.
[0015] According to at least one embodiment of the present disclosure, a summary can be generated for each of a plurality of segments constituting content in various ways.
[0016] According to at least one embodiment of the present disclosure, an optimal viewing method for content can be recommended.
[0017] Further scope of the applicability of the present disclosure will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present disclosure will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present disclosure, are given by way of example only.
[0018] FIG. 1 is a diagram illustrating an image display system according to one embodiment of the present disclosure.
[0019] Figure 2 is an internal block diagram of the image display device of Figure 1.
[0020] Figure 3 is an internal block diagram of the control unit of Figure 2.
[0021] FIG. 4a is a drawing illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.
[0022] Figure 5 is a drawing referenced in the description of the first server of Figure 1.
[0023] FIG. 6 is a block diagram illustrating the configuration of a first server according to one embodiment of the present disclosure.
[0024] FIG. 7 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0025] FIG. 8 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.
[0026] FIGS. 9 and 10 are flowcharts of an operation method of an image display device according to one embodiment of the present disclosure.
[0027] FIGS. 11 to 25 are drawings for reference in explaining the operation of a video display device according to one embodiment of the present disclosure.
[0028] Hereinafter, the present disclosure will be described in detail with reference to the drawings. In the drawings, portions irrelevant to the description are omitted to clearly and concisely describe the present disclosure, and the same reference numerals are used for identical or extremely similar portions throughout the specification.
[0029] The suffixes "module" and "part" used in the following description are given solely for the convenience of writing this specification and do not impart any particularly significant meaning or role to the components themselves. Therefore, the terms "module" and "part" may be used interchangeably.
[0030] In this application, it should be understood that terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Additionally, while terms such as "first" and "second" may be used in this specification to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another.
[0032] FIG. 1 is a diagram illustrating an image display system according to various embodiments of the present invention.
[0033] Referring to FIG. 1, the image display system (10) may include an image display device (100) and / or a remote control device (200).
[0034] The image display device (100) may be a device that processes and outputs an image. The image display device (100) is not particularly limited as long as it can output a screen corresponding to an image signal, such as a TV, a notebook computer, or a monitor.
[0035] The video display device (100) can receive a broadcast signal, process the signal, and output the processed broadcast image. When the video display device (100) receives a broadcast signal, the video display device (100) may correspond to a broadcast receiving device.
[0036] The video display device (100) can receive broadcast signals wirelessly via an antenna, or can receive broadcast signals wired via a cable. For example, the video display device (100) can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, IPTV (Internet Protocol Television) broadcast signals, etc.
[0037] The remote control device (200) can be connected to the image display device (100) by wire and / or wirelessly, and can provide various control signals to the image display device (100). At this time, the remote control device (200) can include a device that establishes a wired or wireless network with the image display device (100), and transmits various control signals to the image display device (100) through the established network, or receives signals related to various operations processed in the image display device (100) from the image display device (100).
[0038] For example, various input devices such as a mouse, keyboard, space remote control, trackball, joystick, etc. can be used as the remote control device (200). The remote control device (200) can be referred to as an external device, and it is to be noted in advance that external devices and remote control devices can be used interchangeably as needed.
[0039] The video display device (100) can be connected to only a single remote control device (200) or can be connected to two or more remote control devices (200) simultaneously, and can change objects displayed on the screen or adjust the status of the screen based on control signals provided from each remote control device (200).
[0040] Meanwhile, the video display system (10) may further include at least one server (300). The video display device (100) may transmit and receive data with at least one server (300). For example, the video display device (100) may transmit and receive data with at least one server (300) via a network such as the Internet.
[0041] According to one embodiment, at least one server (300) may include a first server (400) that performs voice recognition, a second server (500) that processes data using a super-giant artificial intelligence model (hereinafter, super-giant AI), a third server (600) that provides content, etc.
[0042] Figure 2 is an internal block diagram of the image display device of Figure 1.
[0043] Referring to FIG. 2, the video display device (100) may include a broadcast receiving unit (105), an external device interface unit (130), a network interface unit (135), a storage unit (140), a user input interface unit (150), an input unit (160), a control unit (170), a display (180), an audio output unit (185), and / or a power supply unit (190).
[0044] The broadcast receiving unit (105) may include a tuner unit (110) and a demodulator unit (120).
[0045] Meanwhile, unlike the drawing, the image display device (100) may include only the broadcast reception unit (105) and the external device interface unit (130) among the broadcast reception unit (105), the external device interface unit (130), and the network interface unit (135). That is, the image display device (100) may not include the network interface unit (135).
[0046] The tuner unit (110) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit (110) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.
[0047] For example, the tuner unit (110) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (110) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (110) can be directly input to the control unit (170).
[0048] Meanwhile, the tuner unit (110) can sequentially select broadcast signals of all broadcast channels stored through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.
[0049] Meanwhile, the tuner unit (110) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.
[0050] The demodulation unit (120) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (110).
[0051] The demodulation unit (120) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal may be a signal in which a video signal, an audio signal, or a data signal is multiplexed.
[0052] The stream signal output from the demodulation unit (120) can be input to the control unit (170). The control unit (170) can output an image through the display (180) and output an audio through the audio output unit (185) after performing demultiplexing, image / audio signal processing, etc.
[0053] The external device interface unit (130) can transmit or receive data with a connected external device. To this end, the external device interface unit (130) may include an A / V input / output unit (not shown).
[0054] The external device interface unit (130) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.
[0055] In addition, the external device interface unit (130) can establish a communication network with various remote control devices (200) as illustrated in FIG. 1, and receive a control signal related to the operation of the image display device (100) from the remote control device (200) or transmit data related to the operation of the image display device (100) to the remote control device (200).
[0056] The A / V input / output unit can receive video and audio signals from an external device. For example, the A / V input / output unit can include an Ethernet terminal, a USB terminal, a CVBS (Composite Video Banking Sync) terminal, a component terminal, an S-video terminal (analog), a DVI (Digital Visual Interface) terminal, an HDMI (High Definition Multimedia Interface) terminal, an MHL (Mobile High-definition Link) terminal, an RGB terminal, a D-SUB terminal, an IEEE 1394 terminal, an SPDIF terminal, a Liquid HD terminal, etc. Digital signals input through these terminals can be transmitted to the control unit (170). At this time, analog signals input through the CVBS terminal and the S-video terminal can be converted into digital signals through an analog-to-digital converter (not shown) and transmitted to the control unit (170).
[0057] The external device interface unit (130) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit (130) can exchange data with an adjacent mobile terminal. For example, in mirroring mode, the external device interface unit (130) may receive device information, running application information, application images, etc. from the mobile terminal.
[0058] The external device interface unit (130) can perform short-range wireless communication using Bluetooth, RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), UWB (Ultra-Wideband), ZigBee, etc.
[0059] The network interface unit (135) can provide an interface for connecting the video display device (100) to a wired / wireless network including the Internet.
[0060] The network interface unit (135) may include a communication module (not shown) for connection to a wired / wireless network. For example, the network interface unit (135) may include a communication module for WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), etc.
[0061] The network interface unit (135) can transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.
[0062] The network interface unit (135) can receive web content or data provided by a content provider or network operator. That is, the network interface unit (135) can receive content such as movies, advertisements, games, VOD, broadcasts, etc., and information related thereto provided by a content provider or network provider via a network.
[0063] The network interface unit (135) can receive firmware update information and update files provided by the network operator, and transmit data to the Internet or content provider or network operator.
[0064] The network interface unit (135) can select and receive a desired application from among applications open to the public through a network.
[0065] The storage unit (140) may store programs for signal processing and control within the control unit (170), or may store processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (170), and may selectively provide some of the stored application programs upon request from the control unit (170).
[0066] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (170).
[0067] The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (130).
[0068] The storage unit (140) can store information about a specific broadcast channel through a channel memory function such as a channel map.
[0069] Although the storage unit (140) of FIG. 2 is provided separately from the control unit (170), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (170).
[0070] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit (140) and memory may be used interchangeably.
[0071] The user input interface unit (150) can transmit a signal input by the user to the control unit (170) or transmit a signal from the control unit (170) to the user.
[0072] For example, a user input signal such as power on / off, channel selection, screen setting, etc. may be transmitted / received from a remote control device (200), a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. may be transmitted to the control unit (170), a user input signal input from a sensor unit (not shown) that senses a user's gesture may be transmitted to the control unit (170), or a signal from the control unit (170) may be transmitted to the sensor unit.
[0073] The input unit (160) may be provided on one side of the main body of the video display device (100). For example, the input unit (160) may include a touch pad, a physical button, etc.
[0074] The input unit (160) can receive various user commands related to the operation of the video display device (100) and transmit a control signal corresponding to the input command to the control unit (170).
[0075] The input unit (160) may include at least one microphone (not shown) and may receive the user's voice through the microphone.
[0076] The control unit (170) may include at least one processor, and may control the overall operation of the image display device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.
[0077] The control unit (170) can demultiplex a stream input through the tuner unit (110), the demodulator unit (120), the external device interface unit (130), or the network interface unit (135), or process the demultiplexed signals to generate and output a signal for video or audio output.
[0078] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (170) or a video signal, data signal, control signal, etc. received from the external device interface unit (130).
[0079] The display (180) may include a display panel (not shown) having a plurality of pixels.
[0080] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display (180) may convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (170) to generate driving signals for the plurality of pixels.
[0081] The display (180) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display, etc., and may also be a 3D display. The 3D display (180) can be divided into a glasses-free type and a glasses type.
[0082] Meanwhile, the display (180) is configured as a touch screen and can be used as an input device in addition to an output device.
[0083] The audio output unit (185) receives a signal processed by the control unit (170) and outputs it as voice.
[0084] The image signal processed by the control unit (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (170) can also be input to an external output device through the external device interface unit (130).
[0085] The voice signal processed in the control unit (170) can be output as sound to the audio output unit (185). In addition, the voice signal processed in the control unit (170) can be input to an external output device through the external device interface unit (130).
[0086] Although not shown in FIG. 2, the control unit (170) may include a demultiplexing unit, an image processing unit, etc. This will be described later with reference to FIG. 3.
[0087] In addition, the control unit (170) can control the overall operation within the video display device (100). For example, the control unit (170) can control the tuner unit (110) to select (tune) a broadcast corresponding to a channel selected by the user or a previously stored channel.
[0088] In addition, the control unit (170) can control the image display device (100) by a user command or internal program input through the user input interface unit (150).
[0089] Meanwhile, the control unit (170) can control the display (180) to display an image. At this time, the image displayed on the display (180) may be a still image or a moving image, and may be a 2D image or a 3D image.
[0090] Meanwhile, the control unit (170) can cause a predetermined 2D object to be displayed within an image displayed on the display (180). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.
[0091] Meanwhile, the image display device (100) may further include a camera (not shown). The camera can capture images of a user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the image display device (100) above the display (180) or can be separately positioned. Image information captured by the camera can be input to the control unit (170).
[0092] The control unit (170) can recognize the user's location based on the image captured by the camera. For example, the control unit (170) can determine the distance (z-axis coordinate) between the user and the image display device (100). In addition, the control unit (170) can determine the x-axis coordinate and y-axis coordinate within the display (180) corresponding to the user's location.
[0093] The control unit (170) can detect the user's gesture based on an image captured from the camera unit, a signal detected from the sensor unit, or a combination thereof.
[0094] The power supply unit (190) can supply power to the entire image display device (100). In particular, it can supply power to a control unit (170) that can be implemented in the form of a system on chip (SOC), a display (180) for image display, and an audio output unit (185) for audio output.
[0095] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.
[0096] The remote control device (200) can transmit user input to the user input interface unit (150). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (150) and display or output the same as voice on the remote control device (200).
[0097] Meanwhile, the above-described video display device (100) may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.
[0098] Meanwhile, the block diagram of the image display device (100) illustrated in FIG. 2 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the image display device (100) actually implemented.
[0099] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.
[0100] Figure 3 is an internal block diagram of the control unit of Figure 2.
[0101] Referring to FIG. 3, a control unit (170) according to one embodiment of the present invention may include a demultiplexer (310), an image processing unit (320), a processor (330), an OSD generation unit (340), a mixer (345), a frame rate conversion unit (350), and / or a formatter (360). In addition, an audio processing unit (not shown) and a data processing unit (not shown) may be further included.
[0102] The demultiplexer (310) can demultiplex an input stream. For example, when MPEG-2 TS is input, it can be demultiplexed to separate it into video, audio, and data signals, respectively. Here, the stream signal input to the demultiplexer (310) may be a stream signal output from the tuner (110), the demodulator (120), or the external device interface (130).
[0103] The image processing unit (320) can perform image processing of a demultiplexed image signal. To this end, the image processing unit (320) may be equipped with an image decoder (325) and a scaler (335).
[0104] The video decoder (325) can decode a demultiplexed video signal, and the scaler (335) can perform scaling so that the resolution of the decoded video signal can be output on the display (180).
[0105] The video decoder (325) may include decoders of various standards. For example, it may include an MPEG-2, H.264 decoder, a 3D video decoder for color images and depth images, a decoder for multi-view images, etc.
[0106] The processor (330) can control the overall operation within the video display device (100) or the control unit (170). For example, the processor (330) can control the tuner (110) to select (tune) a broadcast corresponding to a channel selected by the user or a pre-stored channel.
[0107] In addition, the processor (330) can control the image display device (100) by a user command or internal program input through the user input interface unit (150).
[0108] Additionally, the processor (330) can perform data transmission control with the network interface unit (135) or the external device interface unit (130).
[0109] Additionally, the processor (330) can control the operation of the demultiplexing unit (310), the image processing unit (320), the OSD generation unit (340), etc. within the control unit (170).
[0110] The OSD generation unit (340) can generate OSD signals based on user input or on its own. For example, based on a user input signal input through the input unit (160), it can generate signals for displaying various information in the form of graphics or text on the screen of the display (180).
[0111] The generated OSD signal may include various data such as the user interface screen of the video display device (100), various menu screens, widgets, icons, etc. In addition, the generated OSD signal may include a 2D object or a 3D object.
[0112] Additionally, the OSD generation unit (340) can generate a pointer that can be displayed on the display (180) based on a pointing signal input from the remote control device (200).
[0113] The OSD generation unit (340) may include a pointing signal processing unit (not shown) that generates a pointer. It is also possible for the pointing signal processing unit (not shown) to be provided separately rather than within the OSD generation unit (240).
[0114] The mixer (345) can mix the OSD signal generated by the OSD generation unit (340) and the decoded image signal processed by the image processing unit (320). The mixed image signal can be provided to the frame rate conversion unit (350).
[0115] The frame rate converter (FRC) (350) can convert the frame rate of an input video. Meanwhile, the frame rate converter (350) can also output the video as is without a separate frame rate conversion.
[0116] The formatter (360) can arrange left-eye image frames and right-eye image frames of a frame rate-converted 3D image. In addition, it can output a synchronization signal (Vsync) for opening the left-eye glasses and right-eye glasses of a 3D viewing device (not shown).
[0117] Meanwhile, the formatter (360) can change the format of the input video signal into a video signal for display on the display (180) and output it.
[0118] Additionally, the formatter (360) can change the format of a 3D video signal. For example, it can change the format to any one of various 3D formats, such as a side-by-side format, a top-down format, a frame sequential format, an interlaced format, and a checker box format.
[0119] Meanwhile, the formatter (360) can also convert a 2D image signal into a 3D image signal. For example, according to a 3D image generation algorithm, an edge or a selectable object can be detected within a 2D image signal, and an object or a selectable object according to the detected edge can be separated and generated as a 3D image signal. At this time, the generated 3D image signal can be separated and aligned into a left-eye image signal (L) and a right-eye image signal (R), as described above.
[0120] Meanwhile, although not shown in the drawing, a 3D processor (not shown) for 3D effect signal processing may be further placed after the formatter (360). This 3D processor may process brightness, tint, and color adjustments of the image signal to improve the 3D effect. For example, it may perform signal processing to make the image clear at a close distance and blurry at a long distance. Meanwhile, the function of this 3D processor may be incorporated into the formatter (360) or incorporated into the image processing unit (320).
[0121] Meanwhile, the audio processing unit (not shown) within the control unit (170) can perform audio processing of the demultiplexed audio signal. For this purpose, the audio processing unit (not shown) can be equipped with various decoders.
[0122] Additionally, the audio processing unit (not shown) within the control unit (170) can process bass, treble, volume control, etc.
[0123] A data processing unit (not shown) within the control unit (170) can perform data processing of a demultiplexed data signal. For example, if the demultiplexed data signal is an encoded data signal, it can be decoded. The encoded data signal may be electronic program guide information (EPG) information that includes broadcast information such as the start time and end time of a broadcast program broadcast on each channel.
[0124] Meanwhile, the block diagram of the control unit (170) illustrated in FIG. 3 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the control unit (170) actually implemented.
[0125] In particular, the frame rate conversion unit (350) and the formatter (360) are not provided within the control unit (170), but may be provided separately, or may be provided separately as one module.
[0126] FIG. 4a is a drawing illustrating a control method of the remote control device of FIG. 2, and FIG. 4b is an example of an internal block diagram of the remote control device of FIG. 2.
[0127] Referring to FIG. 4a, it can be confirmed that a pointer (205) corresponding to a remote control device (200) is displayed on the display (180) of the video display device (100).
[0128] Referring to (a) of Fig. 4a, the user can move or rotate the remote control device (200) up and down, left and right, forward and backward. At this time, the pointer (205) displayed on the display (180) of the image display device (100) can be displayed in response to the movement of the remote control device (200). Since the pointer (205) of this remote control device (200) moves and is displayed according to the movement in 3D space, as shown in the drawing, it can be called a space remote control or a 3D pointing device.
[0129] Referring to (b) of FIG. 4a, when the user moves the remote control device (200) to the left, it can be confirmed that the pointer (205) displayed on the display (180) of the image display device (100) also moves to the left in response to the movement of the remote control device (200).
[0130] Information about the movement of the remote control device (200) detected through the sensor of the remote control device (200) can be transmitted to the image display device (100). The image display device (100) can calculate the coordinates of the pointer (205) from the information about the movement of the remote control device (200). The image display device (100) can display the pointer (205) to correspond to the calculated coordinates.
[0131] Referring to (c) of FIG. 4a, a user can move the remote control device (200) away from the display (180) while pressing a specific button provided on the remote control device (200). As a result, a selection area within the display (180) corresponding to the pointer (205) can be zoomed in and displayed in an enlarged manner. Conversely, when a user moves the remote control device (200) closer to the display (180) while pressing a specific button provided on the remote control device (200), a selection area within the display (180) corresponding to the pointer (205) can be zoomed out and displayed in a reduced manner.
[0132] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.
[0133] Meanwhile, when the user presses a specific button within the remote control device (200), recognition of up, down, left, and right movements may be excluded. That is, when the remote control device (200) moves away from or toward the display (180), up, down, left, and right movements are not recognized, and only forward and backward movements may be recognized. When the user does not press a specific button within the remote control device (200), only up, down, left, and right movements of the remote control device (200) may be recognized, and accordingly, only the pointer (205) may be moved.
[0134] Meanwhile, the movement speed or movement direction of the pointer (205) can correspond to the movement speed or movement direction of the remote control device (200).
[0135] Referring to FIG. 4b, the remote control device (200) may include a wireless communication unit (220), a user input unit (230), a sensor unit (240), an output unit (250), a power supply unit (260), a storage unit (270), and / or a control unit (280).
[0136] The wireless communication unit (220) can transmit and receive signals with the image display device (100).
[0137] In this embodiment, the remote control device (200) may be equipped with an RF module (221) capable of transmitting and receiving signals with the image display device (100) according to RF (Radio frequency) communication standards. In addition, the remote control device (200) may be equipped with an IR module (223) capable of transmitting and receiving signals with the image display device (100) according to IR (Infrared radiation) communication standards.
[0138] The remote control device (200) can transmit a signal including information about the movement of the remote control device (200) to the image display device (100) through the RF module (221). The remote control device (200) can receive the signal transmitted by the image display device (100) through the RF module (221).
[0139] The remote control device (200) can transmit commands for power on / off, channel change, volume change, etc. to the video display device (100) through the IR module (223).
[0140] The user input unit (230) may be composed of a keypad, buttons, a touch pad, a touch screen, etc. The user can input commands related to the video display device (100) to the remote control device (200) by operating the user input unit (230).
[0141] When the user input unit (230) has a hard key button, the user can input a command related to the video display device (100) to the remote control device (200) through a push operation of the hard key button.
[0142] When the user input unit (230) has a touch screen, the user can input commands related to the video display device (100) using the remote control device (200) by touching the soft keys of the touch screen.
[0143] Meanwhile, the user input unit (230) may be equipped with various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the present invention.
[0144] The user input unit (230) may be equipped with a microphone. The user may speak into the microphone provided in the user input unit (230). At this time, the microphone provided in the user input unit (230) may receive the voice spoken by the user.
[0145] The sensor unit (240) may be equipped with a gyro sensor (241) or an acceleration sensor (243). The gyro sensor (241) can sense the movement of the remote control device (200).
[0146] The gyro sensor (241) can sense information about the operation of the remote control device (200) based on the x, y, and z axes. The acceleration sensor (243) can sense information about the movement speed of the remote control device (200). Meanwhile, the sensor unit (240) may further include a distance measuring sensor capable of sensing the distance from the display (180).
[0147] The output unit (250) can output an image or sound corresponding to the operation of the user input unit (230) or to a signal transmitted from the image display device (100). Through the output unit (250), the user can recognize whether the user input unit (230) is being operated or whether the image display device (100) is being controlled.
[0148] The output unit (250) may include an LED module (251) including at least one light-emitting element (e.g., an LED (Light Emitting Diode)), a vibration module (253) that generates vibration, a sound output module (255) that outputs sound, and / or a display module (257) that outputs an image.
[0149] The power supply unit (260) can supply power to each component provided in the remote control device (200). The power supply unit (260) can include at least one battery (not shown).
[0150] The power supply unit (260) can prevent unnecessary power consumption by stopping the power supply to each component provided in the remote control device (200) when movement of the remote control device (200) is not detected for a predetermined period of time through the sensor unit (240).
[0151] The power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when a predetermined event occurs. For example, the power supply unit (260) can resume power supply to each component when a predetermined key equipped in the remote control device (200) is operated. For example, the power supply unit (260) can resume power supply to each component equipped in the remote control device (200) when movement of the remote control device (200) is detected through the sensor unit (240).
[0152] The storage unit (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).
[0153] When the remote control device (200) wirelessly transmits and receives signals through the image display device (100) and the RF module (221), the remote control device (200) and the image display device (100) can transmit and receive signals through a predetermined frequency band. The control unit (280) of the remote control device (200) can store and reference information about the frequency band, etc., through which signals can be wirelessly transmitted and received between the remote control device (200) and the image display device (100) paired therewith in the storage unit (270).
[0154] The control unit (280) may include at least one processor, and may control the overall operation of the remote control device (200) using the processor included therein.
[0155] The control unit (280) can transmit a control signal corresponding to a predetermined key operation of the user input unit (230) or a control signal corresponding to the movement of the remote control device (200) sensed by the sensor unit (240) to the image display device (100) via the wireless communication unit (220).
[0156] The user input interface unit (150) of the video display device (100) may be equipped with a wireless communication unit (151) capable of transmitting and receiving signals wirelessly with a remote control device (200), and a coordinate value calculation unit (155) capable of calculating the coordinate value of a pointer corresponding to the operation of the remote control device (200).
[0157] The user input interface unit (150) can wirelessly transmit and receive signals with the remote control device (200) via the RF module (152). In addition, the user input interface unit (150) can receive signals transmitted by the remote control device (200) according to the IR communication standard via the IR module (153).
[0158] The coordinate value calculation unit (155) can calculate the coordinate values (x, y) of the pointer (205) to be displayed on the display (170) by correcting hand tremors or errors from a signal corresponding to the operation of the remote control device (200) received through the wireless communication unit (151).
[0159] A transmission signal of a remote control device (200) input to a video display device (100) through a user input interface unit (150) can be transmitted to a control unit (170) of the video display device (100). The control unit (170) of the video display device (100) can check information about the operation and key operation of the remote control device (200) from the signal transmitted from the remote control device (200) and control the video display device (100) in response thereto.
[0160] As another example, the remote control device (200) can calculate the pointer coordinate values corresponding to the operation and output them to the user input interface unit (150) of the image display device (100). In this case, the user input interface unit (150) of the image display device (100) can transmit information about the received pointer coordinate values to the control unit (170) without a separate hand shake or error correction process.
[0161] In addition, as another example, the coordinate value calculation unit (155) may be provided inside the control unit (170) rather than the user input interface unit (150), unlike in the drawing.
[0162] Figure 5 is a drawing referenced in the description of the first server of Figure 1.
[0163] Referring to FIG. 5, the first server (400) may include a relay server (410), an STT (Speech To Text) server (420), an NLP (Natural Language Processing) server (430), an AI (Artificial Intelligence) server (440), and / or a database (450). In the present disclosure, the relay server (410), the STT server (420), the NLP server (430), and the AI server (440) are described as being distinct from each other, but are not limited thereto. For example, two or more of the relay server (410), the STT server (420), the NLP server (430), and the AI server (440) may be configured as one server.
[0164] The relay server (410) can communicate with the video display device (100). The relay server (410) can transfer data between the STT server (420), the NLP server (430), and the video display device (100). The relay server (410) can store at least a portion of the data transferred between the STT server (420), the NLP server (430), and the video display device (100).
[0165] The STT server (420) can receive voice data. The STT server (420) can convert the voice data into text data. The STT server (420) can transmit the text data to the video display device (100) via the relay server (410). The STT server (420) may also be referred to as an ASR (Automatic Speech Recognition) server.
[0166] The STT server (420) can improve the accuracy of speech-to-text conversion using a language model. The language model can refer to a model that can calculate the probability of a sentence or the probability of a subsequent word given previous words. For example, the language model can include probabilistic language models such as a unigram model, a bigram model, an N-gram model, etc. In other words, the STT server (420) can use the language model to determine whether text data converted from speech data has been appropriately converted, thereby improving the accuracy of conversion into text data.
[0167] The NLP server (430) can receive text data. Based on the received text data, the NLP server (430) can perform intent analysis on the text data. The NLP server (430) can transmit intent analysis information indicating the results of the intent analysis to the video display device (100) via the relay server (410).
[0168] According to one embodiment, the NLP server (430) may sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate intent analysis information. The morphological analysis step is a step of classifying text data corresponding to speech uttered by a user into morphemes, which are the smallest units having meaning, and determining which part of speech each classified morpheme has. The syntax analysis step is a step of using the results of the morphological analysis step to classify text data into noun phrases, verb phrases, adjective phrases, etc., and to determine what kind of relationship exists between each of the classified phrases. Through the syntax analysis step, the subject, object, and modifiers of the speech uttered by the user can be determined. The speech act analysis step is a step of analyzing the intent of the speech uttered by the user using the results of the syntax analysis step. Specifically, the speech act analysis step is a step of determining the intent of a sentence, such as whether the user is asking a question, making a request, or simply expressing an emotion. The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.
[0169] The AI server (440) may transmit a response according to a request from the NLP server (430) to the relay server (410) or the NLP server (430). For example, when a content search request is received from the NLP server (430), the AI server (440) may transmit data for at least one content corresponding to the received content search request to the NLP server (430). For example, when a recommendation request for a keyword associated with a predetermined keyword (hereinafter, “associated keyword”) is received from the NLP server (430), the AI server (440) may transmit data for at least one associated keyword corresponding to the predetermined keyword to the NLP server (430).
[0170] In one embodiment, the AI server (440) may process data using a super-large AI. For example, the AI server (440) may provide a response to a prompt received from the NLP server (430).
[0171] In one embodiment, the AI server (440) may include multiple AI agent servers. Here, the multiple AI agent servers may each correspond to a type of request from the NLP server (430). For example, the AI server (440) may utilize a first AI agent server when a text-related request is received from the NLP server (430), and may utilize a second AI agent server when a video-related request is received.
[0172] The database (450) can store data used to generate responses from the AI server (440). The database (450) can include multiple sub-databases (451 to 453). For example, the database (450) can include a first database (451) that stores text such as syllables and words, a second database (452) that stores voices, a third database (453) that stores various learning models, etc.
[0173] FIG. 6 is a block diagram illustrating the configuration of a first server according to one embodiment of the present disclosure.
[0174] Referring to FIG. 6, the first server (400) may include a preprocessing unit (460), a controller (470), a communication unit (480), and / or a database (490).
[0175] The preprocessing unit (460) can preprocess voice received through the communication unit (480) or voice stored in the database (490).
[0176] The preprocessing unit (460) may be implemented as a separate chip from the controller (470) or as a chip included in the controller (470).
[0177] The preprocessing unit (460) can receive a voice signal (spoken by a user) and filter out noise signals from the voice signal before converting the received voice signal into text data.
[0178] When a preprocessing unit (460) is provided in the image display device (100), it can recognize a trigger word for activating voice recognition of the image display device (100). The preprocessing unit (460) converts the trigger word received through the user input interface unit (150) into text data, and if the converted text data corresponds to a previously stored trigger word, it can be determined that the trigger word has been recognized.
[0179] The preprocessing unit (460) can convert the noise-removed voice signal into a power spectrum.
[0180] Power spectrum can be a parameter that indicates which frequency components are included in the waveform of a temporally varying voice signal and at what magnitude.
[0181] The power spectrum shows the distribution of the squared amplitude values of a voice signal's waveform according to frequency. This is explained with reference to Figure 7.
[0182] FIG. 7 is a diagram illustrating an example of converting a voice signal into a power spectrum according to one embodiment of the present disclosure.
[0183] Referring to FIG. 7, a voice signal (710) is illustrated. The voice signal (460) may be received from an external device or may be a signal pre-stored in memory (170).
[0184] The x-axis of the voice signal (710) may represent time, and the y-axis may represent the amplitude size.
[0185] The power spectrum processing unit (463) can convert a voice signal (710) whose x-axis is the time axis into a power spectrum (720) whose x-axis is the frequency axis.
[0186] The power spectrum processing unit (463) can convert a voice signal (710) into a power spectrum (720) using a fast Fourier transform (FFT).
[0187] The x-axis of the power spectrum (720) represents frequency, and the y-axis represents the square value of the amplitude.
[0188] Let's explain Figure 6 again.
[0189] The functions of the preprocessing unit (460) and controller (470) described in Fig. 6 can also be performed in the NLP server (430).
[0190] The preprocessing unit (460) may include a wave processing unit (461), a frequency processing unit (462), a power spectrum processing unit (463), a speech to text (STT) conversion unit (464), etc.
[0191] The wave processing unit (461) can extract the waveform of the voice.
[0192] The frequency processing unit (462) can extract the frequency band of the voice.
[0193] The power spectrum processing unit (463) can extract the power spectrum of the voice.
[0194] Power spectrum can be a parameter that indicates which frequency components are included in a given temporally varying waveform and at what magnitude.
[0195] The speech to text (STT) conversion unit (464) can convert speech into text.
[0196] The speech-to-text conversion unit (464) can convert speech in a specific language into text in that language.
[0197] The controller (470) can control the overall operation of the first server (400).
[0198] The controller (470) may include a voice analysis unit (471), a text analysis unit (472), a feature clustering unit (473), a text mapping unit (474), and / or a voice synthesis unit (475).
[0199] The voice analysis unit (471) can extract characteristic information of the voice by using one or more of the voice waveform, voice frequency band, and voice power spectrum preprocessed in the preprocessing unit (460).
[0200] Voice characteristic information may include one or more of the speaker's gender information, the speaker's voice (or tone), pitch, the speaker's speech style, the speaker's speaking rate, and the speaker's emotion.
[0201] Additionally, the characteristic information of the voice may further include the speaker's timbre.
[0202] The text analysis unit (472) can extract key expression phrases from the text converted by the voice text conversion unit (464).
[0203] When the text analysis unit (472) detects a difference in tone between phrases in the converted text, it can extract the phrases with different tones as main expression phrases.
[0204] The text analysis unit (472) can determine that the tone has changed if the frequency band between phrases has changed beyond a preset band.
[0205] The text analysis unit (472) can also extract key words within phrases of the converted text. Key words may be nouns within the phrases, but this is merely an example.
[0206] The feature clustering unit (473) can classify the speaker's speech type using the characteristic information of the voice extracted from the voice analysis unit (471).
[0207] The feature clustering unit (473) can classify the speaker's speech type by assigning weights to each of the type items that constitute the characteristic information of the voice.
[0208] The feature clustering unit (473) can classify the speaker's speech type using the attention technique of the deep learning model.
[0209] The text mapping unit (474) can translate text converted into a first language into text in a second language.
[0210] The text mapping unit (474) can map text translated into a second language to text in a first language.
[0211] The text mapping unit (474) can map the main expression phrases that constitute the text of the first language to the corresponding phrases of the second language.
[0212] The text mapping unit (474) can map utterance types corresponding to the main expression phrases that constitute the text of the first language to phrases of the second language. This is to apply the classified utterance types to the phrases of the second language.
[0213] The voice synthesis unit (475) can generate a synthesized voice by applying the speech type and the speaker's tone classified by the feature clustering unit (473) to the main expression phrases of the text translated into a second language by the text mapping unit (474).
[0214] The controller (470) can determine the user's speech characteristics using one or more of the transmitted text data or power spectrum (720).
[0215] User speech characteristics may include the user's gender, the user's pitch, the user's timbre, the user's speech topic, the user's speech rate, and the user's voice volume.
[0216] The controller (470) can obtain the frequency of the voice signal (710) and the amplitude corresponding to the frequency by using the power spectrum (720).
[0217] The controller (470) can determine the gender of the user who uttered the voice by using the frequency band of the power spectrum (470).
[0218] For example, the controller (470) can determine the user's gender as male if the frequency band of the power spectrum (720) is within a preset first frequency band range.
[0219] The controller (470) can determine the user's gender as female if the frequency band of the power spectrum (720) is within a preset second frequency band range. Here, the second frequency band range may be larger than the first frequency band range.
[0220] The controller (470) can determine the pitch of the voice by using the frequency band of the power spectrum (720).
[0221] For example, the controller (470) can determine the pitch of a sound based on the size of the amplitude within a specific frequency band range.
[0222] The controller (470) can determine the user's tone by using the frequency band of the power spectrum (720). For example, the controller (470) can determine a frequency band with an amplitude greater than a certain size among the frequency bands of the power spectrum (720) as the user's main vocal range, and determine the determined main vocal range as the user's tone.
[0223] The controller (470) can determine the user's speaking speed through the number of syllables spoken per unit time from the converted text data.
[0224] For the converted text data of the controller (470), the user's speech topic can be determined using the Bag-Of-Word Model technique.
[0225] The Bag-Of-Word Model technique extracts frequently used words based on their frequency within a sentence. Specifically, the Bag-Of-Word Model extracts unique words from a sentence, expresses the frequency of each word as a vector, and then determines the topic of the utterance.
[0226] For example, if words such as <running> and <physical strength> frequently appear in the controller (470) text data, the user's speech topic can be classified as exercise.
[0227] The controller (470) can determine the topic of a user's speech from text data using a known text categorization technique. The controller (470) can extract keywords from the text data to determine the topic of the user's speech.
[0228] The controller (470) can determine the user's vocal volume by considering amplitude information across the entire frequency band.
[0229] For example, the user's vocal volume can be determined based on the average or weighted average of the amplitude in each frequency band of the controller (470) power spectrum.
[0230] The communication unit (480) can communicate with an external server via wired or wirelessly. The communication unit (480) can communicate with the image display device (100) via wired or wirelessly.
[0231] The database (490) can store the speech of the first language included in the content.
[0232] The database (490) can store a synthesized voice in which a voice of a first language is converted into a voice of a second language.
[0233] The database (490) can store a first text corresponding to speech in a first language and a second text translated into a second language.
[0234] The database (490) may store various learning models required for voice recognition.
[0235] Meanwhile, the control unit (170) of the image display device (100) illustrated in FIG. 2 may be equipped with the preprocessing unit (460) and controller (470) illustrated in FIG. 6.
[0236] That is, the control unit (170) of the image display device (100) may perform the functions of the preprocessing unit (460) and the controller (470).
[0237] FIG. 8 is a block diagram illustrating the configuration of a control unit for voice recognition and synthesis of an image display device according to one embodiment of the present disclosure.
[0238] That is, the voice recognition and synthesis process of FIG. 8 may be performed by the control unit (170) of the image display device (100) without going through the server.
[0239] Referring to FIG. 8, the processor (180) of the image display device (100) may include an STT engine (810), an NLP engine (820), and a voice synthesis engine (830).
[0240] Each engine can be either hardware or software.
[0241] The STT engine (810) can perform the function of the STT server (420) of FIG. 5. That is, the STT engine (810) can convert voice data into text data.
[0242] The NLP engine (820) can perform the function of the NLP server (430) of FIG. 5. That is, the NLP engine (820) can obtain intent analysis information indicating the speaker's intent from converted text data.
[0243] The voice synthesis engine (830) can perform the function of a voice synthesis server.
[0244] The voice synthesis engine (830) can search for syllables or words corresponding to given text data from a database and synthesize a combination of the searched syllables or words to generate a synthesized voice.
[0245] The speech synthesis engine (830) may include a preprocessing engine (831) and a TTS engine (832).
[0246] The preprocessing engine (831) can preprocess text data before generating synthetic voice.
[0247] Specifically, the preprocessing engine (831) performs tokenization, which divides text data into tokens, which are meaningful units.
[0248] After tokenization is performed, the preprocessing engine (831) can perform a cleansing operation to remove unnecessary characters and symbols to remove noise.
[0249] Afterwards, the preprocessing engine (831) can integrate word tokens with different expression methods to generate the same word token.
[0250] Afterwards, the preprocessing engine (831) can remove meaningless word tokens (stopwords).
[0251] The TTS engine (832) can synthesize voice corresponding to preprocessed text data and generate a synthesized voice.
[0252] FIGS. 9 and 10 are flowcharts illustrating a method of operating a system according to one embodiment of the present disclosure. In the present disclosure, the operations of FIGS. 9 and 10 are exemplified as being performed on a first server (400), but are not limited thereto. For example, at least one of the operations of FIGS. 9 and 10 may be performed on a video display device (100), a second server (500), and / or a third server (600).
[0253] Referring to FIG. 9, the system (10) can generate text corresponding to the content from data about the content (hereinafter, referred to as content data) in operation S910. Here, the content data can include image data for the video of the content, audio data for the audio of the content, subtitle data for the closed caption of the content, synopsis data for the synopsis of the content, etc. For example, the first server (400) can convert a voice signal included in the audio data into text. At this time, the system (10) can generate text corresponding to the content based on the text converted from the audio signal, the closed caption obtained from the subtitle data, the synopsis obtained from the synopsis data, etc.
[0254] Text corresponding to content can be structured based on the playback time of the content. For example, text corresponding to content can include situations, backgrounds, character dialogue, emotions, etc. corresponding to each playback time from the start to the end of the content. Text corresponding to content can be designated as the entire text.
[0255] In operation S920, the system (10) can generate multiple segments from the entire text. Here, the multiple segments may each correspond to multiple sections constituting the content. For example, the system (10) can divide the entire text into multiple segments based on the playback time of the content. This will be described with reference to FIG. 10.
[0256] Referring to FIG. 10, the system (10) can initially segment the entire text into a plurality of temporary segments in operation S1010. The system (10) can segment the entire text into a plurality of temporary segments based on the units by which a language model (LM) recognizes text. For example, the system (10) can segment the entire text into a plurality of temporary segments based on tokens corresponding to the basic units by which a large language model (LLM) recognizes text. Hereinafter, an example of the system (10) processing text using a large language model (LLM) will be described.
[0257] Each of the multiple temporary segments can be divided from the entire text according to the size, number, etc. of tokens. The system (10) can evenly divide the entire text into multiple temporary segments according to the size, number, etc. of tokens.
[0258] The system (10) can, in operation S1020, combine temporary segments that are the subject of analysis. At this time, the system (10) can combine adjacent temporary segments among the temporary segments that are the subject of analysis. Here, the temporary segment that is the subject of analysis may mean a temporary segment in which at least one of the start time and the end time of the temporary segment is not determined as a point that divides an interval (hereinafter, an interval dividing point). A temporary segment in which both the start time and the end time are determined as an interval dividing point may be excluded from the analysis.
[0259] For example, the system (10) can combine a first temporary segment having an end point that is not determined as an interval dividing point among the temporary segments being analyzed, and a second temporary segment having a start point that is not determined as an interval dividing point. In this case, the system (10) can connect the end point of the first temporary segment and the start point of the second temporary segment to each other to form a single temporary segment.
[0260] In operation S1030, the system (10) can determine the section dividing point included in the combined temporary segments by checking the text contents of the combined temporary segments. For example, the system (10) can check the text contents of the combined temporary segments using a large language model (LLM). At this time, the system (10) can determine the section dividing point based on the characters, places, situations, times, etc. that appear in the text contents of the combined temporary segments. For example, the section dividing point can be determined based on cases where the characters that appear change, the content of the conversation change, the place or situation in which the conversation takes place change, the time order in which the conversation takes place change, etc.
[0261] The system (10) may, in operation S1040, re-divide the combined temporary segments according to the determined interval dividing points. For example, if two interval dividing points are determined from the text content of the combined temporary segments, the system (10) may divide the combined temporary segments into three new temporary segments. For example, if, for example, it is determined from the text content of the combined temporary segments that there are no interval dividing points, the system (10) may determine the combined temporary segments as one temporary segment.
[0262] The system (10) can determine, in operation S1050, whether a temporary segment to be analyzed exists.
[0263] In operation S1060, the system (10) can determine multiple temporary segments as multiple segments when there is no temporary segment to be analyzed, i.e., when all of the multiple temporary segments constituting the entire text are excluded from the analysis target.
[0264] Referring to FIG. 11, the system (10) can generate full text (1100) from content data.
[0265] The system (10) can initially divide the entire text (1100) into a plurality of primary temporary segments (1111 to 1120). At this time, the system (10) can divide the entire text (1100) into a plurality of primary temporary segments (1111 to 1120) according to the size, number, etc. of tokens. For example, the system (10) can equally divide the entire text (1100) into a plurality of primary temporary segments (1111 to 1120). At this time, all of the initially divided primary temporary segments (1111 to 1120) can correspond to analysis targets.
[0266] The system (10) can combine a plurality of adjacent primary temporary segments (1111 to 1120). The system (10) can verify the text content of the combined primary temporary segments using a large language model (LLM). The system (10) can determine a segmentation point for each of the combined primary temporary segments. At this time, the combined primary temporary segments can be divided into secondary temporary segments (1121 to 1129) according to the segmentation point determined for the primary temporary segments.
[0267] For example, the first primary temporary segment (1111) and the second primary temporary segment (1112) can be divided into the first secondary temporary segment (1121) and the second secondary temporary segment (1122). At this time, the start and end points of the first secondary temporary segment (1121) may both correspond to interval dividing points. Meanwhile, the start point of the second secondary temporary segment (1122) may correspond to an interval dividing point, but the end point may not correspond to an interval dividing point.
[0268] For example, the fifth primary temporary segment (1115) and the sixth primary temporary segment (1116) may form a fifth secondary temporary segment (1125). That is, the combined fifth primary temporary segment (1115) and the sixth primary temporary segment (1116) may not include an interval dividing point. In this case, both the start and end points of the fifth secondary temporary segment (1125) may not correspond to an interval dividing point.
[0269] The system (10) can combine analysis subjects among a plurality of secondary temporary segments (1121 to 1129). Among the plurality of secondary temporary segments (1121 to 1129), the first secondary temporary segment (1121) and the ninth secondary temporary segment (1129) can be excluded from the analysis subjects. The system (10) can verify the text contents of the combined secondary temporary segments (1122 to 1128) using a large language model (LLM). The system (10) can determine a section dividing point for each of the combined secondary temporary segments (1122 to 1128).
[0270] Meanwhile, the system (10) can determine the plurality of temporary segments (1121, 1129, 1131 to 1135) as the plurality of segments (1101 to 1107) when all of the plurality of temporary segments (1121, 1129, 1131 to 1135) are excluded from the analysis target.
[0271] Referring again to FIG. 9, the system (10) may, in operation S930, generate at least one summary for each of a plurality of segments using a large language model (LLM). The system (10) may generate a database (hereinafter, “summary DB”) of the summaries generated for each of the plurality of segments. The summary DB may be included in the first server (400).
[0272] A segment summary may include information derived from the textual content corresponding to the segment. For example, a segment summary may include the segment type, information about the characters, relationships between characters, character actions, character dialogue, story content, whether there are twists or turns in the overall story, whether there are spoilers for the overall story, and the chronological order within the overall story. Meanwhile, the segment type may be set to an action type if a certain proportion or more of the scenes correspond to action scenes, or a dialogue type if a certain proportion or more of the scenes correspond to dialogue scenes.
[0273] Referring to FIG. 12, the summary DB may include summaries (1201 to 1207) for each of a plurality of segments (1101 to 1107). Each of the summaries (1201 to 1207) for a segment may include a plurality of pieces of information derived from text content corresponding to the segment. For example, the first summary (1201) may correspond to the first segment (1101). At this time, the first summary (1201) may include the type of the first segment, information about characters, relationships between characters, actions of characters, dialogue of characters, content of the story, whether there is a twist in the overall story, whether there is a spoiler for the overall story, chronological order in the overall story, etc., derived from the text content corresponding to the first segment (1101).
[0274] Referring back to FIG. 9, the system (10) may, in operation S940, provide information related to viewing of content based on the summary DB. For example, the system (10) may provide text briefly describing each of the multiple segments, the importance of each of the multiple segments, the popularity of each of the multiple segments, etc.
[0275] According to one embodiment, the system (10) may provide text briefly describing each of a plurality of segments based on a summary database. For example, if the summary for a specific segment includes a predetermined important part of the story, the system (10) may generate text briefly describing the specific segment based on the remaining content excluding the content corresponding to the important part. Here, the important part may include a plot twist, spoiler, etc. Meanwhile, the system (10) may also add words, sentences, etc. that imply the content corresponding to the predetermined important part to the text briefly describing the specific segment.
[0276] According to one embodiment, the system (10) can determine the importance of multiple segments based on the user's preferences, the content of the summary, etc.
[0277] The system (10) can store the user's viewing history. Here, the user's viewing history may include the genre of content viewed by the user, the type of segment viewed by the user, the point in time at which the user stopped viewing, etc. For example, the user's viewing history may be stored in the first server (400). Meanwhile, the user's viewing history may be stored in accordance with the video display device (100) used by the user to view the content.
[0278] The system (10) can determine the importance of multiple segments based on the frequency with which a user has viewed content of a specific genre, the frequency with which a user has viewed segments of a specific type, etc. For example, if the system (10) determines that the user has viewed content of the action genre more than a predetermined standard based on the user's viewing history, the system (10) can determine a segment of the action type among the multiple segments as the main segment. For example, if the system (10) determines that the user has viewed segments of the action type less than a predetermined standard based on the user's viewing history, the system (10) can determine a segment of the conversation type among the multiple segments as the main segment. Here, the predetermined standard can correspond to the number of views, the viewing frequency, etc. Meanwhile, the system (10) can also determine the importance of multiple segments based on the user's preferred genre, type, etc.
[0279] The system (10) can determine the importance of multiple segments based on a summary of each segment. For example, the system (10) can determine a segment whose summary includes a plot twist and / or spoiler as a major segment.
[0280] In one embodiment, the system (10) can determine the popularity of multiple segments based on the viewing frequency of other users. The system (10) can determine the popularity of multiple segments based on the video viewing history of each of the multiple segments by other users. For example, the system (10) can store the video viewing history of each of the multiple segments by other users. In this case, the system (10) can determine a segment among the multiple segments that has been viewed by other users above a predetermined threshold as a popular segment.
[0281] The system (10) can determine, in operation S950, whether a content viewing method is determined. The system (10) can determine, for each of a plurality of segments, whether to provide a video or a summary of the segment. For example, the system (10) can receive user input regarding predetermined conditions that determine a content viewing method. At this time, the system (10) can determine a viewing method corresponding to the predetermined conditions for each of a plurality of segments based on a summary database.
[0282] The system (10) may, in operation S960, provide images or summaries corresponding to each of the plurality of segments, depending on the content viewing method. For example, the image display device (100) may output an image corresponding to the first segment through the display (180). For example, the image display device (100) may output a summary for the second segment through the display (180).
[0283] Referring to FIG. 13, the video display device (100) can output a setting screen (1300) that determines how to view specific content through the display (180).
[0284] The settings screen (1300) may include text (1310) briefly describing each of the plurality of segments, objects (1320) corresponding to a method of viewing each of the plurality of segments, indicators (1331) indicating major segments, indicators (1333) indicating popular segments, objects (1340) determining a viewing method, etc.
[0285] The user can select, for each of the plurality of segments, either an object (1321) corresponding to a summary or an object (1323) corresponding to a video using the pointer (205). When the user selects an object (1340) that determines a method of viewing content using the pointer (205), the system (10) can provide a video or summary corresponding to each of the plurality of segments according to the viewing method set for each of the plurality of segments.
[0286] Referring to FIG. 14, based on the viewing method for a given segment being set to summary, the video display device (100) can output a summary screen including a summary of the given segment through the display (180). For example, the video display device (100) can output a first summary screen (1410) including text (1415) (hereinafter, “basic text”) regarding the basic content of the story corresponding to the first segment.
[0287] According to one embodiment, the summary screen may include objects (1401, 1402) that change the content of the summary for the segment. When a user selects a first object (1401) included in a first summary screen (1410) using a pointer (205), the video display device (100) may output a second summary screen (1420) that includes text (1425) (hereinafter, “detailed text”) about the details of a story corresponding to the first segment. Meanwhile, when a user selects a second object (1402) included in the second summary screen (1420) using a pointer (205), the first summary screen (1410) that includes basic text (1415) corresponding to the first segment may be output again.
[0288] The basic content and / or detailed content of a story corresponding to a given segment may be included in the summary DB of the system (10). For example, the system (10) may provide the user with either the basic content or detailed content of a story corresponding to a first segment included in the summary DB.
[0289] Meanwhile, the system (10) can generate, in real time, the basic content and / or detailed content of a story corresponding to a given segment using a Large Language Model (LLM). For example, when a user selects a first object (1401) included in a first summary screen (1410) using a pointer (205), the system (10) can generate, based on a summary DB, the detailed content of the story corresponding to the first segment using a Large Language Model (LLM).
[0290] According to one embodiment, a summary screen corresponding to a given segment may include a third object (1403) that changes the viewing method for the given segment. When a user selects the third object (1403) included in the summary screen (1410, 1420) using a pointer (205), the image display device (100) may output an image corresponding to the given segment through the display (180).
[0291] Referring to FIG. 15, the image display device (100) can output a video screen including an image corresponding to a predetermined segment through the display (180). For example, based on the viewing method for the first segment being set to a video, the image display device (100) can output a video screen (1500) including an image corresponding to the first segment. For example, when a user selects a third object (1403) included in a summary screen (1410, 1420) corresponding to the first segment using a pointer (205), the image display device (100) can output a video screen (1500) including an image corresponding to the first segment.
[0292] According to one embodiment, a video screen corresponding to a given segment may include an object (1503) that changes the viewing method for the given segment. For example, when a user selects an object (1503) included in a video screen (1500) using a pointer (205), the video display device (100) may output a first summary screen (1410) including a summary of the first segment through the display (180).
[0293] Referring to FIGS. 16 and 17, the video display device (100) can output a summary screen (1600) containing a summary of the fourth segment. The system (10) can receive an input from a user requesting the output of a video corresponding to a specific section of the fourth segment. At this time, the system (10) can determine the playback point in time at which a specific section in the fourth segment starts based on the summary DB using a large language model (LLM).
[0294] For example, the system (10) may receive a user's voice (1601) requesting the output of a specific section corresponding to the conflict between characters included in the fourth segment through the input unit (160) included in the video display device (100) and / or the user input unit (230) included in the remote control device (200). The system (10) may perform voice recognition on the user's voice (1601) and determine that the user's voice (1601) is an input requesting the output of a specific section corresponding to the conflict between characters included in the fourth segment. At this time, the system (10) may determine the playback point in time at which the specific section corresponding to the conflict between characters in the fourth segment starts based on a summary DB using a large language model (LLM). According to one embodiment, the system (10) may determine a segment related to the conflict between characters based on a summary for each of a plurality of segments. At this time, the system (10) can determine a specific section corresponding to the conflict between characters in a segment related to the conflict between characters.
[0295] The video display device (100) can output a video screen (1700) containing a video corresponding to the fourth segment from the playback point in time when a specific section corresponding to the conflict between characters in the fourth segment begins through the display (180). Meanwhile, when an object (1703) included in the video screen (1700) that changes the viewing method is selected by the user, the video display device (100) can re-output a summary screen (1600) containing a summary of the fourth segment.
[0296] Referring to FIGS. 18 to 21, the system (10) may receive an input from a user requesting a recommendation for a viewing method corresponding to a predetermined condition. The system (10) may receive a user's voice requesting a recommendation for a viewing method corresponding to a predetermined condition through an input unit (160) included in a video display device (100) and / or a user input unit (230) included in a remote control device (200). Meanwhile, the system (10) may also receive a key input requesting a recommendation for a viewing method corresponding to a predetermined condition.
[0297] Referring to FIG. 18, the system (10) may receive a user's voice (1801) requesting a recommendation for a viewing method for viewing specific content within one hour. The system (10) may perform voice recognition on the user's voice (1801) and determine that the user's voice (1801) is an input requesting a viewing method for viewing specific content within one hour.
[0298] The system (10) can determine a viewing method for each of the plurality of segments based on the playback time of each of the plurality of segments, whether it is a main segment, whether it is a popular segment, etc. For example, the system (10) can determine a viewing method for the main segments among the plurality of segments as a video. At this time, if the total playback time of the main segments exceeds 1 hour, the system (10) can determine a viewing method for at least one of the main segments as a summary. On the other hand, if the total playback time of the main segments is less than 1 hour, the system (10) can determine a viewing method for the popular segment by considering the playback time of the popular segment. On the other hand, if the total playback time of the main segment and the popular segment is less than 1 hour, the viewing method for the remaining segments can be determined by considering the playback times of the remaining segments that are not the main segment and the popular segment.
[0299] Depending on the viewing method determined for each of the plurality of segments, an object (1320) corresponding to the viewing method for each of the plurality of segments may be displayed on the setting screen (1300). For example, depending on the viewing method for watching specific content within one hour, the viewing method for the first segment introducing the background setting of the specific content, the third, fifth, and seventh segments corresponding to the main segments, and the fourth segment corresponding to the popular segment may be displayed as a video, and the viewing methods for the remaining segments may be displayed as a summary. Meanwhile, the video display device (100) may display text (1800) indicating that it is a recommendation result for a viewing method corresponding to a predetermined condition through the setting screen (1300).
[0300] Meanwhile, referring to FIG. 19, the system (10) may receive a user's voice (1901) requesting a recommendation for a viewing method for watching specific content within 30 minutes. The system (10) may perform voice recognition on the user's voice (1901) and determine that the user's voice (1901) is an input requesting a viewing method for watching specific content within 30 minutes.
[0301] For example, the system (10) may determine the viewing method for major segments among a plurality of segments as a video. At this time, if the total sum of the playback times of the major segments exceeds 30 minutes, the system (10) may determine the viewing method for at least one of the major segments as a summary. On the other hand, if the total sum of the playback times of the major segments is less than 30 minutes, the system (10) may determine the viewing method for the popular segment by considering the playback time of the popular segment. At this time, if the total of the playback times of any one of the popular segments and the major segments exceeds 30 minutes, the system (10) may determine the viewing method for all segments except the major segment as a summary.
[0302] Depending on the viewing method determined for each of the multiple segments, an object (1320) corresponding to the viewing method for each of the multiple segments may be displayed on the settings screen (1300). For example, depending on the viewing method for watching specific content within 30 minutes, the viewing methods for the third, fifth, and seventh segments corresponding to the main segments may be displayed as videos, and the viewing methods for the remaining segments may be displayed as summaries.
[0303] Referring to FIG. 20, the system (10) can provide a viewing method that focuses on scenes corresponding to action scenes. The system (10) can determine at least one segment corresponding to an action scene based on a summary of each of a plurality of segments included in a summary DB. For example, the system (10) can determine the viewing method of a segment that is an action type as a video. For example, the system (10) can determine the viewing method of a segment that corresponds to an action scene for a predetermined portion or more of the playback time as a video.
[0304] Depending on the viewing method determined for each of the plurality of segments, an object (1320) corresponding to the viewing method for each of the plurality of segments may be displayed on the setting screen (1300). For example, the viewing method for the fourth, fifth, and seventh segments corresponding to action scenes may be displayed as a video, and the viewing method for the remaining segments may be displayed as a summary. Meanwhile, the video display device (100) may display text (2000) indicating that it is a recommended result for the viewing method corresponding to the action scene through the setting screen (1300).
[0305] Referring to FIG. 21, the system (10) can provide a viewing method that focuses on scenes corresponding to dialogue scenes. The system (10) can determine at least one segment corresponding to a dialogue scene based on a summary of each of the multiple segments included in the summary DB. For example, the system (10) can determine the viewing method of a segment that is a dialogue type as a video. For example, the system (10) can determine the viewing method of a segment that corresponds to a dialogue scene for a predetermined portion or more of the playback time as a video.
[0306] Depending on the viewing method determined for each of the plurality of segments, an object (1320) corresponding to the viewing method for each of the plurality of segments may be displayed on the setting screen (1300). For example, the viewing method for the second, third, sixth, and seventh segments corresponding to the dialogue scene may be displayed as a video, and the viewing method for the remaining segments may be displayed as a summary. Meanwhile, the video display device (100) may display text (2100) indicating that it is a recommended result for the viewing method corresponding to the dialogue scene through the setting screen (1300).
[0307] According to one embodiment, the system (10) can determine whether a user has previously viewed a specific content based on the user's viewing history when the user wishes to view a specific content. The system (10) can provide various viewing methods for the specific content based on the viewing history of the specific content. For example, if there is a history of viewing the specific content ending at a specific point in time, the system (10) can provide a method of continuing to view the specific content as a video from a specific point in time where viewing ended, a method of providing a summary of the content of the story before a specific point in time and then viewing the specific content as a video from a specific point in time, a method of providing a summary of the content of the story after a specific point in time, a method of viewing the video from a segment including a specific point in time, etc.
[0308] Referring to FIG. 22, in a case where a specific content is composed of multiple segments (1101 to 1107), there may be a history of viewing ending at a specific point in time (2200) of the fourth segment (1104). That is, among the specific content, the user may have already viewed the first segment (2210) up to a specific point in time (2200) of the first to third segments (1101 to 1103) and the fourth segment (1104), and may not have viewed the remaining segments (2220).
[0309] Referring to FIG. 23, if there is a history of viewing of a specific content ending at a specific point in time, the video display device (100) may output a screen (2300) recommending a viewing method based on the history of viewing the specific content. According to one embodiment, the screen (2300) recommending a viewing method may be displayed as a pop-up screen.
[0310] A screen (2300) recommending a viewing method may include an object (2310) for a method of continuing to view specific content as a video from a specific point in time after viewing has ended, an object (2320) for a method of viewing specific content as a video from a specific point in time after providing a summary of the story before a specific point in time, and an object (2330) for a method of providing a summary of the story after a specific point in time. A user may select one of the objects (2310 to 2330) using a pointer (205) to select one of various viewing methods for specific content.
[0311] In one embodiment, the system (10) can generate a summary of a portion of a segment in real time using a Large Language Model (LLM). For example, the system (10) can generate a summary of the content up to a specific point in time (2200) in the fourth segment (1104). For example, the system (10) can generate a summary of the content after a specific point in time (2200) in the fourth segment (1104) using the Large Language Model (LLM).
[0312] Meanwhile, referring to FIG. 24, the video display device (100) may output a settings screen (2400) that determines how to view specific content. The settings screen (2400) may include text (2410) briefly describing each of the plurality of segments. In this case, the text (2410) briefly describing each of the plurality of segments may be arranged according to the order of the plurality of segments.
[0313] The system (10) can receive a user's voice (2401) requesting a recommendation for a viewing method in which the story of a specific content is viewed in chronological order. The system (10) can perform voice recognition on the user's voice (2401) and determine that the user's voice (2401) is an input requesting a viewing method in which the story of a specific content is viewed in chronological order.
[0314] The system (10) can use a large language model (LLM) to determine the temporal order corresponding to each of a plurality of segments based on a summary database. For example, if a first segment corresponds to the present time in the story and a second segment corresponds to a past time in the story, the second segment may precede the first segment in the overall story in temporal order.
[0315] According to one embodiment, the system (10) can use a large language model (LLM) to determine, for each of a plurality of segments, whether it is temporally continuous with other segments. For example, the system (10) can determine, for the first segment (1101), whether it is temporally continuous with the second segment (1102) to the seventh segment (1107). At this time, the system (10) can determine whether the first segment (1101) is not temporally continuous with the other segments, whether the first segment (1101) precedes the other segments if it is temporally continuous, and whether the other segment precedes the first segment (1101).
[0316] Referring to FIG. 25, the video display device (100) may output a setting screen (2500) that determines a method for viewing specific content. The setting screen (2500) may include text (2510) briefly describing each of a plurality of segments. In this case, the text (2510) briefly describing each of the plurality of segments may be arranged in chronological order in the entire story corresponding to each of the plurality of segments. Meanwhile, the video display device (100) may display text (2501) indicating a recommendation result for a viewing method corresponding to a predetermined condition through the setting screen (2500).
[0317] As described above, according to at least one embodiment of the present disclosure, a summary of content can be generated based on full text corresponding to video, audio, subtitles, synopsis, etc. of the content.
[0318] Additionally, according to at least one embodiment of the present disclosure, a summary of content can be generated using a language model.
[0319] Additionally, according to at least one embodiment of the present disclosure, a summary can be generated in various ways for each of a plurality of segments constituting the content.
[0320] Additionally, according to at least one embodiment of the present disclosure, an optimal viewing method for content can be recommended.
[0321] Referring to FIGS. 1 to 25, an operating method of a system (10) including a video display device (100) according to one aspect of the present disclosure may include: an operation of generating an entire text corresponding to a predetermined content from data about the predetermined content; an operation of dividing the entire text into a plurality of segments; an operation of generating at least one summary for each of the plurality of segments through a language model (LM); and an operation of outputting information related to viewing of the predetermined content through the video display device (100) based on the summary for each of the plurality of segments.
[0322] Additionally, according to one aspect of the present disclosure, the data for the predetermined content may include at least one of first data for voice of the predetermined content and second data for closed captions of the predetermined content.
[0323] In addition, according to one aspect of the present disclosure, the operation of dividing the entire text into the plurality of segments may include: an operation of dividing the entire text into a plurality of temporary segments; an operation of combining temporary segments that are adjacent to each other among the plurality of temporary segments; and an operation of determining the plurality of segments based on texts of the combined temporary segments identified through the language model (LM).
[0324] In addition, according to one aspect of the present disclosure, the operation of determining the plurality of segments may include an operation of determining, based on text of the combined temporary segments, a section dividing point that divides a plurality of sections constituting the predetermined content included in the combined temporary segments; an operation of dividing the combined temporary segments based on the section dividing point; and an operation of determining the divided temporary segments into the plurality of segments.
[0325] In addition, according to one aspect of the present disclosure, the operation of determining the plurality of segments may further include an operation of determining the divided temporary segments as the plurality of segments based on the fact that both the start time and the end time of the divided temporary segments correspond to the section dividing point; and an operation of combining some of the divided temporary segments based on the fact that at least one of the start time and the end time of some of the divided temporary segments does not correspond to the section dividing point.
[0326] Additionally, according to one aspect of the present disclosure, the operation of dividing the entire text into the plurality of temporary segments may include an operation of dividing the entire text based on tokens corresponding to basic units by which the language model (LM) recognizes the text.
[0327] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include, if the summary for the predetermined segment includes a preset important part of the story of the content, an operation of outputting text describing the predetermined segment based on the remainder of the summary for the predetermined segment excluding the important part; and if the summary for the predetermined segment does not include the important part, an operation of outputting text describing the predetermined segment based on the entire summary for the predetermined segment.
[0328] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include an operation of determining a main segment among the plurality of segments based on a history of the user viewing the video through the video display device (100) and a summary of each of the plurality of segments; and an operation of outputting an indicator indicating the main segment through the video display device (100).
[0329] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include an operation of determining a popular segment among the plurality of segments based on a history of other users who have viewed the predetermined content viewing each of the plurality of segments as a video; and an operation of outputting an indicator indicating the popular segment through the video display device (100).
[0330] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include an operation of outputting a basic text describing the predetermined segment based on a summary of the predetermined segment; an operation of generating a detailed text describing the predetermined segment based on the summary of the predetermined segment through the language model (LM) when a user input requesting detailed information about the predetermined segment is received; and an operation of outputting the detailed text.
[0331] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include an operation of outputting text describing the predetermined segment based on a summary of the predetermined segment; an operation of determining a playback time point at which the specific section starts based on the summary of the predetermined segment through the language model (LM) when a user input requesting output of a video corresponding to a specific section of the predetermined segment is received; and an operation of outputting a video corresponding to the predetermined segment from a playback time point at which the specific section starts.
[0332] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content includes an operation of recommending at least one viewing method for the predetermined content based on a history of a user ending viewing of the predetermined content at a specific point in time, and the viewing method for the predetermined content may include a first method of providing a first summary of a story of the predetermined content before the specific point in time; and a second method of providing a second summary of a story of the predetermined content after the specific point in time.
[0333] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include an operation of generating the first summary based on a summary for each of the plurality of segments through the language model (LM) when a first user input selecting the first method is received; and an operation of generating the second summary based on a summary for each of the plurality of segments through the language model (LM) when a second user input selecting the second method is received.
[0334] In addition, according to one aspect of the present disclosure, the operation of outputting information related to viewing of the predetermined content may include: receiving a user input for a predetermined condition that determines a viewing method for the predetermined content; determining the viewing method corresponding to the predetermined condition for each of the plurality of segments based on a summary for each of the plurality of segments; and outputting the determined viewing method for each of the plurality of segments through the image display device.
[0335] A system (10) including a video display device (100) according to one aspect of the present disclosure includes a control unit that processes data for a predetermined content, and the control unit generates a full text corresponding to the predetermined content from the data for the predetermined content, divides the full text into a plurality of segments, generates at least one summary for each of the plurality of segments through a language model (LM), and based on the summary for each of the plurality of segments, can output information related to viewing of the predetermined content through the video display device (100).
[0336] The attached drawings are only intended to facilitate understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure.
[0337] Meanwhile, the operating method of the present disclosure can be implemented as processor-readable code on a processor-readable recording medium. A processor-readable recording medium includes all types of recording devices that store data that can be read by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across network-connected computer systems, so that the processor-readable code can be stored and executed in a distributed manner.
[0338] In addition, although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In the operating method of a system including a video display device, An action of generating a full text corresponding to a given content from data about the given content; An action of dividing the entire text into multiple segments; An operation of generating at least one summary for each of the plurality of segments through a language model (LM); and An operating method of a system including an operation of outputting information related to viewing of the predetermined content through the video display device based on a summary for each of the plurality of segments.
2. In paragraph 1, A method of operating a system, characterized in that the data for the predetermined content includes at least one of first data for voice of the predetermined content and second data for closed captions of the predetermined content.
3. In paragraph 1, The operation of dividing the entire text into the plurality of segments is as follows: An action of splitting the entire text into multiple temporary segments; An operation of combining adjacent temporary segments among the above multiple temporary segments; and A method of operating a system, characterized in that it includes an operation of determining the plurality of segments based on the text of the combined temporary segments confirmed through the language model (LM).
4. In paragraph 3, The operation of determining the above multiple segments is: An operation of determining a section dividing point that divides a plurality of sections constituting the predetermined content included in the combined temporary segment based on the text of the combined temporary segments; An operation of dividing the combined temporary segments based on the above section demarcation points; and A method of operating a system, characterized in that it includes an operation of determining the divided temporary segment into the plurality of segments.
5. In paragraph 4, The operation of determining the above multiple segments is: An operation of determining the divided temporary segments into the plurality of segments based on the fact that both the start and end points of the divided temporary segments correspond to the interval dividing points; and A method of operating a system, characterized in that it further includes an operation of combining some of the temporary segments based on at least one of the start time and the end time of some of the divided temporary segments not corresponding to the interval dividing point.
6. In paragraph 3, The operation of dividing the entire text into the above multiple temporary segments is: A method of operating a system, characterized in that the language model (LM) includes an operation of dividing the entire text based on tokens corresponding to basic units for recognizing text.
7. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An action of outputting text describing the given segment based on the remainder of the summary for the given segment, excluding the given important part, if the summary for the given segment includes a predetermined important part of the story of the content; and A method of operating a system, characterized in that it includes an action of outputting text describing the given segment based on the entire summary of the given segment, if the summary for the given segment does not include the important part.
8. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An operation of determining a main segment among the plurality of segments based on a history of the user viewing a video through the video display device and a summary of each of the plurality of segments; and A method of operating a system, characterized by including an operation of outputting an indicator indicating the main segment through the image display device.
9. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An operation of determining a popular segment among the plurality of segments based on the history of other users who have viewed the predetermined content and have viewed each of the plurality of segments as videos; and A method of operating a system, characterized by including an operation of outputting an indicator indicating the popular segment through the video display device.
10. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An action of outputting a basic text describing a given segment based on a summary for the given segment; When a user input requesting details about the given segment is received, an operation of generating detailed text describing the given segment based on a summary of the given segment through the language model (LM); and A method of operating a system, characterized by including an action of outputting the above detailed text.
11. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An action of outputting text describing a given segment based on a summary for the given segment; When a user input requesting the output of a video corresponding to a specific section of the above-mentioned predetermined segment is received, an operation of determining a playback point at which the specific section starts based on a summary of the above-mentioned predetermined segment through the language model (LM); and A method of operating a system, characterized in that it includes an operation of outputting a video corresponding to the predetermined segment from a playback point in time when the specific section is started.
12. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An action of recommending at least one viewing method for the given content based on a history of a user ending viewing of the given content at a specific point in time, The viewing method for the above content is as follows: A first method for providing a first summary of a story prior to the specific point in time of the above-mentioned content; and A method of operating a system characterized by including a second method of providing a second summary of a story after the specific point in time of the predetermined content.
13. In paragraph 12, The action of outputting information related to viewing the above content is as follows: When a first user input selecting the first method is received, an operation of generating the first summary based on a summary for each of the plurality of segments through the language model (LM); and A method of operating a system, characterized in that when a second user input selecting the second method is received, the method includes an operation of generating the second summary based on a summary for each of the plurality of segments through the language model (LM).
14. In paragraph 1, The action of outputting information related to viewing the above content is as follows: An action of receiving user input for a predetermined condition that determines a viewing method for the above-mentioned predetermined content; An operation of determining, for each of the plurality of segments, the viewing method corresponding to the predetermined condition based on the summary for each of the plurality of segments; and An operating method of a system characterized by including an operation of outputting the viewing method determined for each of the plurality of segments through the video display device.
15. In a system including a video display device, Includes a control unit that processes data for a given content, The above control unit, Generate the entire text corresponding to the above-mentioned content from the data for the above-mentioned content, Split the entire text above into multiple segments, Generate at least one summary for each of the plurality of segments using a language model (LM), A system characterized in that, based on a summary for each of the plurality of segments, information related to viewing of the predetermined content is output through the video display device.
Citation Information
Patent Citations
Summary video generation device and program thereof
JP2022181790A
Video generation method and apparatus of the same, and neural network training method and apparatus of the same
JP2023062173A
Highlight moving image generation system, highlight moving image generation method, and program
JP2023137704A
PLC system and method for transmitting control data thereof
KR1020240145273A
Device and method for inputting and transmitting sports data, non-transitory computer-readable storage medium storing program for performing the method
KR102715054B1
Cited By
Medical literature large-scale document self-adaptive block translation method and system
CN121145890A