Multimedia device and control method of multimedia device

By applying a time delay to either audio or video data based on the determined time correction value, the method effectively synchronizes audio and video in multimedia devices, resolving lip sync mismatches and enhancing user experience without increasing memory capacity.

WO2025116063A1PCT designated stage expired Publication Date: 2025-06-05LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/019378
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Multimedia devices face challenges in synchronizing audio and video data, leading to lip sync mismatches, especially when connected wirelessly to external speakers, which can result in poor user experience.

Method used

A method is proposed to synchronize audio and video data by applying a time delay to either the video data or the audio data, using a processor to determine the necessary time correction value and adjust the output timing accordingly.

Benefits of technology

This solution effectively addresses the lip sync mismatch issue without significantly increasing the internal memory capacity of multimedia devices, ensuring improved user experience across various multimedia scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023019378_05062025_PF_FP_ABST
    Figure KR2023019378_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A multimedia device for synchronization between audio data and video data is proposed. The multimedia device may comprise: a processor which processes received audio data and received video data; a display which outputs the processed video data; and a speaker which outputs the processed audio data, wherein the processor determines a time correction value for the video data, and applies the time correction value to delay the timing at which the video data is output on the display.
Need to check novelty before this filing date? Find Prior Art

Description

Multimedia devices and methods for controlling multimedia devices

[0001] The present invention relates to a multimedia device and a method for controlling the multimedia device.

[0002] Recently, new form factors have been discussed for multimedia devices such as mobile phones and TVs. Form factors refer to the structured form of a product.

[0003] The growing importance of form factor innovation in the display industry stems from the increasing mobility of consumers, the rapid advancement of device convergence and smartness, and the growing need for form factors that can be used freely and conveniently regardless of the usage environment, moving away from typical form factors tailored to specific usage environments.

[0004] For example, vertical TVs are expanding, breaking the stereotype of viewing TV horizontally. These TVs, designed to accommodate the MZ generation's familiarity with mobile content, feature screen orientation adjustments. This allows users to conveniently view images from social media or shopping sites, and even read comments while watching videos. Vertical TVs are particularly advantageous when paired with smartphones via Near Field Communication (NFC)-based mirroring. When watching regular TV or movies, the TV can be rotated horizontally.

[0005] As another example, rollable TVs and foldable smartphones are similar in that they both use "flexible displays." A flexible display, as the name suggests, refers to electronic components that are flexible. For these displays to be flexible, they must first be thin. The substrate that receives information and converts it into light must be thin and flexible to ensure long-lasting performance without damage.

[0006] Flexibility means being able to withstand impact without significant impact. While a flexible display is being bent or folded, the joints are constantly under pressure. To prevent internal damage from this pressure, the display must be durable yet still be able to deform smoothly when subjected to pressure.

[0007] Flexible displays, for example, are implemented based on OLEDs. OLEDs are displays that utilize organic light-emitting materials. Organic materials are relatively more flexible than inorganic materials like metals. Furthermore, OLEDs have a thin substrate, giving them a competitive edge over other displays. Previously used LCD substrates required separate liquid crystals and glass, limiting their ability to be reduced in thickness.

[0008] As a new form factor for TVs, demand is growing for TVs that can be easily transported both indoors and outdoors. In particular, the recent coronavirus pandemic has led to increased time spent at home, driving demand for second TVs. Furthermore, the growing number of people going outdoors for activities like camping has driven demand for TVs in new form factors that are easy to carry and transport.

[0009] Unlike conventional TVs, these TVs are connected to various external devices. For example, TVs are connected to external speakers or headsets to output various sound effects. However, the wireless connection between the TV and external speakers creates a problem: the video displayed on the display and the audio output through the speakers are out of sync (i.e., lip sync).

[0010] Meanwhile, the description below applies to all devices that output video and audio, not just TVs, and the term display device or multimedia device is used instead of TV.

[0011] In order to solve the above-described problem, the present invention proposes a method for solving the problem of lip sync mismatch.

[0012] More specifically, the present invention proposes a specific method of applying a time delay to either video data or audio data to match lip sync.

[0013] The problems to be solved by the present invention are not limited to the problems to be solved above, and other problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0014] A multimedia device for synchronizing audio data and video data is proposed, the device comprising: a processor for processing received audio data and received video data; a display for outputting the processed video data; and a speaker for outputting the processed audio data, wherein the processor determines a time correction value for the video data and can delay the output timing of the video data to the display by applying the time correction value.

[0015] A multimedia device for synchronizing audio data and video data is proposed, the device comprising: a processor for processing received audio data and received video data; a display for outputting the processed video data; and a speaker for outputting the processed audio data, wherein the processor requests audio data from a content provider, and after a time for a transmission delay of video data, requests the video data from the content provider, and in response to the request, calculates a time difference between reception points of the audio data and the video data according to reception of the video data, and applies an output time delay to the audio data by the difference between the calculated time difference and the time for the transmission delay.

[0016] A method for synchronizing audio data and video data is proposed, the method comprising: determining a time correction value for the video data; delaying an output timing of the video data to the display by applying the time correction value; and outputting the video data and the audio data according to the output timing delay.

[0017] A method for synchronizing audio data and video data is proposed, the method comprising: requesting audio data from a content provider; requesting video data from the content provider after a time for a transmission delay of the video data; calculating a time difference between reception points of the audio data and the video data in response to the request and upon reception of the video data; and applying an output time delay to the audio data by a difference between the calculated time difference and the time for the transmission delay.

[0018] The above problem solving methods are only some of the embodiments of the present invention, and various embodiments reflecting the technical features of the present invention can be derived and understood by a person having ordinary knowledge in the relevant technical field based on the detailed description of the present invention described below.

[0019] The present invention has the following effects.

[0020] The present invention can solve the lip sync mismatch problem described above.

[0021] Additionally, the present invention can solve the lip sync mismatch problem without significantly increasing the capacity of the internal memory of a multimedia device.

[0022] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present invention, provide embodiments of the present invention and, together with the detailed description, explain the technical idea of ​​the present invention.

[0023] Figure 1 is a block diagram for explaining each component of the display device.

[0024] FIG. 2 is a drawing showing a display device according to one embodiment of the present invention.

[0025] Figure 3 illustrates a processing procedure for audio data and video data in a multimedia device.

[0026] Figure 4 illustrates a procedure for compensating for the difference in processing time between audio data and video data in a multimedia device.

[0027] FIG. 5 illustrates a method for lip sync matching according to one embodiment of the present invention.

[0028] FIG. 6 illustrates a flowchart of a method for lip sync matching according to one embodiment of the present invention.

[0029] FIG. 7 illustrates a method for lip sync matching according to another embodiment of the present invention.

[0030] FIG. 8 illustrates a flowchart of a method for lip sync matching according to another embodiment of the present invention.

[0031] FIG. 9 illustrates a flowchart of a method for lip sync matching according to another embodiment of the present invention.

[0032] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present invention.

[0033] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0034] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0035] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0036] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but should be understood not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0037]

[0038] In the following description, the display device (100) is referred to as such, but the display device may be referred to by various names, such as a TV or a multimedia device, and the scope of the invention is not limited to the name.

[0039]

[0040] FIG. 1 is a block diagram for explaining each component of a display device (100) according to one embodiment of the present invention.

[0041] The display device (100) may include a broadcast receiving unit (110), an external device interface unit (171), a network interface unit (172), a storage unit (140), a user input interface unit (173), an input unit (130), a control unit (180), a display module (150), an audio output unit (160), and / or a power supply unit (190).

[0042] The broadcast receiving unit (110) may include a tuner unit (111) and a demodulator unit (112).

[0043] Meanwhile, unlike the drawing, the display device (100) may include only the external device interface unit (171) and the network interface unit (172) among the broadcast receiving unit (110), the external device interface unit (171), and the network interface unit (172). That is, the display device (100) may not include the broadcast receiving unit (110).

[0044] The tuner unit (111) can select a broadcast signal corresponding to a channel selected by the user or all previously stored channels among broadcast signals received through an antenna (not shown) or a cable (not shown). The tuner unit (111) can convert the selected broadcast signal into an intermediate frequency signal or a baseband video or audio signal.

[0045] For example, the tuner unit (111) can convert the selected broadcast signal into a digital IF signal (DIF) if it is a digital broadcast signal, and can convert it into an analog baseband video or audio signal (CVBS / SIF) if it is an analog broadcast signal. That is, the tuner unit (111) can process a digital broadcast signal or an analog broadcast signal. The analog baseband video or audio signal (CVBS / SIF) output from the tuner unit (111) can be directly input to the control unit (180).

[0046] Meanwhile, the tuner unit (111) can sequentially select broadcast signals of all stored broadcast channels through the channel memory function among the received broadcast signals and convert them into intermediate frequency signals or baseband video or audio signals.

[0047] Meanwhile, the tuner unit (111) may be equipped with multiple tuners to receive broadcast signals of multiple channels. Alternatively, a single tuner that simultaneously receives broadcast signals of multiple channels is also possible.

[0048] The demodulation unit (112) can perform a demodulation operation by receiving a digital IF signal (DIF) converted by the tuner unit (111). The demodulation unit (112) can output a stream signal (TS) after performing demodulation and channel decoding. At this time, the stream signal can be a signal in which a video signal, an audio signal, or a data signal is multiplexed.

[0049] The stream signal output from the demodulation unit (112) can be input to the control unit (180). The control unit (180) can perform demultiplexing, image / audio signal processing, etc., and then output an image through the display module (150) and output an audio through the audio output unit (160).

[0050] The sensing unit (120) refers to a device that detects changes within the display device (100) or detects external changes. For example, it may include at least one of a proximity sensor, an illumination sensor, a touch sensor, an infrared sensor (IR sensor), an ultrasonic sensor, an optical sensor (e.g., a camera), a voice sensor (e.g., a microphone), a battery gauge, and an environmental sensor (e.g., a hygrometer, a thermometer, etc.).

[0051] The control unit (180) can check the status of the display device (100) based on the information collected from the sensing unit (120), and if a problem occurs, can notify the user of it or control it to maintain the best condition by adjusting it on its own.

[0052] In addition, the content, picture quality, size, etc. of the image provided to the display module (150) can be controlled differently depending on the viewer detected by the sensing unit or the surrounding lighting, etc., to provide an optimal viewing environment. As smart display devices advance, the functions installed in the display devices increase, and the sensing unit (20) also increases along with them.

[0053] The input unit (130) may be provided on one side of the main body of the display device (100). For example, the input unit (130) may include a touch pad, a physical button, etc. The input unit (130) may receive various user commands related to the operation of the display device (100) and transmit a control signal corresponding to the input command to the control unit (180).

[0054] Recently, as the size of the bezel of the display device (100) has become smaller, the number of display devices (100) with a minimal input unit (130) in the form of a physical button exposed to the outside has increased. Instead, a minimum number of physical buttons are positioned on the back or side, and user input can be received through a remote control device (200) via a touchpad or a user input interface unit (173) described later.

[0055] The storage unit (140) may store programs for each signal processing and control within the control unit (180), or may store signal-processed video, audio, or data signals. For example, the storage unit (140) may store application programs designed for the purpose of performing various tasks that can be processed by the control unit (180), and may selectively provide some of the stored application programs upon request from the control unit (180).

[0056] Programs stored in the storage unit (140) are not particularly limited as long as they can be executed by the control unit (180). The storage unit (140) may also perform a function for temporarily storing video, audio, or data signals received from an external device through the external device interface unit (171). The storage unit (140) may store information regarding a specific broadcast channel through a channel memory function such as a channel map.

[0057] Although the storage unit (140) of FIG. 1 is provided separately from the control unit (180), the scope of the present invention is not limited thereto, and the storage unit (140) may be included within the control unit (180).

[0058] The storage unit (140) may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.).

[0059] The display module (150) can generate a driving signal by converting a video signal, data signal, OSD signal, control signal, etc. processed by the control unit (180) or a video signal, data signal, control signal, etc. received from the interface unit (171). The display module (150) can include a display panel having a plurality of pixels.

[0060] The plurality of pixels provided on the display panel may have RGB sub-pixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW sub-pixels. The display module (150) can convert image signals, data signals, OSD signals, control signals, etc. processed by the control unit (180) to generate driving signals for the plurality of pixels.

[0061] The display module (150) can be a PDP (Plasma Display Panel), an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diode), a flexible display module, etc., and may also be a 3D display module. The 3D display module (150) can be divided into a glasses-free type and a glasses type.

[0062] The display device (100) includes a display module that occupies most of the front surface area and a case that covers the rear side of the display module and packages the display module.

[0063] Recently, display devices (100) can utilize display modules (150) that can be bent, such as LEDs (Light Emitting Diodes) or OLEDs (Organic Light Emitting Diodes), to implement a curved screen rather than a flat screen.

[0064] Previously, LCDs were primarily used because they were unable to emit light on their own, so they relied on backlight units for light. Backlight units are devices that evenly distribute light from a light source to the liquid crystal display (LCD) located at the front. As backlight units have become thinner, thinner LCDs have become possible. However, it has been difficult to implement backlight units using flexible materials, and when the backlight unit is bent, it becomes difficult to supply light evenly to the LCD, causing screen brightness to fluctuate.

[0065] On the other hand, in the case of LEDs or OLEDs, since each element forming a pixel emits light on its own, it is possible to implement a curved display without using a backlight unit. In addition, since each element emits light on its own, even if the positional relationship with neighboring elements changes, its own brightness is not affected, so a curved display module (150) can be implemented using LEDs or OLEDs.

[0066] OLED (Organic Light-Emitting Diodes) panels made their debut in the mid-2010s and are rapidly replacing LCDs in the small- and medium-sized display market. OLEDs utilize the self-luminous phenomenon, where fluorescent organic compounds emit light when current flows through them. Their image quality response speed is faster than that of LCDs, resulting in virtually no afterimages when displaying video.

[0067] OLED uses three types of fluorescent organic compounds with self-luminous functions: red, green, and blue. It is a luminescent display product that utilizes the phenomenon in which electrons injected from the cathode and anode and positively charged particles combine within the organic material to emit light on their own, so there is no need for a backlight (backlight device) that reduces color.

[0068] LED (Light Emitting Diode) panels are a technology that uses one LED element as one pixel, and can reduce the size of the LED element compared to the past, enabling the implementation of a flexible display module (150). Devices previously referred to as LED TVs used LEDs as a light source for a backlight unit that supplied light to an LCD, and the LEDs themselves did not form a screen.

[0069] The display module includes a display panel, a coupling magnet positioned on a rear surface of the display panel, a first power supply unit, and a first signal module. The display panel may include a plurality of pixels (R, G, B). The plurality of pixels (R, G, B) may be formed in each area where a plurality of data lines and a plurality of gate lines intersect. The plurality of pixels (R, G, B) may be arranged or arranged in a matrix form.

[0070] For example, the plurality of pixels (R, G, B) may include a red (Red, hereinafter referred to as 'R') sub-pixel, a green (Green, 'G') sub-pixel, and a blue (Blue, 'B') sub-pixel. The plurality of pixels (R, G, B) may further include a white (White, hereinafter referred to as 'W') sub-pixel.

[0071] The side of the display module (150) that displays an image can be referred to as the front or front. When the display module (150) displays an image, the side from which the image cannot be observed can be referred to as the rear or back.

[0072] Meanwhile, the display module (150) is configured as a touch screen and can be used as an input device in addition to an output device.

[0073] The audio output unit (160) receives a signal processed by the control unit (180) and outputs it as voice.

[0074] The interface unit (170) serves as a passageway for various types of external devices connected to the display device (100). The interface unit may include not only a wired method for transmitting and receiving data through a cable, but also a wireless method using an antenna.

[0075] The interface unit (170) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port.

[0076] As an example of a wireless method, the broadcast receiving unit (110) described above may be included, and may include not only broadcast signals but also mobile communication signals, short-range communication signals, wireless Internet signals, etc.

[0077] The external device interface unit (171) can transmit or receive data with a connected external device. To this end, the external device interface unit (171) may include an A / V input / output unit (not shown).

[0078] The external device interface unit (171) can be connected to external devices such as a DVD (Digital Versatile Disk), Blu-ray, game device, camera, camcorder, computer (laptop), set-top box, etc., via wired / wireless connection, and can also perform input / output operations with the external devices.

[0079] In addition, the external device interface unit (171) can establish a communication network with various remote control devices (200), and receive a control signal related to the operation of the display device (100) from the remote control device (200), or transmit data related to the operation of the display device (100) to the remote control device (200).

[0080] The external device interface unit (171) may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit (not shown), the external device interface unit (171) can exchange data with an adjacent mobile terminal. In particular, in mirroring mode, the external device interface unit (171) can receive device information, running application information, application images, etc. from the mobile terminal.

[0081] The network interface unit (172) may provide an interface for connecting the display device (100) to a wired / wireless network, including the Internet. For example, the network interface unit (172) may receive content or data provided by the Internet, a content provider, or a network operator via a network. Meanwhile, the network interface unit (172) may include a communication module (not shown) for connection to a wired / wireless network.

[0082] The external device interface unit (171) and / or the network interface unit (172) may include a communication module for short-range communication such as Wi-Fi (Wireless Fidelity), Bluetooth, Bluetooth Low Energy (BLE), Zigbee, and NFC (Near Field Communication), a communication module for cellular communication such as LTE (long-term evolution), LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), and WiBro (Wireless Broadband).

[0083] The user input interface unit (173) can transmit a signal input by the user to the control unit (180), or transmit a signal from the control unit (180) to the user. For example, it can transmit / receive a user input signal such as power on / off, channel selection, screen setting, etc. from a remote control device (200), or transmit a user input signal input from a local key (not shown) such as a power key, a channel key, a volume key, a setting value, etc. to the control unit (180), or transmit a user input signal input from a sensor unit (not shown) that senses a user's gesture to the control unit (180), or transmit a signal from the control unit (180) to the sensor unit.

[0084] The control unit (180) may include at least one processor, and may control the overall operation of the display device (100) using the processor included therein. Here, the processor may be a general processor such as a central processing unit (CPU). Of course, the processor may be a dedicated device such as an ASIC or another hardware-based processor.

[0085] The control unit (180) can demultiplex a stream input through a tuner unit (111), a demodulator unit (112), an external device interface unit (171), or a network interface unit (172), or process demultiplexed signals to generate and output a signal for video or audio output.

[0086] The image signal processed by the control unit (180) can be input to the display module (150) and displayed as an image corresponding to the image signal. In addition, the image signal processed by the control unit (180) can also be input to an external output device through the external device interface unit (171).

[0087] The voice signal processed in the control unit (180) can be output as sound to the audio output unit (160). In addition, the voice signal processed in the control unit (180) can be input to an external output device through the external device interface unit (171). In addition, the control unit (180) can include a demultiplexing unit, an image processing unit, etc.

[0088] In addition, the control unit (180) can control the overall operation within the display device (100). For example, the control unit (180) can control the tuner unit (111) to select (tuning) a broadcast corresponding to a channel selected by the user or a previously stored channel.

[0089] In addition, the control unit (180) can control the display device (100) by a user command or an internal program input through the user input interface unit (173). Meanwhile, the control unit (180) can control the display module (150) to display an image. At this time, the image displayed on the display module (150) may be a still image or a moving image, and may be a 2D image or a 3D image.

[0090] Meanwhile, the control unit (180) can cause a predetermined 2D object to be displayed within an image displayed on the display module (150). For example, the object can be at least one of a connected web screen (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.

[0091] Meanwhile, the control unit (180) may modulate and / or demodulate a signal using the Amplitude Shift Keying (ASK) method. Here, the Amplitude Shift Keying (ASK) method may mean a method of modulating a signal by varying the amplitude of a carrier wave according to a data value, or of restoring an analog signal to a digital data value according to the amplitude of the carrier wave.

[0092] For example, the control unit (180) can modulate a video signal using the amplitude shift keying (ASK) method and transmit it through a wireless communication module.

[0093] For example, the control unit (180) can demodulate and process an image signal received through a wireless communication module using the amplitude shift keying (ASK) method.

[0094] Through this, the display device (100) can easily transmit and receive signals with other adjacently placed video display devices without using a unique identifier such as a MAC address (Media Access Control Address) or a complex communication protocol such as TCP / IP.

[0095] Meanwhile, the display device (100) may further include a camera (not shown). The camera can capture images of the user. The camera can be implemented with a single camera, but is not limited thereto, and may also be implemented with multiple cameras. Meanwhile, the camera can be embedded in the display device (100) above the display module (150) or can be separately positioned. Image information captured by the camera can be input to the control unit (180).

[0096] The control unit (180) can recognize the user's location based on the image captured by the camera. For example, the control unit (180) can determine the distance (z-axis coordinate) between the user and the display device (100). In addition, the control unit (180) can determine the x-axis coordinate and y-axis coordinate within the display module (150) corresponding to the user's location.

[0097] The control unit (180) can detect the user's gesture based on the image captured from the camera unit, or the signal detected from the sensor unit, or a combination thereof.

[0098] The power supply unit (190) can supply power to the entire display device (100). In particular, it can supply power to a control unit (180) that can be implemented in the form of a system on chip (SOC), a display module (150) for displaying images, and an audio output unit (160) for audio output.

[0099] Specifically, the power supply unit (190) may be equipped with a converter (not shown) that converts AC power into DC power and a Dc / Dc converter (not shown) that converts the level of DC power.

[0100] Meanwhile, the power supply unit (190) receives power from an external source and distributes power to each component. The power supply unit (190) may utilize a method of supplying AC power by directly connecting to an external power source, or may include a power supply unit (190) that can be recharged and used, including a battery.

[0101] The former requires a wired cable connection, making movement difficult or limited. The latter offers freedom of movement, but increases weight and volume by the battery. Charging requires a direct connection to a power cable or a charging cradle (not shown) for a set period of time.

[0102] The charging cradle can be connected to the display device through an externally exposed terminal, or the built-in battery can be charged when brought into proximity using a wireless method.

[0103] The remote control device (200) can transmit user input to the user input interface unit (173). To this end, the remote control device (200) can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. In addition, the remote control device (200) can receive images, voices, or data signals output from the user input interface unit (173) and display or output the same as voice on the remote control device (200).

[0104] Meanwhile, the above-described display device (100) may be a digital broadcast receiver capable of receiving fixed or mobile digital broadcasts.

[0105] Meanwhile, the block diagram of the display device (100) illustrated in FIG. 1 is only a block diagram for one embodiment of the present invention, and each component of the block diagram may be integrated, added, or omitted depending on the specifications of the display device (100) actually implemented.

[0106] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0107]

[0108] Figure 2 is a drawing illustrating a display device according to one embodiment of the present invention. Any description that overlaps with the above description will be omitted.

[0109] Referring to FIG. 2, the display device (100) is configured such that a display module (150) is housed inside a housing (210). At this time, the housing (210) includes an upper case (210a) and a lower case (210b), and the upper case (210a) and the lower case (210b) can be made in a structure that opens and closes.

[0110] In one embodiment, an audio output unit (160) may be included in the upper case (210a) of the display device (100), and a control unit (180), such as a main board, a power board, a power supply unit (190), a battery, an interface unit (170), a sensing unit (120), and an input unit (130, including a local key), may be housed in the lower case (210b). Here, the interface unit (170) may include a Wi-Fi module, a Bluetooth module, an NFC module, and the like for communication with an external device, and the sensing unit (120) may include a light sensor and an IR sensor.

[0111] In one embodiment, the display module (150) may include a DC-DC board, a sensor, and an LVDS (low voltage differential signaling) conversion board.

[0112] Additionally, in one embodiment, the display device (100) may further include four detachable legs (220a, 220b, 220c, 220d). Here, the four legs (220a, 220b, 220c, 220d) may be attached to the lower case (210b) to allow the display device (100) to be spaced from the floor.

[0113] The display device according to Fig. 2 has a movable feature.

[0114]

[0115] Figure 3 illustrates a processing procedure for audio data and video data in a multimedia device.

[0116] Multimedia devices typically have separate audio and video data processing units, and the time required for processing video data and audio data differs. If this time difference is significant, the shape or movement of a person's lips when speaking in a video may not align with the audio sound, resulting in an unpleasant experience for the viewer. Figure 3 illustrates the processing time difference between audio and video data.

[0117] This requires multimedia devices to perform separate tasks to synchronize the output timings of video and audio data processing. Synchronization between audio and video data is referred to as lip sync and is managed as a standard feature of multimedia devices.

[0118] In the specification below, the expression “lip sync is not consistent” means that there is no synchronization between audio data and video data.

[0119]

[0120] Figure 4 illustrates a procedure for compensating for the difference in processing time between audio data and video data in a multimedia device.

[0121] Let us explain a conventional method for matching lip sync. A multimedia device (or a processor of the multimedia device) obtains data processing (or delay) times of each of a video data processing unit and an audio data processing unit in advance according to the processing capability of the multimedia device or network status, and manages this in a table. Then, the multimedia device can synchronize the video data output time and the audio data output time by adding “delay logic” (A) to the side with a shorter data processing time among the video data processing unit and the audio data processing unit in the current situation.

[0122] Here, “delay logic” is implemented by storing decoded video data in internal storage (e.g., memory) for a desired delay time and then outputting it in the case of a video data processing unit, and by storing decoded audio data in internal storage (e.g., memory) for a desired delay time and then outputting it in the case of an audio data processing unit.

[0123] Delay logic is implemented by storing data in internal storage for the amount of time you want to delay, and this storage operation uses memory as internal storage, and the use of memory incurs a cost. To delay audio data by 160 msec, 4 * 48000 * 2 (assuming 2 channels) = 384 KBytes of memory are required; whereas, for video data, to delay just 160 msec, 3840 * 2160 (assuming 4K resolution) * 3 * 10 (assuming 60 Hz) = 249 MBytes of memory are required.

[0124] In other words, audio data has a much smaller amount of data compared to video data, so its delay logic does not take up a large portion of the price (cost) of the entire system. However, video data has a large amount of data, so its delay logic can take up a significant portion of the price of the entire system, which can be a problem. Because of this, most multimedia devices have limitations on the maximum time to delay the video, so they often only partially match the lip sync or do not have video delay logic at all, often giving up the lip sync matching operation itself.

[0125] As a solution for lip sync consistency, multimedia devices such as TVs have previously provided audio delay settings based on user input, providing users with a means to improve lip sync mismatch issues on their own. However, even in these cases, only a solution for adding delay to audio data, that is, delaying the output of audio data, was provided, and the reverse was not provided. Fortunately, user scenarios to date have often involved longer video data processing times than audio data processing times, so solutions that inevitably delay audio output have not been a major problem.

[0126] Recently, in order to provide a richer user experience, user scenarios that provide high-quality audio to users by linking TVs with external audio devices such as sound bars, BT (Bluetooth) sound bars, and BT headsets are emerging and expanding. In this case, while the video data processing time remains the same as the existing scenario, the audio data processing time becomes longer than the existing scenario due to the additional delay caused by the multimedia device transmitting audio data to the external audio device, and during this transmission process, various processing operations (decoding / re-encoding, etc.) of the audio data, physical amplification operations of the electrical signal, and wireless transmission and reception operations in the case of wireless communication methods.

[0127] Furthermore, with the recent adoption of high-end sound field effect technologies, user scenarios where audio data processing time exceeds video data processing latency are increasing. In these cases, lip syncing requires adding latency to the video data processing unit. However, conventional audio data delay logic (i.e., additional memory) inevitably incurs increased costs. Therefore, a solution that achieves the same or equivalent results is needed.

[0128]

[0129] FIG. 5 illustrates a method for lip sync matching according to one embodiment of the present invention.

[0130] The method illustrated in Fig. 5 is a method applicable under broadcasting or multicasting conditions, and changes the location where video data processing delay occurs during the video data processing process from after the completion of video decompression to before the completion of video decompression. This is because after the completion of video decompression, the amount of video data increases, requiring a lot of memory, whereas before the completion of video decompression, the amount of video data is small, allowing video delay to be implemented with relatively little memory.

[0131] To implement this, we propose a method that utilizes timestamp information, such as the Presentation Timestamp (PTS) of video frames. This involves adding a time value equal to the desired delay in video data to the original timestamp information, i.e., compensating the timestamp.

[0132] The output point of the video frames that have been decoded by the video decoder is the point in time when the PTS value matches the system time. In addition, the video sink has the characteristic of retrieving the decoded video frames from the video decoder and transmitting them to the video output terminal when the PTS value of the decoded video frames matches the system time.

[0133] Under these operating characteristics, the impact of increasing the PTS value can be directly observed as a delay in video output. GStreamer, widely used today for multimedia content playback, contains a timestamp information element within GstBuffer, which corresponds to the PTS described above.

[0134] Depending on the length of the delay time, the internal video decoder of the multimedia device exhibits the following operating characteristics due to the nature of its operating principle. When the delay time is added, the second buffer, which is the output buffer within the video decoder within the multimedia device, such as the DPB (Decoded Picture Buffer), becomes full first. If the delay time is further increased, the first buffer, which is the input buffer within the video decoder, such as the CPB (Coded Picture Buffer), also becomes full. Since the first buffer stores compressed video data, it can enjoy the advantage of using much less memory to delay the video data. Nevertheless, in order to smoothly perform the proposed method, it would be desirable to increase the capacity of the first or second buffer within the video decoder compared to the existing one.

[0135]

[0136] FIG. 6 illustrates a flowchart of a method for lip sync matching according to one embodiment of the present invention. The method illustrated in FIG. 6 may be performed by a multimedia device (100), and more specifically, by a controller or processor (180) of the multimedia device (100). Hereinafter, the method will be briefly described as being performed by the multimedia device (100).

[0137] A multimedia device (100) can determine a time correction value for video data (S610).

[0138] The multimedia device (100) can delay the output timing of video data to the display (150) by applying a determined time correction value (S620).

[0139] The multimedia device (100) can output video data according to delayed output timing (S630).

[0140]

[0141] Furthermore, further expansions can be considered for the embodiments described with reference to FIGS. 5 and 6. While FIG. 5 was described as being applicable to content transmission and reception using broadcasting or multicasting methods, the embodiments can also be utilized for transmission and reception using the Transmission Control Protocol (TCP) network transmission method. This will be described in more detail below.

[0142]

[0143] A content provider, acting as a network transmitter, can transmit video and audio via separate URLs. A multimedia device (100), acting as a network receiver, can receive audio data using a first URL for audio data and video data using a second URL for video data. In other words, the multimedia device (100) can independently receive video data and audio data using separate URLs.

[0144] Then, the multimedia device (100) implements the delay operation by delaying the output time of video frames that have been decoded (decoded) in the video decoder or the exhaustion time of video frames that have been decoded in the video sync terminal compared to the original time, rather than the conventional memory usage method described above.

[0145] To this end, we propose a method that utilizes timestamp information, such as the Presentation Timestamp (PTS) of video frames, similar to the previously described method. This involves adding a time value equal to the desired delay in video data to the original value of the timestamp information, i.e., compensating the timestamp.

[0146] The output point of the video frames that have been decoded by the video decoder is the point in time when the PTS value matches the system time. In addition, the video sink has an operating characteristic of retrieving the decoded video frames from the video decoder and transmitting them to the video output terminal when the PTS value of the decoded video frames matches the system time. Under these operating characteristics, the effect of an increase in the PTS value can be immediately confirmed as a delay in the video output.

[0147] Depending on the length of the delay time, the video decoder of the multimedia device (100) exhibits the following operating characteristics due to its operating principle. When the delay time is added, the second buffer, which is the output buffer within the video decoder within the multimedia device, such as the DPB (Decoded Picture Buffer), becomes full first. If the delay time is increased further, the first buffer, which is the input buffer within the video decoder, such as the CPB (Coded Picture Buffer), also becomes full. If the delay time is increased further, the network receiver buffer also becomes full. When the data in the network receiver buffer becomes full, the network transmitter stops transmitting video data due to the TCP flow control operation. Accordingly, the multimedia device (100) stops receiving video data. Thereafter, when the video decoder outputs the decoded video frames and the internal buffer is not full, the network transmitter transmits video data again. Since the TCP flow control operates independently for each URL (or port), this delay operation of the video data does not affect the transmission of audio data. That is, the multimedia device (100) can continue to receive audio data even if reception of video data is stopped.

[0148]

[0149] FIG. 7 illustrates a method for lip sync matching according to another embodiment of the present invention.

[0150] Figure 7 is a network transmission and reception method, and is a method that can be used in a situation where IP (Internet Protocol) is adopted. Since IP does not support the “data flow control” function supported by TCP, the method described above cannot be used.

[0151] Even in this manner, the content provider, which is the network transmitter, can transmit video and audio via separate URLs. However, the multimedia device (100), which is the network receiver, can be configured to first request audio data transmission and then wait a sufficient amount of time before requesting video data transmission when requesting content transmission from the content provider.

[0152] When requesting data transmission from the content provider's server, the multimedia device (100) may recognize that the current state of the multimedia device (100) is such that the processing time for audio data is longer than the processing time for video data. Accordingly, the multimedia device (100) may first request the content provider to transmit audio data, and then request the transmission of video data after a sufficiently long period of time has elapsed.

[0153] Here, a sufficiently long time refers to the time for transmission delay of video data. Due to this, even if the processing time of audio data is longer than the processing time of video data, the output order of video data and audio data may be reversed due to the time difference between the video data and audio data input or arriving at the multimedia device (100).

[0154] This proposal aims to resolve the lip sync matching problem caused by the delay in processing audio data compared to processing video data, and provides a means for delaying the processing of video data or its output to a display (150). As mentioned above, the time for the transmission delay of video data must be sufficiently long, which is to ensure that the video data arrives or is input to the audio data processing unit or video data processing unit of the multimedia device (100) later than the audio data.

[0155] However, it should be noted that even if a multimedia device (100), which is a network receiver, requests transmission of audio data from a content provider, which is a network transmitter, and then requests transmission of video data after a time T, it is not guaranteed that the multimedia device (100) receives the video data exactly after a time T after receiving the audio data. This is due to the operating characteristics of the network, and there is no guarantee that the time difference between the reception of two responses to two requests with an interval of time T through the network is maintained exactly as much as the time T.

[0156] As illustrated in FIG. 7, in a situation where the input or arrival of (compressed) video data to a video decoder of a multimedia device (100) is delayed compared to the input or arrival of (compressed) audio data to an audio decoder of the multimedia device (100), the multimedia device (100) can obtain a time difference T' between the video data and the audio data. This can be achieved by detecting a timestamp value of the video data at the moment when the video data is first input or arrived, and simultaneously detecting a timestamp value of the audio data being input or arrived, and calculating T' through the difference between the two timestamp values.

[0157] That is, a situation occurs where audio data is ahead of video data by T'. This is due to the time difference (T) between the transmission request of audio data and the transmission request of video data. Of course, the transmission and reception delay of the network may also have an effect. The multimedia device (100) may additionally delay the audio data using the existing memory method (i.e., the “delay logic” described above) to eliminate the time difference T' or adjust for the lip sync gap (time difference) to match the lip sync. Since the compensation (time T) for the expected lip sync gap has been performed in advance at the data request stage, the “delay logic” is expected not to incur a large cost for the additional hardware.

[0158] When applying an additional delay to audio data using delay logic, the additional delay time can be obtained based on the information managed in the table described above. The table may include information on the processing time delay between audio data and video data according to the format of the audio data or video data, the performance (capability) of each data processing unit, etc., assuming that the audio data and the video data are input or arrived at the audio data processing unit and the video data processing unit at the same time. Accordingly, the multimedia device (100) may calculate the expected time delay (T) when there is a difference in arrival or input time between the audio data and the video data by the time T'. est ) can be produced. T est If is a positive value (i.e., processing of audio data is fast), the multimedia device (100) can further delay the processing time of audio data using delay logic. Conversely, the difference (T) in the request times of audio data and video data is also T est It should be determined so that it can have a positive value.

[0159] Additionally, the following two options may be considered:

[0160] A1) As an additional explanation, when determining the difference (T) in request times between audio data and video data, the multimedia device (100) may consider the difference in processing times between audio data and video data processed by the audio data processing unit and the video data processing unit. In this case, the difference (T) in request times may be a value that prevents lip sync from mismatching, i.e., ensures that lip sync matches. However, even in this case, since T' may be greater than T depending on the network status, the multimedia device (100) may remove the time difference corresponding to T'-T using delay logic.

[0161] A2) Similar to the A1 method, however, the multimedia device (100) can additionally consider delays according to network conditions when determining the difference (T) between the request times of audio data and video data. That is, the A2 method considers not only the audio data processing time of the audio data processing unit and the video data processing time of the video data processing unit, but also the network conditions.

[0162] To this end, the difference in request times between audio data and video data can be determined to have a smaller value than T determined in the A1 method. That is, the difference in request times between audio data and video data can be determined so that the time difference between audio data and video data input or arriving at the multimedia device (100) or its audio / video data processing unit becomes T. Ideally, this method does not require the use of delay logic, but since the network condition or the data processing performance of the multimedia device (100) may not be ideal, the use of delay logic is not completely ruled out.

[0163]

[0164] Fig. 8 illustrates a flowchart of a method for lip sync matching according to another embodiment of the present invention. Fig. 8 is a flowchart for the method described in Fig. 7. To avoid duplication, the description given with reference to Fig. 7 will be omitted with respect to Fig. 8.

[0165] The method illustrated in FIG. 8 can be performed by a multimedia device (100), and more specifically, by a controller or processor (180) of the multimedia device (100). Hereinafter, the method will be briefly described as being performed by the multimedia device (100).

[0166] A multimedia device (100) requests audio data from a content provider (S810).

[0167] The multimedia device (100) requests video data from the content provider after a time (T) for transmission delay of video data (S820).

[0168] The multimedia device (100) initiates processing of audio data and video data as video data is received in response to a request for video data (S830).

[0169] The multimedia device (100) can determine whether a time delay is required for audio data (S840). More specifically, the multimedia device (100) can calculate the time difference (T') between the reception of audio data and video data when video data is received in response to a request for video data. This is to determine whether there is a difference in the output times of the processing results of audio data and video data, i.e., whether lip sync mismatch occurs.

[0170] When it is determined that a time delay is required for audio data, the multimedia device (100) applies an output time delay to the audio data (S850). The amount of the output time delay may correspond to the difference between the previously calculated time difference (T') and the time for transmission delay (T), but may not be the exact value (T'-T).

[0171] Meanwhile, the time (T) for transmission delay may be determined by considering the state of the network for transmission of audio data and video data (i.e., the communication state between the multimedia device (100) and the content provider) or the processing capability or performance of audio data and video data of the multimedia device (100).

[0172] As it is determined that no time delay is required for audio data, the multimedia device (100) can output processed audio data and video data.

[0173]

[0174] FIG. 9 illustrates a flowchart of a method for lip sync matching according to another embodiment of the present invention.

[0175] The method illustrated in FIG. 9 can be performed by a multimedia device (100), and more specifically, by a controller or processor (180) of the multimedia device (100). Hereinafter, the method will be briefly described as being performed by the multimedia device (100).

[0176] The method described in Figure 9 utilizes general content data (i.e., original content) and content data with a time delay applied by the content provider. Here, the time delay refers to the time delay between the audio and video data of the content, and refers to the time delay applied to the video data relative to the audio data. In this case, the general content data and the delayed content data are indicated through different URLs, and each URL can be used to request or receive the content data.

[0177] A multimedia device (100) can request general or delayed content data from a content provider (S910).

[0178] If the multimedia device (100) determines that the current situation of the multimedia device (100) is one in which the audio data processing time is longer than the video data processing time, the multimedia device (100) requests the content provider to transmit content in which the video data is delayed compared to the audio data. As a result, even though the audio data processing time is longer than the video data processing time, the audio decoder of the multimedia device (100) can output the audio data first, and the video decoder can output the video data later, because the video data input or arriving at the multimedia device (100) is delayed compared to the audio data.

[0179] The multimedia device (100) can receive general or delayed content data in response to a content request (S920).

[0180] Furthermore, even in this situation, the multimedia device (100) may further delay the output audio data of the audio decoder using the existing memory usage method. To this end, the multimedia device (100) may determine whether a time delay is required for the audio data (S930).

[0181] As it is determined that a time delay for audio data is required, the multimedia device (100) applies an output time delay to the audio data (S940).

[0182] As it is determined that no time delay is required for audio data, the multimedia device (100) can output the received and processed audio data and video data.

[0183] If the multimedia device (100) determines that the current situation of the multimedia device (100) is one in which the audio data processing time is longer than the video data processing time, the multimedia device (100) can request the content provider to transmit general content.

[0184] The method according to FIG. 9 has the advantage of reducing the processing load on the multimedia device (100) compared to the method described in FIG. 5 or FIG. 6 by selectively requesting general content or delayed content depending on the processing status of audio data and video data of the multimedia device (100).

[0185]

[0186] In addition, as another aspect of the present invention, the operation of the proposal or invention described above may be implemented, performed or executed by a “computer” (a comprehensive concept including a system on chip (SoC) or a (micro) processor, etc.), or may be provided as a code or a computer-readable storage medium storing or including the code or a computer program product, and the scope of the present invention may be extended to the code or the computer-readable storage medium storing or including the code or the computer program product.

[0187]

[0188] The detailed description of the preferred embodiments of the present invention disclosed above has been provided to enable those skilled in the art to implement and practice the present invention. While the above description has been made with reference to preferred embodiments of the present invention, those skilled in the art will appreciate that various modifications and variations of the present invention, as defined by the following claims, are possible. Accordingly, the present invention is not intended to be limited to the embodiments disclosed herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. As a multimedia device for synchronizing audio data and video data, A processor for processing received audio data and received video data; A display for outputting the processed video data; and Includes a speaker for outputting the processed audio data, The above processor: Determining a time correction value for the above video data, and applying the time correction value to delay the output timing of the video data to the display. , multimedia devices.

2. In paragraph 1, the processor: Receive said video data independently from said audio data by using a URL for said video data from a content provider, Adding the time correction value to the timestamp value of the received video data , multimedia devices.

3. In the second paragraph, the received video data or the processed video data is stored in the internal buffer according to the addition of the time correction value, and when the internal buffer becomes full, the reception of the video data is stopped. , multimedia devices.

4. In the third paragraph, the processor: Even if the reception of the above video data is stopped, the above audio data is received. , multimedia devices.

5. In paragraph 1, the processor: Applying the time correction value to the received video data before decompression processing of the received video data. , multimedia devices.

6. In paragraph 1, the processor: Outputting the video data to the display as the timestamp value of the above video data matches the system time. , multimedia devices.

7. As a multimedia device for synchronizing audio data and video data, A processor for processing received audio data and received video data; A display for outputting the processed video data; and Includes a speaker for outputting the processed audio data, The above processor: Request audio data from a content provider, After the time for delaying the transmission of video data, requesting the video data from the content provider, In response to the above request, upon receipt of the above video data, the time difference between the reception times of the above audio data and the above video data is calculated, Apply an output time delay to the audio data equal to the difference between the calculated time difference and the time for the transmission delay. , multimedia devices.

8. In the 7th paragraph, the time for the transmission delay is determined by considering the status of the network for transmission of the audio data and the video data or the processing capability of the audio data and the video data of the multimedia device. , multimedia devices.

9. A method for synchronizing audio data and video data, A step of determining a time correction value for the above video data; A step of delaying the output timing of the video data to the display by applying the above time correction value; and A step of outputting the video data and the audio data according to the output timing delay. , method.

10. In paragraph 9, A step of receiving the video data independently from the audio data using a URL for the video data from a content provider; and A step of adding the time correction value to the time stamp value of the received video data is included. , method.

11. In the 10th paragraph, the received video data or the processed video data is stored in the internal buffer according to the addition of the time correction value, and when the internal buffer becomes full, the reception of the video data is stopped. , method.

12. In paragraph 11, A step of receiving the audio data even when the reception of the video data is stopped. , method.

13. In paragraph 9, A step of applying the time correction value to the received video data prior to decompression processing of the received video data. , method.

14. In paragraph 9, A step of outputting the video data to the display when the timestamp value of the video data matches the system time. , method.

15. A method for synchronizing audio data and video data, Step of requesting audio data from a content provider; A step of requesting the video data to the content provider after a time for delaying transmission of the video data; In response to the request, a step of calculating a time difference between the reception time of the audio data and the video data, upon reception of the video data; and A step of applying an output time delay to the audio data by the difference between the calculated time difference and the time for the transmission delay. , method.

Citation Information

Patent Citations

  • Streaming playback apparatus, streaming distribution playback system, streaming playback method and streaming playback program

    JP2010021867A

  • Distribution of IP broadcast streaming services using a file distribution method.

    JP2014517558A

  • Audio-video synchronization system and monitoring equipment

    JP4182437B2

  • Apparatus and method for signal control of hometheater system

    KR101703862B1

  • Electronic device for performing synchronization of video data and audio data, and control method therefor

    US20230232063A1