3D Lyric Video Display With Depth-Based Background Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing music video lyrics are monotonous and separated from the background, leading to a poor user experience.

Innovation Solution

A method and device that integrate lyrics into the background by determining target objects and adjusting display effects based on depth information, embedding lyrics into the real space of the image data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If lyrics are simply superimposed with the video using basic effects, then the implementation is simple, but the display form of the lyrics is monotonous and the lyrics are completely separated from the background

Engineering Contradiction:
Improveimplementation simplicityVSAvoidlyrics-background integration
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces depth information (z-axis dimension) to position lyrics in three-dimensional space relative to the background. Lyrics are no longer flatly superimposed but are positioned at different depths, allowing them to appear in front of or behind background elements, creating a spatial relationship that integrates lyrics with the video background while maintaining implementation feasibility through depth map processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent applies different display effects to lyrics based on their local spatial relationship with background elements. By analyzing depth information at specific positions, the system adjusts lyric attributes (such as opacity, size, or position) locally to match the background context, creating natural integration without requiring complete redesign of the entire lyric display system.

Inventive Principle:
Principle #3Local quality

2Device complexity

If lyrics are completely separated from the video background with no correlation, then the processing is simple, but the user experience is poor

Engineering Contradiction:
Improveprocessing complexityVSAvoiduser experience quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces depth information as an intermediary element that mediates between the lyrics and the video background. This depth map serves as a bridge that provides spatial relationship data without requiring complex real-time interaction processing, enabling natural lyric-background integration while keeping the processing architecture relatively simple and reliable.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If basic mechanical effects are applied to lyrics, then the implementation is straightforward, but the display form is monotonous resulting in poor user experience

Engineering Contradiction:
Improveimplementation straightforwardnessVSAvoiddisplay form diversity
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent makes the lyric display dynamic by continuously adjusting lyric attributes based on real-time depth information and background analysis. Instead of static superimposition, lyrics adapt their position, size, and transparency dynamically as they move through the video, creating diverse and engaging display forms while building upon straightforward basic effects rather than replacing them entirely.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12549801B2Lyric video display method and device, electronic apparatus and computer-readable medium
Publication Date: 2026.02.10 BEIJING ZITIAO NETWORK TECH CO LTD
  • US12549801B2 patent drawing
  • US12549801B2 patent drawing
  • US12549801B2 patent drawing

AI summary

The present disclosure provides a lyric video display method and device, an electronic apparatus, and a computer-readable medium. The method includes: playing, based on a lyric video display operation of a user, multimedia data and data about music to be displayed, the multimedia data including image data, and the music data including audio data and lyrics; determining a target time point, determining a target object in the image data corresponding to the target time point, and determining target lyrics in the lyrics corresponding to the target time point; and displaying the target lyrics within a preset range of a position of the target object in the target image, and adjusting display effects of the target lyrics based on depth information of the target object, while playing a part of the audio data corresponding to the target lyrics.