Customized Lyric Captions Using User-Aware Generative AI

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Media platforms face challenges in providing dynamic and personalized lyric captions that account for individual user preferences, incorporating non-lyric context, and offering interactive content, often requiring significant manual effort and resources.

Innovation Solution

Utilizing machine learning models, particularly large language models (LLMs) and large multi-modal models (LMMs), to generate customized lyric captions based on user data, identify and describe non-lyric context, and provide interactive content, reducing the need for manual curation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual curation is used to create personalized lyric captions, then customization quality can be maintained, but significant time and resources are required

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidmanual curation time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system enables automated self-service generation of personalized lyric captions through machine learning models. The model automatically processes audio data, identifies non-lyric context, and generates customized captions based on user profiles without requiring manual human intervention for each caption set.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual curation process with an automated machine learning system. The ML model substitutes human operators by performing audio analysis, context identification, and caption generation tasks that were previously done manually, significantly reducing time and resource requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If machine learning models are used to generate customized lyric captions, then manual labor is reduced, but model complexity and computational resources increase

Engineering Contradiction:
Improvecaption generation efficiencyVSAvoidmodel complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex caption generation task into distinct functional components: audio data processing, non-lyric context identification, user profile analysis, and caption synthesis. Each component can be independently optimized and managed, reducing overall system complexity while maintaining high productivity.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If lyric captions are customized for individual users, then user experience is enhanced, but system complexity increases

Engineering Contradiction:
Improveuser experienceVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The machine learning model is designed with universal functionality to handle multiple tasks: processing audio data, analyzing diverse user profiles with different preferences and language proficiencies, identifying various types of non-lyric context, and generating customized captions. This multi-functionality allows the system to serve multiple user needs through a single unified platform.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of information

If non-lyric context is incorporated into lyric captions, then information completeness is improved, but processing complexity increases

Engineering Contradiction:
Improvecontext information completenessVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The machine learning model acts as an intermediary that automatically identifies and extracts non-lyric context from audio data, such as musical instruments, background sounds, and emotional tone. This intermediary process simplifies the incorporation of contextual information by handling the complex analysis and integration tasks automatically, reducing manual processing complexity while maintaining information completeness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260044669A1Generating customized lyric captions using machine learning models
Publication Date: 2026.02.12 GOOGLE LLC
  • US20260044669A1 patent drawing
  • US20260044669A1 patent drawing
  • US20260044669A1 patent drawing

AI summary

A media stream comprising audio data and first lyric data associated with the audio data is received by a processing device. A set of user data associated with a user of a client device is identified. The first lyric data and the set of user data are provided as input to a generative machine learning model. An output of the generative machine learning model is obtained. The output comprises second lyric data. The second lyric data is a version of the first lyric data that is customized for the user. The second lyric data and the media stream are caused to be presented in a graphical user interface on the client device.