CGR Content Generation from Lyric and Audio Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audiovisual experiences, such as music videos and algorithmic audio visualizations, are not truly immersive and are not tailored to a user's environment, particularly due to the inherent ambiguity of language in song lyrics.

Innovation Solution

The generation of computer-generated reality (CGR) content based on both natural language analysis and semantic analysis of audio files, which considers the audio data's key, tempo, rhythm, mood, and vocal timbre to determine the meaning of lyrics and create immersive CGR environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If natural language analysis alone is used to generate CGR content from lyrics, then the content generation process is simple, but the resulting content lacks accuracy and relevance due to language ambiguity

Engineering Contradiction:
Improvelyric interpretation accuracyVSAvoidcontent generation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the lyric analysis process into two distinct stages: natural language analysis to generate candidate meanings, and semantic analysis to select the most appropriate meaning based on audio characteristics. This segmentation allows the system to handle complexity in a structured manner while improving accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces audio characteristics (key, tempo, rhythm, mood, vocal timbre) as an intermediary element that mediates between the ambiguous lyrics and the CGR content generation. This intermediary provides additional context to disambiguate lyric meanings and guide content creation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If generic audio visualizations are used, then the implementation is straightforward, but the user experience lacks immersion and personalization

Engineering Contradiction:
Improveexperience personalizationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by tailoring the CGR content to match specific segments of the audio file. Each portion of the audio (with its unique key, tempo, rhythm, mood, and vocal timbre) generates customized visual content, rather than applying a uniform visualization throughout.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by making the CGR content adaptive and responsive to changing audio characteristics. As the audio progresses and its characteristics change, the generated visual content dynamically adjusts to maintain relevance and immersion.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If multiple analysis methods are combined to improve CGR content quality, then the content relevance increases, but the processing time and computational resources increase

Engineering Contradiction:
Improvecontent generation qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by extracting and analyzing audio characteristics (key, tempo, rhythm, mood, vocal timbre) before generating the CGR content. This preprocessing step organizes the audio data in a way that facilitates more efficient and accurate content generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback by using the audio characteristics as a feedback mechanism to guide and constrain the CGR content generation process. This feedback loop ensures that the generated content remains aligned with the audio's emotional and thematic content, reducing the need for iterative adjustments.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12293759B2Method and device for presenting a CGR environment based on audio data and lyric data
Publication Date: 2025.05.06 APPLE INC
  • US12293759B2 patent drawing
  • US12293759B2 patent drawing
  • US12293759B2 patent drawing

AI summary

In one implementation, a method of generating CGR content to accompany an audio file including audio data and lyric data based on semantic analysis of the audio data and the lyric data is performed by a device including a processor, non-transitory memory, a speaker, and a display. The method includes obtaining an audio file including audio data and lyric data associated with the audio data. The method includes performing natural language analysis of at least a portion of the lyric data to determine a plurality of candidate meanings of the portion of the lyric data. The method includes performing semantic analysis of the portion of the lyric data to determine a meaning of the portion of the lyric data by selecting, based on a corresponding portion of the audio data, one of the plurality of candidate meanings as the meaning of the portion of the lyric data. The method includes generating CGR content associated with the portion of the lyric data based on the meaning of the portion of the lyric data.