Neural Network Lyrics Alignment via Probability Matrix

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing media content systems face challenges in providing time-aligned lyrics for audio data, as manually transcribed lyrics are often unavailable or costly, and lack temporal alignment with the music, limiting user interaction and content enhancement.

Innovation Solution

A neural network is used to generate a probability matrix from audio data, identifying textual units and their timing information, allowing for the alignment of lyrics with the audio, enabling features like karaoke, search functionality, and explicit content removal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manually transcribed lyrics are used, then lyrics content is available, but time alignment information is lost and cost increases

Engineering Contradiction:
Improvetime alignment informationVSAvoidtranscription cost
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The system uses the audio data itself to generate both the lyrics and timing information through neural network processing. The audio waveform is directly transformed into a probability matrix that identifies textual units and their temporal positions, making the system self-sufficient without requiring external transcription services.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical transcription process with an automated neural network system. The neural network directly processes audio waveforms to generate time-aligned lyrics, substituting human labor with an automated computational system that provides both text and timing information simultaneously.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If manually transcribed lyrics are obtained, then lyrics content is available, but availability becomes limited and cost increases

Engineering Contradiction:
Improvelyrics contentVSAvoidlyrics availability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The system extracts lyrics directly from the audio data using neural network processing, making it self-sufficient and eliminating dependence on external transcription sources. This approach enables universal application to any audio content regardless of whether manual transcriptions exist.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The neural network system provides a universal solution that can generate time-aligned lyrics for any audio content, making the system adaptable to diverse media types and sources without requiring different transcription approaches for different content.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If neural network processing is applied, then time-aligned lyrics are generated, but processing complexity increases

Engineering Contradiction:
Improvetime alignment accuracyVSAvoidneural network processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential features needed for time-aligned lyric generation from the audio data. The neural network focuses on transforming the audio waveform directly into a probability matrix for textual units, extracting and processing only the relevant temporal and spectral information necessary for accurate lyric alignment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms the audio waveform into a probability matrix through neural network processing, changing the parameter representation from raw audio signals to probabilistic textual unit predictions with temporal information. This parameter transformation enables accurate time alignment while managing processing complexity through efficient network architecture.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11475887B2Systems and methods for aligning lyrics using a neural network
Publication Date: 2022.10.18 SPOTIFY
  • US11475887B2 patent drawing
  • US11475887B2 patent drawing
  • US11475887B2 patent drawing

AI summary

An electronic device receives audio data for a media item. The electronic device generates, from the audio data, a plurality of samples, each sample having a predefined maximum length. The electronic device, using a neural network trained to predict textal unit probabilities, generates a probability matrix of textual units for a first portion of a first sample of the plurality of samples. The probability matrix includes information about textual units, timing information, and respective probabilities of respective textual units at respective times. The electronic device identifies, for the first portion of the first sample, a first sequence of textual units based on the generated probability matrix.