Neural Network Lyrics Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing media content systems face challenges in providing time-aligned lyrics for audio data, as manually transcribed lyrics are often unavailable or costly, and lack synchronization with the audio, limiting user interaction and functionality such as karaoke and content editing.

Innovation Solution

A neural network is used to generate a probability matrix from audio data, predicting character probabilities and timing information, allowing for the identification and alignment of lyrics within the audio, enabling efficient extraction and synchronization of lyrics with the audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manually transcribed lyrics are obtained, then lyrical content is available, but time alignment information is lacking and costs increase

Engineering Contradiction:
Improvetime alignment informationVSAvoidcost of obtaining lyrics
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The neural network model automatically extracts and aligns lyrics with audio data without requiring manual transcription services. The system self-services by processing audio input directly to generate time-aligned lyrics output, eliminating dependency on external manual transcription resources while capturing both lyrical content and temporal information simultaneously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual transcription with an automated neural network-based acoustic model. This substitution transforms the lyric extraction process from a labor-intensive manual operation to an automated computational process that simultaneously retrieves lyrics and their temporal alignment information

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manually transcribed lyrics are used, then lyrical content is provided, but user interaction functionality is limited

Engineering Contradiction:
Improveuser interaction functionalityVSAvoidtiming information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent adds the temporal dimension to lyric data by generating time alignment information that maps each lyric segment to its corresponding position in the audio. This transforms static lyric text into dynamically timed lyric events, enabling dimensionally richer user interactions such as synchronized display, karaoke timing, and navigation to specific song portions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If a neural network is used to generate probability matrices for all samples, then accurate lyric identification is achieved, but processing time increases

Engineering Contradiction:
Improvelyric identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the audio data into multiple samples or segments that can be processed independently or in parallel. By segmenting the input audio into manageable portions, the system maintains high lyric identification accuracy through comprehensive neural network analysis while reducing overall processing time through parallel computation and incremental processing of individual segments

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11308943B2Systems and methods for aligning lyrics using a neural network
Publication Date: 2022.04.19 SPOTIFY
  • US11308943B2 patent drawing
  • US11308943B2 patent drawing
  • US11308943B2 patent drawing

AI summary

An electronic device receives audio data for a media item. The electronic device generates, from the audio data, a plurality of samples, each sample having a predefined maximum length. The electronic device, using a neural network trained to predict character probabilities, generates a probability matrix of characters for a first portion of a first sample of the plurality of samples. The probability matrix includes character information, timing information, and respective probabilities of respective characters at respective times. The electronic device identifies, for the first portion of the first sample, a first sequence of characters based on the generated probability matrix.