Crowd-Sourced Vocal Pitch Track Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating high-quality vocal pitch tracks for karaoke-style applications is challenging due to the constant change in popular music releases, with existing technologies lacking efficient automated or semi-automated techniques to produce pitch tracks for mass-market use.

Innovation Solution

Employing digital signal processing and machine learning techniques to computationally generate vocal pitch tracks from crowd-sourced audio performances captured against a common temporal baseline, using a geographically distributed network of devices and a service platform to estimate, aggregate, and align pitch sequences, with confidence ratings and predictive models for accurate pitch tracking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated or semi-automated techniques are used to generate pitch tracks, then productivity is improved, but measurement precision may deteriorate

Engineering Contradiction:
Improvepitch track generation efficiencyVSAvoidpitch tracking accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent combines multiple pitch estimation results from different crowd-sourced vocal performances into a single aggregated pitch track. By merging numerous individual pitch estimates through computational aggregation, the system achieves both high productivity (automated generation from many sources) and high measurement precision (accurate pitch tracking through statistical consolidation of multiple measurements)

Inventive Principle:
Principle #5Merging (Combining)

2Quantity of substance

If crowd-sourced vocal performances are aggregated, then quantity of data is improved, but device complexity increases

Engineering Contradiction:
Improvenumber of vocal performancesVSAvoidprocessing system complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system enables crowd-sourced vocal performances to self-organize and self-correct through automated aggregation algorithms. Each user contribution automatically becomes part of the collective dataset, and the computational system automatically processes and aggregates these contributions without requiring manual intervention, thus managing large quantities of data while keeping operational complexity manageable

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If pitch tracks are generated for ever-changing popular music, then adaptability is improved, but loss of time increases

Engineering Contradiction:
Improvecontent library update capabilityVSAvoidpitch track production time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously maintaining a library of pitch tracks from crowd-sourced performances. When new popular music emerges, the system can quickly generate pitch tracks by aggregating new user performances, rather than waiting for manual processing. This preliminary aggregation capability reduces the time loss when adapting to changing music trends

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11900904B2Crowd-sourced technique for pitch track generation
Publication Date: 2024.02.13 SMULE INC
  • US11900904B2 patent drawing
  • US11900904B2 patent drawing
  • US11900904B2 patent drawing

AI summary

Digital signal processing and machine learning techniques can be employed in a vocal capture and performance social network to computationally generate vocal pitch tracks from a collection of vocal performances captured against a common temporal baseline such as a backing track or an original performance by a popularizing artist. In this way, crowd-sourced pitch tracks may be generated and distributed for use in subsequent karaoke-style vocal audio captures or other applications. Large numbers of performances of a song can be used to generate a pitch track. Computationally determined pitch trackings from individual audio signal encodings of the crowd-sourced vocal performance set are aggregated and processed as an observation sequence of a trained Hidden Markov Model (HMM) or other statistical model to produce an output pitch track.