Demographic Prediction Using Sparse Vector Media Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for predicting demographics and taste preferences from music or media data rely on limited source data, leading to less accurate predictions.

Innovation Solution

A system that constructs a sparse vector from a set of terms and corresponding weights associated with an artist or song title, which is then input into a training model to predict demographic data, using a larger data set to enhance predictive accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a relatively small set of source data (e.g., web browsing history) is used for prediction, then the system complexity is reduced, but the prediction accuracy deteriorates

Engineering Contradiction:
Improvedata set sizeVSAvoidprediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple data sources including web browsing history, music listening data, demographic data, and media consumption data into a unified comprehensive data set. This merging of diverse data types allows the system to maintain low complexity while achieving high prediction accuracy through the synergistic use of multiple data dimensions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a universal prediction model that can handle multiple types of source data (browsing history, music preferences, demographic information, media consumption) through a single unified approach. This multi-functional capability allows the system to process diverse data types without increasing operational complexity, while significantly improving prediction accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If a larger data set with multiple data sources is used for prediction, then the prediction accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata set size
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive data set into distinct modular components: browsing history data, music listening data, demographic data, and media consumption data. Each segment is processed independently through specialized modules, which then feed into the unified prediction model. This segmentation reduces system complexity by breaking down the handling of large data sets into manageable, independent tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary processing modules that act as mediators between the raw data sources and the final prediction model. These intermediary modules pre-process and structure the diverse data types into standardized formats, reducing the complexity burden on the main prediction system while preserving the benefits of using comprehensive data sets.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9547679B2Demographic and media preference prediction using media content data analysis
Publication Date: 2017.01.17 SPOTIFY
  • US9547679B2 patent drawing
  • US9547679B2 patent drawing
  • US9547679B2 patent drawing

AI summary

Methods, systems and computer program products are provided for predicting data. A name or title is obtained from a taste profile. There is an index into a data set based on the name or title, and a set of terms and corresponding term weights associated with the name or title are retrieved. A sparse vector is constructed based on the set of terms and term weights. The sparse vector is input to a training model including target data. The target data includes a subset of test data which has a correspondence to a predetermined target metric of data. A respective binary value and confidence level is output for each term, corresponding to an association between the term and the target metric.