Expectation-Maximization Algorithm for Biased Population Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing population estimation methods, such as the capture-recapture method, face challenges in accurately estimating population sizes due to incomplete data and bias in single-source data, which can lead to inaccurate population composition analysis.

Innovation Solution

The use of an expectation-maximization algorithm to adjust user data from biased sources, generating weighted population composition data that accounts for bias, resulting in more accurate estimates of user populations accessing a resource.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If capture-recapture method or single-source panel approach is used to estimate population size, then population estimation can be performed with limited data, but the accuracy of population composition estimates deteriorates due to incomplete data and bias

Engineering Contradiction:
Improveaccuracy of population composition estimatesVSAvoidincomplete data from single source
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces an expectation-maximization (EM) algorithm as an intermediary computational framework that processes biased single-source data. The EM algorithm iteratively refines population composition estimates by alternating between estimating missing data patterns (E-step) and optimizing parameters (M-step), effectively mediating between incomplete observed data and accurate population composition estimates without requiring multiple independent data sources

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by iteratively adjusting weights and probabilities in the EM algorithm. The algorithm changes parameters (population segment proportions, detection probabilities) across iterations, converging to optimal estimates that account for bias in the single-source data. This iterative parameter refinement enables accurate composition estimates despite incomplete information

Inventive Principle:
Principle #35Parameter changes

2Productivity

If single-source biased data is used for population estimation, then data collection is simplified and faster, but the accuracy of user demographic representation deteriorates

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidaccuracy of user demographic representation
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The EM algorithm serves as a computational intermediary that processes biased single-source data efficiently. It takes the simplified single-source data as input and produces corrected population composition estimates as output, maintaining the productivity advantage of single-source collection while eliminating the accuracy penalty through iterative statistical correction

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback through the iterative nature of the EM algorithm. Each iteration uses the results from the previous iteration to refine estimates further, with the algorithm continuously feeding corrected estimates back into the calculation process until convergence. This feedback mechanism progressively improves demographic representation accuracy while working with fixed single-source data

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8954577B2Estimating a composition of a population
Publication Date: 2015.02.10 GOOGLE LLC
  • US8954577B2 patent drawing
  • US8954577B2 patent drawing
  • US8954577B2 patent drawing

AI summary

A method performed by one or more processing devices includes receiving data indicative of amounts of users in population segments that access a resource; applying an expectation-maximization algorithm to the data received; generating, based on applying, estimates of weights indicative of an accuracy of the amounts of users; wherein the expectation-maximization algorithm is applied and the estimates are generated until the estimates reach an asymptotic approximation of the weights; adjusting the amounts of the users in accordance with the estimates of the weights; and generating, based on the amounts of the users adjusted, an estimate of a composition of a population of users that access the resource.