Expectation-Maximization Algorithm for Biased Population Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing population estimation methods, such as the capture-recapture method, face challenges in accurately estimating population sizes due to incomplete data and bias in single-source data, which can lead to inaccurate population composition analysis.
Innovation Solution
The use of an expectation-maximization algorithm to adjust user data from biased sources, generating weighted population composition data that accounts for bias, resulting in more accurate estimates of user populations accessing a resource.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If capture-recapture method or single-source panel approach is used to estimate population size, then population estimation can be performed with limited data, but the accuracy of population composition estimates deteriorates due to incomplete data and bias
Solution Approach 1:
The patent introduces an expectation-maximization (EM) algorithm as an intermediary computational framework that processes biased single-source data. The EM algorithm iteratively refines population composition estimates by alternating between estimating missing data patterns (E-step) and optimizing parameters (M-step), effectively mediating between incomplete observed data and accurate population composition estimates without requiring multiple independent data sources
Solution Approach 2:
The patent applies parameter changes by iteratively adjusting weights and probabilities in the EM algorithm. The algorithm changes parameters (population segment proportions, detection probabilities) across iterations, converging to optimal estimates that account for bias in the single-source data. This iterative parameter refinement enables accurate composition estimates despite incomplete information
2Productivity
If single-source biased data is used for population estimation, then data collection is simplified and faster, but the accuracy of user demographic representation deteriorates
Solution Approach 1:
The EM algorithm serves as a computational intermediary that processes biased single-source data efficiently. It takes the simplified single-source data as input and produces corrected population composition estimates as output, maintaining the productivity advantage of single-source collection while eliminating the accuracy penalty through iterative statistical correction
Solution Approach 2:
The patent implements feedback through the iterative nature of the EM algorithm. Each iteration uses the results from the previous iteration to refine estimates further, with the algorithm continuously feeding corrected estimates back into the calculation process until convergence. This feedback mechanism progressively improves demographic representation accuracy while working with fixed single-source data
Data Source
AI summary
A method performed by one or more processing devices includes receiving data indicative of amounts of users in population segments that access a resource; applying an expectation-maximization algorithm to the data received; generating, based on applying, estimates of weights indicative of an accuracy of the amounts of users; wherein the expectation-maximization algorithm is applied and the estimates are generated until the estimates reach an asymptotic approximation of the weights; adjusting the amounts of the users in accordance with the estimates of the weights; and generating, based on the amounts of the users adjusted, an estimate of a composition of a population of users that access the resource.


