De-biasing Mobile App Usage Data via ML Demographic Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile application usage data is often biased due to a lack of demographic information for users, making it difficult to generate accurate metrics for the true desired population.
Innovation Solution
A de-biasing module utilizing a machine learning model, integrated with a VPN or utility application, collects user attribute data through consented questionnaires or ad targeting criteria, and weights usage data to reflect the demographics of the desired population by comparing aggregate user attributes with census data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a VPN or utility application is used to collect usage data from mobile applications, then the quantity of usage data collected increases, but the demographic information accuracy deteriorates because users of the utility application do not reflect the demographics of the true population
Solution Approach 1:
The patent uses an intermediary machine learning model that acts as a mediator between the collected usage data and the true population demographics. The model is trained on data from users who provided demographic information and applies this learned mapping to infer demographics for users without such information, thereby recovering accurate demographic representation from the biased utility application user base.
Solution Approach 2:
The patent changes the parameter representation by transforming the biased user sample into unbiased population metrics through the machine learning model. The model learns to map from the observable utility application user characteristics to the underlying true population demographic parameters, effectively changing how the data represents the population.
2Measurement precision
If demographic information is collected through questionnaires to improve accuracy, then the measurement precision improves, but the device complexity and user burden increase
Solution Approach 1:
The patent applies partial action by collecting demographic information from only a subset of users through questionnaires, rather than requiring all users to complete surveys. This partial collection is sufficient to train the machine learning model, which then infers demographics for the remaining users, reducing overall complexity while maintaining accuracy.
Solution Approach 2:
The machine learning model serves itself by automatically learning demographic patterns from the questionnaire responses and applying this knowledge to infer demographics for all users. The system becomes self-sufficient in generating demographic information without requiring continuous manual data collection from every user.
3Measurement precision
If the machine learning model is trained on a subset of users with known demographics, then the model accuracy improves, but the loss of information increases for users without demographic data
Solution Approach 1:
The patent creates a copy of the demographic information pattern learned from users with known demographics and applies this copied knowledge to users without demographic data. The machine learning model captures the demographic distribution patterns from the training subset and replicates this understanding across the entire user population, recovering information that would otherwise be lost.
Data Source
AI summary
A utility application for a mobile device inspects data packets from other mobile applications running on the device to gather and record usage data about those applications. Since users of the utility application may not reflect the true population for which the usage data is desired, a system de-biases the data reported from the utility applications using a machine learning model to predict demographics of the users of the utility application. To determine a training data set for the model, the system requests a user to provide a desired user attribute by way of an in-app questionnaire. This enables labeling utility usage data with the demographics, which can be weighted and extrapolated to determine usage across the population as a whole.


