Propensity Score Matching for Unsampled Computing Device Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Service providers face challenges in collecting and analyzing usage data from millions of computing devices due to resource and cost limitations, as well as privacy regulations, leading to inaccurate representation of user experience across a larger group of devices.
Innovation Solution
A system using propensity score matching to identify a sampled computing device that best represents an unsampled device based on configuration data, allowing for the prediction of metrics of interest on unsampled devices without requiring explicit user consent for telemetry data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If telemetry data is collected from a small sampled group of computing devices, then resource and cost limitations are addressed, but the accuracy of representing user experience across the larger population deteriorates
Solution Approach 1:
The patent creates synthetic copies of telemetry data by training a machine learning model on sampled device data and applying it to generate predicted telemetry data for unsampled devices. This copying approach allows the system to estimate user experience metrics for the entire population without actually collecting data from every device, thus maintaining measurement precision while addressing resource constraints
Solution Approach 2:
The patent performs preliminary training of a machine learning model using telemetry data from sampled devices before deployment. This preliminary action creates a predictive model that can be applied to unsampled devices, enabling accurate estimation of user experience metrics without requiring real-time data collection from the entire population
2Measurement precision
If telemetry data is collected from all computing devices, then measurement precision of user experience is improved, but resource and cost limitations are exceeded
Solution Approach 1:
The patent applies partial action by collecting telemetry data from only a sampled subset of devices (those that have opted in) rather than all devices. The machine learning model then extrapolates this partial data to represent the entire population, achieving sufficient measurement precision without the excessive resource cost of universal data collection
3Loss of energy
If a sampling policy is implemented to address resource limitations, then resource and cost constraints are satisfied, but sampling bias occurs leading to inaccurate representation of the population
Solution Approach 1:
The patent incorporates feedback by using configuration data from both sampled and unsampled devices to train and refine the machine learning model. The model learns from the characteristics of unsampled devices and adjusts its predictions accordingly, creating a feedback loop that reduces sampling bias and improves representation accuracy while maintaining resource efficiency
Solution Approach 2:
The patent changes parameters by incorporating additional configuration data features (device type, OS version, location) into the machine learning model training process. These parameter changes enable the model to account for population characteristics and reduce sampling bias, improving reliability without requiring increased resource expenditure
Data Source
AI summary
Disclosed herein is a system for leveraging telemetry data representing usage of a component installed on a group of sampled computing devices to confidently infer the quality of a user experience and/or the behavior of the component (e.g., an operating system) on a larger group of unsampled computing devices. The system is configured to use a propensity score matching approach to identify a sampled computing device that best represents an unsampled computing device using configuration data that is collected from both the sampled and unsampled computing devices. The quality of the user experience and/or the behavior of the component may be captured by a metric of interest (e.g., a QoS value). Accordingly, the system is configured to use the known metric of interest, determined from the telemetry data collected for the sampled computing device, to determine or predict the metric of interest for the unsampled computing device.


