Synthetic Respondent Data Generation Using Constrained Markov Chains
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audience measurement methods lack detailed respondent-level data, particularly demographics and viewing data, in return path data, which limits the accuracy of media audience size and demographic analysis.
Innovation Solution
The generation of synthetic respondent-level data using constrained Markov chains, where a seed panel is created based on monitored panelists' data and adjusted to satisfy target reach constraints, generating viewing data that corresponds to return path data to enhance the value of return path data for advertisers and media providers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audience measurement methods are used, then media consumers can be enlisted to provide demographic data, but the data collection process is time-consuming and requires manual cooperation from panelists
Solution Approach 1:
The patent creates synthetic panelists that replicate the viewing behavior patterns of real panelists. These synthetic panelists are generated by copying and extrapolating from the viewing data of actual panelists, allowing the system to maintain accurate demographic representation without requiring continuous manual participation from real individuals.
Solution Approach 2:
The patent replaces manual data collection processes with automated computer-generated synthetic panelists. Instead of relying on mechanical processes like manual surveys and direct observer cooperation, the system uses computational algorithms to generate and manage panelist data automatically, significantly reducing time requirements.
2Loss of information
If return path data is collected from media presentation devices, then tuning data can be obtained, but detailed respondent-level demographic information is missing
Solution Approach 1:
The patent introduces synthetic panelists as intermediary entities that bridge the gap between aggregate return path data and detailed demographic information. These synthetic panelists serve as mediators that can be assigned demographic characteristics and viewing behaviors, allowing the system to infer respondent-level data from device-level measurements without directly collecting sensitive personal information.
Solution Approach 2:
The patent segments the audience measurement process into distinct components: return path data collection from devices, synthetic panelist generation, and data assignment. This segmentation allows each component to be optimized independently, with return path data providing the structural framework and synthetic panelists filling in the demographic details through computational assignment.
3Measurement precision
If synthetic panelists are generated to represent audiences, then detailed demographic data can be created, but the generation process requires complex algorithms and computing resources
Solution Approach 1:
The patent changes the parameters of panelist representation by generating synthetic panelists with adjusted demographic and viewing characteristics. These synthetic panelists are created by modifying and extrapolating from real panelist data, allowing the system to achieve high measurement precision while managing computational resources through efficient parameter transformation rather than brute-force processing.
4Duration of action of moving object
If panelists are monitored over extended periods, then comprehensive viewing data can be collected, but the cooperation and retention of panelists becomes difficult
Solution Approach 1:
The patent creates synthetic panelists that copy and replicate the viewing behavior patterns of real panelists. These synthetic representations can be monitored indefinitely without requiring human cooperation, as they are generated through computational processes that extrapolate from initial panelist data, eliminating the retention and cooperation challenges associated with long-term human panel studies.
Data Source
AI summary
Methods, apparatus, systems, and articles of manufacture are disclosed to generate synthetic respondent level data. Example apparatus disclosed herein include means for generating a synthetic panel corresponding to a duration of time, the means for generating the synthetic panel to: generate a transition matrix corresponding to a first sub-duration of the duration of time and a second sub-duration of the duration of time; generate, based on the transition matrix, a plurality of synthetic panelists and associated viewing data; remove first ones of the synthetic panelists associated with one or more weights that do not satisfy a threshold to generate the synthetic panel corresponding to the duration of time, the synthetic panel representative of audiences of media presented by a plurality of media devices during the duration of time; and generate synthetic respondent level data based on the viewing data associated with remaining second ones of the synthetic panelists.


