Web Crawler Verification for Measurement Tag Placement Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for verifying direct page view measurement tag placement on websites are inadequate, particularly for large websites, as they lack robustness and adaptability to various data sources and verification techniques, leading to issues like poor tagging results, missed page views, and incorrect visit counts.
Innovation Solution
A method involving web crawling, using multiple data sources such as clickstream and search result data, to verify measurement code placement by analyzing page view data from both panel and direct measurement sources, identifying variance, and applying threshold-based analysis to ensure accurate tag placement across domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional tag verification methods are used, then verification process is simple, but measurement precision and reliability are insufficient
Solution Approach 1:
The verification process is divided into multiple independent stages: data collection from multiple sources (clickstream, search results, panel data), web crawling and tag detection, variance analysis, and threshold-based verification. Each stage processes specific data types and produces intermediate results that feed into the next stage, enabling comprehensive verification without requiring a single complex verification system.
Solution Approach 2:
The verification system integrates multiple data sources and verification techniques into a single platform that can handle diverse website types and tag placements. The system universally processes panel data, direct measurement data, clickstream data, and search result data through a common variance analysis framework, making it adaptable to different verification scenarios while maintaining consistent precision standards.
2Reliability
If comprehensive verification strategies are implemented, then measurement accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The system applies variance thresholds to determine the level of verification needed for each website. Websites with low variance across multiple data sources require minimal verification, while those exceeding thresholds trigger more intensive crawling and analysis. This partial action approach maintains high reliability for problematic sites while reducing processing time for well-behaved sites.
Solution Approach 2:
The verification system dynamically adjusts processing parameters based on website characteristics and data quality. Threshold values, crawl depths, and sample sizes are modified according to the variance observed in initial data analysis, allowing the system to allocate computational resources efficiently while maintaining consistent reliability standards across different website types and sizes.
3Adaptability or versatility
If multiple data sources are integrated, then verification robustness improves, but system complexity and data processing requirements increase
Solution Approach 1:
The system introduces standardized data interfaces and normalization layers that act as intermediaries between diverse data sources (panel data collectors, web crawlers, search result analyzers) and the core variance analysis engine. These intermediaries translate different data formats into a unified structure, enabling the system to integrate multiple sources without creating unmanageable complexity in the core verification logic.
4Measurement precision
If thorough tag verification is performed across all pages, then measurement accuracy improves, but productivity and coverage speed decrease
Solution Approach 1:
The system performs thorough verification only on pages that exceed variance thresholds or are identified as high-risk through initial sampling. For the majority of pages that show consistent tag placement across multiple data sources, the system accepts the verification results without exhaustive checking, thereby maintaining high measurement precision for problematic pages while preserving overall verification productivity.
Data Source
AI summary
Disclosed herein are strategies for verifying placement of a direct measurement tag useful for measuring Internet traffic of a plurality of users at a website. For example, a method may include receiving web page identification data that is derived from user clickstream data, determining a URL associated with a domain based on the webpage identification information, and providing a measurement code verification web crawler with the URL and the depth to which to explore the domain for verifying measurement code placement with the web crawler.


