Machine Learning URL Parsing for Anonymous Identifier Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for documenting and tracking electronic devices using permanent identifiers have raised concerns about user privacy, as these identifiers can be exploited for profiling and marketing purposes without user consent, leading to issues with anonymous identifier management in cellular networks.
Innovation Solution
A machine learning algorithm is employed to verify and identify network devices using anonymous advertising identifiers by parsing Uniform Resource Locators (URLs) and associating potential identifiers with device identifiers, thereby distinguishing actual from unverified identifiers and mitigating malicious traffic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If permanent device identifiers are used to track electronic devices, then device tracking capability is improved, but user privacy is compromised due to profiling and marketing without consent
Solution Approach 1:
The patent introduces an intermediary verification system that sits between device identifiers and advertising networks. This system parses URLs, extracts potential advertising identifiers, and verifies them against a database of known valid identifiers before allowing data sharing. The intermediary prevents direct exposure of user identifiers to marketing companies without verification, thus protecting privacy while enabling legitimate tracking.
Solution Approach 2:
The patent implements a feedback mechanism where the verification system continuously learns from parsed URLs and identified advertising identifiers. By analyzing patterns in valid identifiers and their usage contexts, the system refines its verification capabilities over time, improving accuracy in distinguishing legitimate advertising identifiers from malicious ones while maintaining privacy protection.
2Ease of operation
If advertising identifiers are made changeable by users, then user control and privacy are improved, but device identification reliability deteriorates
Solution Approach 1:
The verification system acts as an intermediary that validates changeable advertising identifiers against a trusted database. Even though users can change their identifiers, the system verifies each identifier's legitimacy by parsing URLs and checking against known valid identifiers, maintaining reliability despite user-controlled changes.
Solution Approach 2:
The system performs preliminary verification of advertising identifiers before they are used for tracking or marketing purposes. By pre-validating identifiers through URL parsing and database verification, the system ensures that even changeable identifiers maintain their reliability for legitimate device identification.
3Measurement precision
If machine learning algorithms are used to verify advertising identifiers, then identifier verification accuracy is improved, but computational complexity and processing time increase
Solution Approach 1:
The verification process is segmented into distinct stages: URL parsing, extraction of potential identifiers, verification against database, and machine learning-based validation. This segmentation allows the system to apply computational resources selectively at each stage, improving accuracy through multiple layers of verification while managing complexity by breaking down the overall process.
Solution Approach 2:
The system applies partial verification to all identifiers (basic parsing and database checking) and reserves the more computationally intensive machine learning verification for cases where partial verification is inconclusive or for high-priority identifiers. This approach maintains high accuracy while reducing overall computational complexity by not applying full verification uniformly to all cases.
Data Source
AI summary
Techniques for identifying certain types of network activity are disclosed, including parsing of a Uniform Resource Locator (URL) to identify a plurality of key-value pairs in a query string of the URL. The plurality of key-value pairs may include one or more potential anonymous identifiers. In an example embodiment, a machine learning algorithm is trained on the URL to determine whether the one or more potential anonymous identifiers are actual anonymous identifiers (i.e., advertising identifiers) that provide advertisers a method to identify a user device without using, for example, a permanent device identifier. In this embodiment, a ranking threshold is used to verify the URL. A verified URL associate the one or more potential anonymous identifiers with the user device as actual anonymous identifiers. Such techniques may be used to identify and eliminate malicious and/or undesirable network traffic.


