Anonymous Log Entry Generation via Data Field Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Search engines collect user-specific data that raises privacy concerns, as it can be used to track user behavior and compromise security, while also being necessary for improving search results and delivering targeted advertisements.
Innovation Solution
Generating anonymous log entries by adding non-user-specific data fields and deleting user-specific data fields from original log entries, ensuring complete user privacy and allowing analysis for improving search results without tracking individual users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If user-specific data fields are retained in log entries, then search engine can improve search results quality and deliver targeted advertisements, but user privacy is compromised and security risks increase
Solution Approach 1:
The patent extracts and removes user-specific data fields (IP addresses, cookie identifications, device identifiers) from log entries while retaining non-user-specific data fields (search queries, timestamps, result counts). This extraction principle resolves the contradiction by eliminating the harmful tracking capability while preserving the useful search analysis functionality.
Solution Approach 2:
The patent changes the parameter state of log entries by transforming identifiable user data into anonymized data. Specifically, it removes parameters that uniquely identify users (IP addresses, cookies) while maintaining parameters that describe search behavior patterns. This parameter transformation allows the system to maintain productivity for search improvement while eliminating privacy harms.
2Object-affected harmful factors
If user-specific data fields are deleted from log entries, then user privacy is protected, but the ability to link search behaviors of the same user together is lost
Solution Approach 1:
The patent selectively extracts only the harmful user-identifying components (IP addresses, cookie IDs) while preserving the useful search behavior data (queries, timestamps, result metadata). This selective extraction maintains privacy protection while avoiding loss of information needed for analyzing search patterns and improving results.
Solution Approach 2:
The patent segments log entry data into two distinct categories: user-specific identifiers (removed for privacy) and search behavior data (retained for analysis). This segmentation allows the system to protect user privacy by removing identifying segments while maintaining the informational segments needed for improving search results and delivering relevant advertisements.
3Object-affected harmful factors
If anonymous log entries are generated by adding and deleting data fields, then complete user anonymity is achieved, but processing complexity increases
Solution Approach 1:
The patent applies a straightforward extraction process that automatically identifies and removes user-specific data fields based on predefined criteria (IP address, cookie identification, device identifiers). This automated extraction approach achieves complete user anonymity through systematic field removal without requiring complex manual processing or sophisticated algorithms.
Solution Approach 2:
The patent implements parameter changes by systematically removing specific data fields from log entries according to a defined anonymization rule set. This parameter transformation approach achieves user anonymity through consistent application of removal criteria, maintaining processing simplicity while ensuring complete anonymity of user identifiers.
Data Source
AI summary
Assigning session identifications to log entries and generating anonymous log entries are provided. In order to balance users' privacy concerns with the need for analysis of the log entries to provide high quality search results, non-user-specific data fields, such as a user's location (e.g., city, state, and latitude/longitude) and connection speed, are inserted into the log entries, and user-specific data fields, such as the IP address and cookie identifications, are deleted from the log entries. In addition or alternatively, prior to anonymization of the log entries, session identifications are assigned to identified groups of log entries. The groups are identified based on factors such as the user's identification, the IP address, the time of search, and differences between the search terms used in the search queries.


