Personal Information Aggregation via Pre-crawled Profile Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack an efficient method to aggregate and identify personal information dispersed across multiple Internet sources accurately, making it difficult to determine the ownership or association of webpages with specific individuals.
Innovation Solution
A distributed personal information aggregation system that receives and matches public data with pre-profiled individuals, using data-matching techniques and confidence scores to associate URLs and information with the correct user profiles, creating a comprehensive directory of online presence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual searching and examination of each webpage source is performed to identify personal information, then completeness of information gathering may be improved, but time consumption and operational complexity increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-crawling and storing information from multiple web sources in a database before user queries are submitted. Personal information, URLs, and associated data are collected and organized in advance, so when a user searches for information about a specific person, the results are immediately available from pre-processed data rather than requiring real-time manual searching of numerous websites.
Solution Approach 2:
The system creates copies of personal information and associated web data by storing them in a centralized database. Instead of requiring users to access and examine original webpages manually, the system maintains replicated copies of the information in structured format, allowing rapid retrieval and comparison without needing to revisit source websites.
2Loss of information
If multiple independent searches are conducted to gather personal information from different sources, then information completeness improves, but system complexity and difficulty of operation increase
Solution Approach 1:
The system merges multiple independent search functions into a single integrated platform. Instead of requiring users to conduct separate searches across different websites and manually compile results, the system combines data from numerous web sources into one unified database that can be queried through a single interface, automatically aggregating personal information from diverse sources.
Solution Approach 2:
The system provides universal access to personal information aggregation through a single interface that handles multiple functions: searching by name, analyzing URLs, comparing data profiles, and displaying comprehensive results. This multi-functional approach eliminates the need for users to navigate multiple independent search systems and perform manual data compilation.
3Loss of information
If data from multiple public sources is aggregated and merged, then comprehensiveness of personal information profile improves, but accuracy and reliability may deteriorate due to data quality variations
Solution Approach 1:
The system replaces manual data verification processes with automated computational methods. Instead of relying on human judgment to assess data quality and authenticity, the system uses algorithms to compare data profiles, analyze patterns, calculate confidence scores, and automatically determine the likelihood that information refers to the same individual, thereby maintaining accuracy through systematic automated analysis.
Solution Approach 2:
The system implements feedback mechanisms by calculating and displaying confidence scores that indicate the reliability of matched information. The comparison engine analyzes data from multiple sources and provides feedback on the strength of matches, allowing users to assess the reliability of aggregated information while the system continuously refines its matching algorithms based on accumulated data patterns.
Data Source
AI summary
A method of aggregating personal information available from public sources over a network. The method includes the steps of receiving at a computer server, data associated with a person, the data being publicly available over a network, and including at least a first name and a last name; using a processor to compare the received data to a plurality of data profiles stored in a database of one or more memory devices, each profile corresponding to a previously-profiled person and containing data associated with the previously-profiled person; determining whether the received data sufficiently matches data associated with the previously-profiled person of the data profile; and merging the received data with the data associated with the previously-profiled person.


