Visual Data Association System for Web Content Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The vast and disorganized nature of information on the internet makes it difficult for researchers and professionals to collect and process relevant data, as information is spread across trillions of webpages in various formats, often including irrelevant content.
Innovation Solution
A data fusion system that transforms data sources into an object model using an ontology and schema map, allowing users to visually define data associations and automatically collect relevant information from documents by identifying and storing objects with specific characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If information is collected from trillions of webpages in different formats, then the quantity of information increases, but the relevance and quality of collected data deteriorates
Solution Approach 1:
The patent extracts only the relevant objects and data from webpages by allowing users to visually select specific elements (images, text, videos) that match desired characteristics. The system parses HTML code to identify and extract only these selected objects, filtering out irrelevant information while collecting data from trillions of webpages.
Solution Approach 2:
The patent applies different selection criteria and characteristics to different parts of the webpage content. Users can specify local characteristics such as object type, size, color, text content, or hierarchical position to selectively extract relevant data from specific regions or elements of webpages, ensuring high relevance while collecting large quantities of information.
2Quantity of substance
If all information from selected webpages is downloaded, then the completeness of data collection improves, but the amount of irrelevant information increases
Solution Approach 1:
Instead of downloading all information from webpages, the system extracts only the specific objects and data elements that users have visually selected and defined as relevant. The HTML parser identifies and extracts only these selected elements, completely filtering out irrelevant information while maintaining completeness of the desired data.
Solution Approach 2:
The system performs partial action by selecting and extracting only the necessary portion of webpage content rather than downloading everything. Users define specific characteristics and select particular objects, causing the system to extract only that partial set of relevant data, avoiding the excessive collection of irrelevant information.
3Adaptability or versatility
If data is collected in different formats from multiple sources, then the versatility of data collection improves, but the complexity of data processing increases
Solution Approach 1:
The system uses a universal HTML parser that can handle multiple data formats and webpage structures through a single unified interface. The visual selection mechanism and characteristic-based filtering work consistently across different webpage formats, sources, and content types, providing multi-functional data collection without requiring separate processing pipelines for each format.
Solution Approach 2:
The system manages complexity by allowing users to define and modify parameters such as object characteristics, selection criteria, and data formats through a visual interface. These parameter changes enable flexible adaptation to different data sources and formats without increasing underlying system complexity, as the HTML parser adjusts its extraction based on the specified parameters.
4Measurement precision
If manual selection of relevant data is performed, then the precision of data selection improves, but the time required for data collection increases
Solution Approach 1:
The system performs preliminary action by pre-defining object characteristics, selection criteria, and filtering rules before actual data collection begins. Users visually select and define the characteristics of desired objects in advance, allowing the HTML parser to automatically and precisely extract matching data from webpages without requiring manual review during the collection process, thus maintaining high precision while reducing time.
Data Source
AI summary
Systems and methods are disclosed for visual definitions of data associations. In accordance with one implementation, a method is provided for visual definitions of data associations. The method includes obtaining and displaying a first sample document, receiving a first input indicating selection of one or more objects within the first sample document, and determining a first set of one or more characteristics shared by the selected objects. The method also includes identifying, within one or more target documents, one or more target objects characterized by the first set of one or more characteristics, and storing object data associated with the target objects.


