Visual Data Association System for Web Content Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The vast and disorganized nature of information on the internet makes it difficult for researchers and professionals to collect and process relevant data, as information is spread across trillions of webpages in various formats, often including irrelevant content.

Innovation Solution

A data fusion system that transforms data sources into an object model using an ontology and schema map, allowing users to visually define data associations and automatically collect relevant information from documents by identifying and storing objects with specific characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If information is collected from trillions of webpages in different formats, then the quantity of information increases, but the relevance and quality of collected data deteriorates

Engineering Contradiction:
Improvequantity of informationVSAvoidrelevance of data
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent extracts only the relevant objects and data from webpages by allowing users to visually select specific elements (images, text, videos) that match desired characteristics. The system parses HTML code to identify and extract only these selected objects, filtering out irrelevant information while collecting data from trillions of webpages.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different selection criteria and characteristics to different parts of the webpage content. Users can specify local characteristics such as object type, size, color, text content, or hierarchical position to selectively extract relevant data from specific regions or elements of webpages, ensuring high relevance while collecting large quantities of information.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If all information from selected webpages is downloaded, then the completeness of data collection improves, but the amount of irrelevant information increases

Engineering Contradiction:
Improvecompleteness of dataVSAvoidirrelevant information
Core Design Contradiction:
Quantity of substanceVSLoss of substance

Solution Approach 1:

Instead of downloading all information from webpages, the system extracts only the specific objects and data elements that users have visually selected and defined as relevant. The HTML parser identifies and extracts only these selected elements, completely filtering out irrelevant information while maintaining completeness of the desired data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial action by selecting and extracting only the necessary portion of webpage content rather than downloading everything. Users define specific characteristics and select particular objects, causing the system to extract only that partial set of relevant data, avoiding the excessive collection of irrelevant information.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If data is collected in different formats from multiple sources, then the versatility of data collection improves, but the complexity of data processing increases

Engineering Contradiction:
Improvedata collection versatilityVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system uses a universal HTML parser that can handle multiple data formats and webpage structures through a single unified interface. The visual selection mechanism and characteristic-based filtering work consistently across different webpage formats, sources, and content types, providing multi-functional data collection without requiring separate processing pipelines for each format.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages complexity by allowing users to define and modify parameters such as object characteristics, selection criteria, and data formats through a visual interface. These parameter changes enable flexible adaptation to different data sources and formats without increasing underlying system complexity, as the HTML parser adjusts its extraction based on the specified parameters.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If manual selection of relevant data is performed, then the precision of data selection improves, but the time required for data collection increases

Engineering Contradiction:
Improveprecision of data selectionVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-defining object characteristics, selection criteria, and filtering rules before actual data collection begins. Users visually select and define the characteristics of desired objects in advance, allowing the HTML parser to automatically and precisely extract matching data from webpages without requiring manual review during the collection process, thus maintaining high precision while reducing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10901583B2Systems and methods for visual definition of data associations
Publication Date: 2021.01.26 PALANTIR TECHNOLOGIES INC
  • US10901583B2 patent drawing
  • US10901583B2 patent drawing
  • US10901583B2 patent drawing

AI summary

Systems and methods are disclosed for visual definitions of data associations. In accordance with one implementation, a method is provided for visual definitions of data associations. The method includes obtaining and displaying a first sample document, receiving a first input indicating selection of one or more objects within the first sample document, and determining a first set of one or more characteristics shared by the selected objects. The method also includes identifying, within one or more target documents, one or more target objects characterized by the first set of one or more characteristics, and storing object data associated with the target objects.