Hypothesis Generation from Unstructured Data via Search Engine Intermediary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning and big data analysis face challenges in utilizing unstructured data from sources like the World Wide Web, as computers struggle to automatically extract relevant information from this messy data, limiting the effectiveness of prediction models and insights.

Innovation Solution

A method that generates and validates hypotheses using unstructured data by querying search engines, where attributes from dataset instances are used to obtain additional information, and new features are defined based on search results, enabling the use of unstructured data in machine learning models without requiring explicit data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If unstructured data from the World Wide Web is used as a data source, then the quantity of available data increases, but the difficulty of detecting and measuring relevant information increases

Engineering Contradiction:
Improvequantity of dataVSAvoiddifficulty of extracting relevant information
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces search engines as intermediary tools between the unstructured web data and the machine learning system. The search engine automatically queries the web, retrieves relevant information, and structures it into usable features, bridging the gap between messy unstructured data and structured machine learning inputs without requiring manual intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables computers to automatically perform tasks that previously required human researchers. The machine learning system self-services by automatically generating search queries, executing them against search engines, extracting relevant information, and incorporating it as new features, eliminating the need for manual data extraction and annotation

Inventive Principle:
Principle #25Self-service

2Loss of information

If manual web search and knowledge attachment is performed, then relevant information can be extracted, but the productivity of the process decreases

Engineering Contradiction:
Improveextraction of relevant informationVSAvoidproductivity of data extraction
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system enables computers to automatically perform tasks that previously required human researchers. The machine learning system self-services by automatically generating search queries, executing them against search engines, extracting relevant information, and incorporating it as new features, eliminating the need for manual data extraction and annotation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual web searching and information attachment with an automated computational system. Instead of human researchers manually browsing the web and attaching knowledge, the system uses search engines and automated processing to extract and structure information, dramatically increasing productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If more properties and examples of entities are collected, then the effectiveness of prediction models improves, but the complexity of data collection and processing increases

Engineering Contradiction:
Improveeffectiveness of prediction modelVSAvoidcomplexity of data collection system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal data collection mechanism that can automatically gather diverse types of information across different domains through a single integrated system. The search engine-based approach can retrieve various entity properties (revenue, location, industry, etc.) using the same underlying infrastructure, reducing complexity compared to specialized collection methods for each data type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The search engine acts as an intermediary that handles the complexity of data collection. Instead of building complex specialized systems for each type of entity property, the patent uses the search engine to mediate between the need for diverse data and the simplicity of the collection process, automatically retrieving structured information about entity properties through standardized queries

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11182441B2Hypotheses generation using searchable unstructured data corpus
Publication Date: 2021.11.23 SPARKBEYOND
  • US11182441B2 patent drawing
  • US11182441B2 patent drawing
  • US11182441B2 patent drawing

AI summary

Methods, products and apparatus are provided for hypotheses generation using searchable unstructured data corpus. In one method, a query is generated based on at least one attribute of at least one instance in a dataset. The query is provided to a search engine searching in an unstructured data corpus. An hypothesis for the database is based on a new attribute whose value is defined based on the one or more results. Another method comprises obtaining a set of keywords from a plurality of hypotheses extracted from a database. A query is generated based on an attribute of an instance in the dataset, where the attribute corresponds to an hypothesis. A search engine executes the query to provide results which are used to augment an instance with a new attribute, where a value of the new attribute is computed based on the one or more results.