Hypothesis Generation from Unstructured Data via Search Engine Intermediary
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning and big data analysis face challenges in utilizing unstructured data from sources like the World Wide Web, as computers struggle to automatically extract relevant information from this messy data, limiting the effectiveness of prediction models and insights.
Innovation Solution
A method that generates and validates hypotheses using unstructured data by querying search engines, where attributes from dataset instances are used to obtain additional information, and new features are defined based on search results, enabling the use of unstructured data in machine learning models without requiring explicit data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If unstructured data from the World Wide Web is used as a data source, then the quantity of available data increases, but the difficulty of detecting and measuring relevant information increases
Solution Approach 1:
The patent introduces search engines as intermediary tools between the unstructured web data and the machine learning system. The search engine automatically queries the web, retrieves relevant information, and structures it into usable features, bridging the gap between messy unstructured data and structured machine learning inputs without requiring manual intervention
Solution Approach 2:
The system enables computers to automatically perform tasks that previously required human researchers. The machine learning system self-services by automatically generating search queries, executing them against search engines, extracting relevant information, and incorporating it as new features, eliminating the need for manual data extraction and annotation
2Loss of information
If manual web search and knowledge attachment is performed, then relevant information can be extracted, but the productivity of the process decreases
Solution Approach 1:
The system enables computers to automatically perform tasks that previously required human researchers. The machine learning system self-services by automatically generating search queries, executing them against search engines, extracting relevant information, and incorporating it as new features, eliminating the need for manual data extraction and annotation
Solution Approach 2:
The patent replaces the mechanical process of manual web searching and information attachment with an automated computational system. Instead of human researchers manually browsing the web and attaching knowledge, the system uses search engines and automated processing to extract and structure information, dramatically increasing productivity
3Reliability
If more properties and examples of entities are collected, then the effectiveness of prediction models improves, but the complexity of data collection and processing increases
Solution Approach 1:
The patent creates a universal data collection mechanism that can automatically gather diverse types of information across different domains through a single integrated system. The search engine-based approach can retrieve various entity properties (revenue, location, industry, etc.) using the same underlying infrastructure, reducing complexity compared to specialized collection methods for each data type
Solution Approach 2:
The search engine acts as an intermediary that handles the complexity of data collection. Instead of building complex specialized systems for each type of entity property, the patent uses the search engine to mediate between the need for diverse data and the simplicity of the collection process, automatically retrieving structured information about entity properties through standardized queries
Data Source
AI summary
Methods, products and apparatus are provided for hypotheses generation using searchable unstructured data corpus. In one method, a query is generated based on at least one attribute of at least one instance in a dataset. The query is provided to a search engine searching in an unstructured data corpus. An hypothesis for the database is based on a new attribute whose value is defined based on the one or more results. Another method comprises obtaining a set of keywords from a plurality of hypotheses extracted from a database. A query is generated based on an attribute of an instance in the dataset, where the attribute corresponds to an hypothesis. A search engine executes the query to provide results which are used to augment an instance with a new attribute, where a value of the new attribute is computed based on the one or more results.


