Information Mining Using Domain Taxonomies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining methods are labor-intensive and time-consuming, often requiring significant manual processing to identify insights from large volumes of data, and lack the capability to effectively employ domain-specific knowledge to filter and analyze information.
Innovation Solution
The method involves categorizing documents using a taxonomy, generating a contingency table to identify relationships between categories, and incorporating domain-specific knowledge to refine taxonomies, enabling the identification of representative documents and deeper relationships within the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual processing and googling are used to search for insights, then comprehensive information can be gathered, but the process becomes labor-intensive and time-consuming
Solution Approach 1:
The patent pre-processes and structures information into taxonomies and categories before actual search needs arise. By organizing data into hierarchical structures with defined relationships in advance, the system enables rapid querying and information retrieval without requiring manual browsing or googling at the time of information seeking.
Solution Approach 2:
The patent introduces automated information mining tools and structured taxonomies as intermediaries between the user and the vast amount of available information. These intermediaries automatically process, categorize, and present relevant information, eliminating the need for manual searching while maintaining comprehensive coverage.
2Measurement precision
If manual processing is used to make sense of search results, then detailed analysis can be performed, but significant manual effort is required
Solution Approach 1:
The patent enables the information mining system to perform detailed analysis automatically without requiring manual intervention. The system self-organizes information into structured taxonomies, automatically identifies relationships between categories, and presents analyzed results, thereby maintaining high analysis precision while eliminating manual processing effort.
Solution Approach 2:
The patent replaces manual mechanical processing with automated computational methods. Instead of human analysts manually examining and categorizing information, the system uses automated algorithms to process, structure, and analyze data, achieving the same detailed analysis capability with minimal human effort.
3Adaptability or versatility
If standard text mining techniques are used, then general patterns can be identified, but domain-specific insights are missed
Solution Approach 1:
The patent applies different levels of quality and specificity to different parts of the information structure. While maintaining general taxonomic structures that capture broad patterns, the system incorporates domain-specific terminology, relationships, and categorizations at local levels within the hierarchy, ensuring both general adaptability and domain-specific precision.
Solution Approach 2:
The patent creates dynamic taxonomies that can adapt to different domains and contexts. The structured framework allows for domain-specific customization while maintaining the overall architectural integrity, enabling the system to flexibly adjust to various domains without losing either general pattern recognition capability or domain-specific insights.
Data Source
AI summary
A method and analytics tools for information mining incorporating domain specific knowledge and conceptual structures are disclosed, the method including: providing a first set of documents related to a first topic of interest; using a first taxonomy to categorize the first set of documents into a set of categories; providing a second set of documents related to a second topic of interest; categorizing the second set of documents according to the set of categories of the first set of documents; using an element of domain knowledge to re-categorize the first set of documents; and examining a category to identify a document of interest.


