Domain Ontology Extraction from Reference Papers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data analysis systems face inefficiencies in identifying and connecting relevant data and relationships between different theories and ontologies, leading to wasted time and difficulty in interpreting complex scenarios due to unclear interdependencies between variables in domains like behavioral science.
Innovation Solution
A method and system for data extraction that involves collecting and processing reference papers using crawlers, classifying relevant sections, identifying candidate sentences with relation terms, and extracting qualitative and quantitative relations to create domain dictionaries and ontologies, utilizing a relation miner module to analyze and preprocess data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a single database contains many theories and ontologies, then the quantity of available data increases, but it becomes difficult to identify relevant data and relationships between different theories
Solution Approach 1:
The patent segments the large database into multiple ontologies, each representing a specific research area or domain. This segmentation allows researchers to navigate and query specific domains without being overwhelmed by the entire database, while still maintaining the ability to discover relationships across ontologies through standardized schemas and mapping mechanisms.
Solution Approach 2:
The patent introduces an intermediary layer consisting of ontology mapping mechanisms and relationship extraction algorithms that mediate between the stored ontologies and user queries. This intermediary automatically identifies and presents relationships between different ontologies, eliminating the need for users to manually search through all data.
2Reliability
If researchers manually review multiple theories and data sources, then comprehensive understanding may be achieved, but significant time is wasted
Solution Approach 1:
The patent performs preliminary actions by automatically organizing data into structured ontologies, pre-computing relationships between concepts, and indexing data according to multiple schemas before user queries are submitted. This preliminary structuring enables rapid retrieval and comprehensive analysis without requiring manual review of all source materials.
Solution Approach 2:
The system provides feedback to researchers by automatically presenting relevant relationships, conflicts, and connections between theories based on their queries. This feedback mechanism guides researchers through the complex data landscape, highlighting important findings without requiring them to manually examine all underlying data sources.
3Measurement precision
If detailed processing of all reference papers is performed, then extraction accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial processing by focusing computational resources on extracting and validating only the most critical relationship terms and ontology structures from reference papers. Rather than analyzing every sentence in detail, the system identifies key passages and relationships, achieving sufficient accuracy for research purposes while maintaining high processing throughput.
Data Source
AI summary
Method and system to extract domain concepts to create domain dictionaries and ontologies comprises collecting a plurality of reference papers and further classifying the collected plurality of reference papers as relevant and irrelevant. Each of the ‘relevant’ reference papers is further processed by the system, during which the system identifies relevant sections from each document and further processes data in the relevant sections to extract required information and also to identify a relationship between different extracted information, which is further used to create domain dictionaries and ontologies.


