Domain Ontology Extraction from Reference Papers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data analysis systems face inefficiencies in identifying and connecting relevant data and relationships between different theories and ontologies, leading to wasted time and difficulty in interpreting complex scenarios due to unclear interdependencies between variables in domains like behavioral science.

Innovation Solution

A method and system for data extraction that involves collecting and processing reference papers using crawlers, classifying relevant sections, identifying candidate sentences with relation terms, and extracting qualitative and quantitative relations to create domain dictionaries and ontologies, utilizing a relation miner module to analyze and preprocess data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a single database contains many theories and ontologies, then the quantity of available data increases, but it becomes difficult to identify relevant data and relationships between different theories

Engineering Contradiction:
Improvequantity of dataVSAvoiddifficulty of identifying relationships
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the large database into multiple ontologies, each representing a specific research area or domain. This segmentation allows researchers to navigate and query specific domains without being overwhelmed by the entire database, while still maintaining the ability to discover relationships across ontologies through standardized schemas and mapping mechanisms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer consisting of ontology mapping mechanisms and relationship extraction algorithms that mediate between the stored ontologies and user queries. This intermediary automatically identifies and presents relationships between different ontologies, eliminating the need for users to manually search through all data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If researchers manually review multiple theories and data sources, then comprehensive understanding may be achieved, but significant time is wasted

Engineering Contradiction:
Improvecomprehensive understandingVSAvoidtime for data review
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by automatically organizing data into structured ontologies, pre-computing relationships between concepts, and indexing data according to multiple schemas before user queries are submitted. This preliminary structuring enables rapid retrieval and comprehensive analysis without requiring manual review of all source materials.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides feedback to researchers by automatically presenting relevant relationships, conflicts, and connections between theories based on their queries. This feedback mechanism guides researchers through the complex data landscape, highlighting important findings without requiring them to manually examine all underlying data sources.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If detailed processing of all reference papers is performed, then extraction accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies partial processing by focusing computational resources on extracting and validating only the most critical relationship terms and ontology structures from reference papers. Rather than analyzing every sentence in detail, the system identifies key passages and relationships, achieving sufficient accuracy for research purposes while maintaining high processing throughput.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11599580B2Method and system to extract domain concepts to create domain dictionaries and ontologies
Publication Date: 2023.03.07 TATA CONSULTANCY SERVICES LTD
  • US11599580B2 patent drawing
  • US11599580B2 patent drawing
  • US11599580B2 patent drawing

AI summary

Method and system to extract domain concepts to create domain dictionaries and ontologies comprises collecting a plurality of reference papers and further classifying the collected plurality of reference papers as relevant and irrelevant. Each of the ‘relevant’ reference papers is further processed by the system, during which the system identifies relevant sections from each document and further processes data in the relevant sections to extract required information and also to identify a relationship between different extracted information, which is further used to create domain dictionaries and ontologies.