Automated Ontology Generation via UMLS Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing ontologies from structured and semi-structured knowledge sources are inefficient in automatically generating domain-specific relationships, making it difficult to infer and apply reasoning to unstructured medical knowledge for diagnosis-specific questions.
Innovation Solution
A method for automatically creating ontology by identifying primary subjects, extracting concept descriptors, and validating relationships between them using pre-existing information sources like UMLS, to generate a structured ontology that can infer relationships not explicitly stated in the sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated ontology generation from structured and semi-structured knowledge sources is implemented, then productivity in creating domain-specific ontologies is improved, but manufacturing precision in capturing accurate medical relationships deteriorates
Solution Approach 1:
The patent introduces an intermediary validation layer that uses pre-existing authoritative medical knowledge sources (UMLS, SNOMED CT, ICD-10) to verify relationships extracted from unstructured text. This mediator system bridges the gap between automated extraction speed and medical accuracy by filtering and validating extracted relationships against established medical ontologies before incorporation.
Solution Approach 2:
The patent replaces manual ontology construction (mechanical human effort) with an automated NLP-based extraction system. The system uses computational methods to identify relationships in unstructured medical text, replacing the traditional manual curation process while maintaining precision through validation against authoritative sources.
2Quantity of substance
If comprehensive relationship extraction from unstructured medical text is performed, then quantity of extracted relationships is improved, but measurement precision in identifying valid medical relationships deteriorates
Solution Approach 1:
The validation layer acts as an intermediary that processes the large volume of extracted relationships through multiple filtering stages. It uses pattern matching, semantic analysis, and cross-referencing with authoritative sources to distinguish valid medical relationships from spurious extractions, maintaining precision while handling high quantities of data.
Solution Approach 2:
The system performs excessive extraction initially, capturing all potential relationships from unstructured text regardless of validity. This excessive action ensures no valid relationship is missed, followed by a validation filtering process that removes false positives, thereby achieving both high quantity and high precision.
3Reliability
If manual validation of extracted relationships is performed, then reliability of ontology relationships is improved, but loss of time in the ontology generation process increases
Solution Approach 1:
The system implements self-service validation where the extraction process automatically validates its own outputs against pre-loaded authoritative knowledge sources. The validation rules and reference ontologies are pre-configured, enabling the system to autonomously verify relationships without requiring manual expert review for each extraction, thereby maintaining reliability while reducing time loss.
4Ease of operation
If domain-specific ontology generation is automated, then ease of operation in creating medical ontologies is improved, but device complexity in the NLP processing system increases
Solution Approach 1:
The complex NLP system is segmented into distinct functional modules: text preprocessing, relationship extraction, validation filtering, and ontology generation. Each module handles a specific aspect of the process independently, making the overall system more manageable and easier to operate despite the inherent complexity. The modular architecture allows users to interact with discrete functions rather than a monolithic complex system.
Data Source
AI summary
Exemplary embodiments of the present invention relate to a solution for the extraction of information from structured and semi-structured knowledge sources. Further, ontological relationships are inferred between the extracted information by parsing the layout structure of the structured and semi-structured knowledge sources. The inferred ontological relationships are verified in relation to existing ontological sources.


