Automated Domain Model Extraction from Semi-Structured Documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Building detailed domain models from scratch is time-consuming and cumbersome, especially in complex domains where no single person is an authoritative expert, and existing approaches fail to leverage fragmented domain models in documents created by word processors.
Innovation Solution
The system automatically harvests documents to separate content from presentation, identifies candidate model elements within and across document types, consolidates these to learn a global model, and allows for manual review to resolve domain-specific ambiguities, using indicative words for classification and concept extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If domain models are built from scratch manually, then model quality and accuracy can be ensured, but the time and effort required becomes excessively large
Solution Approach 1:
The system performs preliminary actions by automatically harvesting and analyzing existing documentation before model construction. It pre-processes documents to extract domain concepts, relationships, and terminology, creating a foundation that significantly reduces the time required for subsequent manual model building while maintaining quality through structured extraction methods
Solution Approach 2:
The patent introduces an intermediary automated system that acts as a bridge between existing documentation and the final domain model. This intermediary tooling harvests information from documents, extracts domain concepts and relationships, and presents structured data to modelers, reducing both time and manual effort while preserving model quality
2Manufacturing precision
If domain experts are consulted to build models, then domain accuracy is improved, but the complexity of coordinating multiple experts and locations increases
Solution Approach 1:
The system segments the model building process into independent, automated stages: document harvesting, concept extraction, relationship identification, and model generation. This segmentation allows different teams to work on separate document sets independently, with the system automatically integrating results, thereby reducing coordination complexity while maintaining domain accuracy through systematic processing
Solution Approach 2:
The patent creates copies of domain knowledge from existing documentation across multiple locations and teams. By harvesting and extracting information from various document sources into a unified model structure, it replicates domain expertise without requiring constant expert coordination, reducing complexity while preserving accuracy
3Productivity
If automated document processing is used, then productivity is improved, but the precision of domain concept extraction may deteriorate
Solution Approach 1:
The system incorporates feedback mechanisms where extracted domain concepts and relationships are validated against the source documentation and refined iteratively. Automated processing identifies candidate concepts and relationships, which are then reviewed and refined based on feedback from the documentation context, maintaining high precision while achieving productivity through automation
Solution Approach 2:
The patent replaces manual mechanical analysis of documents with automated text processing and information extraction systems. These systems use computational methods to identify domain concepts, relationships, and terminology with high precision, achieving both productivity improvement through automation and accuracy maintenance through sophisticated NLP techniques
Data Source
AI summary
Systems and associated methods for automated and semi-automated building of domain models for documents are described. Embodiments provide an approach to discover an information model by mining documentation about a particular domain captured in the documents. Embodiments classify the documents into one or more types corresponding to concepts using indicative words, identify candidate model elements (concepts) for document types, identify relationships both within and across document types, and consolidate and learn a global model for the domain.


