Automatic IT Ontology Generation via Document Path Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The generation of reliable IT ontologies is complex, resource-intensive, and requires significant human effort, making it costly and time-consuming.
Innovation Solution
An automatic method that analyzes digital documents to generate ontologies by associating path information with content-derived keywords, using first-level keywords as macro-categories and second-level keywords as sub-categories, reducing the need for human intervention and improving categorization precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual methods are used to create ontologies with skilled personnel defining and associating terms, then reliability and precision of ontology are improved, but loss of time and economic cost increase significantly
Solution Approach 1:
The system automatically generates ontologies by extracting and analyzing terms from digital documents itself, without requiring external skilled personnel. The computer executes algorithms that autonomously identify, extract, and associate terms to create the ontology structure, making the system self-sufficient in ontology creation.
Solution Approach 2:
The patent replaces the manual mechanical process of skilled personnel defining and associating terms with an automated computer-based system. The computer executes structured algorithms that automatically perform term extraction, association, and ontology generation, substituting human intellectual labor with computational processes.
2Reliability
If manual methods are used to create ontologies with skilled personnel defining and associating terms, then reliability and precision of ontology are improved, but economic cost increases significantly
Solution Approach 1:
The system automatically generates ontologies by extracting and analyzing terms from digital documents itself, without requiring external skilled personnel. The computer executes algorithms that autonomously identify, extract, and associate terms to create the ontology structure, making the system self-sufficient in ontology creation.
Solution Approach 2:
The patent replaces the manual mechanical process of skilled personnel defining and associating terms with an automated computer-based system. The computer executes structured algorithms that automatically perform term extraction, association, and ontology generation, substituting human intellectual labor with computational processes.
3Ease of operation
If traditional automatic ontology generation methods are used independently of storage organization, then ease of operation is improved, but manufacturing precision and categorization accuracy deteriorate
Solution Approach 1:
The patent segments the ontology generation process into distinct levels: first-level keywords (macro-categories) derived from path information and second-level keywords (sub-categories) derived from content analysis. This segmentation allows the system to utilize both structural and semantic information for more precise categorization.
Solution Approach 2:
The patent adds a new dimension to ontology generation by incorporating path information from the digital document storage structure. Instead of relying solely on content analysis, the system utilizes the hierarchical path structure as an additional dimension for extracting macro-categories, thereby improving categorization precision.
Data Source
Figure 1
AI summary
The invention is a method (1) for the generation of an IT ontology including the following steps: extracting the string related to the path of the digital document and analysing it in such a way as to identify the identification name of each individual folder of the path; indexing the folder names with an increasing integer numeric value from 1 to n; associating each one of the folder names with value m with the corresponding folder name with value m+1, being 1<= m <= n-1 and being m progressively increased by one unit; associating the folder names with a pre-defined number x of stored sentences present in said document; storing the result of these associations in a third memory allocation (300).