Ontology-Based Classifier for Unstructured Occupational Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Occupational data collected by organizations is often unstructured and inconsistent, making it difficult for analysis and comparison across different organizations due to disparate standards and conventions, and existing methodologies for normalization are inefficient and prone to errors.
Innovation Solution
A method and system for classifying unstructured occupational data using semantic analysis and ontology-based classification systems, which interprets and structures the data to generate enhanced, standardized datasets suitable for computer-based processing, enabling deeper analysis and insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If keyword-based search approach is used for occupational data classification, then the search capability is provided, but the contextual understanding and accuracy are insufficient
Solution Approach 1:
The patent introduces an intermediary layer between keyword search and classification results: a context-aware semantic analysis system that uses ontology-based concept extraction and relationship mapping. This intermediary processes the semantic meaning of occupational data, expanding simple keyword matches into contextually accurate classifications by understanding relationships between occupations, skills, and industries.
Solution Approach 2:
The patent replaces the mechanical keyword-matching system with a semantic analysis system that processes language understanding. Instead of simple string comparison, the system uses natural language processing, ontology reasoning, and contextual analysis to achieve accurate classification, substituting mechanical operations with intelligent processing.
2Measurement precision
If manual classification steps are implemented, then the classification accuracy can be improved, but the processing time and cost increase significantly
Solution Approach 1:
The patent implements a self-service classification system where the computational model automatically performs semantic analysis, concept extraction, and classification without requiring manual intervention. The system serves itself by using trained AI models to process occupational data, eliminating the need for human experts to manually classify each data point while maintaining high accuracy through automated reasoning.
Solution Approach 2:
The patent performs preliminary actions by pre-training the classification model on extensive occupational data and ontology relationships before actual classification tasks. This preliminary training enables the system to quickly and accurately classify new data without manual intervention, as the model has already learned the complex relationships and patterns during the pre-processing phase.
3Adaptability or versatility
If existing normalization methodologies are used, then some standardization is achieved, but the results are ineffective and inefficient
Solution Approach 1:
The patent fundamentally changes the parameters of data normalization by transitioning from rule-based string matching to semantic parameter extraction. Instead of normalizing based on fixed conventions, the system extracts semantic parameters such as occupation concepts, skills, responsibilities, and industry contexts, then normalizes these parameters using ontology-based relationships, achieving both compatibility and efficiency.
Solution Approach 2:
The patent creates a universal classification system that can handle multiple types of occupational data (job titles, descriptions, skills, responsibilities) through a single semantic analysis framework. The ontology-based approach provides multi-functionality by accommodating various data formats and classification standards, making the system adaptable to different organizational needs while maintaining consistent processing efficiency.
4Loss of information
If semantic analysis and ontology-based classification are implemented, then the data structure and meaning are enhanced, but the computational complexity increases
Solution Approach 1:
The patent segments the complex semantic analysis process into distinct modular components: text preprocessing, entity recognition, concept extraction, relationship mapping, and classification. Each module handles a specific aspect of semantic analysis, reducing overall computational complexity by breaking down the monolithic task into manageable segments that can be processed independently and efficiently.
Data Source
AI summary
Disclosed herein are systems and methods for classifying unstructured datasets according to a classification system and generating an enhanced, classified and structured data-set enabling efficient supplemental computer-based processing. The exemplary computer-implemented classification algorithms involve, for each entry in the input dataset, semantically interpreting a text-based occupation description, analyzing the description according to an ontology of interrelated “concepts” and identifying semantically relevant concept(s) and any associated descriptors specific to the classification system. The system is also configured to expand the list of relevant concepts to include concepts that bear a relationship thereto, scoring the various concepts and associated descriptors and identifying the concept(s) and descriptors that most accurately correspond to the input data. Further, the system is configured to generate the new structured and classified occupation dataset by selectively combining certain input data and augmenting each entry with supplemental information inferred through the classification process.


