Hybrid Ontology Creation via Collocation Extraction and Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of online content and the time-consuming process of developing website ontologies make it difficult for users, publishers, and marketers to efficiently acquire and categorize information, leading to challenges in keyword-based searches and ad placement on websites.
Innovation Solution
Automated systems and methods for creating a combined domain ontology by extracting collocations from webpages, performing semantic and statistical analyses, and aggregating ontologies to determine topics of interest and webpages relevant to users, using techniques like CUR matrix decomposition and WordNet data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual categorization of website content is performed, then accuracy of topic classification is improved, but time consumption and labor effort increase significantly
Solution Approach 1:
The patent segments the ontology creation process into multiple independent components: collocation extraction from webpages, statistical analysis of collocations, semantic analysis using WordNet, and aggregation of results. This segmentation allows automated processing of large volumes of content while maintaining classification accuracy through specialized analysis at each stage.
Solution Approach 2:
The patent introduces collocations as an intermediary concept between raw webpage text and final topic classification. By extracting and analyzing collocations (frequently co-occurring word pairs) as an intermediate step, the system bridges the gap between unstructured text and structured classification, enabling automated processing with high accuracy.
2Adaptability or versatility
If complete ontology coverage is achieved through manual methods, then comprehensiveness of topic hierarchy is improved, but complexity and cost of development increase
Solution Approach 1:
The patent enables the ontology creation system to be self-service by automatically extracting collocations from webpages, performing statistical analysis to identify significant patterns, conducting semantic analysis using WordNet relationships, and aggregating results into a comprehensive ontology hierarchy without requiring manual intervention at each step.
Solution Approach 2:
The patent changes the fundamental parameters of ontology development by transitioning from manual text analysis to automated collocation-based analysis. This parameter change involves using statistical measures of collocation strength and semantic relationships rather than manual categorization, thereby reducing complexity while maintaining or improving comprehensiveness.
3Ease of operation
If traditional keyword search is used, then simplicity of search operation is maintained, but effectiveness of information retrieval deteriorates without topic structure
Solution Approach 1:
The patent performs preliminary action by automatically creating a comprehensive topic ontology hierarchy before the search operation. This pre-established structure organizes all website content into meaningful categories and relationships, enabling subsequent search operations to leverage this structure for improved effectiveness while maintaining user-friendly interfaces.
4Measurement precision
If marketers manually review all webpages for ad placement, then precision of ad placement is improved, but productivity and speed of campaign deployment decrease
Solution Approach 1:
The patent extracts relevant features and topics from webpage content through automated collocation extraction and analysis, identifying key characteristics that determine appropriate ad placement. This extraction process enables automated matching of advertisements to suitable webpages based on extracted topic information, maintaining precision while dramatically improving productivity.
Data Source
AI summary
Systems and methods are discussed to automatically create a domain ontology that is a combination of ontologies. Some embodiments include systems and methods for developing a combined ontology for a website that includes extracting collocations for each webpage within the website, creating first and second ontologies from the collocations, and then aggregating the ontologies into a combined ontology. Some embodiments of the invention include unique ways to calculate collocations, to develop a smaller yet meaningful document sample from a large sample, to determine webpages of interest to users interacting with a website, and to determine topics of interest of users interacting with a website. Various other embodiments of the invention are disclosed.


