Enterprise Collective Term Index Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current term and phrase indices are user and application specific, not shared across multiple enterprise applications or users, and are limited to specific file repositories, failing to capture collective knowledge content across an entire enterprise.
Innovation Solution
A method for automatically generating a collective term and phrase index (corporate dictionary) by selecting knowledge elements, deriving term vectors, identifying key terms, extracting n-grams, scoring them based on frequency and probability, and adding them to an index, enabling enterprise-wide sharing and updating.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If term and phrase indices are generated for local content only, then the index is specific to individual users and applications, but the index cannot be shared across multiple enterprise applications or users
Solution Approach 1:
The patent merges multiple local term and phrase indices from different users and applications into a single centralized corporate dictionary. The system collects terms and phrases from various knowledge repositories, applications, and users, then consolidates them into a shared index that can be accessed enterprise-wide, resolving the contradiction between index specificity and shareability.
Solution Approach 2:
The corporate dictionary is designed as a universal index that serves multiple functions across different applications and users. It provides a common vocabulary and search capability that works across diverse enterprise systems, making the index adaptable and versatile while maintaining a single manageable structure.
2Loss of information
If term and phrase indices are limited to specific file repositories, then the index is easy to manage, but it fails to capture collective knowledge content across the entire enterprise
Solution Approach 1:
The system segments the enterprise knowledge base into multiple knowledge repositories while maintaining a unified corporate dictionary. Each repository can be independently managed, but their term and phrase indices are aggregated into the centralized structure, allowing comprehensive knowledge capture without overwhelming complexity in any single component.
Solution Approach 2:
The corporate dictionary acts as an intermediary layer between specific file repositories and the enterprise-wide search functionality. It collects and standardizes terms from various repositories, then provides a unified access point that captures collective knowledge without requiring direct integration of all underlying repositories.
3Manufacturing precision
If n-grams are extracted without probability filtering, then all possible term combinations are captured, but the index includes low-quality or irrelevant terms
Solution Approach 1:
The system performs preliminary probability calculations for adjacent terms before finalizing n-gram extraction. By calculating the probability that terms appear together in meaningful contexts beforehand, the system filters out low-quality term combinations early in the process, ensuring high precision without requiring exhaustive analysis of all possible combinations.
Solution Approach 2:
The system uses probability thresholds as a parameter to control the quality of extracted n-grams. By adjusting the minimum probability threshold, the system can balance between capturing comprehensive term combinations and filtering out irrelevant terms, optimizing both precision and productivity based on specific needs.
Data Source
AI summary
Knowledge automation techniques may include selecting a knowledge element from a knowledge corpus of an enterprise for extraction of n-grams, and deriving a term vector comprising terms in the knowledge element. Based at least on a frequency of occurrence of each term in the knowledge element, key terms are identified in the term vector. Thereafter, the identified key terms are used to extract one or more n-grams from the knowledge element. Each of the extracted n-grams is scored as a function of at least a frequency of occurrence of each of the n-grams across the knowledge corpus of the enterprise, and based on the scoring, one or more of the n-grams is added to a collective term and phrase index.


