Canonical Rule Sets for Cross-Corpus Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems fail to effectively identify data files with common characteristics across various contexts and sub-domains, as they struggle to distinguish between unique and common predictive terms, leading to inefficiencies in data classification and analysis.
Innovation Solution
The development of a system that generates canonical rule sets by comparing cross-corpus and dimensional rule sets, using dimensional differentiators and common nodes to visually distinguish between contexts, allowing for the identification of unique and common predictive terms across multiple contexts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data mining techniques are used to locate and categorize data files, then data classification can be performed, but the system fails to effectively distinguish between unique and common predictive terms across different contexts and sub-domains, leading to poor classification accuracy
Solution Approach 1:
The patent segments the data files into different contexts and sub-domains, allowing the system to analyze and compare predictive terms within specific segments rather than treating all data uniformly. This segmentation enables the system to identify which terms are unique to specific contexts versus those that are common across multiple contexts, thereby improving classification accuracy while preserving contextual information.
Solution Approach 2:
The patent introduces a new dimension of analysis by comparing rule sets across multiple contexts and sub-domains simultaneously. Instead of analyzing data in a single context, the system operates in a multi-dimensional space where contexts and sub-domains are distinguished dimensions, allowing for the identification of predictive terms that vary across these dimensions and improving overall classification precision.
2Adaptability or versatility
If the system analyzes data across multiple contexts and sub-domains, then more comprehensive classification can be achieved, but the complexity of distinguishing between unique and common predictive terms increases
Solution Approach 1:
The patent creates a universal rule set comparison mechanism that can handle multiple contexts and sub-domains through a single integrated framework. The system uses common nodes and dimensional differentiators as universal structures that can represent relationships across any combination of contexts, eliminating the need for separate analysis procedures for each context and reducing overall system complexity despite the multi-context capability.
Solution Approach 2:
The patent generates canonical rule sets by copying and comparing rule structures across different contexts. By creating standardized representations (copies) of rules from multiple contexts and comparing them systematically, the system reduces the complexity of original multi-context analysis into a more manageable canonical form, making the versatile cross-context classification capability more manageable.
3Measurement precision
If the system generates canonical rule sets with visual distinction between dimensional differentiators and common nodes, then predictive term identification improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary actions by pre-identifying and marking dimensional differentiators and common nodes in the rule sets before the final comparison and canonical generation. This preliminary structuring and visual distinction of elements reduces the computational burden during the main comparison process, as the system only needs to process already-organized rule sets rather than creating all distinctions from scratch, thereby improving accuracy while reducing processing time.
Data Source
AI summary
Systems and methods for performing analyses on data sets to display canonical rules sets with dimensional targets are disclosed. A cross-corpus rule set for a given Topic can be generated based on the entire corpus of data. A first dimensional rule set can be generated based on a first context (e.g., based on the same Topic but using a first sub-domain of the corpus of data). A second dimensional rule set can be generated based on a second context (e.g., based on the same Topic but using a second sub-domain of the corpus of data). Key dimensional differentiators (e.g., for each dimension, or context, of the Topic) can be determined based on a comparison of the general rule set, the first dimensional rule set, and the second dimensional rule set. A canonical rule set visualization can be displayed. The visualization can highlight the dimensional selectors (e.g., those tokens, or nodes, that differ between the first dimensional rule set and the second dimensional rule set).


