Automated Rule Generation for Text Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to efficiently identify data files with a common characteristic from large datasets, as they often require manual intervention and are not effective in automatically categorizing or locating relevant data without human intervention.
Innovation Solution
A system and method that generates a rule set by selecting key terms from a list of data files, evaluating potential rules using term and rule evaluation metrics to determine relevance and applicability, and iteratively adding rules until a stopping criterion is met, allowing for the identification of data files with a common characteristic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual intervention methods are used to identify data files with common characteristics, then accuracy and reliability of identification can be maintained, but productivity and efficiency deteriorate due to time-consuming manual processes
Solution Approach 1:
The system enables automated self-service by allowing the data mining system to independently generate rules, evaluate them against criteria, and identify data files without requiring manual human intervention. The automated rule generation and evaluation process maintains identification accuracy while dramatically improving processing efficiency and productivity.
Solution Approach 2:
The patent replaces manual mechanical processes with automated computer-based systems. The mechanical action of manually reviewing and categorizing data files is substituted by an automated system that uses computational algorithms to generate and evaluate rules, thereby maintaining reliability while enhancing productivity.
2Productivity
If automated rule generation methods are implemented, then productivity and efficiency improve, but device complexity increases due to multiple evaluation metrics and iterative processes
Solution Approach 1:
The system segments the complex rule generation process into distinct manageable components: rule generation module, term evaluation metric module, rule evaluation metric module, and data file identification module. This segmentation reduces overall system complexity by making each component independent and easier to implement while maintaining high productivity through automated operation.
Solution Approach 2:
The system performs preliminary actions by pre-defining evaluation metrics and criteria before the main processing begins. The term evaluation metric and rule evaluation metric are established in advance, allowing the automated system to efficiently process data without requiring complex real-time decision-making, thus reducing operational complexity while maintaining high productivity.
3Measurement precision
If iterative rule evaluation processes are used, then measurement precision and rule relevancy improve, but loss of time increases due to repeated generation and evaluation cycles
Solution Approach 1:
The system implements feedback mechanisms where the rule evaluation metric continuously monitors and provides feedback on rule performance. This feedback allows the system to iteratively refine rules and stop when criteria are met, improving measurement precision while minimizing unnecessary processing time through efficient feedback-driven iteration.
Solution Approach 2:
The system maintains continuity of useful action by continuously generating and evaluating rules until the stopping criterion is satisfied. The iterative process ensures that time is not wasted on unnecessary iterations, as the system efficiently progresses through meaningful evaluation cycles to achieve high precision while optimizing processing time.
Data Source
AI summary
Systems and methods for identifying data files that have a common characteristic are provided. A plurality of data files including one or more data files having a common characteristic are received. A potential rule is generated by selecting key terms from a list that satisfy a term evaluation metric, and the potential rule is evaluated using a rule evaluation metric. The potential rule is added to the rule set if the rule evaluation metric is satisfied. Based upon the potential rule being added to the rule set, data files covered by the potential rule are removed from the plurality of data files. The potential rule generation and evaluation steps are repeated until a stopping criterion is met. After the stopping criterion has been met, the rule set is used to identify other data files having the common characteristic.


