Unified Text Analytics Annotator Combining Rule-Based and Machine Learning Techniques
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional information extraction systems face challenges in developing unified tools for iteratively developing, training, and testing extraction flows, as they are typically independent and disconnected, leading to labor-intensive manual design, training, and maintenance, with ad-hoc connections that complicate scaling and lack unified tooling.
Innovation Solution
A methodology for developing a combined annotator that integrates rule-based and machine learning techniques, using a development tool to create an evaluation dataset, training a machine-learning annotator, and combining it with a rule-based annotator to form a unified annotator that exceeds a predetermined performance threshold, facilitating efficient extraction of target concepts across datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If rule-based and machine learning based information extraction systems are developed independently, then each system can be optimized for its specific approach, but the overall development process becomes labor-intensive and requires manual design, training, and maintenance of separate systems
Solution Approach 1:
The patent combines rule-based annotators and machine learning-based annotators into a single unified annotator system. The development tool integrates both approaches, allowing rules and machine learning models to work together within one system, eliminating the need for separate independent systems and reducing integration complexity.
Solution Approach 2:
The unified annotator system performs multiple functions: it can apply rule-based extraction, machine learning-based extraction, or both simultaneously through a single development tool. This multi-functional approach allows the system to handle diverse information extraction tasks without requiring separate specialized systems.
2Adaptability or versatility
If rule-based and machine learning based systems are connected in an ad-hoc manner, then flexibility in connecting different components is achieved, but scaling becomes complicated and unified tooling is lacking
Solution Approach 1:
The development tool provides a universal platform that handles both rule-based and machine learning-based annotator development within a single interface. This unified tooling maintains flexibility in configuring different extraction approaches while enabling efficient scaling through consistent, standardized processes rather than ad-hoc connections.
3Adaptability or versatility
If manual design, training, and maintenance of information extraction systems is performed, then customization to specific use-cases is achieved, but the process becomes labor-intensive and time-consuming
Solution Approach 1:
The system enables automated development of both rule-based and machine learning-based annotators through the unified development tool. The tool automatically handles dataset development, model training, and annotator creation, reducing manual labor and development time while maintaining customization capabilities through configurable parameters and user input.
Solution Approach 2:
The development tool performs preliminary actions by automatically generating evaluation datasets and training data based on user specifications. This preliminary automation of data preparation and model training significantly reduces the time required for customization while maintaining adaptability to specific use-cases.
4Productivity
If a unified tool for developing both rule-based and machine learning based annotators is created, then development efficiency is improved, but the complexity of the development tool itself increases
Solution Approach 1:
The development tool achieves universality by providing a single integrated platform that handles both rule-based and machine learning-based annotator development. This unified approach improves development efficiency by eliminating the need for separate tools while managing complexity through a standardized, cohesive interface that handles multiple functions within one system.
Data Source
AI summary
One embodiment provides a method for developing a text analytics program for extracting at least one target concept including: utilizing at least one processor to execute computer code that performs the steps of: initiating a development tool that accepts user input to develop rules for extraction of features of the at least one target concept within a dataset comprising textual information; developing, using the rules for feature extraction, an evaluation dataset comprising at least one document annotated with the at least one target concept to be extracted by the text analytics program; creating, using the rules for feature extraction, a rule-based annotator to extract the at least one target concept; training, using the evaluation dataset, a machine-learning annotator to extract the at least one target concept within the dataset; combining the rule-based annotator and the machine learning annotator to form a combined annotator; evaluating, using the evaluation dataset, extraction performance of the combined annotator against a predetermined threshold; and publishing, when the extraction performance of the combined annotator exceeds the predetermined threshold, the combined annotator for use in an application that extracts the at least one target concept from a plurality of datasets.


